| Literature DB >> 28381821 |
Hongyan Cao1, Zhi Li2, Haitao Yang3, Yuehua Cui4,5, Yanbo Zhang6.
Abstract
Longitudinal genetic data provide more information regarding genetic effects over time compared with cross-sectional data. Coupled with next-generation sequencing technologies, it becomes reality to identify important genes containing both rare and common variants in a longitudinal design. In this work, we adopted a weighted sum statistic (WSS) to collapse multiple variants in a gene region to form a gene score. When multiple genes in a pathway were considered together, a penalized longitudinal model under the quadratic inference function (QIF) framework was applied for efficient gene selection. We evaluated the estimation accuracy and model selection performance under different model settings, then applied the method to a real dataset from the Genetic Analysis Workshop 18 (GAW18). Compared with the unpenalized QIF method, the penalized QIF (pQIF) method achieved better estimation accuracy and higher selection efficiency. The pQIF remained optimal even when the working correlation structure was mis-specified. The real data analysis identified one important gene, angiotensin II receptor type 1 (AGTR1), in the Ca2+/AT-IIR/α-AR signaling pathway. The estimated effect implied that AGTR1 may have a protective effect for hypertension. Our pQIF method provides a general tool for longitudinal sequencing studies involving large numbers of genetic variants.Entities:
Mesh:
Year: 2017 PMID: 28381821 PMCID: PMC5429681 DOI: 10.1038/s41598-017-00712-9
Source DB: PubMed Journal: Sci Rep ISSN: 2045-2322 Impact factor: 4.379
Figure 1Performance of the pQIF for different sample sizes and different dimensions. (a) p = 20, (b) p = 40. The horizontal axis represents the variables, where 1 represents covariate age, 2–4 represent the three gene variables (MAP4, TNN, and NRF1) when p = 20, and 2–6 represent the five gene variables (MAP4, TNN, LEPR, FLT3, and NRF1) when p = 40, others represent the noise variables. The true and working correlation structures were set as AR(1). The title of each subfigure (e.g., “n = 142” in the top left panel) refers to the sample size. Since the pQIF did not converge well for n = 142, p = 40 in some simulations runs, the estimation results were not listed in the figure.
Estimation accuracy of parameters for the QIF and pQIF under different model conditions.
| Sample | Method |
|
| ||||||
|---|---|---|---|---|---|---|---|---|---|
|
|
|
|
| ||||||
| TMSEa | NMSEb | TMSE | NMSE | TMSE | NMSE | TMSE | NMSE | ||
|
| oQIFc | 0.0068 | — | 0.0093 | — | — | — | — | — |
| QIF | 0.0711 | 0.0523 | 0.2458 | 0.1837 | — | — | — | — | |
| pQIF | 0.0191 | 0.0033 | 0.0329 | 0.0056 | — | — | — | — | |
|
| oQIF | 0.0029 | — | 0.0043 | — | 0.0029 | — | 0.0035 | — |
| QIF | 0.0211 | 0.0191 | 0.0305 | 0.0277 | 0.0659 | 0.0462 | 0.301 | 0.2033 | |
| pQIF | 0.0062 | 0.0022 | 0.0098 | 0.0041 | 0.0067 | 0.0020 | 0.0123 | 0.0028 | |
|
| oQIF | 0.0015 | — | 0.0022 | — | 0.0012 | — | 0.0017 | — |
| QIF | 0.0076 | 0.0069 | 0.0110 | 0.0099 | 0.0123 | 0.0102 | 0.018 | 0.0149 | |
| pQIF | 0.0023 | 0.0009 | 0.0035 | 0.0015 | 0.0021 | 0.0010 | 0.0036 | 0.0020 | |
Since the pQIF did not converge well for n = 142, p = 40 in some simulations runs, the estimation results were not listed in the table. The notation “—” indicates that the results were not available for n = 142, p = 40, and so does for the NMSE of oQIF.
aTMSE = total mean squared error for all the variables in the model.
bNMSE = mean squared error for all the noisy gene variants in the model.
coQIF = Oracle QIF.
Figure 2Model selection performance of the pQIF under three different working correlation structures. The true correlation structures are assumed as (a) AR(1) structure, (b) EXCH (exchangeable) structure, (c) INDEP (independent) structure. AR(1), EXCH (exchangeable), and INDEP (independent) in each sub-part are the working correlation structures. The bootstrapped sample size is 300, p=20, and ρ=0.5. The horizontal axis represents the variables, where 1 represents covariate age and 2–4 represent the three gene variables (MAP4, TNN, and NRF1).
Estimation accuracy of the pQIF method under three types of working correlations, AR(1), EXCH (exchangeable), INDEP (independent).
| True correlation | Working correlation | MSE1a | NMSEb | TMSEc |
|---|---|---|---|---|
| AR(1) | AR(1) | 0.0182 | 0.002 | 0.0052 |
| EXCH | 0.0225 | 0.0027 | 0.0067 | |
| INDEP | 0.0164 | 0.0018 | 0.0047 | |
| EXCH | AR(1) | 0.0215 | 0.0025 | 0.0063 |
| EXCH | 0.0260 | 0.0029 | 0.0075 | |
| INDEP | 0.0183 | 0.0019 | 0.0052 | |
| INDEP | AR(1) | 0.0090 | 0.0002 | 0.0020 |
| EXCH | 0.0097 | 0.0003 | 0.0022 | |
| INDEP | 0.0078 | 0.0001 | 0.0017 |
aMSE1 = mean squared error of the four nonzero coefficients for age, MAP4, TNN, and NRF1.
bNMSE = mean squared error for all the noisy gene variants in the model.
cTMSE = total mean squared error for all the variables in the model.
Distribution of age, sex, smoking, and hypertension in the GAW18 real dataset at different exam stages.
| Variables | values | Exam 1 | Exam 2 | Exam 3 |
|---|---|---|---|---|
| Age (years) | (20.3~96.72) | 53.30 ± 15.91a | 56.38 ± 12.96 | 58.24 ± 11.89 |
| Sex | 1 = Male, 2 = Female | 61:81b | 61:81 | 61:81 |
| Smoking | 1 = Smoking, 0 = Non-smoking | 33:98 | 14:67 | 14:78 |
| Hypertension | 1 = Hypertension, 0 = Non-Hypertension | 41:92 | 49:40 | 50:42 |
a for age in each exam, where is the average and S is the standard deviation.
bRatios for male: female, smoking: non-smoking, and hypertension: non-hypotension.
The coefficients estimated by the QIF and pQIF methods.
| Variables | QIF (S.E)a | pQIF |
|---|---|---|
| Intercept | 3.438 (1.458) | 0.054 |
| Age | 6.903 (2.391) | 2.199 |
| Gender | 2.829 (1.082) | — |
| Smoking | 1.170 (0.861) | — |
| AGTR1 | −4.546 (1.697) | −0.889 |
| ADRA1B | −0.916 (0.627) | — |
| GNAQ | −3.311 (1.335) | — |
| PLCB3 | −0.008 (0.611) | — |
| PRKCA | −1.598 (0.953) | — |
| PPP1CA | 1.906 (0.857) | — |
| CAMK2A | −0.535 (0.532) | — |
| CALM3 | −1.593 (0.690) | — |
| RYR2 | −0.883 (0.645) | — |
| PPP2CA | 4.688 (1.826) | — |
| CREB3L2 | −0.240 (0.367) | — |
AGTR1, ADRA1B, GNAQ, PLCB3, PRKCA, PPP1CA, CAMK2A, CALM3, RYR2, PPP2CA, and CREB3L2 were genes in the Ca2+/AT-IIR/α-AR signaling pathway. Other covariates in the analysis were age, gender, and tobacco smoking. The notation “—” indicates that the coefficient of the related variable was penalized to 0.
aS.E = standard error of the coefficient estimate.