| Literature DB >> 25886991 |
Sang He1, Yusheng Zhao2, M Florian Mette3, Reiner Bothe4, Erhard Ebmeyer5, Timothy F Sharbel6, Jochen C Reif7, Yong Jiang8.
Abstract
BACKGROUND: The main goal of our study was to investigate the implementation, prospects, and limits of marker imputation for quantitative genetic studies contrasting map-independent and map-dependent algorithms. We used a diversity panel consisting of 372 European elite wheat (Triticum aestivum L.) varieties, which had been genotyped with SNP arrays, and performed intensive simulation studies.Entities:
Mesh:
Substances:
Year: 2015 PMID: 25886991 PMCID: PMC4364688 DOI: 10.1186/s12864-015-1366-y
Source DB: PubMed Journal: BMC Genomics ISSN: 1471-2164 Impact factor: 3.969
Accuracies of imputing measured as average correlations (cor) between observed and estimated marker genotypes
|
|
|
|
|
|
|---|---|---|---|---|
|
|
|
|
| |
| Low to high marker density | ||||
| Beagle | 0.61 | 0.70 | 0.75 | 0.78 |
| FImpute | 0.68 | 0.73 | 0.77 | 0.80 |
| IMPUTE2 | 0.74 | 0.77 | 0.81 | 0.84 |
| Random Forest | 0.56 | 0.61 | 0.66 | 0.69 |
| Genotyping-by-sequencing-like | ||||
| Beagle | 0.76 | 0.85 | 0.92 | 0.95 |
| FImpute | 0.59 | 0.79 | 0.91 | 0.95 |
| IMPUTE2 | 0.68 | 0.82 | 0.91 | 0.95 |
| Random Forest | 0.54 | 0.64 | 0.75 | 0.83 |
Map- dependent (Beagle, FImpute, and IMPUTE2) and map-independent (Random Forest) algorithms were applied with reference population sizes of 50, 100, 200, and 300 lines out of 371, and imputing was performed for a low to high marker density and for a GBS-like data scenario.
*For GBS-like imputation scenarios, Ref 50, Ref 100, Ref 200, and Ref 300 refer to missing value rates 72.8%; 61.5%; 38.8%; 16.1% for all lines of the population, corresponding to scenarios with reference population sizes of 50, 100, 200, and 300, of the total of 371 lines.
Figure 1Linkage disequilibrium influences the accuracy of imputing missing values. The relationship between linkage disequilibrium (as measured by r2 between 90 k SNPs and the respective most closely linked 9 k SNPs) and the average correlation between observed and imputed genotypic data, as calculated using map-dependent (Beagle, FImpute, and IMPUTE2) and map-independent (Random Forest) imputation algorithms, for a reference population size of 50 out of 371 lines. Trends are shown as boxplot displays separately for three minor allele frequency (MAF) classes.
Correlations between Rogers’ distance matrices of the individual lines of the test population
|
|
|
|
|
|
|---|---|---|---|---|
|
|
|
|
| |
| 9 k panel | 0.95 | 0.95 | 0.95 | 0.95 |
| Beagle | 0.83 | 0.92 | 0.95 | 0.96 |
| FImpute | 0.95 | 0.96 | 0.97 | 0.97 |
| IMPUTE2 | 0.96 | 0.97 | 0.98 | 0.98 |
| Random Forest | 0.61 | 0.61 | 0.61 | 0.66 |
Estimates are based solely on imputed parts of data sets (90 k SNP minus 9 k SNP data) and the original 90 k SNP data set, as well as the correlation between Rogers’ distance matrices of the original 9 k and original 90 k SNP data sets. Different imputed low to high marker density data sets were generated by map- dependent (Beagle, FImpute, and IMPUTE2) and map-independent (Random Forest) imputation algorithms for reference populations of 50, 100, 200, and 300 out of 371 lines. All correlations were significantly larger than zero (P < 0.01) according to a Mantel test.
Figure 2Imputing from low to high density has a limited effect on the accuracy of genomic selection. Correlation between results of genomic selection based on true and predicted genotypic values applying genomic selection for the original 90 k (Total-90 k) and 9 k SNP data sets (Total-9 k) for the total 371 lines, as well as for imputing low to high density marker data applying map-dependent (Beagle, FImpute, and IMPUTE2) and map-independent (Random Forest) imputation algorithms, for reference population sizes 50, 100, 200, and 300 out of 371 lines.
Figure 3Imputing improves accuracy of prediction of genomic selection based on GBS-like data sets. Correlation between true and predicted genotypic values applying genomic selection for the original 90 k SNP data set (GBLUP, 0%), as well as for genotype-by-sequencing-like data sets where missing values were not imputed (GBLUP) or were imputed with map-dependent (Beagle, FImpute and IMPUTE2) and map-independent (Random Forest) algorithms for rates of missing values of 72.8%, 61.5%, 38.8% and 16.1% in the total population of 371 lines.
Figure 4Imputing from low to high marker density increases the power of association mapping. Detection frequency of a major QTL explaining 10% of the genotypic variance in the total population based on the 90 k SNP data set (Total-90 k) and the SNP present in the 9 k data set most closely linked to the SNP in the 90 k data set (Total-9 k). Detection frequency of a major QTL in the reference population with sizes from 50 to 300 individuals fingerprinted with the 90 k SNP array (Reference-90 k). Detection frequency of major QTL with varying degrees of linkage disequilibrium (r2) between the QTL and closest linked 9 k SNP marker estimated in the total panel for which depleted 90 k SNP array data have been imputed for the test population with map- dependent (Beagle, FImpute, and IMPUTE2) and map-independent (Random Forest) algorithms for the reference population sizes of 50, 100, 200, and 300 out of 371 lines.
Figure 5Strategy for molecular breeding of wheat. We suggest combining high-density genotyping of a core set of wheat lines with low density and low cost fingerprinting of the entire breeding population followed by imputing of missing marker data for the subsequent quantitative genetic analyses.