Literature DB >> 22686303

Imputing missing genotypes with weighted k nearest neighbors.

Holger Schwender1.   

Abstract

Missing values are a common problem in genetic association studies concerned with single-nucleotide polymorphisms (SNPs). Since many statistical methods cannot handle missing values, such values need to be removed prior to the actual analysis. Considering only complete observations, however, often leads to an immense loss of information. Therefore, procedures are required that can be used to impute such missing values. In this study, an imputation procedure based on a weighted k nearest neighbors algorithm is presented. This approach, called KNNcatImpute, searches for the k SNPs that are most similar to the SNP whose missing values need to be replaced and uses these k SNPs to impute the missing values. Alternatively, KNNcatImpute can search for the k nearest subjects. In this situation, the missing values of an individual are imputed by considering subjects showing a DNA pattern similar to the one of this individual. In a comparison to other imputation approaches, KNNcatImpute shows the lowest rates of falsely imputed genotypes when applied to the SNP data from the GENICA study, a candidate SNP study dedicated to the identification of genetic and gene-environment interactions associated with sporadic breast cancer. Moreover, KNNcatImpute can also be applied to data from genome-wide association studies, as an application to a subset of the HapMap data demonstrates.

Entities:  

Mesh:

Substances:

Year:  2012        PMID: 22686303     DOI: 10.1080/15287394.2012.674910

Source DB:  PubMed          Journal:  J Toxicol Environ Health A        ISSN: 0098-4108


  14 in total

1.  Efficient imputation of missing markers in low-coverage genotyping-by-sequencing data from multiparental crosses.

Authors:  B Emma Huang; Chitra Raghavan; Ramil Mauleon; Karl W Broman; Hei Leung
Journal:  Genetics       Date:  2014-02-28       Impact factor: 4.562

2.  KLFDAPC: a supervised machine learning approach for spatial genetic structure analysis.

Authors:  Xinghu Qin; Charleston W K Chiang; Oscar E Gaggiotti
Journal:  Brief Bioinform       Date:  2022-07-18       Impact factor: 13.994

3.  A Multi-Marker Genetic Association Test Based on the Rasch Model Applied to Alzheimer's Disease.

Authors:  Wenjia Wang; Jonas Mandel; Jan Bouaziz; Daniel Commenges; Serguei Nabirotchkine; Ilya Chumakov; Daniel Cohen; Mickaël Guedj
Journal:  PLoS One       Date:  2015-09-17       Impact factor: 3.240

4.  LinkImpute: Fast and Accurate Genotype Imputation for Nonmodel Organisms.

Authors:  Daniel Money; Kyle Gardner; Zoë Migicovsky; Heidi Schwaninger; Gan-Yuan Zhong; Sean Myles
Journal:  G3 (Bethesda)       Date:  2015-09-15       Impact factor: 3.154

5.  Genome-wide association mapping of partial resistance to Aphanomyces euteiches in pea.

Authors:  Aurore Desgroux; Virginie L'Anthoëne; Martine Roux-Duparque; Jean-Philippe Rivière; Grégoire Aubert; Nadim Tayeh; Anne Moussart; Pierre Mangin; Pierrick Vetel; Christophe Piriou; Rebecca J McGee; Clarice J Coyne; Judith Burstin; Alain Baranger; Maria Manzanares-Dauleux; Virginie Bourion; Marie-Laure Pilet-Nayel
Journal:  BMC Genomics       Date:  2016-02-20       Impact factor: 3.969

6.  Prenatal Glucocorticoid Exposure Modifies Endocrine Function and Behaviour for 3 Generations Following Maternal and Paternal Transmission.

Authors:  Vasilis G Moisiadis; Andrea Constantinof; Alisa Kostaki; Moshe Szyf; Stephen G Matthews
Journal:  Sci Rep       Date:  2017-09-18       Impact factor: 4.379

7.  Effect of Co-segregating Markers on High-Density Genetic Maps and Prediction of Map Expansion Using Machine Learning Algorithms.

Authors:  Amidou N'Diaye; Jemanesh K Haile; D Brian Fowler; Karim Ammar; Curtis J Pozniak
Journal:  Front Plant Sci       Date:  2017-08-23       Impact factor: 5.753

8.  Robustification of GWAS to explore effective SNPs addressing the challenges of hidden population stratification and polygenic effects.

Authors:  Zobaer Akond; Md Asif Ahsan; Munirul Alam; Md Nurul Haque Mollah
Journal:  Sci Rep       Date:  2021-06-22       Impact factor: 4.379

9.  Large-scale multiple testing in genome-wide association studies via region-specific hidden Markov models.

Authors:  Jian Xiao; Wensheng Zhu; Jianhua Guo
Journal:  BMC Bioinformatics       Date:  2013-09-25       Impact factor: 3.169

10.  Association Mapping in Scandinavian Winter Wheat for Yield, Plant Height, and Traits Important for Second-Generation Bioethanol Production.

Authors:  Andrea Bellucci; Anna Maria Torp; Sander Bruun; Jakob Magid; Sven B Andersen; Søren K Rasmussen
Journal:  Front Plant Sci       Date:  2015-11-26       Impact factor: 5.753

View more

北京卡尤迪生物科技股份有限公司 © 2022-2023.