Han Zhang1, Hsin-Chou Yang, Yaning Yang. 1. Department of Statistics and Finance, University of Science and Technology of China, Anhui, Taiwan. zhanghan@mail.ustc.edu.cn
Abstract
MOTIVATION: Pooling DNA is a cost-effective alternative to individual genotyping method. It is often used for initial screening in genome-wide association analysis. In some studies, large pools with sizes up to several hundreds were applied in order to significantly reduce genotyping cost. However, method for estimating haplotype frequencies from large DNA pools has not been available due to computational complexity involved. METHODS: We propose a novel constrained EM algorithm, PoooL, to estimate frequencies of single-nucleotide polymorphism (SNP) haplotypes from DNA pools. A quantity called importance factor is introduced to measure the contribution of a haplotype to the likelihood. Under the assumption of asymptotic normality of the estimated allele frequencies and a system of linear constraints on haplotype frequencies the importance factor remains a constant in the iterative maximization process. The maximization problem in the EM algorithm is then formulated into a constrained maximum entropy model and solved by the improved iterative scaling method. RESULTS: Simulation study shows that our algorithm can efficiently estimate haplotype frequencies from DNA pools with arbitrarily large sizes. The algorithm works equally well for large pools with sizes up to hundreds or thousands and for pools with sizes as small as one or two individuals. The computational complexity of the PoooL algorithm is independent of pool sizes, and the computational efficiency for large pools is thus substantially improved over existing estimating methods. Simulation results also show that the proposed method is robust to genotype errors and population admixture.
MOTIVATION: Pooling DNA is a cost-effective alternative to individual genotyping method. It is often used for initial screening in genome-wide association analysis. In some studies, large pools with sizes up to several hundreds were applied in order to significantly reduce genotyping cost. However, method for estimating haplotype frequencies from large DNA pools has not been available due to computational complexity involved. METHODS: We propose a novel constrained EM algorithm, PoooL, to estimate frequencies of single-nucleotide polymorphism (SNP) haplotypes from DNA pools. A quantity called importance factor is introduced to measure the contribution of a haplotype to the likelihood. Under the assumption of asymptotic normality of the estimated allele frequencies and a system of linear constraints on haplotype frequencies the importance factor remains a constant in the iterative maximization process. The maximization problem in the EM algorithm is then formulated into a constrained maximum entropy model and solved by the improved iterative scaling method. RESULTS: Simulation study shows that our algorithm can efficiently estimate haplotype frequencies from DNA pools with arbitrarily large sizes. The algorithm works equally well for large pools with sizes up to hundreds or thousands and for pools with sizes as small as one or two individuals. The computational complexity of the PoooL algorithm is independent of pool sizes, and the computational efficiency for large pools is thus substantially improved over existing estimating methods. Simulation results also show that the proposed method is robust to genotype errors and population admixture.
Authors: Jamie E Craig; Alex W Hewitt; Amy E McMellon; Anjali K Henders; Lingjun Ma; Leanne Wallace; Shiwani Sharma; Kathryn P Burdon; Peter M Visscher; Grant W Montgomery; Stuart MacGregor Journal: Genome Res Date: 2009-10-03 Impact factor: 9.043
Authors: Charleston W K Chiang; Zofia K Z Gajdos; Joshua M Korn; Johannah L Butler; Rachel Hackett; Candace Guiducci; Thutrang T Nguyen; Rainford Wilks; Terrence Forrester; Katherine D Henderson; Loic Le Marchand; Brian E Henderson; Christopher A Haiman; Richard S Cooper; Helen N Lyon; Xiaofeng Zhu; Colin A McKenzie; Mark R Palmert; Joel N Hirschhorn Journal: Hum Genet Date: 2011-03-22 Impact factor: 4.132
Authors: Charleston W K Chiang; Zofia K Z Gajdos; Joshua M Korn; Finny G Kuruvilla; Johannah L Butler; Rachel Hackett; Candace Guiducci; Thutrang T Nguyen; Rainford Wilks; Terrence Forrester; Christopher A Haiman; Katherine D Henderson; Loic Le Marchand; Brian E Henderson; Mark R Palmert; Colin A McKenzie; Helen N Lyon; Richard S Cooper; Xiaofeng Zhu; Joel N Hirschhorn Journal: PLoS Genet Date: 2010-03-05 Impact factor: 5.917