| Literature DB >> 27540268 |
Solaiappan Manimaran1,2, Heather Marie Selby3, Kwame Okrah4, Claire Ruberman5, Jeffrey T Leek5, John Quackenbush6,7, Benjamin Haibe-Kains8,9,10, Hector Corrada Bravo11, W Evan Johnson1,2,3.
Abstract
Sequencing and microarray samples often are collected or processed in multiple batches or at different times. This often produces technical biases that can lead to incorrect results in the downstream analysis. There are several existing batch adjustment tools for '-omics' data, but they do not indicate a priori whether adjustment needs to be conducted or how correction should be applied. We present a software pipeline, BatchQC, which addresses these issues using interactive visualizations and statistics that evaluate the impact of batch effects in a genomic dataset. BatchQC can also apply existing adjustment tools and allow users to evaluate their benefits interactively. We used the BatchQC pipeline on both simulated and real data to demonstrate the effectiveness of this software toolkit.Entities:
Mesh:
Year: 2016 PMID: 27540268 PMCID: PMC5167063 DOI: 10.1093/bioinformatics/btw538
Source DB: PubMed Journal: Bioinformatics ISSN: 1367-4803 Impact factor: 6.937
Fig. 1.Examples from the BatchQC interface. (top) Boxplots from the simulated dataset showing clear distributional differences between batches. (bottom) The first two PCA components from the signature dataset shows strong batch effects