Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Classifying next-generation sequencing data using a zero-inflated Poisson model.

Literature DB >> 29186294

Classifying next-generation sequencing data using a zero-inflated Poisson model.

Yan Zhou¹, Xiang Wan², Baoxue Zhang³, Tiejun Tong⁴.

Abstract

Motivation: With the development of high-throughput techniques, RNA-sequencing (RNA-seq) is becoming increasingly popular as an alternative for gene expression analysis, such as RNAs profiling and classification. Identifying which type of diseases a new patient belongs to with RNA-seq data has been recognized as a vital problem in medical research. As RNA-seq data are discrete, statistical methods developed for classifying microarray data cannot be readily applied for RNA-seq data classification. Witten proposed a Poisson linear discriminant analysis (PLDA) to classify the RNA-seq data in 2011. Note, however, that the count datasets are frequently characterized by excess zeros in real RNA-seq or microRNA sequence data (i.e. when the sequence depth is not enough or small RNAs with the length of 18-30 nucleotides). Therefore, it is desired to develop a new model to analyze RNA-seq data with an excess of zeros.
Results: In this paper, we propose a Zero-Inflated Poisson Logistic Discriminant Analysis (ZIPLDA) for RNA-seq data with an excess of zeros. The new method assumes that the data are from a mixture of two distributions: one is a point mass at zero, and the other follows a Poisson distribution. We then consider a logistic relation between the probability of observing zeros and the mean of the genes and the sequencing depth in the model. Simulation studies show that the proposed method performs better than, or at least as well as, the existing methods in a wide range of settings. Two real datasets including a breast cancer RNA-seq dataset and a microRNA-seq dataset are also analyzed, and they coincide with the simulation results that our proposed method outperforms the existing competitors. Availability and implementation: The software is available at http://www.math.hkbu.edu.hk/∼tongt. Contact: xwan@comp.hkbu.edu.hk or tongt@hkbu.edu.hk. Supplementary information: Supplementary data are available at Bioinformatics online.

Entities: Chemical

Mesh：

Substances：
MicroRNAs

Year: 2018 PMID： 29186294 DOI： 10.1093/bioinformatics/btx768

Source DB: PubMed Journal: Bioinformatics ISSN： 1367-4803 Impact factor: 6.937

Keyword Cloud
Cited

3 in total

Classifying next-generation sequencing data using a zero-inflated Poisson model.

1. Quantile regression for challenging cases of eQTL mapping.

2. Modelling RNA-Seq data with a zero-inflated mixture Poisson linear model.

3. scDLC: a deep learning framework to classify large sample single-cell RNA-seq data.