Literature DB >> 34341986

Effect of high variation in transcript expression on identifying differentially expressed genes in RNA-seq analysis.

Weitong Cui1, Huaru Xue1, Yifan Geng1,2, Jing Zhang1, Yajun Liang1, Xuewen Tian3, Qinglu Wang1,3.   

Abstract

Great efforts have been made on the algorithms that deal with RNA-seq data to enhance the accuracy and efficiency of differential expression (DE) analysis. However, no consensus has been reached on the proper threshold values of fold change and adjusted p-value for filtering differentially expressed genes (DEGs). It is generally believed that the more stringent the filtering threshold, the more reliable the result of a DE analysis. Nevertheless, by analyzing the impact of both adjusted p-value and fold change thresholds on DE analyses, with RNA-seq data obtained for three different cancer types from the Cancer Genome Atlas (TCGA) database, we found that, for a given sample size, the reproducibility of DE results became poorer when more stringent thresholds were applied. No matter which threshold level was applied, the overlap rates of DEGs were generally lower for small sample sizes than for large sample sizes. The raw read count analysis demonstrated that the transcript expression of the same gene in different samples, whether in tumor groups or in normal groups, showed high variations, which resulted in a drastic fluctuation in fold change values and adjustedp-values when different sets of samples were used. Overall, more stringent thresholds did not yield more reliable DEGs due to high variations in transcript expression; the reliability of DEGs obtained with small sample sizes was more susceptible to these variations. Therefore, less stringent thresholds are recommended for screening DEGs. Moreover, large sample sizes should be considered in RNA-seq experimental designs to reduce the interfering effect of variations in transcript expression on DEG identification.
© 2021 John Wiley & Sons Ltd/University College London.

Entities:  

Keywords:  Differential expression; RNA-seq; false discovery rate; fold change; sample size; threshold

Mesh:

Substances:

Year:  2021        PMID: 34341986     DOI: 10.1111/ahg.12441

Source DB:  PubMed          Journal:  Ann Hum Genet        ISSN: 0003-4800            Impact factor:   1.670


  2 in total

1.  Transcriptomic and proteomic retinal pigment epithelium signatures of age-related macular degeneration.

Authors:  Anne Senabouth; Maciej Daniszewski; Grace E Lidgerwood; Helena H Liang; Damián Hernández; Mehdi Mirzaei; Stacey N Keenan; Ran Zhang; Xikun Han; Drew Neavin; Louise Rooney; Maria Isabel G Lopez Sanchez; Lerna Gulluyan; Joao A Paulo; Linda Clarke; Lisa S Kearns; Vikkitharan Gnanasambandapillai; Chia-Ling Chan; Uyen Nguyen; Angela M Steinmann; Rachael A McCloy; Nona Farbehi; Vivek K Gupta; David A Mackey; Guy Bylsma; Nitin Verma; Stuart MacGregor; Matthew J Watt; Robyn H Guymer; Joseph E Powell; Alex W Hewitt; Alice Pébay
Journal:  Nat Commun       Date:  2022-07-26       Impact factor: 17.694

2.  NETest: serial liquid biopsies in gastroenteropancreatic NET surveillance.

Authors:  Mark J C van Treijen; Catharina M Korse; Wieke H Verbeek; Margot E T Tesselaar; Gerlof D Valk
Journal:  Endocr Connect       Date:  2022-09-07       Impact factor: 3.221

  2 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.