Chuan Lu1, Ross D King. 1. Department of Computer Science, Aberystwyth University, Ceredigion SY23 3DB, UK.
Abstract
MOTIVATION: Distribution analysis is one of the most basic forms of statistical analysis. Thanks to improved analytical methods, accurate and extensive quantitative measurements can now be made of the mRNA, protein and metabolite from biological systems. Here, we report a large-scale analysis of the population abundance distributions of the transcriptomes, proteomes and metabolomes from varied biological systems. RESULTS: We compared the observed empirical distributions with a number of distributions: power law, lognormal, loglogistic, loggamma, right Pareto-lognormal (PLN) and double PLN (dPLN). The best-fit for mRNA, protein and metabolite population abundance distributions was found to be the dPLN. This distribution behaves like a lognormal distribution around the centre, and like a power law distribution in the tails. To better understand the cause of this observed distribution, we explored a simple stochastic model based on geometric Brownian motion. The distribution indicates that multiplicative effects are causally dominant in biological systems. We speculate that these effects arise from chemical reactions: the central-limit theorem then explains the central lognormal, and a number of possible mechanisms could explain the long tails: positive feedback, network topology, etc. Many of the components in the central lognormal parts of the empirical distributions are unidentified and/or have unknown function. This indicates that much more biology awaits discovery.
MOTIVATION: Distribution analysis is one of the most basic forms of statistical analysis. Thanks to improved analytical methods, accurate and extensive quantitative measurements can now be made of the mRNA, protein and metabolite from biological systems. Here, we report a large-scale analysis of the population abundance distributions of the transcriptomes, proteomes and metabolomes from varied biological systems. RESULTS: We compared the observed empirical distributions with a number of distributions: power law, lognormal, loglogistic, loggamma, right Pareto-lognormal (PLN) and double PLN (dPLN). The best-fit for mRNA, protein and metabolite population abundance distributions was found to be the dPLN. This distribution behaves like a lognormal distribution around the centre, and like a power law distribution in the tails. To better understand the cause of this observed distribution, we explored a simple stochastic model based on geometric Brownian motion. The distribution indicates that multiplicative effects are causally dominant in biological systems. We speculate that these effects arise from chemical reactions: the central-limit theorem then explains the central lognormal, and a number of possible mechanisms could explain the long tails: positive feedback, network topology, etc. Many of the components in the central lognormal parts of the empirical distributions are unidentified and/or have unknown function. This indicates that much more biology awaits discovery.
Authors: Brian T Helfand; Victor P Andreev; Nazema Y Siddiqui; Gang Liu; Bradley A Erickson; Margaret E Helmuth; Susan K Lutgendorf; H Henry Lai; Ziya Kirkali Journal: Urology Date: 2019-03-25 Impact factor: 2.649
Authors: Maciej Dobrzyński; Lan K Nguyen; Marc R Birtwistle; Alexander von Kriegsheim; Alfonso Blanco Fernández; Alex Cheong; Walter Kolch; Boris N Kholodenko Journal: J R Soc Interface Date: 2014-09-06 Impact factor: 4.118
Authors: Margot Fournier; Carina Ferrari; Philipp S Baumann; Andrea Polari; Aline Monin; Tanja Bellier-Teichmann; Jacob Wulff; Kirk L Pappan; Michel Cuenod; Philippe Conus; Kim Q Do Journal: Schizophr Bull Date: 2014-03-31 Impact factor: 9.306
Authors: Mohammad R Nezami Ranjbar; Mahlet G Tadesse; Yue Wang; Habtom W Ressom Journal: IEEE/ACM Trans Comput Biol Bioinform Date: 2015 Jul-Aug Impact factor: 3.710
Authors: Daniel Hebenstreit; Miaoqing Fang; Muxin Gu; Varodom Charoensawan; Alexander van Oudenaarden; Sarah A Teichmann Journal: Mol Syst Biol Date: 2011-06-07 Impact factor: 11.429
Authors: Jiangang Liu; Robert A Jolly; Aaron T Smith; George H Searfoss; Keith M Goldstein; Vladimir N Uversky; Keith Dunker; Shuyu Li; Craig E Thomas; Tao Wei Journal: PLoS One Date: 2011-09-15 Impact factor: 3.240