Literature DB >> 26668004

Positive and negative forms of replicability in gene network analysis.

W Verleyen1, S Ballouz1, J Gillis1.   

Abstract

MOTIVATION: Gene networks have become a central tool in the analysis of genomic data but are widely regarded as hard to interpret. This has motivated a great deal of comparative evaluation and research into best practices. We explore the possibility that this may lead to overfitting in the field as a whole.
RESULTS: We construct a model of 'research communities' sampling from real gene network data and machine learning methods to characterize performance trends. Our analysis reveals an important principle limiting the value of replication, namely that targeting it directly causes 'easy' or uninformative replication to dominate analyses. We find that when sampling across network data and algorithms with similar variability, the relationship between replicability and accuracy is positive (Spearman's correlation, rs ∼0.33) but where no such constraint is imposed, the relationship becomes negative for a given gene function (rs ∼ -0.13). We predict factors driving replicability in some prior analyses of gene networks and show that they are unconnected with the correctness of the original result, instead reflecting replicable biases. Without these biases, the original results also vanish replicably. We show these effects can occur quite far upstream in network data and that there is a strong tendency within protein-protein interaction data for highly replicable interactions to be associated with poor quality control.
AVAILABILITY AND IMPLEMENTATION: Algorithms, network data and a guide to the code available at: https://github.com/wimverleyen/AggregateGeneFunctionPrediction CONTACT: jgillis@cshl.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
© The Author 2015. Published by Oxford University Press. All rights reserved. For Permissions, please e-mail: journals.permissions@oup.com.

Mesh:

Year:  2015        PMID: 26668004     DOI: 10.1093/bioinformatics/btv734

Source DB:  PubMed          Journal:  Bioinformatics        ISSN: 1367-4803            Impact factor:   6.937


  7 in total

1.  EGAD: ultra-fast functional analysis of gene networks.

Authors:  Sara Ballouz; Melanie Weber; Paul Pavlidis; Jesse Gillis
Journal:  Bioinformatics       Date:  2017-02-15       Impact factor: 6.937

2.  Using predictive specificity to determine when gene set analysis is biologically meaningful.

Authors:  Sara Ballouz; Paul Pavlidis; Jesse Gillis
Journal:  Nucleic Acids Res       Date:  2017-02-28       Impact factor: 16.971

3.  Functional networks inference from rule-based machine learning models.

Authors:  Nicola Lazzarini; Paweł Widera; Stuart Williamson; Rakesh Heer; Natalio Krasnogor; Jaume Bacardit
Journal:  BioData Min       Date:  2016-09-05       Impact factor: 2.522

4.  Extracting replicable associations across multiple studies: Empirical Bayes algorithms for controlling the false discovery rate.

Authors:  David Amar; Ron Shamir; Daniel Yekutieli
Journal:  PLoS Comput Biol       Date:  2017-08-18       Impact factor: 4.475

5.  Strength of functional signature correlates with effect size in autism.

Authors:  Sara Ballouz; Jesse Gillis
Journal:  Genome Med       Date:  2017-07-07       Impact factor: 11.117

6.  Ligand Similarity Complements Sequence, Physical Interaction, and Co-Expression for Gene Function Prediction.

Authors:  Matthew J O'Meara; Sara Ballouz; Brian K Shoichet; Jesse Gillis
Journal:  PLoS One       Date:  2016-07-28       Impact factor: 3.240

7.  Dynamic rewiring of the human interactome by interferon signaling.

Authors:  Craig H Kerr; Michael A Skinnider; Daniel D T Andrews; Angel M Madero; Queenie W T Chan; R Greg Stacey; Nikolay Stoynov; Eric Jan; Leonard J Foster
Journal:  Genome Biol       Date:  2020-06-15       Impact factor: 13.583

  7 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.