Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 An algorithm for direct causal learning of influences on patient outcomes.

Literature DB >> 28363452

An algorithm for direct causal learning of influences on patient outcomes.

Chandramouli Rathnam¹, Sanghoon Lee¹, Xia Jiang².

Abstract

OBJECTIVE: This study aims at developing and introducing a new algorithm, called direct causal learner (DCL), for learning the direct causal influences of a single target. We applied it to both simulated and real clinical and genome wide association study (GWAS) datasets and compared its performance to classic causal learning algorithms.
METHOD: The DCL algorithm learns the causes of a single target from passive data using Bayesian-scoring, instead of using independence checks, and a novel deletion algorithm. We generate 14,400 simulated datasets and measure the number of datasets for which DCL correctly and partially predicts the direct causes. We then compare its performance with the constraint-based path consistency (PC) and conservative PC (CPC) algorithms, the Bayesian-score based fast greedy search (FGS) algorithm, and the partial ancestral graphs algorithm fast causal inference (FCI). In addition, we extend our comparison of all five algorithms to both a real GWAS dataset and real breast cancer datasets over various time-points in order to observe how effective they are at predicting the causal influences of Alzheimer's disease and breast cancer survival.
RESULTS: DCL consistently outperforms FGS, PC, CPC, and FCI in discovering the parents of the target for the datasets simulated using a simple network. Overall, DCL predicts significantly more datasets correctly (McNemar's test significance: p<<0.0001) than any of the other algorithms for these network types. For example, when assessing overall performance (simple and complex network results combined), DCL correctly predicts approximately 1400 more datasets than the top FGS method, 1600 more datasets than the top CPC method, 4500 more datasets than the top PC method, and 5600 more datasets than the top FCI method. Although FGS did correctly predict more datasets than DCL for the complex networks, and DCL correctly predicted only a few more datasets than CPC for these networks, there is no significant difference in performance between these three algorithms for this network type. However, when we use a more continuous measure of accuracy, we find that all the DCL methods are able to better partially predict more direct causes than FGS and CPC for the complex networks. In addition, DCL consistently had faster runtimes than the other algorithms. In the application to the real datasets, DCL identified rs6784615, located on the NISCH gene, and rs10824310, located on the PRKG1 gene, as direct causes of late onset Alzheimer's disease (LOAD) development. In addition, DCL identified ER category as a direct predictor of breast cancer mortality within 5 years, and HER2 status as a direct predictor of 10-year breast cancer mortality. These predictors have been identified in previous studies to have a direct causal relationship with their respective phenotypes, supporting the predictive power of DCL. When the other algorithms discovered predictors from the real datasets, these predictors were either also found by DCL or could not be supported by previous studies.
CONCLUSION: Our results show that DCL outperforms FGS, PC, CPC, and FCI in almost every case, demonstrating its potential to advance causal learning. Furthermore, our DCL algorithm effectively identifies direct causes in the LOAD and Metabric GWAS datasets, which indicates its potential for clinical applications.

Entities: Chemical Disease Gene Mutation Species

Keywords: Bayesian-score based learning; Causal discovery; Clinical decision support; Constraint-based learning; Predictive medicine; Simulated data

Mesh：

Year: 2016 PMID： 28363452 PMCID： PMC5415921 DOI： 10.1016/j.artmed.2016.10.003

Source DB: PubMed Journal: Artif Intell Med ISSN： 0933-3657 Impact factor: 5.326

25 in total

1. Using Bayesian networks to analyze expression data.

Authors: N Friedman; M Linial; I Nachman; D Pe'er
Journal: J Comput Biol Date: 2000 Impact factor: 1.479

2. Optimizing exact genetic linkage computations.

Authors: Ma'ayan Fishelson; Dan Geiger
Journal: J Comput Biol Date: 2004 Impact factor: 1.479

3. The TETRAD Project: Constraint Based Aids to Causal Model Specification.

Authors: R Scheines; P Spirtes; C Glymour; C Meek; T Richardson
Journal: Multivariate Behav Res Date: 1998-01-01 Impact factor: 5.923

4. Evaluation of a two-stage framework for prediction using big genomic data.

Authors: Xia Jiang; Richard E Neapolitan
Journal: Brief Bioinform Date: 2015-03-18 Impact factor: 11.622

5. Prognostic significance of HER-2/neu expression in breast cancer and its relationship to other prognostic factors.

Authors: F Rilke; M I Colnaghi; N Cascinelli; S Andreola; M T Baldini; R Bufalino; G Della Porta; S Ménard; M A Pierotti; A Testori
Journal: Int J Cancer Date: 1991-08-19 Impact factor: 7.396

6. GWAS reveals new recessive loci associated with non-syndromic facial clefting.

Authors: Mauricio Camargo; Dora Rivera; Lina Moreno; Andrew C Lidral; Ursula Harper; Marypat Jones; Benjamin D Solomon; Erich Roessler; Jorge I Vélez; Ariel F Martinez; Settara C Chandrasekharappa; Mauricio Arcos-Burgos
Journal: Eur J Med Genet Date: 2012-06-27 Impact factor: 2.708

7. A bayesian method for evaluating and discovering disease loci associations.

Authors: Xia Jiang; M Michael Barmada; Gregory F Cooper; Michael J Becich
Journal: PLoS One Date: 2011-08-10 Impact factor: 3.240

8. Data mining of high density genomic variant data for prediction of Alzheimer's disease risk.

Authors: Natalia Briones; Valentin Dinu
Journal: BMC Med Genet Date: 2012-01-25 Impact factor: 2.103

9. Genomewide association study for onset age in Parkinson disease.

Authors: Jeanne C Latourelle; Nathan Pankratz; Alexandra Dumitriu; Jemma B Wilk; Stefano Goldwurm; Gianni Pezzoli; Claudio B Mariani; Anita L DeStefano; Cheryl Halter; James F Gusella; William C Nichols; Richard H Myers; Tatiana Foroud
Journal: BMC Med Genet Date: 2009-09-22 Impact factor: 2.103

10. Mining pure, strict epistatic interactions from high-dimensional datasets: ameliorating the curse of dimensionality.

Authors: Xia Jiang; Richard E Neapolitan
Journal: PLoS One Date: 2012-10-12 Impact factor: 3.240

3 in total

Review 1. Role of Nischarin in the pathology of diseases: a special emphasis on breast cancer.

Authors: Samuel C Okpechi; Hassan Yousefi; Khoa Nguyen; Thomas Cheng; Nikhilesh V Alahari; Bridgette Collins-Burow; Matthew E Burow; Suresh K Alahari
Journal: Oncogene Date: 2022-01-22 Impact factor: 9.867

2. Prevalence of hyperlipidemia in Shanxi Province, China and application of Bayesian networks to analyse its related factors.

Authors: Jinhua Pan; Zeping Ren; Wenhan Li; Zhen Wei; Huaxiang Rao; Hao Ren; Zhuang Zhang; Weimei Song; Yuling He; Chenglian Li; Xiaojuan Yang; LiMin Chen; Lixia Qiu
Journal: Sci Rep Date: 2018-02-28 Impact factor: 4.379

3. From hype to reality: data science enabling personalized medicine.

Authors: Holger Fröhlich; Rudi Balling; Niko Beerenwinkel; Oliver Kohlbacher; Santosh Kumar; Thomas Lengauer; Marloes H Maathuis; Yves Moreau; Susan A Murphy; Teresa M Przytycka; Michael Rebhan; Hannes Röst; Andreas Schuppert; Matthias Schwab; Rainer Spang; Daniel Stekhoven; Jimeng Sun; Andreas Weber; Daniel Ziemek; Blaz Zupan
Journal: BMC Med Date: 2018-08-27 Impact factor: 8.775

3 in total