| Literature DB >> 31292179 |
Holly Tibble1,2, Athanasios Tsanas1,2, Elsie Horne1,2, Robert Horne2,3, Mehrdad Mizani1,2, Colin R Simpson2,4, Aziz Sheikh1,2.
Abstract
INTRODUCTION: Asthma is a long-term condition with rapid onset worsening of symptoms ('attacks') which can be unpredictable and may prove fatal. Models predicting asthma attacks require high sensitivity to minimise mortality risk, and high specificity to avoid unnecessary prescribing of preventative medications that carry an associated risk of adverse events. We aim to create a risk score to predict asthma attacks in primary care using a statistical learning approach trained on routinely collected electronic health record data. METHODS AND ANALYSIS: We will employ machine-learning classifiers (naïve Bayes, support vector machines, and random forests) to create an asthma attack risk prediction model, using the Asthma Learning Health System (ALHS) study patient registry comprising 500 000 individuals across 75 Scottish general practices, with linked longitudinal primary care prescribing records, primary care Read codes, accident and emergency records, hospital admissions and deaths. Models will be compared on a partition of the dataset reserved for validation, and the final model will be tested in both an unseen partition of the derivation dataset and an external dataset from the Seasonal Influenza Vaccination Effectiveness II (SIVE II) study. ETHICS AND DISSEMINATION: Permissions for the ALHS project were obtained from the South East Scotland Research Ethics Committee 02 [16/SS/0130] and the Public Benefit and Privacy Panel for Health and Social Care (1516-0489). Permissions for the SIVE II project were obtained from the Privacy Advisory Committee (National Services NHS Scotland) [68/14] and the National Research Ethics Committee West Midlands-Edgbaston [15/WM/0035]. The subsequent research paper will be submitted for publication to a peer-reviewed journal and code scripts used for all components of the data cleaning, compiling, and analysis will be made available in the open source GitHub website (https://github.com/hollytibble). © Author(s) (or their employer(s)) 2019. Re-use permitted under CC BY. Published by BMJ.Entities:
Keywords: asthma; asthma attacks; machine learning; prediction; primary care
Year: 2019 PMID: 31292179 PMCID: PMC6624024 DOI: 10.1136/bmjopen-2018-028375
Source DB: PubMed Journal: BMJ Open ISSN: 2044-6055 Impact factor: 2.692
Metadata for clinical data sources in derivation dataset (ALHS)
| Data Source | Number of Records | Number of Individuals | Extraction Date | Data Specification Date Range |
| Primary Care Prescribing* | 4 709 231 | 47 095 | October 2018 | January 2009–April 2017 |
| Primary Care Encounters* | 11 766 100 | 49 307 | March 2018 | January 2000–November 2017 |
| Accident & Emergency | 1 831 789 | 500 321 | November 2017 | June 2007–September 2017 |
| Hospital Inpatient Admissions | 1 668 957 | 342 838 | August 2018 | January 2000–March 2017 |
| Mortality |
| 91 758 | May 2018 | January 2000–March 2017 |
*Records available for subset of study population with asthma diagnosis only.
Metadata for clinical data sources in external dataset (SIVE II)
| Data Source | Number of Records | Number of Individuals | Extraction Date | Data Specification Date Range |
| Primary Care Prescribing | 29 360 448 | 1 073 377 | May 2017 | January 2003–March 2017 |
| Primary Care Encounters | 31 878 423 | 1 887 957 | May 2017 | January 2000*–March 2017 |
| Accident & Emergency | 4 116 561 | 1 247 314 | April 2017 | June 2007– August 2016 |
| Hospital Inpatient Admissions | 3 549 174 | 794 937 | April 2017 | January 2000–March 2017 |
| Mortality |
| 215 466 | April 2017 | January 2000– March 2017 |
*Diagnosis codes entered in this period, but post-dated from 1940 onwards retained.
Figure 1Process of selecting the highest performing model from the validation data and the average performance of this model across iterations in the testing dataset. In the foreground, we have the first iteration. We will use 100 iterations for statistical confidence, randomly permuting the data into training, validation, and testing subsets in each iteration.