Literature DB >> 32411828

Translating evidence into practice: eligibility criteria fail to eliminate clinically significant differences between real-world and study populations.

Amelia J Averitt1, Chunhua Weng1, Patrick Ryan1,2, Adler Perotte1.   

Abstract

Randomized controlled trials (RCTs) are regarded as the most reputable source of evidence. In some studies, factors beyond the intervention itself may contribute to the measured effect, an occurrence known as heterogeneity of treatment effect (HTE). If the RCT population differs from the real-world population on factors that induce HTE, the trials effect will not replicate. The RCTs eligibility criteria should identify the sub-population in which its evidence will replicate. However, the extent to which the eligibility criteria identify the appropriate population is unknown, which raises concerns for generalizability. We compared reported data from RCTs with real-world data from the electronic health records of a large, academic medical center that was curated according to RCT eligibility criteria. Our results show fundamental differences between the RCT population and our observational cohorts, which suggests that eligibility criteria may be insufficient for identifying the applicable real-world population in which RCT evidence will replicate.
© The Author(s) 2020.

Entities:  

Keywords:  Communication and replication; Epidemiology; Outcomes research; Randomized controlled trials

Year:  2020        PMID: 32411828      PMCID: PMC7214444          DOI: 10.1038/s41746-020-0277-8

Source DB:  PubMed          Journal:  NPJ Digit Med        ISSN: 2398-6352


Introduction

Generalizability closes the gap between biomedical research and clinical practice[1]. When research is translated into the healthcare setting, the application of biomedical evidence to clinical care is known as evidence-based medicine (EBM). Since its inception in the 1990s, EBM has become the standard of operation for many clinicians[2-5]. EBM encourages clinicians to seek the most reputable evidence for any patient, according to a hierarchy of study quality in which randomized controlled trials (RCTs) are the best single study design[5]. RCTs are most often used to unbiasedly assess the effect of an intervention, such as a drug or procedure, on an outcome. Although EBM may be employed successfully for many different clinical decisions, challenges remain. Underlying EBM’s success is the assumption that the effect shown in RCTs will replicate in real-world populations[6,7]. However, research has shown that factors beyond the intervention itself, such as age, sex, or medical history, may modify the measured effect, a phenomenon known as heterogeneity of treatment effect (HTE)[8]. If the RCT population differs from the real-world population based on factors that induce HTE, RCT results will not be replicated in real-world application. Realistically, clinicians cannot evaluate HTE on a case-by-case basis and must assume that HTE is not a significant factor. However, when applying evidence from RCTs, this assumption is likely unmet. Research has shown that HTE is often found to exist[9,10]. This raises concerns for reproducibility of studies in the presence of additional heterogeneity in real-world populations. The RCT is well-regarded for many reasons, but randomization is the most important. Randomization ensures the highest possible internal validity, which speaks to whether the true effect is biased by systematic error[11,12]. The notion of internal validity does not speak to how well the causal relationship will generalize, only how unbiased it is for the study population. The patients for which the effect estimate is internally valid are nominally defined by eligibility criteria. These criteria both stipulate the characteristics that all study patients must share and nominally identify the real-world population for which the effect is internally valid. When operationalized, the eligibility criteria are represented as inclusion and exclusion criteria[13-15]; and with every addition of a criterion to a study population, a different sub-population is identified with increasingly controlled conditions[16]. If HTE exists, then application of eligibility criteria to a population may identify a sub-population of patients for which there is a more homogeneous effect of the intervention. RCTs often employ very restrictive eligibility criteria and are often cited as poorly representative of the real-world, as many subpopulations may be excluded. This may result in poor external validity. External validity refers to the extent to which the treatment effect estimate applies those outside of the study with potentially different patient and treatment setting characteristics[11]. External validity always poses a concern, except in the circumstance in which HTE is known to be absent. With poor external validity, replication of the study effect can be challenging[17-21]. Replication of trial evidence with real-world data, ideally, requires that the right persons, in the right treatment setting, exist in the right proportions. In the context of treating a population that differs significantly from the clinical trial population, it can be unclear how appropriate the evidence is for this new population. Presumably, the eligibility criteria of a study should be sufficient to identify the population in which the effect will replicate, which we call the applicable population. To address this knowledge gap, we leverage observational data to assesses if RCT populations and real-world populations after application of eligibility criteria differ. If the populations differ, the evidence may not apply due to HTE. If HTE exists in observational populations, it may impede the replication of RCT effect estimates. These methods will contribute (i) a means to determine if the eligibility criteria are adequate for identifying the applicable population; (ii) a framework for evaluating the external validity of studies; and (iii) highlight tensions between assumptions of EBM practice and qualities of reputable evidence. This research may encourage clinicians to reconsider the assumptions made when practicing EBM, and whether these assumptions are valid. Furthermore, the empirical evidence put forth by this study highlights the limitations of the current system of clinical knowledge generation. The current system sacrifices external validity in favor of internal validity, through the selection of the experimental population. Such a decision impedes the ability of experimental evidence to translate to the general population, resulting in non-optimal or damaging clinical care. This problem motivates the use of study populations that are more representative of the real-world and is only truly optimized when study populations and the populations targeted for treatment are one in the same. Such an analysis is called real-world evidence (RWE) generation, in which clinical knowledge is learned from the analysis of routinely collected, real-world data[22]. The results of this research identify the need for RWE in clinical medicine and underscore how RWE may improve the practice of EBM.

Results

Experimental vs. observational populations

The results of this study are presented in Tables 1–4 and Fig. 1.
Table 1

Results for sitagliptin vs. glimepiride trial.

Sitagliptin vs. glimepirideHartley[42] Columbia University Irving Medical Center (CUIMC)
Baseline characterticsSitagliptinGlimepiridePooledIndication onlyWith eligibility criteria
n = 197n = 191n = 388σn = 5942ΔRCTn = 3056ΔRCT
Age70.670.870.74.8569.03−0.26068.98−0.275
Sex
 Male937743.8%35.87%−0.07931.41%−0.124
 Female10411456.2%64.11%0.07968.55%0.124
 Unknown000.0%0.02%0.0000.03%0.000
Race
 White12110357.7%16.62%−0.41116.10%−0.416
 Multi-racial486128.1%33.03%0.04934.29%0.062
 Native American/Alaska Native18158.5%0.09%−0.0840.07%−0.084
 Asian5124.4%1.17%−0.0321.44%−0.029
 African American401.0%11.51%0.10511.32%0.103
 Native Hawaiian/Pacific Islander100.3%0.35%0.0010.29%0.000
 Unknown000.0%37.23%0.37236.49%0.365
Body weight76.975.376.1176.810.02875.39−0.030
BMI29.729.729.74.5430.350.06430.190.055
Duration of DM (years)89.48.696.433.97−0.5493.30−0.668
HbA1c % mean7.87.87.80.77.52−0.1676.81−0.120
Min6.45.76.063.87−1.3054.29−1.059
Max10.69.910.2515.83.30715.83.3301
HbA1c
<8.0%13112566.0%59.61%−0.06459.00%−0.070
≥8.0%666634.0%33.74%−0.00334.20%0.002
Unknown000.00%6.65%0.0666.81%0.068
FPG168.4169.7169.0433.21140.35−0.448141.55−0.440

ΔRCT = difference from observational cohort and reported RCT data; standardized difference in the means for continuous variables; difference in percentage points for discrete variables.

BMI body mass index, DM diabetes mellitus, yrs years, HbA1c hemoglobin A1c, Min minimum, Max maximum, FPG fasting plasma glucose.

Table 4

Results for benazepril-amlodipine vs. benazepril and hydrochlorothiazide (HCTZ) trial (ACCOMPLISH).

The ACCOMPLISH TrialNEJM[40] Columbia University Irving Medical Center (CUIMC)
Baseline characteristicsBenazepril-amlodipineBenazepril– HCTZ GroupPooledIndication onlyWith eligibility criteria
n = 5744n = 5762n = 11,506σn = 36,854ΔRCTn = 4198ΔRCT
Age
 ≥65 years3813382766.40%17.98%−0.45160.05%−0.063
 ≥70 years2363234040.87%9.59%−0.29543.22%0.023
Gender
 Female2296224639.48%67.81%0.28370.41%0.309
 Male3448351560.52%32.18%−0.28329.56%−0.310
 Unknown000.00%0.01%0.0000.02%0.000
Race
 White4817479583.54%25.31%−0.59510.65%−0.729
 Black69771912.31%14.38%0.01012.51%0.002
 Hispanic3003235.41%30.25%0.23036.45%0.310
 Other2302474.15%19.41%0.16730.12%0.260
 Unknown000.00%7.25%0.13410.260.103
Weight88.788.588.6018.9578.01−0.34674.65−0.514
Waist circumference103.9103.8103.8515.30NEDNED
Body mass index313131.006.2030.13−0.06129.95−0.096
Blood pressure
 Systolic145.3145.4145.3518.25129.75−0.704133.41−0.537
 Diastolic80.180.180.1010.7576.78−0.25173.85−0.479
 Pulse70.570.370.4011.0079.330.55277.950.496
 eGFR78.97978.9521.35NED*NED*
Serum values
 Creatinine (mg/dL)1.001.001.000.301.080.0981.330.308
 Glucose (mg/dL)127.9127.0127.4546.60149.550.336165.770.581
 Potassium (mmol/L)4.34.34.300.404.28−0.0314.360.107
 Total cholesterol (mg/dL)184.9184.1184.5039.90187.360.053168.80−0.282
 HDL (mg/dL)49.649.549.5514.1050.310.03846.87−0.140
Previous AHT treatments
 01691532.80%75.42%0.7262.28%−0.006
 11312127922.52%10.10%−0.1246.60%−0.159
 22116204736.18%7.38%−0.28812.97%−0.232
 ≥32147228338.50%7.11%−0.31478.21%0.397
Lipid lowering agents3851397167.98%12.31%−0.55779.75%0.118
Beta blockers2675280747.64%13.18%−0.34573.56%0.259
Antiplatlet agents3710373564.71%17.48%−0.47287.71%0.230
Characteristics
 Previous MI1337137223.54%2.98%−0.20616.76%−0.068
 Previous Stroke76273613.02%1.94%−0.11110.53%−0.025
 Previous hospitalization for unstable angina65367111.51%2.12%−0.09411.78%0.003
 Diabetes mellitus3478346860.37%22.68%−0.37785.76%0.254
 Renal disease3523536.13%7.25%0.01134.69%0.286
 eGFR < 601047103018.05%0.47%−0.17616.59%−0.015
 Previous coronary revasc.2044207335.78%1.56%−0.3427.99%−0.278
 Coronary artery bypass grafting1248119721.25%0.53%−0.2071.98%−0.193
 Percutaneous coronary intervention1055112318.93%1.08%−0.1796.52%−0.124
 Left ventricular hypertrophy76375813.22%0.21%−0.1301.32%−0.119
 Current smoking64165811.29%1.87%−0.0947.48%−0.038
 Dyslipidemia4221431974.22%18.01%−0.56277.45%0.032
 AFib3764036.77%3.67%−0.03113.63%0.069

NED = not enough data for measurement.

NED* = eGFR is incomplete in a biased manner due to lack of reporting of values greater than 60.

Fig. 1

Summary of ∆RCT for baseline characteristics of Indication Only vs RCT and ∆RCT Indication + Eligibilty Criteria vs. RCT.

a ACCOMPLISH trial b NCT01189890 trial (sitagliptin vs. glimepiride), c PROVE-IT trial d RENAAL trial. The shape of the marker corresponds to the data type. Circles (●) denote the standardized difference in the mean of continuous data. Pluses (+) denote the difference in percentage points of discrete data.

Results for sitagliptin vs. glimepiride trial. ΔRCT = difference from observational cohort and reported RCT data; standardized difference in the means for continuous variables; difference in percentage points for discrete variables. BMI body mass index, DM diabetes mellitus, yrs years, HbA1c hemoglobin A1c, Min minimum, Max maximum, FPG fasting plasma glucose. Results for atorvastatin vs. pravastatin trial (PROVE-IT). Results for losartan vs. placebo trial (RENAAL). Results for benazepril-amlodipine vs. benazepril and hydrochlorothiazide (HCTZ) trial (ACCOMPLISH). NED = not enough data for measurement. NED* = eGFR is incomplete in a biased manner due to lack of reporting of values greater than 60.

Summary of ∆RCT for baseline characteristics of Indication Only vs RCT and ∆RCT Indication + Eligibilty Criteria vs. RCT.

a ACCOMPLISH trial b NCT01189890 trial (sitagliptin vs. glimepiride), c PROVE-IT trial d RENAAL trial. The shape of the marker corresponds to the data type. Circles (●) denote the standardized difference in the mean of continuous data. Pluses (+) denote the difference in percentage points of discrete data.

Sitagliptin vs. glimepiride

The sitagliptin vs. glimepiride trial in elderly patients with Type 2 Diabetes Mellitus is given in Table 1. Application of eligibility criteria to the Indication Only cohort identified the Indication + Eligibility criteria cohort that was more similar to the RCT with regard to BMI, Fasting plasma glucose, and HbA1c % (mean); and less similar to the RCT with regard to age, years since diabetes diagnosis, gender, HbA1c > 8%, race/ethnicity, and weight. Indication + Eligibility Criteria patients did not significantly differ from the trial in regards to BMI, weight, and HbA1c % (mean), all other baseline characteristics metrics did significantly differ. These results highlight that the indicated real-world population and the real-world population that meets the stringent eligibility criteria have generally less progressed diabetes than those patients in the trial. This is exemplified by (i) years since diabetes diagnosis, which is 3.97 for the Indication Only cohort and 3.30 in the Indication + Eligibility Criteria cohort, but is 8.69 in the trial (p = 0.007) and (ii) fasting plasma glucose, which is 140.35 in the Indication Only cohort and 141.55 in the Indication + Eligibility Criteria cohort, but is 169.04 in the trial (p = 0.007). With regard to these two baseline characteristics metrics, the application of the eligibility criteria to the Indication Only cohort identified a subset of patients with a fasting plasma glucose that was more similar to the trial and a years since diabetes diagnosis that was less similar to the trial.

PROVE-IT

The atorvastatin vs. pravastatin trial in patients with a history of ACS (PROVE-IT Trial) is given in Table 2. Application of eligibility criteria to the Indication Only cohort identified the Indication + Eligibility Criteria cohort that was more similar to the RCT with regard to age, race/ethnicity, diabetes, hypertension, prior MI, peripheral artery disease, and prior statin therapy, and less similar to the RCT with regard to sex, current smoker, percutaneous coronary intervention, index event, and median lipid values. Indication + Eligibility Criteria patients differed significantly from the trial in regards to all baseline characteristics.
Table 2

Results for atorvastatin vs. pravastatin trial (PROVE-IT).

The PROVE-IT TrialCannon et al.[41] Columbia University Irving Medical Center (CUIMC)
Baseline characterticsPravastatinAtorvastatinPooledIndication onlyWith eligibility criteria
n = 2063n = 2099n = 4162σn = 3972ΔRCTn = 3180ΔRCT
Age58.358.158.2011.2560.370.13759.950.111
Sex
Male1617163478.11%45.92%−0.32245.88%−0.323
Female44546521.89%54.08%0.32254.09%0.323
Unknown0.00%0.03%0.0000.03%0.000
Race
White1865191190.73%28.23%−0.61171.42%−0.604
Other1981889.27%71.77%0.61128.58%0.604
Diabetes36137317.64%29.82%26.57%
Hypertension1014107750.24%60.72%0.10557.64%0.074
Current smoker76676336.74%4.48%−0.3234.18%−0.326
Prior MI39537418.48%34.42%0.15934.40%0.159
PCI
Prior to index event32032215.43%10.31%−0.04810.31%−0.051
After index event1426144268.91%15.30%−0.53615.16%−0.538
Coronary bypass surgery22123310.91%4.00%−0.0691.38%−0.095
Peripheral artery disease1361055.79%15.17%0.09413.33%0.075
Prior statin therapy51453525.20%42.73%0.17537.30%0.121
Index event
Unstable angina61460429.26%48.47%0.19250.88%0.046
MI without ST segment elevation (NSTEMI)75774736.14%19.80%−0.16315.22%−0.209
MI with ST segment elevation (STEMI)69074834.55%31.73%−0.02833.90%0.163
Median lipid values
 Total cholesterol180181180.50171.67−0.151169.54−0.194
 LDL cholesterol106106106.00100.41−0.11099.18−0.138
 HDL cholesterol393838.5045.070.36445.060.370
 Triglycerides154158156.02141.95−0.110137.95−0.145
The results for this trial show that patients that meet either the Indication or the Indication subject to all criteria, have less severe cardiovascular lipid measurements than patients in the trial. This is demonstrated in the median lipid values, where in total cholesterol, LDL, HDL, and triglycerides are 171.67, 100.41, 45.07, and 141.95, respectively, in the Indication Only cohort and 169.55, 99.19, 45.07, and 138.00, respectively, in the Indication + Eligibility Criteria. This is compared to the 180.50, 106.00, 38.50, and 156.02, respectively, that is reported in the trial.

RENAAL

The losartan vs. placebo trial in patients with diabetic nephropathy (RENAAL Trial) is given in Table 3. Application of eligibility criteria to the Indication Only cohort identified the Indication + Eligibility Criteria cohort that was more similar to the RCT with regard to age, pulse, angina pectoris, coronary revascularization, stroke, lipid disorder, total cholesterol, serum triglycerides, hemoglobin, and glycosylated hemoglobin, and less similar to the RCT with regard to sex, race/ethnicity, blood pressure measurements, use of antihypertensive drugs, myocardial infarction, amputation, neuropathy, retinopathy, current smoking, laboratory values, LDL and HDL. Indication + Eligibility Criteria patients significantly differ from the trial in regards to angina pectoris, stroke, amputation, lipid disorder, glycosylated hemoglobin % all other baseline characteristics metrics significantly differ. Significance of median urinary alb:creatinine ratio measurements could not be assessed due to insufficient reporting in the EHR.
Table 3

Results for losartan vs. placebo trial (RENAAL).

The RENAAL TrialBrenner et al.[39] Columbia University Irving Medical Center (CUIMC)
Baseline characteristicsLosartanPlaceboPooledIndication onlyWith eligibility criteria
n = 751n = 762n = 1513σn = 3818ΔRCTn = 72ΔRCT
Age606060.007.0063.720.257−0.095
Sex
 Male46249463.19%40.86%−0.22340.28%−0.229
 Female28626836.62%59.11%0.22559.72%0.231
 Unknown000.00%0.03%0.0000.00%0.000
Race
 Asian11713516.66%0.58%−0.1570.00%−0.153
 Black12510515.20%15.82%0.00613.89%−0.013
 White35837848.65%0.92%−0.4811.39%−0.486
 Hispanic14013618.24%36.14%0.17941.67%0.234
 Other1181.26%27.50%0.26218.06%0.168
 Unknown000.00%19.04%0.19025.00%0.250
BMI30.02929.506.0030.560.08434.000.386
Blood pressure (mmHg)
 Systolic152.0123137.3919.50136.95−0.017137.780.015
 Diastolic82.08282.0010.5071.01−0.79671.94−0.741
 Mean arterial105.5106105.7511.25104.01−0.109104.86−0.055
 Pulse69.470.870.1117.7579.650.45477.560.359
Medical history
Use of antihypertension drugs69372193.46%18.91%−0.7454.17%−0.893
Angina pectoris65759.25%14.14%0.0495.56%−0.037
Myocardial infarction759411.17%17.89%0.0672.78%−0.084
Coronary revasc.110.13%2.02%0.0190.00%−0.001
Stroke010.07%8.64%0.0860.005−0.001
Lipid disorder23427133.38%58.15%0.24843.06%0.097
Amputation65698.86%1.60%−0.0680.00%−0.089
Neuropathy37539751.02%19.83%−0.31211.11%−0.399
Retinopathy49447063.71%5.40%−0.5834.17%−0.595
Current smoking14713018.31%6.47%−0.1182.78%−0.155
Laboratory values
 Median urinary alb:creat ratio123712611249.09NEDNED-
 Serum creatinine (mg/dL)1.91.91.900.501.89−0.0042.450.282
Serum cholsterol (mg/dL)
 Total227229228.0155.50164.98−0.926171.11−0.908
 LDL142142142.0045.99132.18−0.00598.99−0.837
 HDL454545.0015.5043.86−0.05643.02−0.112
 Serum triglycerides (mg/dL)213225219.04190.07154.29−0.310156.21−0.308
Hemoglobin12.512.512.501.8511.53−0.47011.92−0.243
Glycosylated hemoglobin (%)8.58.48.451.658.35−0.3398.24−0.080
Similar to the trial results previously mentioned, patients enrolled in the RCT demonstrate hallmarks of advanced disease. A greater proportion of trial patients had a medical history of amputation (8.86%), neuropathy (51.02%), and retinopathy (63.71%), than compared to either the Indication Only cohort (1.60%, 19.83%, 5.40%, respectively) or the Indication + Eligibility Criteria cohort (0.00%, 11.11%, 4.17%).

ACCOMPLISH

The benazepril-amlopidine vs. benazepril-hydocholorothiazide trial in patients with systolic hypertension (ACCOMPLISH Trial) is given in Table 4. Application of eligibility criteria to the Indication Only cohort identified the Indication + Eligibility Criteria cohort that was more similar to the RCT with regard to age, potassium, lipid lowering agents, beta blockers, antiplatelet agents; history of MI, stroke, hospitalization for unstable angina, diabetes mellitus, eGFR < 60, coronary revascularization, CABG, PCI, left ventricular hypertrophy, current smoking, dyslipidemia, and AFib, and less similar to the RCT with regard to sex, race/ethnicity, weight, blood pressure measurements, pulse, creatinine, glucose, total cholesterol, HDL, and history of renal disease. Indication + Eligibility Criteria patients significantly differ from the trial in regards to all baseline characteristics, except for history of previous hospitalization for unstable angina. Significance of waist circumference and eGFR could not be assessed due to data availability and insufficient reporting in the EHR. The results of the four trials are summarized in Fig. 1. In this figure, each quadrant of the plot corresponds to a trial. For each trial, the ΔRCT for baseline characteristics are plotted for Indication Only vs. RCT and indication + Eligibility Criteria vs. RCT. The minimum and maximum HbA1c measurements for the NCT01189890 trial were excluded in this plot due to biologically implausible values that were likely transcription errors.

Discussion

This research suggests that eligibility criteria are insufficient for identifying the applicable real-world population in which experimental treatment effects will replicate with confidence. The comparison between the trial and the Indication + Eligibility Criteria cohorts highlights that the RCT and real-world cohorts are not similar. This result suggests that the eligibility criteria may not identify the applicable patients if HTE exists. In some cases, application of the eligibility criteria to the Indication Only cohort encouraged the mean feature to be more like that of the RCT. For example, the distribution of gender in the PROVE-IT trial. However, much more commonly, the application of the eligibility criteria to the Indication Only cohort also results in (i) an exacerbation of the difference between the Indication Only cohort and the RCT, as was seen with gender in the RENAAL trial; or (ii) an over-correction of the bias between the Indication Only cohort and the RCT, as was seen with gender in the ACCOMPLISH trial. This evaluation, with something as fundamental as gender, demonstrates that the eligibility criteria do not strictly encourage the data to be more like that reported in the RCT baseline characteristics data. Often, the eligibility criteria identified a subset a patient that was less like the trial on certain baseline characteristics. This suggests that the eligibility criteria applied in a different setting may actually increase confounding and introduce new biases in such an analysis. This assertion is additionally supported by the summarization of results in Fig. 1. A clustering of points near the center of the plot (0.0, 0.0) indicates that the observational cohorts differ very little from the RCT. Points that lie on the 45-degree line are indicative of baseline characteristics that are unaffected by the addition of eligibility criteria to the Indication Only. Deviations from this perfect correlation highlight the extent to which application of eligibility criteria encourage or worsen representativeness to the RCT. The sitagliptin vs. glimepiride trial (NCT01189890) and PROVE-IT show the least impact of the RCT criteria. The high linearity of points in these plots suggests that the eligibility criteria do not identify a subset of patients that are meaningfully different from the Indication Only cohort. This is contrary to the ACCOMPLISH and RENAAL trials. In these plots, there is more variance in the distribution of points along the 45-degree line with certain features improving representativeness and others worsening. The creation of these observational cohorts permits the comparison of real-world populations to summaries of clinical trials. A number of studies have previously examined the of misalignment between experimental populations and real-world populations by quantifying the discrepancy between these two data sources[23-34]. Despite this ongoing conversation regarding the lack of representativeness and generalizability of clinical trials, the relationship between the eligibility criteria and HTE and how they may contribute to poor external validity, remains poorly addressed. This research makes a thorough assessment of these two populations by comparing experimental cohorts with observational cohorts that were curated by carefully operationalizaed of eligibility criteria; it is highly rigorous and encourages confidence in our assessment of the inconsistencies between the trial and the real-world. We believe our approach is limited in practice because the task is complex and requires many components to line up. Our methods require not only a substantial amount of observational data, but standardization of that data into a common data model. The use of the Observational Health Data Sciences and Informatics’ common data model (OHDSI CDM) in this research facilitated the normalization of medical concepts to a single code and made simplified the definition of the cohorts. This study contributes a systematic evaluation of cohort characteristics under eligibility criteria. We show that discrepancies may appear as differences in aggregate features in the baseline characteristics. If these features have a meaningful effect on the outcome, the differences or imbalance between the cohorts, may result in confounding of the observational treatment effect. In this circumstance, the RCT evidence would not applicable to this real-world cohort. Computational methods could assist in identifying patients that match the RCT cohort for applicability, and perhaps such methods should be applied given the results shown here. Experimental trial participants are not only an inherently poor representation of the target population, but this research suggests that factors beyond eligibility criteria may introduce new hidden bias. Furthermore, the importance of HTE and potential for feature imbalance, even under careful cohort curation, highlight the current methodological gap in trial replication. Through our replication efforts, we were also able to articulate a framework of external validity. As noted earlier, external validity refers to the extent to which the trial results can be applied outside of the experimental setting[13]. Underlying the results of this research is the inherent tension that exists between the practice of EBM and what is regarded as credible evidence. The RCT, which is the most reputable source of biomedical knowledge, employs highly discriminative eligibility criteria that serve to identify a targeted effect of the intervention. This research uses real-world evidence to demonstrate that EBM practitioners cannot reasonably assume that trial participants are representative of real-world eligible patients. This raises the question as whether the RCT evidence is, therefore, applicable. This consideration is complicated by the inability of clinicians to determine who the applicable patients are. This is despite the publication of eligibility criteria, which are often incompletely or insufficiently reported in the modes of evidence most often consumed by clinicians[35]. This research does have limitations. Most importantly, the trials presented in this research were selected according to a set of criteria that enabled their analysis using the tools described. These criteria included an active intervention and comparator, published eligibility criteria, and ease of operationalization of concepts. The trials that were investigated as part of this research represent common indications. It is possible that the results presented here are specific to trials of common conditions and may not be representative of rare condition trials. The translation of clinical trial eligibility criteria to operationalized and computable queries may be prone to subjectivity. Although we sought to represent the criteria as unbiasedly as possible and consulted with clinicians to ensure accuracy, there is inherent ambiguity in the criteria themselves, which make perfect RCT representation impossible. Furthermore, information regarding the eligibility criteria may be found within the clinical note, which was not used when constructing the cohorts in this research. Additionally, when subjecting an observational cohort to many criteria, the resultant cohort may become very small, leading to a lack of power for the detection of relevant differences. In our evaluation of the external validity of trials, we compare aggregate metrics rather than a full distribution of features, which would be preferable. This comparison is the best we can do with the data that is available to us. However, such a comparison may fail to capture meaningful differences between the trial and real-world populations, as distributions with greatly differing functional forms may still have similar means. Lastly, and most notably, experimental data and EHR data are fundamentally different, which makes comparison between these two sources difficult. Though differences to experimental data may be inherent, the EHR houses the information that is available to clinicians at the time when treatment decisions are made. Furthermore, it is a valuable resource for identifying the applicable patients to support the practice of EBM. We believe that discrepancies between experimental data and EHR data are necessary to study so that we may develop methodologies to ensure appropriate applicability at the point of care. Based on the results of the research presented, the eligibility criteria, that nominally should be sufficient for effect replication, may not actually be sufficient if HTE exists. If HTE exists and the differences we observed in our cohort are common, factors beyond eligibility criteria may be necessary to identify applicable patients. This finding has significant implications on how we create and apply biomedical evidence. The expectation of EBM is that the population of patients that a single clinician sees, is an applicable population, and will mirror the population in the RCT in all ways, including the distribution of the treatments effect. This assumption does not take into account variation undocumented factors that affect HTE. If factors that induce HTE are not accounted for in the eligibility criteria but exist, a clinician cannot reasonably assume that the treatment effect will be seen in his treated patient population. The discrepancies between experimental and real-world populations that are presented here may be due to a number of sources, including overly restrictive eligibility criteria, insufficient documentation of eligibility criteria, or the self-selection of trial participants. When seeking to rectify this gap and improve generalizability of RCT findings, these issues may be addressed by the relaxation of trial eligibility criteria, a thorough and accurate description of eligibility criteria (perhaps recorded in a codified manner), or the active recruitment of a representative experimental population. Regardless of the source of this discrepancy, until addressed, careful consideration beyond who is eligible for the trial is necessary to determine whether results of a given RCT are an appropriate source of evidence when considering the care of a given patient.

Methods

The comparison of experimental and observational populations

We hypothesized that significant baseline characteristic differences exist between clinical trial populations and observational cohorts that meet all eligibility criteria. Such differences could be the source of poor external validity in the presence of HTE. The presence of differences could be confirmed by comparing empirical distributions of features between the RCT data and real-world observational data. However, patient-level RCT data is rarely released, so such as assessment is infeasible for most published RCTs. The best available proxy is to compare the real-world observational cohort to the summary of baseline characteristics of RCTs, as commonly presented in Table 1 of RCT publications. We will refer to these summary statistics as baseline characteristics. The baseline characteristics summarize the baseline demographic and clinical characteristics for each arm of the study[36]. The intent of publishing this table is to describe the clinical trial population in detail and report the similarity of arms in the RCT post-randomization. This data can also be used to evaluate external validity, and by association, replicability[37]. To examine how potential differences between experimental and observational cohorts may contribute to poor replicability, we compared RCT baseline characteristics with the same metrics from observational EHR data.

Data

Observational clinical data was obtained from the Columbia University Irving Medical Center (CUIMC) clinical data warehouse (CDW). Data elements evaluated in this study include laboratory measurements, diagnosis codes, and medications. This database is comprises predominantly of emergency and inpatient visits with a smaller number of outpatient visits at the hospital’s teaching clinics. The data used for this research was formatted according to the OHDSI (http://www.ohdsi.org) CDM to support downstream interoperability within the OHDSI community and to support replication and extension by OHDSI collaborators. This research was approved by the Columbia University Institutional Review Board. Informed consent was waived as this the research could not practicably be carried out without the waiver. The code to collect and query the data is freely available at https://github.com/ameliaaveritt/Translating_Evidence_Into_Practice.

Cohort creation

Corresponding to each RCT, observational cohorts were curated from EHR data according to two approaches. The first approach curated based on only the indication of the drug (Indication Only), e.g., diabetes or heart failure. This cohort represents the most basic assessment that clinicians can make when considering a treatment for a patient, per EBM. The second approach curated based on both the indication of the drug and all published eligibility criteria (Indication + Eligibility Criteria). This cohort represents the most thorough assessment that clinicians can make under EBM. Both the Indication Only and Indication + Eligibility Criteria cohorts were constructed using OHDSI’s ATLAS tool. ATLAS is an analytics platform used to support the design and execution of observational analyses. Part of this platform includes the ability to create cohort definitions. Cohort definitions identify a set of patients that satisfy one or more criteria for a duration of time. The Indication Only and Indication + Eligibility Criteria cohorts were defined using this tool. The Indication and Eligibility Criteria that were extracted from published RCT documentation were operationalized using the Observational Medical Outcomes Partnership (OMOP) CDM and served as criteria for cohort definitions. This was a rigorously done procedure, in which medical doctors were consulted to ensure the accuracy of the operationalization and faithfulness to the original criteria. To operationalize the criteria, we created concept sets, which enumerate both the medical concepts that should be included in the definition of our criteria and excludes the concepts that should not be included. This procedure often employed the hierarchical relationships that exist in the OMOP CDM ontology, where in all descendants of a single concept could be selected as part of a concept set and selectively removed, if needed. This procedure is outlined in Fig. 2.
Fig. 2

Pipeline to operationalize eligibility criteria using OHDSI tools.

The process begins by identifying the resources (e.g., an RCT protocol) that detail the eligibility criteria of a trial. Each criterion is then extracted and mapped to codified concepts in a controlled vocabulary. The concept is then mapped to the OHDSI common data model (CDM), which aggregates the same concepts from different vocabularies, into a single standardized concept. This concept is then refined to best define the eligibility criterion.

Pipeline to operationalize eligibility criteria using OHDSI tools.

The process begins by identifying the resources (e.g., an RCT protocol) that detail the eligibility criteria of a trial. Each criterion is then extracted and mapped to codified concepts in a controlled vocabulary. The concept is then mapped to the OHDSI common data model (CDM), which aggregates the same concepts from different vocabularies, into a single standardized concept. This concept is then refined to best define the eligibility criterion.

Cohort comparisons

For each RCT under study, we calculated the pooled baseline characteristics using the metrics reported for both the intervention and comparator arms. Discrete data was summed across both arms and is presented as a percent. Continuous data was taken as the average of each arm’s reported metrics, weighted by the proportion of patients in that arm. The Indication Only and Indication + Eligibility Criteria cohorts were queried to obtain metrics that corresponds to the RCT baseline characteristics. To explore the differences that exist between the observational patient cohorts and the RCT patient cohort, we calculated (i) the standardized difference in the means for continuous variables and (ii) percentage point differences between discrete variables (ΔRCT). If ΔRCT evaluates to zero, this indicates that the observational cohort does not differ from the trial cohort. If ΔRCT does not equal zero, this indicates that observational and trial cohorts differ, with greater magnitudes corresponding to greater discrepancies between the cohorts.

Trial selection

For this research, we purposefully picked landmark clinical trials, which are highly influential studies that are noted to change the practice of medicine. We began with a list of landmark trials, and after application of criteria that are outlined below, we decided on a small number. Our primary focus was landmark trials, but to increase the diversity of studies and to demonstrate applicability outside of efficacy trials, we evaluated a safety trial that met our criteria as well. When selecting candidate trials for this research, there were practical considerations that informed our choice of trials[38]. The RCT must have an active intervention and comparator drug, as we would be unable to sufficiently codify a cohort exposed to a placebo. Additionally, the intervention cannot be a new investigatory drug, as it would not exist in our EHR. The eligibility criteria for the RCT must be published and accessible; and most of the eligibility criteria must be hard criteria that are easily operationalized into concept codes (e.g., “age of at least 55 years”). While most trials have inescapable soft criteria that are not easily operationalized (e.g., “no contraindications” or “no current participation in another clinical trial”), it is important that our chosen trials have few of these. Consider, for example, the soft criteria “expected survival of at least 2 years”, which embodies a judgment call by a healthcare practitioner that cannot reasonably replicated with data. Finally, we sought trials that detailed a patient population that exists within the CUIMC EHR. This would ensure that a sufficient number of patients remain in our cohorts after application of the eligibility criteria. As we are interested in comparing the RCT Table 1 metrics with the same metrics from our observational cohort, it is important that our observational data contain as many patients as possible, as greater number of patients will increase confidence that our reported data is truly representative of the CUIMC population. To that end, we investigated four trials (1) the RENAAL Trial, which compared the effect of losartan and placebo on diabetic nephropathy[39]; (2) the ACCOMPLISH Trial[40], which compared benazepril-amlodipine to benazepril and hydrochlorothiazide on CV-related mortality, (3) the PROVE-IT Trial[41], which compared atorvastatin and pravastatin in patients with a history of acute coronary syndrome (ACS); and (4) the sitagliptin and glimepiride trial[42], which compared sitagliptin and glimepiride in elderly, diabetic patients. RENAAL, ACCOMPLISH, and PROVE-IT are Landmark RCTs with efficacy endpoints, and the sitagliptin vs. glimepiride trial is a smaller trial with a safety endpoint. Details on how the Indication Only and Indication + Eligibility Criteria cohorts were created can be found in the Supplementary Information Tables 1–10.
  33 in total

1.  Epistemologic inquiries in evidence-based medicine.

Authors:  Benjamin Djulbegovic; Gordon H Guyatt; Richard E Ashcroft
Journal:  Cancer Control       Date:  2009-04       Impact factor: 3.302

2.  Medicine. Life cycle of translational research for medical interventions.

Authors:  Despina G Contopoulos-Ioannidis; George A Alexiou; Theodore C Gouvias; John P A Ioannidis
Journal:  Science       Date:  2008-09-05       Impact factor: 47.728

Review 3.  Assessing the quality of randomized controlled trials. Current issues and future directions.

Authors:  D Moher; A R Jadad; P Tugwell
Journal:  Int J Technol Assess Health Care       Date:  1996       Impact factor: 2.188

4.  Comparison of baseline characteristics, management and outcome of patients with non-ST-segment elevation acute coronary syndrome in versus not in clinical trials.

Authors:  Adam B Hutchinson-Jaffe; Shaun G Goodman; Raymond T Yan; Ron Wald; Basem Elbarouni; Barry Rose; Kim A Eagle; Christopher C Lai; Carolyn Baer; Anatoly Langer; Andrew T Yan
Journal:  Am J Cardiol       Date:  2010-11-15       Impact factor: 2.778

5.  Evidence based medicine: what it is and what it isn't.

Authors:  D L Sackett; W M Rosenberg; J A Gray; R B Haynes; W S Richardson
Journal:  BMJ       Date:  1996-01-13

Review 6.  Representation of women in randomized clinical trials of cardiovascular disease prevention.

Authors:  Chiara Melloni; Jeffrey S Berger; Tracy Y Wang; Funda Gunes; Amanda Stebbins; Karen S Pieper; Rowena J Dolor; Pamela S Douglas; Daniel B Mark; L Kristin Newby
Journal:  Circ Cardiovasc Qual Outcomes       Date:  2010-02-16

7.  Acute heart failure: perspectives from a randomized trial and a simultaneous registry.

Authors:  Justin A Ezekowitz; Jia Hu; Diego Delgado; Adrian F Hernandez; Padma Kaul; Rolland Leader; Guy Proulx; Sean Virani; Michel White; Shelley Zieroth; Christopher O'Connor; Cynthia M Westerhout; Paul W Armstrong
Journal:  Circ Heart Fail       Date:  2012-10-02       Impact factor: 8.790

8.  Assessing the generalizability of randomized trial results to target populations.

Authors:  Elizabeth A Stuart; Catherine P Bradshaw; Philip J Leaf
Journal:  Prev Sci       Date:  2015-04

9.  Participant demographics reported in "Table 1" of randomised controlled trials: a case of "inverse evidence"?

Authors:  John Furler; Parker Magin; Marie Pirotta; Mieke van Driel
Journal:  Int J Equity Health       Date:  2012-03-19

10.  The older the better: are elderly study participants more non-representative? A cross-sectional analysis of clinical trial and observational study samples.

Authors:  Beatrice A Golomb; Virginia T Chan; Marcella A Evans; Sabrina Koperski; Halbert L White; Michael H Criqui
Journal:  BMJ Open       Date:  2012-12-14       Impact factor: 2.692

View more
  18 in total

1.  Real-world evidence of age-independent electroconvulsive therapy efficacy: A retrospective cohort study.

Authors:  James Luccarelli; Thomas H McCoy; Stephen J Seiner; Michael E Henry
Journal:  Acta Psychiatr Scand       Date:  2021-10-25       Impact factor: 6.392

2.  Treatment of Atopic Dermatitis with Baricitinib: First Real-life Experience.

Authors:  Danielle Rogner; Tilo Biedermann; Felix Lauffer
Journal:  Acta Derm Venereol       Date:  2022-03-22       Impact factor: 3.875

3.  Secular and Regional Trends among Pulmonary Arterial Hypertension Clinical Trial Participants.

Authors:  Jeff Min; Dina H Appleby; Robyn L McClelland; Jasleen Minhas; John H Holmes; Ryan J Urbanowicz; Steven C Pugliese; Jeremy A Mazurek; K Akaya Smith; Jason S Fritz; Harold I Palevsky; Jude Moutchia Suh; Nadine Al-Naamani; Steven M Kawut
Journal:  Ann Am Thorac Soc       Date:  2022-06

4.  Leveraging Real-World Data for the Selection of Relevant Eligibility Criteria for the Implementation of Electronic Recruitment Support in Clinical Trials.

Authors:  Georg Melzer; Tim Maiwald; Hans-Ulrich Prokosch; Thomas Ganslandt
Journal:  Appl Clin Inform       Date:  2021-01-13       Impact factor: 2.342

5.  Treatment of Open-Angle Glaucoma and Ocular Hypertension with Preservative-Free Tafluprost/Timolol Fixed-Dose Combination Therapy: UK and Ireland Results from the VISIONARY Study.

Authors:  Ejaz Ansari; Jasna Pavicic-Astalos; Filis Ayan; Anthony J King; Matthew Kinsella; Eugene Ng; Anca Nita
Journal:  Adv Ther       Date:  2021-04-22       Impact factor: 3.845

6.  Recruitment and enrollment of participants in an online diabetes self-management intervention in a virtual environment.

Authors:  Allison Vorderstrasse; Louise Reagan; Gail D'Eramo Melkus; Sarah Y Nowlin; Stacia B Birdsall; Andrew Burd; Yoon Hee Cho; Myoungock Jang; Constance Johnson
Journal:  Contemp Clin Trials       Date:  2021-04-20       Impact factor: 2.261

7.  Protocol for data extraction: how real-world data have been used in the National Institute for Health and Care Excellence appraisals of cancer therapy.

Authors:  Jiyeon Kang; John Cairns
Journal:  BMJ Open       Date:  2022-01-06       Impact factor: 2.692

8.  Clinical comparison between trial participants and potentially eligible patients using electronic health record data: A generalizability assessment method.

Authors:  James R Rogers; George Hripcsak; Ying Kuen Cheung; Chunhua Weng
Journal:  J Biomed Inform       Date:  2021-05-25       Impact factor: 8.000

9.  Post-Trial Enhanced Deployment and Technical Performance with the MISTIE Procedure per Lessons Learned.

Authors:  Ali Mansour; Andrea Loggini; Faten El Ammar; Ronald Alvarado-Dyer; Sean Polster; Agnieszka Stadnik; Paramita Das; Peter C Warnke; Bakhtiar Yamini; Christos Lazaridis; Christopher Kramer; W Andrew Mould; Meghan Hildreth; Matthew Sharrock; Daniel F Hanley; Fernando D Goldenberg; Issam A Awad
Journal:  J Stroke Cerebrovasc Dis       Date:  2021-07-22       Impact factor: 2.677

10.  Demographic and Health Behavior Factors Associated With Clinical Trial Invitation and Participation in the United States.

Authors:  Courtney P Williams; Nicole Senft Everson; Nonniekaye Shelburne; Wynne E Norton
Journal:  JAMA Netw Open       Date:  2021-09-01
View more

北京卡尤迪生物科技股份有限公司 © 2022-2023.