OBJECTIVE: In this study, we aimed to evaluate interrater agreement statistics (IRAS) for use in research on low base rate clinical diagnoses or observed behaviors. Establishing and reporting sufficient interrater agreement is essential in such studies. Yet the most commonly applied agreement statistic, Cohen's κ, has a well known sensitivity to base rates that results in a substantial penalization of interrater agreement when behaviors or diagnoses are very uncommon, a prevalent and frustrating concern in such studies. METHOD: We performed Monte Carlo simulations to evaluate the performance of 5 of κ's alternatives (Van Eerdewegh's V, Yule's Y, Holley and Guilford's G, Scott's π, and Gwet's AC₁), alongside κ itself. The simulations investigated the robustness of these IRAS to conditions that are common in clinical research, with varying levels of behavior or diagnosis base rate, rater bias, observed interrater agreement, and sample size. RESULTS: When the base rate was 0.5, each IRAS provided similar estimates, particularly with unbiased raters. G was the least sensitive of the IRAS to base rates. CONCLUSIONS: The results encourage the use of the G statistic for its consistent performance across the simulation conditions. We recommend separately reporting the rates of agreement on the presence and absence of a behavior or diagnosis alongside G as an index of chance corrected overall agreement.
OBJECTIVE: In this study, we aimed to evaluate interrater agreement statistics (IRAS) for use in research on low base rate clinical diagnoses or observed behaviors. Establishing and reporting sufficient interrater agreement is essential in such studies. Yet the most commonly applied agreement statistic, Cohen's κ, has a well known sensitivity to base rates that results in a substantial penalization of interrater agreement when behaviors or diagnoses are very uncommon, a prevalent and frustrating concern in such studies. METHOD: We performed Monte Carlo simulations to evaluate the performance of 5 of κ's alternatives (Van Eerdewegh's V, Yule's Y, Holley and Guilford's G, Scott's π, and Gwet's AC₁), alongside κ itself. The simulations investigated the robustness of these IRAS to conditions that are common in clinical research, with varying levels of behavior or diagnosis base rate, rater bias, observed interrater agreement, and sample size. RESULTS: When the base rate was 0.5, each IRAS provided similar estimates, particularly with unbiased raters. G was the least sensitive of the IRAS to base rates. CONCLUSIONS: The results encourage the use of the G statistic for its consistent performance across the simulation conditions. We recommend separately reporting the rates of agreement on the presence and absence of a behavior or diagnosis alongside G as an index of chance corrected overall agreement.
Authors: Stephanie J Wilson; Lisa M Jaremka; Christopher P Fagundes; Rebecca Andridge; Juan Peng; William B Malarkey; Diane Habash; Martha A Belury; Janice K Kiecolt-Glaser Journal: Psychoneuroendocrinology Date: 2017-02-16 Impact factor: 4.905
Authors: Janice K Kiecolt-Glaser; Stephanie J Wilson; Michael L Bailey; Rebecca Andridge; Juan Peng; Lisa M Jaremka; Christopher P Fagundes; William B Malarkey; Bryon Laskowski; Martha A Belury Journal: Psychoneuroendocrinology Date: 2018-08-04 Impact factor: 4.905
Authors: Stephanie J Wilson; Juan Peng; Rebecca Andridge; Lisa M Jaremka; Christopher P Fagundes; William B Malarkey; Martha A Belury; Janice K Kiecolt-Glaser Journal: Psychoneuroendocrinology Date: 2020-06-17 Impact factor: 4.905
Authors: M Magill; Timothy R Apodaca; Justin Walthers; Jacques Gaume; Ayla Durst; Richard Longabaugh; Robert L Stout; Kathleen M Carroll Journal: J Subst Abuse Treat Date: 2016-07-29
Authors: Lisa M Jaremka; Martha A Belury; Rebecca R Andridge; Monica E Lindgren; Diane Habash; William B Malarkey; Janice K Kiecolt-Glaser Journal: Clin Psychol Sci Date: 2015-07-29