Mathew Vithayathil1, Scott Smith1, Sergey Goryachev2, Jennifer Nayor3, Mingyang Song4,5,6. 1. Departments of Epidemiology and Nutrition, Harvard T. H. Chan School of Public Health, 667 Huntington Avenue, Kresge 906A, Boston, MA, 02115, USA. 2. Research Information Science and Computing (RISC), Partners Healthcare, Boston, MA, USA. 3. Division of Gastroenterology, Emerson Hospital, Concord, MA, USA. 4. Departments of Epidemiology and Nutrition, Harvard T. H. Chan School of Public Health, 667 Huntington Avenue, Kresge 906A, Boston, MA, 02115, USA. mingyangsong@mail.harvard.edu. 5. Clinical and Translational Epidemiology Unit, Massachusetts General Hospital and Harvard Medical School, Boston, MA, USA. mingyangsong@mail.harvard.edu. 6. Division of Gastroenterology, Massachusetts General Hospital and Harvard Medical School, Boston, MA, USA. mingyangsong@mail.harvard.edu.
Abstract
BACKGROUND AND AIMS: Conventional adenomas (CAs) and serrated polyps (SPs) are precursors to colorectal cancer (CRC). Understanding metachronous cancer risk is poor due to lack of accurate large-volume datasets. We outline the use of natural language processing (NLP) in forming the Partners Colonoscopy Cohort, an integrated longitudinal cohort of patients undergoing colonoscopies. METHODS: We identified endoscopy quality data from endoscopy reports for colonoscopies performed from 2007 to 2018 in a large integrated healthcare system, Mass General Brigham). Through modification of an established NLP pipeline, we extracted histopathological data (polyp location, histology and dysplasia) from corresponding pathology reports. Pathology and endoscopy data were merged by polyp location using a four-stage algorithm. NLP and merging procedures were validated by manual review of 500 pathology reports. RESULTS: 305,656 colonoscopies in 213,924 patients were identified. After merging, 76,137 patients had matched polyp data for 334,750 polyps. CAs and SPs were present in 86,707 (28.5%) and 55,373 (18.2%) colonoscopies. Among patients with polyps at index screening colonoscopy, 14,931 (33.4%) had follow-up colonoscopy (median 46.4, interquartile range 33.8-62.4 months); 91 (0.2%) and 1127 (2.5%) patients developed metachronous CRC and high-risk polyps (polyps ≥ 10 mm or CAs having high-grade dysplasia/villous/tublovillous histology or SPs with dysplasia). Genetic data were available for 23,787 (31.7%) patients with polyps from the Partners Biobank. The validation study showed a positive predictive value of 100% for polyp histology and locations. CONCLUSION: We created the Partners Colonoscopy Cohort providing essential infrastructure for future studies to better understand the natural history of CRC and improve screening and post-polypectomy strategies.
BACKGROUND AND AIMS: Conventional adenomas (CAs) and serrated polyps (SPs) are precursors to colorectal cancer (CRC). Understanding metachronous cancer risk is poor due to lack of accurate large-volume datasets. We outline the use of natural language processing (NLP) in forming the Partners Colonoscopy Cohort, an integrated longitudinal cohort of patients undergoing colonoscopies. METHODS: We identified endoscopy quality data from endoscopy reports for colonoscopies performed from 2007 to 2018 in a large integrated healthcare system, Mass General Brigham). Through modification of an established NLP pipeline, we extracted histopathological data (polyp location, histology and dysplasia) from corresponding pathology reports. Pathology and endoscopy data were merged by polyp location using a four-stage algorithm. NLP and merging procedures were validated by manual review of 500 pathology reports. RESULTS: 305,656 colonoscopies in 213,924 patients were identified. After merging, 76,137 patients had matched polyp data for 334,750 polyps. CAs and SPs were present in 86,707 (28.5%) and 55,373 (18.2%) colonoscopies. Among patients with polyps at index screening colonoscopy, 14,931 (33.4%) had follow-up colonoscopy (median 46.4, interquartile range 33.8-62.4 months); 91 (0.2%) and 1127 (2.5%) patients developed metachronous CRC and high-risk polyps (polyps ≥ 10 mm or CAs having high-grade dysplasia/villous/tublovillous histology or SPs with dysplasia). Genetic data were available for 23,787 (31.7%) patients with polyps from the Partners Biobank. The validation study showed a positive predictive value of 100% for polyp histology and locations. CONCLUSION: We created the Partners Colonoscopy Cohort providing essential infrastructure for future studies to better understand the natural history of CRC and improve screening and post-polypectomy strategies.
Authors: Aasma Shaukat; Tonya Kaltenbach; Jason A Dominitz; Douglas J Robertson; Joseph C Anderson; Michael Cruise; Carol A Burke; Samir Gupta; David Lieberman; Sapna Syngal; Douglas K Rex Journal: Gastroenterology Date: 2020-11-04 Impact factor: 22.682
Authors: Kevin J Spring; Zhen Zhen Zhao; Rozemary Karamatic; Michael D Walsh; Vicki L J Whitehall; Tanya Pike; Lisa A Simms; Joanne Young; Michael James; Grant W Montgomery; Mark Appleyard; David Hewett; Kazutomo Togashi; Jeremy R Jass; Barbara A Leggett Journal: Gastroenterology Date: 2006-08-18 Impact factor: 22.682
Authors: Ann G Zauber; Sidney J Winawer; Michael J O'Brien; Iris Lansdorp-Vogelaar; Marjolein van Ballegooijen; Benjamin F Hankey; Weiji Shi; John H Bond; Melvin Schapiro; Joel F Panish; Edward T Stewart; Jerome D Waye Journal: N Engl J Med Date: 2012-02-23 Impact factor: 91.245
Authors: Paul C Schroy; John B Wong; Michael J O'Brien; Clara A Chen; John L Griffith Journal: Am J Gastroenterol Date: 2015-05-26 Impact factor: 10.864
Authors: S M Powell; N Zilz; Y Beazer-Barclay; T M Bryan; S R Hamilton; S N Thibodeau; B Vogelstein; K W Kinzler Journal: Nature Date: 1992-09-17 Impact factor: 49.962
Authors: Anne F Peery; Evan S Dellon; Jennifer Lund; Seth D Crockett; Christopher E McGowan; William J Bulsiewicz; Lisa M Gangarosa; Michelle T Thiny; Karyn Stizenberg; Douglas R Morgan; Yehuda Ringel; Hannah P Kim; Marco Dacosta DiBonaventura; Charlotte F Carroll; Jeffery K Allen; Suzanne F Cook; Robert S Sandler; Michael D Kappelman; Nicholas J Shaheen Journal: Gastroenterology Date: 2012-08-08 Impact factor: 22.682
Authors: Henrik Munch Roager; Josef K Vogt; Mette Kristensen; Lea Benedicte S Hansen; Sabine Ibrügger; Rasmus B Mærkedahl; Martin Iain Bahl; Mads Vendelbo Lind; Rikke L Nielsen; Hanne Frøkiær; Rikke Juul Gøbel; Rikard Landberg; Alastair B Ross; Susanne Brix; Jesper Holck; Anne S Meyer; Morten H Sparholt; Anders F Christensen; Vera Carvalho; Bolette Hartmann; Jens Juul Holst; Jüri Johannes Rumessen; Allan Linneberg; Thomas Sicheritz-Pontén; Marlene D Dalgaard; Andreas Blennow; Henrik Lauritz Frandsen; Silas Villas-Bôas; Karsten Kristiansen; Henrik Vestergaard; Torben Hansen; Claus T Ekstrøm; Christian Ritz; Henrik Bjørn Nielsen; Oluf Borbye Pedersen; Ramneek Gupta; Lotte Lauritzen; Tine Rask Licht Journal: Gut Date: 2017-11-01 Impact factor: 23.059