| Literature DB >> 35854729 |
Linh Hoanga1, Lan Jiang1, Halil Kilicoglu1.
Abstract
Lack of large quantities of annotated data is a major barrier in developing effective text mining models of biomedical literature. In this study, we explored weak supervision to improve the accuracy of text classification models for assessing methodological transparency of randomized controlled trial (RCT) publications. Specifically, we used Snorkel, a framework to programmatically build training sets, and UMLS-EDA, a data augmentation method that leverages a small number of labeled examples to generate new training instances, and assessed their effect on a BioBERT-based text classification model proposed for the task in previous work. Performance improvements due to weak supervision were limited and were surpassed by gains from hyperparameter tuning. Our analysis suggests that refinements to the weak supervision strategies to better deal with multi-label case could be beneficial. Our code and data are available at https://github.com/kilicogluh/CONSORT-TM/tree/master/weakSupervision. ©2022 AMIA - All rights reserved.Entities:
Mesh:
Year: 2022 PMID: 35854729 PMCID: PMC9285178
Source DB: PubMed Journal: AMIA Annu Symp Proc ISSN: 1559-4076