| Literature DB >> 23055626 |
Daksha Shukla1, Valadi K Jayaraman.
Abstract
Protein Glycosylation is an important post translational event that plays a pivotal role in protein folding and protein is trafficking. We describe a dictionary based and a rule based approach to mine 'mentions' of protein glycosylation in text. The dictionary based approach relies on a set of manually curated dictionaries specially constructed to address this task. Abstracts are then screened for the 'mentions' of words from these dictionaries which are further scored followed by classification on the basis of a threshold. The rule based approaches also relies on the words in the dictionary to arrive at the features which are used for classification. The performance of the system using both the approaches has been evaluated using a manually curated corpus of 3133 abstracts. The evaluation suggests that the performance of the Rule based approach supersedes that of the Dictionary based approach.Entities:
Keywords: Dictionary -based approach; Glycosylation; Rule-based approach; Text mining
Year: 2012 PMID: 23055626 PMCID: PMC3449393 DOI: 10.6026/97320630008758
Source DB: PubMed Journal: Bioinformation ISSN: 0973-2063
Figure 1Summary of Methodology