Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Detecting Gene Symbols and Names in Biological Texts: A First Step toward Pertinent Information Extraction.

Literature DB >> 11072323

Detecting Gene Symbols and Names in Biological Texts: A First Step toward Pertinent Information Extraction.

Abstract

Gathering data on molecular interactions to be fed into a specialized database has motivated the development of a computer system to help extracting pertinent information from texts, relying on advanced linguistic tools, completed with object-oriented knowledge modeling capabilities. As a first step toward this challenging objective, a program for the identification of gene symbols and names inside sentences has been devised. The main difficulty is that these names and symbols do not appear to follow construction rules. The program is thus made up of a series of sieves of different natures, lexical, morphological and semantic, to distinguish among the words of a sentence those which can only be potential gene symbols or names. Its performance has been evaluated, in terms of coverage and precision ratios, on a corpus of texts concerning D. melanogaster for which the list of names of known genes is available for checking.

Entities: Species

Year: 1998 PMID： 11072323

Source DB: PubMed Journal: Genome Inform Ser Workshop Genome Inform

Keyword Cloud
Cited

20 in total

1. Mining the bibliome: searching for a needle in a haystack? New computing tools are needed to effectively scan the growing amount of scientific literature for useful information.

Authors: Les Grivell
Journal: EMBO Rep Date: 2002-03 Impact factor: 8.807

Detecting Gene Symbols and Names in Biological Texts: A First Step toward Pertinent Information Extraction.

1. Mining the bibliome: searching for a needle in a haystack? New computing tools are needed to effectively scan the growing amount of scientific literature for useful information.

2. Automatic extraction of gene and protein synonyms from MEDLINE and journal articles.

3. A method for finding communities of related genes.

4. A simple and practical dictionary-based approach for identification of proteins in Medline abstracts.

5. Identification of related gene/protein names based on an HMM of name variations.

6. NLProt: extracting protein names and sequences from papers.

7. Using co-occurrence network structure to extract synonymous gene and protein names from MEDLINE abstracts.

8. Unsupervised biomedical named entity recognition: experiments with clinical and biological texts.

9. Textpresso: an ontology-based information retrieval and extraction system for biological literature.

10. BIOADI: a machine learning approach to identifying abbreviations and definitions in biological literature.