| Literature DB >> 25377257 |
Julien Rey1, Patrick Deschavanne2, Pierre Tuffery3.
Abstract
With the recent progress in complete genome sequencing, mining the increasing amount of genomic information available should in theory provide the means to discover new classes of peptides. However, annotation pipelines often do not consider small reading frames likely to be expressed. BactPepDB, available online at http://bactpepdb.rpbs.univ-paris-diderot.fr, is a database that aims at providing an exhaustive re-annotation of all complete prokaryotic genomes-chromosomal and plasmid DNA-available in RefSeq for coding sequences ranging between 10 and 80 amino acids. The identified peptides are classified as (i) previously identified in RefSeq, (ii) entity-overlapping (intragenic) or intergenic, and (iii) potential pseudogenes-intergenic sequences corresponding to a portion of a previously annotated larger gene. Additional information is related to homologs within order, predicted signal sequence, transmembrane segments, disulfide bonds, secondary structure, and the existence of a related 3D structure in the Protein Databank. As a result, BactPepDB provides insights about candidate peptides, and provides information about their conservation, together with some of their expected biological/structural features. The BactPepDB interface allows to search for candidate peptides in the database, or to search for peptides similar to a query, according to the multiple properties predicted or related to genomic localization. Database URL: http://www.yeastgenome.org/Entities:
Mesh:
Substances:
Year: 2014 PMID: 25377257 PMCID: PMC4221844 DOI: 10.1093/database/bau106
Source DB: PubMed Journal: Database (Oxford) ISSN: 1758-0463 Impact factor: 3.451
Figure 1.BactPepDB flowchart.
Figure 2.BactPepDB entries according to peptide size.
BactPepDB entries by categories
| Small | Medium | Large | |
|---|---|---|---|
| RefSeq | 3946 (0.6) | 143 157 (30.4) | 362 395 (63.2) |
| Potential pseudogenes | 201 228 (28.6) | 51 752 (11.0) | 11 743 (2.1) |
| Intergenic | 324 533 (46.2) | 189 573 (40.2) | 121 897 (21.3) |
| Entity-overlapping | 173 270 (24.6) | 86 907 (18.4) | 77 012 (13.4) |
The three categories correspond to small (<30 amino acids), medium (30–50 amino acids) and large (>50 amino acids) peptide sizes. Peptides already annotated in RefSeq are distinguished from newcomers of BactPepDB categorized as potential pseudogenes, intergenic and entity-overlapping. Fractions in % within brackets.
Conserved SCSs
| Small | Medium | Large | ||
|---|---|---|---|---|
| Intra-genus | RefSeq | 978 (25) | 40 913 (29) | 152 429 (42) |
| New intergenic SCSs | 84 276 (26) | 57 584 (30) | 41 447 (34) | |
| Intra-order | RefSeq | 750 (19) | 20 041 (14) | 60 194 (17) |
| New intergenic SCSs | 61 662 (19) | 28 047 (15) | 21 940 (18) | |
Numbers and fractions (% within brackets) of peptide entries that are conserved across species of a genus (intra-genus) and across families of an order (intra-order). Fractions are relative to the total number of entries in each category (see Table 1). The three categories correspond to small (<30 amino acids), medium (30–50 amino acids) and large (>50 amino acids) peptide sizes. Peptides already annotated in RefSeq are distinguished from the new intergenic SCSs of BactPepDB.