Literature DB >> 11917143

Use of covariance analysis for the prediction of structural domain boundaries from multiple protein sequence alignments.

Daniel J Rigden1.   

Abstract

Current methods for identification of domains within protein sequences require either structural information or the identification of homologous domain sequences in different sequence contexts. Knowledge of structural domain boundaries is important for fold recognition experiments and structural determination by X-ray crystallography or nuclear magnetic resonance spectroscopy using the divide-and-conquer approach. Here, a new and conceptually simple method for the identification of structural domain boundaries in multiple protein sequence alignments is presented. Analysis of covariance at positions within the alignment is first used to predict 3D contacts. By the nature of the domain as an independent folding unit, inter-domain predicted contacts are fewer than intra-domain predicted contacts. By analysing all possible domain boundaries and constructing a smoothed profile of predicted contact density (PCD), true structural domain boundaries are predicted as local profile minima associated with low PCD. A training data set is constructed from 52 non-homologous two-domain protein sequences of known 3D structure and used to determine optimal parameters for the profile analysis. The alignments in the training data set contained 48 +/- 17 (mean +/- SD) sequences and lengths of 257 +/- 121 residues. Of the 47 alignments yielding predictions, 35% of true domain boundaries are predicted to within 15 amino acids by the local profile minimum with the lowest profile value. Including predictions from the second- and third-lowest local minima increases the correct domain boundary coverage to 60%, whereas the lowest five local minima cover 79% of correct domain boundaries. Through further profile analysis, criteria are presented which reliably identify subsets of more accurate predictions. Retrospective analysis of CASP3 targets shows predictions of sufficient accuracy to enable dramatically improved fold recognition results. Finally, a prediction is made for geminivirus AL1 protein which is in full agreement with biochemical data, yielding a plausible, novel threading result.

Entities:  

Mesh:

Substances:

Year:  2002        PMID: 11917143     DOI: 10.1093/protein/15.2.65

Source DB:  PubMed          Journal:  Protein Eng        ISSN: 0269-2139


  13 in total

1.  Prediction of protein domain boundaries from inverse covariances.

Authors:  Michael I Sadowski
Journal:  Proteins       Date:  2012-10-16

2.  Identification of putative domain linkers by a neural network - application to a large sequence database.

Authors:  Satoshi Miyazaki; Yutaka Kuroda; Shigeyuki Yokoyama
Journal:  BMC Bioinformatics       Date:  2006-06-27       Impact factor: 3.169

3.  Residue contacts predicted by evolutionary covariance extend the application of ab initio molecular replacement to larger and more challenging protein folds.

Authors:  Felix Simkovic; Jens M H Thomas; Ronan M Keegan; Martyn D Winn; Olga Mayans; Daniel J Rigden
Journal:  IUCrJ       Date:  2016-06-15       Impact factor: 4.769

Review 4.  Co-evolution techniques are reshaping the way we do structural bioinformatics.

Authors:  Saulo de Oliveira; Charlotte Deane
Journal:  F1000Res       Date:  2017-07-25

Review 5.  Folding by numbers: primary sequence statistics and their use in studying protein folding.

Authors:  Brent Wathen; Zongchao Jia
Journal:  Int J Mol Sci       Date:  2009-04-08       Impact factor: 6.208

6.  Ab initio and homology based prediction of protein domains by recursive neural networks.

Authors:  Ian Walsh; Alberto J M Martin; Catherine Mooney; Enrico Rubagotti; Alessandro Vullo; Gianluca Pollastri
Journal:  BMC Bioinformatics       Date:  2009-06-26       Impact factor: 3.169

7.  Domain selection combined with improved cloning strategy for high throughput expression of higher eukaryotic proteins.

Authors:  Yunjia Chen; Shihong Qiu; Chi-Hao Luan; Ming Luo
Journal:  BMC Biotechnol       Date:  2007-07-30       Impact factor: 2.563

8.  CATHEDRAL: a fast and effective algorithm to predict folds and domain boundaries from multidomain protein structures.

Authors:  Oliver C Redfern; Andrew Harrison; Tim Dallman; Frances M G Pearl; Christine A Orengo
Journal:  PLoS Comput Biol       Date:  2007-11       Impact factor: 4.475

9.  Potential DNA binding and nuclease functions of ComEC domains characterized in silico.

Authors:  James A Baker; Felix Simkovic; Helen M C Taylor; Daniel J Rigden
Journal:  Proteins       Date:  2016-07-01

Review 10.  Applications of contact predictions to structural biology.

Authors:  Felix Simkovic; Sergey Ovchinnikov; David Baker; Daniel J Rigden
Journal:  IUCrJ       Date:  2017-04-18       Impact factor: 4.769

View more

北京卡尤迪生物科技股份有限公司 © 2022-2023.