Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Poisson, compound Poisson and process approximations for testing statistical significance in sequence comparisons.

Literature DB >> 1638260

Poisson, compound Poisson and process approximations for testing statistical significance in sequence comparisons.

Abstract

DNA and protein sequence comparisons are performed by a number of computational algorithms. Most of these algorithms search for the alignment of two sequences that optimizes some alignment score. It is an important problem to assess the statistical significance of a given score. In this paper we use newly developed methods for Poisson approximation to derive estimates of the statistical significance of k-word matches on a diagonal of a sequence comparison. We require at least q of the k letters of the words to match where 0 less than q less than or equal to k. The distribution of the number of matches on a diagonal is approximated as well as the distribution of the order statistics of the sizes of clumps of matches on the diagonal. These methods provide an easily computed approximation of the distribution of the longest exact matching word between sequences. The methods are validated using comparisons of vertebrate and E. coli protein sequences. In addition, we compare two HLA class II transplantation antigens by this method and contrast the results with a dynamic programming approach. Several open problems are outlined in the last section.

Entities: Species

Mesh：

Substances：
Proteins

Year: 1992 PMID： 1638260 DOI： 10.1007/bf02459930

Source DB: PubMed Journal: Bull Math Biol ISSN： 0092-8240 Impact factor: 1.758

15 in total

Poisson, compound Poisson and process approximations for testing statistical significance in sequence comparisons.

1. Basic local alignment search tool.

2. Rapid and sensitive protein similarity searches.

3. Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes.

4. Tutorial on large deviations for the binomial distribution.

5. Improved tools for biological sequence comparison.

6. A general method applicable to the search for similarities in the amino acid sequence of two proteins.

7. The statistical distribution of nucleic acid similarities.

8. Identification of common molecular subsequences.

9. New approaches for computer analysis of nucleic acid sequences.

10. Rapid similarity searches of nucleic acid and protein data banks.

1. A local algorithm for DNA sequence alignment with inversions.

2. Rapid and accurate estimates of statistical significance for sequence data base searches.