Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Marginalized kernels for biological sequences.

Literature DB >> 12169556

Marginalized kernels for biological sequences.

Abstract

MOTIVATION: Kernel methods such as support vector machines require a kernel function between objects to be defined a priori. Several works have been done to derive kernels from probability distributions, e.g., the Fisher kernel. However, a general methodology to design a kernel is not fully developed.
RESULTS: We propose a reasonable way of designing a kernel when objects are generated from latent variable models (e.g., HMM). First of all, a joint kernel is designed for complete data which include both visible and hidden variables. Then a marginalized kernel for visible data is obtained by taking the expectation with respect to hidden variables. We will show that the Fisher kernel is a special case of marginalized kernels, which gives another viewpoint to the Fisher kernel theory. Although our approach can be applied to any object, we particularly derive several marginalized kernels useful for biological sequences (e.g., DNA and proteins). The effectiveness of marginalized kernels is illustrated in the task of classifying bacterial gyrase subunit B (gyrB) amino acid sequences.

Mesh：

Substances：
DNA Gyrase

Year: 2002 PMID： 12169556 DOI： 10.1093/bioinformatics/18.suppl_1.s268

Source DB: PubMed Journal: Bioinformatics ISSN： 1367-4803 Impact factor: 6.937

Keyword Cloud
Cited

10 in total

10. A Distance-Based Framework for the Characterization of Metabolic Heterogeneity in Large Sets of Genome-Scale Metabolic Models.

Authors: Andrea Cabbia; Peter A J Hilbers; Natal A W van Riel
Journal: Patterns (N Y) Date: 2020-08-06

10 in total

Marginalized kernels for biological sequences.

Review 1. Genomic similarity and kernel methods II: methods for genomic information.

Review 2. Kernel methods for large-scale genomic data analysis.

3. Protein-ligand interaction prediction: an improved chemogenomics approach.

Review 4. Machine learning for in silico virtual screening and chemical genomics: new strategies.

5. Virtual screening of GPCRs: an in silico chemogenomics approach.

6. An efficient algorithm for de novo predictions of biochemical pathways between chemical compounds.

7. The distance-profile representation and its application to detection of distantly related protein families.

8. Classification of heterogeneous microarray data by maximum entropy kernel.

Review 9. Support vector machines and kernels for computational biology.

10. A Distance-Based Framework for the Characterization of Metabolic Heterogeneity in Large Sets of Genome-Scale Metabolic Models.