Literature DB >> 34750626

MathFeature: feature extraction package for DNA, RNA and protein sequences based on mathematical descriptors.

Robson P Bonidia1, Douglas S Domingues2, Danilo S Sanches3, André C P L F de Carvalho1.   

Abstract

One of the main challenges in applying machine learning algorithms to biological sequence data is how to numerically represent a sequence in a numeric input vector. Feature extraction techniques capable of extracting numerical information from biological sequences have been reported in the literature. However, many of these techniques are not available in existing packages, such as mathematical descriptors. This paper presents a new package, MathFeature, which implements mathematical descriptors able to extract relevant numerical information from biological sequences, i.e. DNA, RNA and proteins (prediction of structural features along the primary sequence of amino acids). MathFeature makes available 20 numerical feature extraction descriptors based on approaches found in the literature, e.g. multiple numeric mappings, genomic signal processing, chaos game theory, entropy and complex networks. MathFeature also allows the extraction of alternative features, complementing the existing packages. To ensure that our descriptors are robust and to assess their relevance, experimental results are presented in nine case studies. According to these results, the features extracted by MathFeature showed high performance (0.6350-0.9897, accuracy), both applying only mathematical descriptors, but also hybridization with well-known descriptors in the literature. Finally, through MathFeature, we overcame several studies in eight benchmark datasets, exemplifying the robustness and viability of the proposed package. MathFeature has advanced in the area by bringing descriptors not available in other packages, as well as allowing non-experts to use feature extraction techniques.
© The Author(s) 2021. Published by Oxford University Press.

Entities:  

Keywords:  GUI-based platform; biological sequences; feature extraction; mathematical descriptors; package; python

Mesh:

Substances:

Year:  2022        PMID: 34750626      PMCID: PMC8769707          DOI: 10.1093/bib/bbab434

Source DB:  PubMed          Journal:  Brief Bioinform        ISSN: 1467-5463            Impact factor:   11.622


  70 in total

1.  repRNA: a web server for generating various feature vectors of RNA sequences.

Authors:  Bin Liu; Fule Liu; Longyun Fang; Xiaolong Wang; Kuo-Chen Chou
Journal:  Mol Genet Genomics       Date:  2015-06-18       Impact factor: 3.291

2.  protr/ProtrWeb: R package and web server for generating various numerical representation schemes of protein sequences.

Authors:  Nan Xiao; Dong-Sheng Cao; Min-Feng Zhu; Qing-Song Xu
Journal:  Bioinformatics       Date:  2015-01-24       Impact factor: 6.937

3.  iLearn: an integrated platform and meta-learner for feature engineering, machine-learning analysis and modeling of DNA, RNA and protein sequence data.

Authors:  Zhen Chen; Pei Zhao; Fuyi Li; Tatiana T Marquez-Lago; André Leier; Jerico Revote; Yan Zhu; David R Powell; Tatsuya Akutsu; Geoffrey I Webb; Kuo-Chen Chou; A Ian Smith; Roger J Daly; Jian Li; Jiangning Song
Journal:  Brief Bioinform       Date:  2020-05-21       Impact factor: 11.622

4.  Seq2Feature: a comprehensive web-based feature extraction tool.

Authors:  Rahul Nikam; M Michael Gromiha
Journal:  Bioinformatics       Date:  2019-11-01       Impact factor: 6.937

5.  SubFeat: Feature subspacing ensemble classifier for function prediction of DNA, RNA and protein sequences.

Authors:  H M Fazlul Haque; Muhammod Rafsanjani; Fariha Arifin; Sheikh Adilina; Swakkhar Shatabda
Journal:  Comput Biol Chem       Date:  2021-04-24       Impact factor: 2.877

6.  Characterization and identification of lysine crotonylation sites based on machine learning method on both plant and mammalian.

Authors:  Rulan Wang; Zhuo Wang; Hongfei Wang; Yuxuan Pang; Tzong-Yi Lee
Journal:  Sci Rep       Date:  2020-11-24       Impact factor: 4.379

7.  periodicDNA: an R/Bioconductor package to investigate k-mer periodicity in DNA.

Authors:  Jacques Serizay; Julie Ahringer
Journal:  F1000Res       Date:  2021-02-24

8.  PlncRNA-HDeep: plant long noncoding RNA prediction using hybrid deep learning based on two encoding styles.

Authors:  Jun Meng; Qiang Kang; Zheng Chang; Yushi Luan
Journal:  BMC Bioinformatics       Date:  2021-05-12       Impact factor: 3.169

9.  PyBioMed: a python library for various molecular representations of chemicals, proteins and DNAs and their interactions.

Authors:  Jie Dong; Zhi-Jiang Yao; Lin Zhang; Feijun Luo; Qinlu Lin; Ai-Ping Lu; Alex F Chen; Dong-Sheng Cao
Journal:  J Cheminform       Date:  2018-03-20       Impact factor: 5.514

10.  LncFinder: an integrated platform for long non-coding RNA identification utilizing sequence intrinsic composition, structural information and physicochemical property.

Authors:  Siyu Han; Yanchun Liang; Qin Ma; Yangyi Xu; Yu Zhang; Wei Du; Cankun Wang; Ying Li
Journal:  Brief Bioinform       Date:  2019-11-27       Impact factor: 11.622

View more
  3 in total

1.  BioAutoML: automated feature engineering and metalearning to predict noncoding RNAs in bacteria.

Authors:  Robson P Bonidia; Anderson P Avila Santos; Breno L S de Almeida; Peter F Stadler; Ulisses N da Rocha; Danilo S Sanches; André C P L F de Carvalho
Journal:  Brief Bioinform       Date:  2022-07-18       Impact factor: 13.994

2.  iFeatureOmega: an integrative platform for engineering, visualization and analysis of features from molecular sequences, structural and ligand data sets.

Authors:  Zhen Chen; Xuhan Liu; Pei Zhao; Chen Li; Yanan Wang; Fuyi Li; Tatsuya Akutsu; Chris Bain; Robin B Gasser; Junzhou Li; Zuoren Yang; Xin Gao; Lukasz Kurgan; Jiangning Song
Journal:  Nucleic Acids Res       Date:  2022-05-07       Impact factor: 19.160

3.  WalkIm: Compact image-based encoding for high-performance classification of biological sequences using simple tuning-free CNNs.

Authors:  Saeedeh Akbari Rokn Abadi; Amirhossein Mohammadi; Somayyeh Koohi
Journal:  PLoS One       Date:  2022-04-15       Impact factor: 3.752

  3 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.