Literature DB >> 29850783

RIFRAF: a frame-resolving consensus algorithm.

Kemal Eren1, Ben Murrell2.   

Abstract

Motivation: Protein coding genes can be studied using long-read next generation sequencing. However, high rates of indel sequencing errors are problematic, corrupting the reading frame. Even the consensus of multiple independent sequence reads retains indel errors. To solve this problem, we introduce Reference-Informed Frame-Resolving multiple-Alignment Free template inference algorithm (RIFRAF), a sequence consensus algorithm that takes a set of error-prone reads and a reference sequence and infers an accurate in-frame consensus. RIFRAF uses a novel structure, analogous to a two-layer hidden Markov model: the consensus is optimized to maximize alignment scores with both the set of noisy reads and with a reference. The template-to-reads component of the model encodes the preponderance of indels, and is sensitive to the per-base quality scores, giving greater weight to more accurate bases. The reference-to-template component of the model penalizes frame-destroying indels. A local search algorithm proceeds in stages to find the best consensus sequence for both objectives.
Results: Using Pacific Biosciences SMRT sequences from an HIV-1 env clone, NL4-3, we compare our approach to other consensus and frame correction methods. RIFRAF consistently finds a consensus sequence that is more accurate and in-frame, especially with small numbers of reads. It was able to perfectly reconstruct over 80% of consensus sequences from as few as three reads, whereas the best alternative required twice as many. RIFRAF is able to achieve these results and keep the consensus in-frame even with a distantly related reference sequence. Moreover, unlike other frame correction methods, RIFRAF can detect and keep true indels while removing erroneous ones. Availability and implementation: RIFRAF is implemented in Julia, and source code is publicly available at https://github.com/MurrellGroup/Rifraf.jl. Supplementary information: Supplementary data are available at Bioinformatics online.

Entities:  

Mesh:

Year:  2018        PMID: 29850783      PMCID: PMC6223371          DOI: 10.1093/bioinformatics/bty426

Source DB:  PubMed          Journal:  Bioinformatics        ISSN: 1367-4803            Impact factor:   6.937


  22 in total

1.  EMBOSS: the European Molecular Biology Open Software Suite.

Authors:  P Rice; I Longden; A Bleasby
Journal:  Trends Genet       Date:  2000-06       Impact factor: 11.639

2.  MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform.

Authors:  Kazutaka Katoh; Kazuharu Misawa; Kei-ichi Kuma; Takashi Miyata
Journal:  Nucleic Acids Res       Date:  2002-07-15       Impact factor: 16.971

3.  Aligning two sequences within a specified diagonal band.

Authors:  K M Chao; W R Pearson; W Miller
Journal:  Comput Appl Biosci       Date:  1992-10

4.  Accurate sampling and deep sequencing of the HIV-1 protease gene using a Primer ID.

Authors:  Cassandra B Jabara; Corbin D Jones; Jeffrey Roach; Jeffrey A Anderson; Ronald Swanstrom
Journal:  Proc Natl Acad Sci U S A       Date:  2011-11-30       Impact factor: 11.205

5.  Degenerate Primer IDs and the birthday problem.

Authors:  Daniel J Sheward; Ben Murrell; Carolyn Williamson
Journal:  Proc Natl Acad Sci U S A       Date:  2012-04-19       Impact factor: 11.205

Review 6.  De novo assembly of short sequence reads.

Authors:  Konrad Paszkiewicz; David J Studholme
Journal:  Brief Bioinform       Date:  2010-08-19       Impact factor: 11.622

7.  Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data.

Authors:  Chen-Shan Chin; David H Alexander; Patrick Marks; Aaron A Klammer; James Drake; Cheryl Heiner; Alicia Clum; Alex Copeland; John Huddleston; Evan E Eichler; Stephen W Turner; Jonas Korlach
Journal:  Nat Methods       Date:  2013-05-05       Impact factor: 28.547

8.  MAFFT multiple sequence alignment software version 7: improvements in performance and usability.

Authors:  Kazutaka Katoh; Daron M Standley
Journal:  Mol Biol Evol       Date:  2013-01-16       Impact factor: 16.240

9.  Real-time DNA sequencing from single polymerase molecules.

Authors:  John Eid; Adrian Fehr; Jeremy Gray; Khai Luong; John Lyle; Geoff Otto; Paul Peluso; David Rank; Primo Baybayan; Brad Bettman; Arkadiusz Bibillo; Keith Bjornson; Bidhan Chaudhuri; Frederick Christians; Ronald Cicero; Sonya Clark; Ravindra Dalal; Alex Dewinter; John Dixon; Mathieu Foquet; Alfred Gaertner; Paul Hardenbol; Cheryl Heiner; Kevin Hester; David Holden; Gregory Kearns; Xiangxu Kong; Ronald Kuse; Yves Lacroix; Steven Lin; Paul Lundquist; Congcong Ma; Patrick Marks; Mark Maxham; Devon Murphy; Insil Park; Thang Pham; Michael Phillips; Joy Roy; Robert Sebra; Gene Shen; Jon Sorenson; Austin Tomaney; Kevin Travers; Mark Trulson; John Vieceli; Jeffrey Wegener; Dawn Wu; Alicia Yang; Denis Zaccarin; Peter Zhao; Frank Zhong; Jonas Korlach; Stephen Turner
Journal:  Science       Date:  2008-11-20       Impact factor: 47.728

10.  HMM-FRAME: accurate protein domain classification for metagenomic sequences containing frameshift errors.

Authors:  Yuan Zhang; Yanni Sun
Journal:  BMC Bioinformatics       Date:  2011-05-24       Impact factor: 3.169

View more
  2 in total

1.  Long-read amplicon denoising.

Authors:  Venkatesh Kumar; Thomas Vollbrecht; Mark Chernyshev; Sanjay Mohan; Brian Hanst; Nicholas Bavafa; Antonia Lorenzo; Nikesh Kumar; Robert Ketteringham; Kemal Eren; Michael Golden; Michelli F Oliveira; Ben Murrell
Journal:  Nucleic Acids Res       Date:  2019-10-10       Impact factor: 16.971

2.  A new full-length circular DNA sequencing method for viral-sized genomes reveals that RNAi transgenic plants provoke a shift in geminivirus populations in the field.

Authors:  Devang Mehta; Matthias Hirsch-Hoffmann; Mariam Were; Andrea Patrignani; Syed Shan-E-Ali Zaidi; Hassan Were; Wilhelm Gruissem; Hervé Vanderschuren
Journal:  Nucleic Acids Res       Date:  2019-01-25       Impact factor: 16.971

  2 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.