| Literature DB >> 35165687 |
Tavis K Anderson1, Blake Inderski1, Diego G Diel2,3,4, Benjamin M Hause2,3, Elizabeth G Porter5,6, Travis Clement2,3, Eric A Nelson2,3, Jianfa Bai5,6, Jane Christopher-Hennings2,3, Phillip C Gauger7,8, Jianqiang Zhang7,8, Karen M Harmon7,8, Rodger Main7,8, Kelly M Lager1, Kay S Faaberg1.
Abstract
Veterinary diagnostic laboratories derive thousands of nucleotide sequences from clinical samples of swine pathogens such as porcine reproductive and respiratory syndrome virus (PRRSV), Senecavirus A and swine enteric coronaviruses. In addition, next generation sequencing has resulted in the rapid production of full-length genomes. Presently, sequence data are released to diagnostic clients but are not publicly available as data may be associated with sensitive information. However, these data can be used for field-relevant vaccines; determining where and when pathogens are spreading; have relevance to research in molecular and comparative virology; and are a component in pandemic preparedness efforts. We have developed a centralized sequence database that integrates private clinical data using PRRSV data as an exemplar, alongside publicly available genomic information. We implemented the Tripal toolkit, a collection of Drupal modules that are used to manage, visualize and disseminate biological data stored within the Chado database schema. New sequences sourced from diagnostic laboratories contain: genomic information; date of collection; collection location; and a unique identifier. Users can download annotated genomic sequences using a customized search interface that incorporates data mined from published literature; search for similar sequences using BLAST-based tools; and explore annotated reference genomes. Additionally, custom annotation pipelines have determined species, the location of open reading frames and nonstructural proteins and the occurrence of putative frame shifts. Eighteen swine pathogens have been curated. The database provides researchers access to sequences discovered by veterinary diagnosticians, allowing for epidemiological and comparative virology studies. The result will be a better understanding on the emergence of novel swine viruses and how these novel strains are disseminated in the USA and abroad. Database URLhttps://swinepathogendb.org. Published by Oxford University Press 2021. This work is written by (a) US Government employee(s) and is in the public domain in the US.Entities:
Mesh:
Year: 2021 PMID: 35165687 PMCID: PMC8903347 DOI: 10.1093/database/baab078
Source DB: PubMed Journal: Database (Oxford) ISSN: 1758-0463 Impact factor: 3.451
Figure 2.Genome annotation for the United States Swine Pathogen Database. Genome annotation begins with preprocessing, which requires nucleotide multiple sequence alignment files (MSA) representing genomic features as input (①). The products of preprocessing (②) are a nucleotide profile hidden Markov model (HMM) and a structured file containing regular expression patterns representative of diversity within sections of translated input MSAs. Following preprocessing, query nucleotide sequences in FASTA format (③) may be supplied to the annotation pipeline. If species identification is necessary (e.g. differentiating type 1 and 2 PRRSV), BLAST is performed (④) using preprocessing input files (①) as a reference. Once the query sequence species is known, relevant preprocessing files (②) are selected. Regular expressions are matched against three reading frames (⑤) of query sequence (③) to determine the location of genomic features. If more processing is necessary due to frame changes or uncertainty in the start or stop position, a profile HMM alignment is performed (⑥). These steps produce genome annotation and additional information with high confidence. The output produced by the pipeline is general feature format (GFF3) file, a standard nucleic acid or protein feature file format (⑦). The annotation pipeline is available at https://github.com/us-spd/.
Figure 1.Conceptual model describing the automated pipeline implemented in the US Swine Pathogen Database that takes raw sequence data to fully anonymized and annotated virus sequence record in the relational database.
Figure 3.Observed frequency of the top 15 most commonly detected restriction fragment length polymorphism (RFLP) patterns in Type 2 porcine reproductive and respiratory syndrome virus sampled in the USA from 2000 to present (n = 16 403). Less common RFLP patterns were grouped together and labeled as “Other”.
Figure 4.Phylogenetic network of porcine reproductive and respiratory syndrome virus collected in the Midwest of the USA from 2014 to present (a) Putative recombination events are indicated by verticle lines within the network, annotated by red stars. Panel (b) strain IA30788-R (GB043798), Panel (c) strain 23199-S4-L001 (GB043498) and Panel (d) strain 7705R-S1 (GB043536) are visualized separately to demonstrate evolutionary relationships and recombination nodes. Phylogenetic network with tip labels, recombination layers 1 through 20 and gene tree embeddings is provided at https://github.com/us-spd/.