| Literature DB >> 35438138 |
Dietmar Fernández-Orth1, Manuel Rueda1, Babita Singh1, Mauricio Moldes1, Aina Jene1, Marta Ferri1, Claudia Vasallo1, Lauren A Fromont1, Arcadi Navarro1, Jordi Rambla1.
Abstract
Since its launch in 2008, the European Genome-Phenome Archive (EGA) has been leading the archiving and distribution of human identifiable genomic data. In this regard, one of the community concerns is the potential usability of the stored data, as of now, data submitters are not mandated to perform any quality control (QC) before uploading their data and associated metadata information. Here, we present a new File QC Portal developed at EGA, along with QC reports performed and created for 1 694 442 files [Fastq, sequence alignment map (SAM)/binary alignment map (BAM)/CRAM and variant call format (VCF)] submitted at EGA. QC reports allow anonymous EGA users to view summary-level information regarding the files within a specific dataset, such as quality of reads, alignment quality, number and type of variants and other features. Researchers benefit from being able to assess the quality of data prior to the data access decision and thereby, increasing the reusability of data (https://ega-archive.org/blog/data-upcycling-powered-by-ega/).Entities:
Keywords: European Genome-Phenome Archive (EGA); Fastq; binary alignment map (BAM); quality control; variant call format (VCF)
Mesh:
Year: 2022 PMID: 35438138 PMCID: PMC9116225 DOI: 10.1093/bib/bbac136
Source DB: PubMed Journal: Brief Bioinform ISSN: 1467-5463 Impact factor: 13.994
Figure 1EGA website. Primary file information and how to access the QC report. (A) List with the EGA ID files composing the dataset. (B) Link to the File QC for each specific ID (https://ega-archive.org/datasets/EGAD00001004220/files).
Figure 2QC File Information section for a BAM file from H3AFRICA TRYPANOGEN2. (A) File Information section with general data about the bam file. (B) Link to bam header and plots generated by bamstats plot plugin from SAM tools (https://filesportal.ega-archive.org/EGAF00002051993).
Figure 3Left. Detailed QC plots for BAM files. (A) Base coverage distribution and base quality plots. (B) Example description for forward strand plot and pie chart showing % of proper pairs found in the H3AFRICA TRYPANOGEN2 BAM file (https://filesportal.ega-archive.org/EGAF00002051993).
Figure 4Pie chart showing number and percentages of NGS files at the EGA (update November 2021). Source: https://ega-archive.org/about/ega-statistics.