| Literature DB >> 30610921 |
Stephen E Lincoln1, Rebecca Truty2, Chiao-Feng Lin3, Justin M Zook4, Joshua Paul2, Vincent H Ramey2, Marc Salit5, Heidi L Rehm6, Robert L Nussbaum7, Matthew S Lebo8.
Abstract
Orthogonal confirmation of next-generation sequencing (NGS)-detected germline variants is standard practice, although published studies have suggested that confirmation of the highest-quality calls may not always be necessary. The key question is how laboratories can establish criteria that consistently identify those NGS calls that require confirmation. Most prior studies addressing this question have had limitations: they have been generally of small scale, omitted statistical justification, and explored limited aspects of underlying data. The rigorous definition of criteria that separate high-accuracy NGS calls from those that may or may not be true remains a crucial issue. We analyzed five reference samples and over 80,000 patient specimens from two laboratories. Quality metrics were examined for approximately 200,000 NGS calls with orthogonal data, including 1662 false positives. A classification algorithm used these data to identify a battery of criteria that flag 100% of false positives as requiring confirmation (CI lower bound, 98.5% to 99.8%, depending on variant type) while minimizing the number of flagged true positives. These criteria identify false positives that the previously published criteria miss. Sampling analysis showed that smaller data sets resulted in less effective criteria. Our methodology for determining test- and laboratory-specific criteria can be generalized into a practical approach that can be used by laboratories to reduce the cost and time burdens of confirmation without affecting clinical accuracy.Entities:
Mesh:
Year: 2019 PMID: 30610921 PMCID: PMC6629256 DOI: 10.1016/j.jmoldx.2018.10.009
Source DB: PubMed Journal: J Mol Diagn ISSN: 1525-1578 Impact factor: 5.568
Summary of the Data Sets
| Source | Variant type | Samples | Unique variants | Variant calls | TPs | FPs | FDR, % | Total calls | Total FPs | FP sensitivity, % | CI lower bound, % |
|---|---|---|---|---|---|---|---|---|---|---|---|
| This study: Lab 1 | SNVs | GIAB | 27,202 | 136,146 | 135,945 | 201 | 0.15 | ||||
| Patients | 2840 | 3699 | 3689 | 10 | 0.27 | 139,845 | 211 | 100 | 98.9 | ||
| Indels | GIAB | 3715 | 15,574 | 14,594 | 980 | 6.29 | |||||
| Patients | 1749 | 2274 | 2262 | 12 | 0.53 | 17,848 | 992 | 100 | 99.8 | ||
| This study: Lab 2 | SNVs | GIAB | 5816 | 29,148 | 29,110 | 38 | 0.13 | ||||
| Patients | 4359 | 4934 | 4804 | 130 | 2.63 | 34,082 | 168 | 100 | 98.5 | ||
| Indels | GIAB | 1185 | 3617 | 3343 | 274 | 7.58 | |||||
| Patients | 267 | 389 | 372 | 17 | 4.37 | 4006 | 291 | 100 | 99.1 | ||
| Strom et al | SNVs | Patients | - | 108 | 107 | 1 | 0.93 | 108 | 1 | 100 | 5.1 |
| Baudhuin et al | SNVs | Patients | 380 | 797 | 797 | 0 | 0 | ||||
| 1KG | 736 | 736 | 736 | 0 | 0 | 1533 | 0 | N/A | N/A | ||
| Indels | Patients | 63 | 122 | 122 | 0 | 0 | |||||
| 1KG | 26 | 26 | 26 | 0 | 0 | 148 | 0 | N/A | N/A | ||
| Mu et al | SNVs | Patients | - | 6912 | 6818 | 94 | 1.36 | 6912 | 94 | 100 | 97.4 |
| Indels | Patients | - | 933 | 928 | 5 | 0.54 | 933 | 5 | 100 | 62.1 | |
| van den Akker et al | SNVs | Patients | 3044 | 5829 | 5524 | 305 | 5.23 | 5829 | 305 | 100 | 99.2 |
| Indels | Patients | 526 | 1350 | 1142 | 208 | 15.41 | 1350 | 208 | 100 | 98.8 |
1KG, 1000 Genomes Project; FDR, false discovery rate [calculated as FPs/(FPs + TPs)]; FP, false positive; FP sensitivity, the fraction of FPs captured using the study's proposed criteria; GIAB, Genome in a Bottle Consortium; Indel, insertion or deletion; SNV, single-nucleotide variant; unique variant, a particular alteration which may be present in one or more individuals; TP, true positive.
Clinical and GIAB data were combined in the final analysis. GIAB data included on- and off-target calls.
Manual review removed certain FPs (particularly indels) and thus reduced FDRs in the Lab 1 clinical data.
Many of the clinical FPs were systematic errors in the OTOA and CFTR genes, which were tested in many patients.
The authors did not provide a count of unique variants. For the van den Akker study, it was calculated from the data provided.
CIs were calculated based on data from the publication. No such statistics were provided by the study authors.
The lack of FPs may have been a result of aggressive filtering, which can remove clinical TPs as well as FPs.
Only unique 1KG variants were analyzed. The results from the updated 1KG data mentioned in this article are described.
Most of the indels in patients were intronic and homopolymer associated. These are generally not clinically significant.
The relatively high FDRs in this data set may have been a result of under-filtering, which can also affect CIs.
Figure 1Study methodology. A: Variant calls can be classified as high-confidence true positive (green) and intermediate-confidence (yellow), using strict thresholds intended to maximize the specificity of the high-confidence set. The intermediate set will contain a mixture of true- and false-positive calls. The study objective was to rigorously determine test-specific criteria that distinguish these two categories. Variant calls that are confidently false positives (red) are typically filtered out using different criteria that emphasize sensitivity. The analogy to a traffic light is illustrative. B: Process for collecting true positive and false positive variant calls for both the clinical and Genome in a Bottle Consortium (GIAB) specimens. Each laboratory's clinically validated next-generation sequencing (NGS) assays, bioinformatics pipelines, and filtering criteria were used for both specimen types. Single-nucleotide variants and indels were collected as shown. Copy number and structural variants were excluded from this study, as were any variants with an unknown confirmation status. Filtering and manual review processes were designed to remove clearly false variant calls but not those considered even potentially true. Manual review was used only with the Lab 1 clinical data. HC, high-confidence.
Key Aspects of Methodology Used in This Study
| Aspect of methodology used | Rationale (see text for details) |
|---|---|
| Large data sets used for both SNVs and indels | Provides confidence in the resulting criteria and helps minimize overfitting. Appropriate sizes determined by CI calculations (below). |
| Both clinical and reference (GIAB) samples used | Greatly increases the size and diversity of the data sets, particularly in FPs. |
| Same quality filtering thresholds as used in clinical practice | Confirmation criteria can depend on filtering criteria. Using lower filters would add many FPs to the study but could result in the selection of ineffective or biased confirmation criteria. |
| Separate filtering and confirmation thresholds used | Allows high sensitivity (by keeping variants of marginal quality) and high specificity (by subjecting these variants to confirmation). |
| On- and off-target variant calls analyzed in GIAB samples | Further increases the data set size and diversity. Off-target calls were subject to the same quality filters as were on-target calls. |
| Indels and SNVs analyzed separately | Indels and SNVs can have different quality determinants. An adequate population of each was required to achieve statistical significance. |
| Partial matches considered FPs | Zygosity errors and incorrect diploid genotypes do occur and can be as clinically important as “pure” FPs. |
| Algorithm selects criteria primarily by their ability to flag FPs | Other algorithms equally value classification of TPs, which may result in biased criteria, particularly as TPs greatly outnumber FPs. |
| Multiple quality metrics used to flag possible FPs | Various call-specific metrics (eg, quality scores, read depth, allele balance, strand bias) and genomic annotations (eg, repeats, segmental duplications) proved crucial. |
| Key metric: fraction of FPs flagged (FP sensitivity) | Other metrics, including test PPV and overall classifier accuracy, can be uninformative or misleading in the evaluation of confirmation criteria. |
| Requirement of 100% FP sensitivity on training data | Clinically appropriate. Resulting criteria will be effective on any subset of the training data (clinical or GIAB, on-target or off-target, etc.) |
| Statistical significance metric: CI on FP sensitivity | Rigorously indicates validity of the resulting criteria: eg, flagging 100% of 125 FPs demonstrates ≥98% FP sensitivity at |
| Separate training and test sets used (cross validation) | In conjunction with large data sets, cross-validation is a crucial step to avoid overfitting, which can otherwise result in ineffective criteria. |
| Prior confirmation of a variant was not used as a quality metric | Successful confirmations of a particular variant can indicate little about whether future calls of that same variant are true or false. |
| All variants outside of GIAB-HC regions require confirmation | Outside of these regions, too few confirmatory data are available to prove whether the criteria are effective. |
| Laboratory- and workflow-specific criteria | Effective confirmation criteria can vary based on numerous details of a test's methodology and its target genes. Changes can necessitate revalidation of confirmation criteria. |
FP, false positive; GIAB, Genome in a Bottle; GIAB-HC, regions in which high-confidence truth data are available from the GIAB specimens (unrelated to confidence in our own calls or to on/off-target regions); indel, insertion or deletion; PPV, positive predictive value; SNV, single-nucleotide variant; TP, true positive.
Figure 2Individual quality metric results and statistics. A: Variant call quality scores (QUAL, y axis) for true-positive (TP; blue) and false-positive (FP; red) single-nucleotide variant (SNV) calls in the Lab 1 Genome in a Bottle Consortium (GIAB) data set. The x axis position of each point is randomly assigned. To make density changes visible, a random selection of one-fifth of the TP calls was plotted along with all FPs. In the large data set, some FPs had quite high QUAL scores, demonstrating that this metric is inadequate alone. Corresponding histograms using the full data set without down-sampling are shown in Supplemental Figure S1. B: Random sample of 1000 data points from the same data set plotted in A. All points are displayed. Arrows indicate the two FPs present. One thousand such random samples were generated, and compared with the full data set in A, many would lead to quite different conclusions about the effectiveness of QUAL thresholds. C: Lower bound of the 95% CI on the fraction of FPs flagged (y axis) as a function of the number of FPs used to determine criteria (x axis) assuming that 100% success is observed. The y axis range is 90% to 100%. This calculation used the Jeffreys method (blue), the Wilson score method (red), and the tolerance interval method (gray). All methods produce generally similar results and indicate the validity of any study such as this. For example, using the Jeffreys method, flagging 49 of 49 FPs shows 100% effectiveness, with a CI of 95% to 100%. Many prior studies did not achieve this level of statistical significance (Table 1). Consistent with these CI calculations, small data sets indeed resulted in ineffective criteria (see Results). D: Histogram of per-variant false-discovery rates (FDRs; x axis) for all variants that were observed more than once in the Lab 1 data set, and for which one or more of those calls was an FP. SNVs and insertions and deletions (indels) are combined. An FDR of 100% indicates a fully systematic FP (insofar as we can measure); an FDR of 0% indicates a consistent TP (not shown in this graph). Each unique variant (ie, a genetic alteration that may be present in multiple individuals) is counted once. The y axis range is 0% to 50% of variants. Approximately half of all variants that were FPs were also correctly called as TPs in a different specimen(s) or run(s). Examples of this were observed in both the clinical and the GIAB specimens and included both SNVs and indels. Lab 2 results were similar. Many of these variants have low per-variant FDRs, which usually but not always are correctly called. Repeated TP observations of such a variant provide little information about the accuracy of any following observation of that same variant. This study was underpowered to measure FDRs near 0% or 100%, and many more of these variants may exist than are shown here.
Figure 3Combining flags. These plots show the cumulative effect (top to bottom) of sequentially combining flags chosen by our algorithm. A and B: Single-nucleotide variants (SNVs) (A) and insertions and deletions (indels; B) in the Lab 1 data set. Red indicates the fraction of all false positives (FPs) captured; blue, the fraction of clinical true positives (TPs); gray, the fraction of FPs captured by two or more flags. The dashed lines illustrate the flags needed to capture 100% of the FPs using at least one flag each. To be conservative, the full set of flags shown here was used (maximizing double coverage) and, in particular, required confirmation of all repeat-associated calls. Note that in the indel analysis, higher QUAL and QD thresholds were required to maximize double coverage than were required to achieve 100% capture of variants by a single flag each. The flags include: DP, read depth, specifically GT_DP; DP/DP, ratio of GT_DP to INFO_DP; SB5 and SB25, strand bias metrics; QUAL, quality score; QD, quality-depth score; ABnorm, allele balance for heterozygotes, normalized to be within 0.0 to 0.5; ABhom, allele balance for homozygotes; HetAlt, heterozygous call for which neither allele is in the GRCh37 reference genome; Repeat, variant call within a homopolymer or short tandem repeat.