| Literature DB >> 31727106 |
Arne De Roeck1,2, Wouter De Coster1,2, Liene Bossaerts1,2, Rita Cacace1,2, Tim De Pooter3, Jasper Van Dongen1,2, Svenn D'Hert3, Peter De Rijk3, Mojca Strazisar3, Christine Van Broeckhoven1,2, Kristel Sleegers4,5.
Abstract
Technological limitations have hindered the large-scale genetic investigation of tandem repeats in disease. We show that long-read sequencing with a single Oxford Nanopore Technologies PromethION flow cell per individual achieves 30× human genome coverage and enables accurate assessment of tandem repeats including the 10,000-bp Alzheimer's disease-associated ABCA7 VNTR. The Guppy "flip-flop" base caller and tandem-genotypes tandem repeat caller are efficient for large-scale tandem repeat assessment, but base calling and alignment challenges persist. We present NanoSatellite, which analyzes tandem repeats directly on electric current data and improves calling of GC-rich tandem repeats, expanded alleles, and motif interruptions.Entities:
Keywords: ATP-binding cassette; Alzheimer’s disease; Dynamic time warping (DTW); Long-read whole genome sequencing; Member 7 (ABCA7); Sub-family A; Variable number tandem repeat (VNTR)
Year: 2019 PMID: 31727106 PMCID: PMC6857246 DOI: 10.1186/s13059-019-1856-3
Source DB: PubMed Journal: Genome Biol ISSN: 1474-7596 Impact factor: 13.583
Fig. 1Tandem repeat analysis methods. To extract TR length and sequence information from PromethION data, we used a base calling (red) and our squiggle-based NanoSatellite (blue) approach. Consecutive steps are shown in bold with below the names of the used bioinformatics tools or squiggle illustrations. The “Raw PromethION data” illustration corresponds to a partial PromethION squiggle from a single read spanning the ABCA7 VNTR. The “TR consensus sequence” figure was obtained from De Roeck et al. [22]. The height of each nucleotide corresponds to its frequency on that position [22]. In the “Reference Squiggles” figure, three ABCA7 VNTR units (alternating colors) are shown based on Scrappie current estimation. After DTW with raw PromethION data and reference squiggles, the TR is delineated from the flanking sequence and segmented into individual TR units (alternating colors in the “Delineation and segmentation” figure). The final “TR unit clustering” figure depicts the DTW process between two TR units. Each current measurement from a TR unit (red or black) is matched (gray lines) to a current measurement from the other TR unit. More detail of the NanoSatellite process is shown in Additional file 1: Figure S6
Summary of individuals included in this study
| Individual | DNA source | PromethION device | Sequencing chemistry | Number of flow cells | Yield (Gb) | Read length N50 (kb) | ||
|---|---|---|---|---|---|---|---|---|
| Subject01 | 5.1 | 4.5 | QIAmp | α | 108 | 6 + 7d | 74.4 | 7.0 |
| Subject02 | 1.8 | 1.8 | QIAmp | α | 108 | 1 | 61.3 | 11.2 |
| Subject03 | 4.7 | 2.3 | QIAmp | α | 109 | 1 | 72.5 | 16.3 |
| Subject04 | 9.7 | 0.8 | QIAmp | α | 109 | 1 | 85.5 | 11.5 |
| Subject05 | 10.7 | 1.7 | QIAmp | α | 109 | 1 | 88.5 | 11.9 |
| Subject06 | 2.3 | 0.3 | QIAmp | α | 109 | 1 | 98.0 | 16.2 |
| Subject07 | 3.3 | 3.3 | Magtration | α | 109 | 1 | 77.7 | 10.6 |
| Subject08 | 2.3 | 0.4 | Magtration | β | 109 | 1 | 56.2 | 14.7 |
| Subject09a | 5.5 | 4.3 | Magtration | β | 109 | 1 | 43.2 | 29.3 |
| Subject10b | 8.7 | 2.8 | Magtration | β | 109 | 1 | 48.9 | 9.1 |
| NA19240 | 4.5 | 2.2 | Magtration | β | 109c | 5 | 220.0 | 15.8 |
DNA was either extracted with QIAmp (Qiagen) or on a Magtration robotic platform (PSS). An alpha (α) and beta (β) PromethION device were used. Two chemistries were applied: SQK-LSK108 (108) or SQK-LSK109 (109). Yield and mean read length were calculated with NanoPack [23] after Albacore or Guppy base calling. kb kilobases, Gb gigabases, N50 50% of the total sequencing dataset is contained in reads equal or larger than this value. aDNA was not sheared for this individual. bAn 8-year-old DNA extraction was used. cSQK-LSK109 was used for four flow cells and for one flow cell was run with SQK-LSK108. dSix debug PromethION flow cells and seven MinION flow cells were used
Fig. 2ABCA7 VNTR length estimates. TR length estimates (the number of TR units is depicted on the y-axis) per positive strand (red) or negative strand (blue) PromethION sequencing reads (dots) are shown in comparison to the Southern blotting lengths (dashed lines). a Comparison of the five methods (Albacore + tandem-genotypes (tg), Scrappie (S.) events + tg, S. raw + tg, Guppy “flip-flop” + tg, and NanoSatellite) for NA19240, the individual with most sequencing reads. b Comparison of the five methods for Subject05, the individual with the largest expanded ABCA7 VNTR allele. c NanoSatellite ABCA7 VNTR length estimations for all individuals
Evaluation of tandem repeat analysis methods on the ABCA7 VNTR
| Method | Accuracy (%) | Relative standard deviation (%) | Number of spanning reads | Expanded read detection (%) | Sequence composition |
|---|---|---|---|---|---|
| Albacore + tandem-genotypes | 67.3 | 36.0 | 154 | 63 | Low consistency |
| Scrappie events + tandem-genotypes | 83.3 | 14.8 | 177 | 88 | Low consistency |
| Scrappie raw + tandem-genotypes | 87.7 | 5.5 | 181 | 25 | Low consistency |
| Guppy “flip-flop” + tandem-genotypes | 91.2 | 3.1 | 176 | 75 | Low consistency |
| NanoSatellite | 90.5 | 5.6 | 194 | 100 | High consistency |
Accuracy corresponds to the degree of resemblance of the average length estimation and Southern blotting length. The relative standard deviation depicts the spread of length estimates to the mean. The total number of ABCA7 VNTR spanning reads detected per method is shown under “Number of spanning reads”; shown in more detail in Additional file 1: Table S2. “Expanded read detection” corresponds to the proportion of expanded reads that were detected using the accompanying method compared to the method that detected most expanded reads, with 100% corresponding to the detection of all expanded reads. Sequence composition is based on the resemblance between base called sequences and the ABCA7 VNTR reference sequence for conventional tools (Additional file 1: Table S1), and the reoccurrence of alternative TR unit patterns in independent sequencing reads for NanoSatellite (Fig. 4)
Fig. 4ABCA7 VNTR sequence reconstruction based on squiggle clusters. Alignments are shown with individual PromethION sequencing reads (narrow segments) and a consensus (broad segments), as annotated in panels a and b. Each rectangle corresponds to a TR unit. Colors correspond to the TR unit cluster as assigned in Fig. 3. a Negative reads originating from the two NA19240 alleles. For both alleles of NA19240, two reads with deviating TR length were removed for clearer visualization. b Negative reads corresponding to both alleles of Subject02, for whom a single VNTR length is observed in Southern blotting. c Positive reads corresponding to the expanded alleles of Subject04, Subject05, and Subject10
Fig. 3ABCA7 VNTR squiggle clustering by NanoSatellite. Centroids are shown which were extracted from hierarchical ABCA7 VNTR squiggle unit clusters originating from positive (a) or negative (b) DNA strands. Each cluster is shown in a different color. We compared these centroids to positive (c) and negative (d) reference squiggles with corresponding sequence motifs shown below. The top sequences (blue and purple) correspond to the expected VNTR motifs, while the sequences below (orange and green) contain nucleotide differences (in bold and italic). The alternative (orange) cluster, observed in panel a, contains two alternative alleles: a guanine insertion (solid orange line in panel c) and a cytosine to adenine substitution (dashed orange line in panel c)
Fig. 5Application of NanoSatellite to other tandem repeats. Comparison of NanoSatellite and Albacore + tandem-genotypes is shown (panels) for two TRs other than the ABCA7 VNTR, with hg19 genomic coordinates and TR unit motif denoted below. The estimated number of TR units is depicted on the y-axis per sequencing read (dots) originating from positive (red) or negative strands (blue). In both examples, NanoSatellite provides a “preferred” outcome as a more concise length call can be made for the largest allele in particular (i.e., read lengths cluster more closely, and more information from both DNA strands is present)