| Literature DB >> 25788829 |
Hatice Gulcin Ozer1, Aisulu Usubalieva1, Adrienne Dorrance2, Ayse Selen Yilmaz1, Michael Caligiuri2, Guido Marcucci2, Kun Huang1.
Abstract
The genome-wide discoveries such as detection of copy number alterations (CNA) from high-throughput whole-genome sequencing data enabled new developments in personalized medicine. The CNAs have been reported to be associated with various diseases and cancers including acute myeloid leukemia. However, there are multiple challenges to the use of current CNA detection tools that lead to high false-positive rates and thus impede widespread use of such tools in cancer research. In this paper, we discuss these issues and propose possible solutions. First, since the entire genome cannot be mapped due to some regions lacking sequence uniqueness, current methods cannot be appropriately adjusted to handle these regions in the analyses. Thus, detection of medium-sized CNAs is also being directly affected by these mappability problems. The requirement for matching control samples is also an important limitation because acquiring matching controls might not be possible or might not be cost efficient. Here we present an approach that addresses these issues and detects medium-sized CNAs in cancer genomes by (1) masking unmappable regions during the initial CNA detection phase, (2) using pool of a few normal samples as control, and (3) employing median filtering to adjust CNA ratios to its surrounding coverage and eliminate false positives.Entities:
Keywords: acute myeloid leukemia (AML); copy number alteration (CNA); whole-genome sequencing
Year: 2015 PMID: 25788829 PMCID: PMC4356486 DOI: 10.4137/CIN.S14023
Source DB: PubMed Journal: Cancer Inform ISSN: 1176-9351
Percentage of unmappable regions of the mouse (mm9) and human (hg19) reference genomes.
| SEQUENCE READ LENGTH | MOUSE REFERENCE GENOME | HUMAN REFERENCE GENOME | ||
|---|---|---|---|---|
| GEM SCORE <1 | UNMAPPABLE | GEM SCORE <1 | UNMAPPABLE | |
| 36 mer | 28% | 55% | 29% | 64% |
| 40 mer | 26% | 52% | 27% | 60% |
| 50 mer | 23% | 46% | 23% | 51% |
| 75 mer | 18% | 33% | 16% | 29% |
| 100 mer | 16% | 27% | 14% | 20% |
Notes: Second and fourth columns show the percentages of the mouse and human genome with a GEM score less than 1 (this means the region is not unique). If an unmappable region covered more than 100 bp, we extended such region by 1 kb and consolidated regions. This aligns with our 90% cutoff on mappability percentage. Total size of these extended regions gives a more realistic idea about overall unmappability of the genome. We report percentages of unmappable regions in the third and fifth columns for mouse and human genomes, respectively.
Figure 1Complete CNA identification workflow.
Figure 2An example of highly unmappable region from chromosome 12. The selected region contains about 3 million bases. It is highly unmappable and also contains segmental duplications. The dark gray track shows mappability of mouse genome (mm9) with 75-bp reads. The red tracks show coverage of 3 PTD/ITD (Mll [PTD/wt]: Flt3 [ITD/ITD] double knock-in) mouse samples. The green tracks show coverage of 3 PTD (nonleukemic Mll [PTD/wt]) mouse samples, while the blue tracks show the coverage of ITD (Flt3 [ITD/wt] single knock-in) mouse samples.
Figure 3At the top: Integrated Genomic Viewer screenshot of a region in chromosome 1 with 3-kb candidate CNA. The dark gray track shows mappability of mouse genome (mm9) with 75-bp reads. The red and blue tracks show coverage of 3 PTD/ITD (Mll [PTD/wt]: Flt3 [ITD/ITD] double knock-in) and 2 WT mouse samples in 100-bp resolution, respectively. Green intervals show 1-kb regions with unmappable bases. At the bottom: Log2 ratio for the first PTD/ITD sample versus the first WT sample (black solid line) and its median filtering (red dotted line) in 100-bp resolution across the region. Size of the detected CNA is denoted with w.
Figure 4The region of the chromosome 9 that has Mll1 gene. A slight enrichment is observed around PTD region in the coverage tracks of PTD/ITD and PTD samples.
Figure 5A false-positive example. Genomic coverage, original log2 ratios, and median filtered log2 ratios are depicted for a 3-kb candidate CNA on chromosome 1. Slight genomic coverage change originally reported as a candidate CNA suggesting 49% loss. Genomic coverage upstream and downstream of the candidate was taken into account by median filter adjustment, and adjusted ratios suggest only 15% loss.
Combinations of various bin size and lambda parameters used in the detection of CNAs with BIC-seq toolBIC-seq parameters.
| AVERAGE # OF CNAs | % ON COMPLETELY MAPPABLE REGIONS | AVERAGE SIZE OF CNAs | ||
|---|---|---|---|---|
| Bin size = 1000 | Lambda = 10 | 283 | 4.6% | 19,363 |
| Bin size = 100 | Lambda = 10 | 330 | 5.3% | 7, 7 8 9 |
| Bin size = 100 | Lambda = 4 | 928 | 6.6% | 3,714 |
Notes: On average only about 5.5% of CNAs were coming from securely mappable regions, while the rest was from unmappable regions of the genome.