| Literature DB >> 30258547 |
Mucheng Zhang1, Deli Liu1, Jie Tang1, Yuan Feng1, Tianfang Wang1, Kevin K Dobbin2, Paul Schliekelman3, Shaying Zhao1.
Abstract
As next-generation sequencing technology advances and the cost decreases, whole genome sequencing (WGS) has become the preferred platform for the identification of somatic copy number alteration (CNA) events in cancer genomes. To more effectively decipher these massive sequencing data, we developed a software program named SEG, shortened from the word "segment". SEG utilizes mapped read or fragment density for CNA discovery. To reduce CNA artifacts arisen from sequencing and mapping biases, SEG first normalizes the data by taking the log2-ratio of each tumor density against its matching normal density. SEG then uses dynamic programming to find change-points among a contiguous log2-ratio data series along a chromosome, dividing the chromosome into different segments. SEG finally identifies those segments having CNA. Our analyses with both simulated and real sequencing data indicate that SEG finds more small CNAs than other published software tools.Entities:
Keywords: Cancer; SEG; Somatic Copy Number Alteration; Whole Genome Sequencing
Year: 2018 PMID: 30258547 PMCID: PMC6154469 DOI: 10.1016/j.csbj.2018.09.001
Source DB: PubMed Journal: Comput Struct Biotechnol J ISSN: 2001-0370 Impact factor: 7.271
Fig. 1The algorithm of SEG. SEG will: 1)normalize the data and exclude the log2-ratio outliers (smooth data); 2)identify change-points; and 3)find CNAs (label segments).For change-point detection, SEG first depends upon the user's input to assign initial change-points, and then loops through the SSE (sum of squared error) to remove insignificant change-points using dynamic programming (see text).The program is implemented in C and can be downloaded from GitHub at https://github.com/ZhaoS-Lab/SEG.
Fig. 2CNAs identified by SEG, BICseq, FREEC and CBS in 10 simulated samples of chromosome 22.A.Amplifications and deletions of ground truth, and those identified by SEG or other software tools drew as described18 for one simulated sample. B. Heatmaps showing the overall sensitivity and specificity of CNA detection in each of the 10 simulated samples by SEG or other software tools. C. Heatmaps showing the overall sensitivity of CNA detection based on the size by SEG or other software tools. D. Heatmaps showing the overall sensitivity of CNA detection for each category indicated by SEG or other software tools.
Fig. 3Data normalization in the three canine mammary cancer genomes. A.The distribution of average mapped fragment density, d, of 100 bp tilting window of the tumor and normal genome of the cancer cases with ID indicated. B. The distribution of the normalized density against its genome wide average by. C. The distribution of the final normalized density of the tumor against the matching normal data by (equation).
Fig. 4Large CNAs identified with WGS (A)and aCGH (B)by SEG.Each line represents a dog chromosome with its chromosome number indicated on the left.Red (amplifications) and blue (deletions) vertical lines shown above the chromosomes are drew as previously described4.Only CNAs of >8.5 kb were plotted, as 8.5 kb is the minimal size of CNAs found by aCGH.
Small CNAs of ≤3 kb identified by SEG from WGS data.
| Tumor ID | Amplification | Deletion | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Total Amount | Average size | Exon content | GC content | Repeats content | Total Amount | Average size | Exon content | GC content | Repeats content | |
| 32,510 | 8.7 Mb | 418 bp | 1/4.9kb | 47.0% | 28.7% | 12.7 Mb | 443 bp | 1/8.5 kb | 40.8% | 36.40% |
| 76 | 36.2 Mb | 308 bp | 1/6.9 kb | 42.4% | 33.0% | 44.5 Mb | 318 bp | 1/5.5 kb | 44.6% | 27.5% |
| 406,434 | 32.1 Mb | 621 bp | 1/7.3 kb | 40.0% | 33.9% | 56.6 Mb | 673 bp | 1/10.5 kb | 40.0% | 32.7% |
The calculations are based on the canFam2 genome assembly, Ensembl gene annotation release-65 (exon content), and RepeatMasker 4.0.5 with repeats database Dfam_2.0.
One exon every 4.5 kb on average. Genome wide: 1/11.7 kb.
Large CNAs of >3 kb identified by SEG from WGS data.
| Tumor ID | Amplification | Deletion | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Total Amount | Average size | Exon content | GC content | Repeatscontent | Total Amount | Average size | Exon content | GC content | Repeats content | |
| 32,510 | None | None | ||||||||
| 76 | 9.2 Mb | 74,656 bp | 1/8.0 kb | 43.5% | 35.6% | None | ||||
| 406,434 | 67.5 Mb | 13,575 | 1/13.2 kb | 40.3% | 35.3% | 1.7 Mb | 4017 bp | 1/17.7 kb | 37.9% | 36.1% |