| Literature DB >> 31610622 |
Pora Kim1, Ye Eun Jang1,2, Sanghyuk Lee1,2,3.
Abstract
Identification of fusion gene is of prominent importance in cancer research field because of their potential as carcinogenic drivers. RNA sequencing (RNA-Seq) data have been the most useful source for identification of fusion transcripts. Although a number of algorithms have been developed thus far, most programs produce too many false-positives, thus making experimental confirmation almost impossible. We still lack a reliable program that achieves high precision with reasonable recall rate. Here, we present FusionScan, a highly optimized tool for predicting fusion transcripts from RNA-Seq data. We specifically search for split reads composed of intact exons at the fusion boundaries. Using 269 known fusion cases as the reference, we have implemented various mapping and filtering strategies to remove false-positives without discarding genuine fusions. In the performance test using three cell line datasets with validated fusion cases (NCI-H660, K562, and MCF-7), FusionScan outperformed other existing programs by a considerable margin, achieving the precision and recall rates of 60% and 79%, respectively. Simulation test also demonstrated that FusionScan recovered most of true positives without producing an overwhelming number of false-positives regardless of sequencing depth and read length. The computation time was comparable to other leading tools. We also provide several curative means to help users investigate the details of fusion candidates easily. We believe that FusionScan would be a reliable, efficient and convenient program for detecting fusion transcripts that meet the requirements in the clinical and experimental community. FusionScan is freely available at http://fusionscan.ewha.ac.kr/.Entities:
Keywords: RNA-Seq; chromosomal translocation; fusion transcript; gene fusion; transcriptome sequencing
Year: 2019 PMID: 31610622 PMCID: PMC6808644 DOI: 10.5808/GI.2019.17.3.e26
Source DB: PubMed Journal: Genomics Inform ISSN: 1598-866X
Fig. 1.Overview of FusionScan algorithm. Computational pipeline is shown with programs used in each step. The statistics for processing K562 RNA sequencing data illustrates the effect of each procedure on reducing the candidates of split reads and fusion gene pairs. ENCODE, Encyclopedia of DNA Elements.
Comparison of RNA-Seq alignment programs
| Mapping program | No. of correct alignments out of 269 known fusion transcripts[ | ||
|---|---|---|---|
| 50 bp | 75 bp | 100 bp | |
| GMAP | 59 | 28 | 3 |
| SSAHA2 | 237 | 248 | 252 |
| Bowtie2 | 242 | 245 | 248 |
| BWA | 1 | 238 | 244 |
| BLAT | 218 | 225 | 226 |
| TopHat2 | 227 | 228 | 226 |
All alignment tools were run with default options.
RNA-Seq, RNA sequencing.
Two hundred sixty-nine known fusion transcripts were collected from TICdb and ChimerDB 2.0.
Fig. 2.Alignment and coverage plots. BCR-ABL1 gene fusion detected from RNA sequencing data of K562 cell line is shown as an example. (A) Fusion alignment view is the read alignment of seed and support reads on hypothetical fusion transcript. (B) Genome alignment view shows the alignment of split reads on the University of California Santa Cruz (UCSC) genome browser for head and tail genes obtained by BLAT alignment tool. (C) Coverage plots on transcript coordinate show abrupt change in read depth at the fusion boundary for both head and tail genes. Blue vertical lines indicate the exon boundaries in each gene.
Identification of known fusion genes by various fusion detection tools
| Sample | Known (Gold) fusion genes | FS | SF | dF | FH | FM | THF |
|---|---|---|---|---|---|---|---|
| NCI-H660 (2) | TMPRSS2-ERG | ● | ● | ● | ● | ● | ● |
| EEF2-SLC25A42 | ● | ● | ● | ● | - | ● | |
| TP/FP | 2/0 | 2/16 | 2/11 | 2/1 | 1/1 | 2/1 | |
| Precision | 1.0 | 0.11 | 0.15 | 0.67 | 0.50 | 0.67 | |
| Recall | 1.0 | 1.0 | 1.0 | 1.0 | 0.50 | 1.0 | |
| K562 (3) | BCR-ABL1 | ● | ● | ● | - | ● | ● |
| NUP214-XKR3 | ● | ● | ○ | ● | ● | ● | |
| BAT3-SLC44A4 | ● | ● | ○ | - | ● | ● | |
| TP/FP | 3/1 | 3/7 | 3/27 | 1/1 | 3/12 | 3/0 | |
| Precision | 0.75 | 0.30 | 0.10 | 1.0 | 0.20 | 1.0 | |
| Recall | 1.0 | 1.0 | 1.0 | 0.33 | 1.0 | 1.0 | |
| MCF-7 (23) | USP31-CRYL1 | ● | ● | ○ | ○ | ○ | ● |
| ARFGEF2-SULF2 | ● | ● | ○ | - | ○ | ● | |
| TXLNG-SYAP1 | ● | ● | ● | - | ○ | ● | |
| DEPDC1B-ELOVL7 | ● | ● | ● | ○ | ● | ● | |
| SYTL2-PICALM | ● | ● | ● | ○ | ● | ● | |
| RPS6KB1-DIAPH3 | - | ● | ● | - | - | - | |
| AHCYL1-RAD51C | - | ● | - | ● | ● | ● | |
| TAF4-BRIP1 | - | ● | - | - | ○ | - | |
| POP1-MATN2 | ● | ● | ● | - | ○ | - | |
| GCN1L1-MSI1 | ● | ● | ● | - | - | ● | |
| ESR1-CCDC170 | ● | ● | ● | ● | ○ | ● | |
| SMARCA4-CARM1 | ● | ● | ● | - | ○ | ● | |
| MYO6-SENP6 | ● | ● | ● | ● | ○ | ● | |
| ADAMTS19-SLC27A6 | ● | ● | ● | ● | ○ | ● | |
| GATAD2B-NUP210L | - | ● | ● | - | ○ | ● | |
| SLC25A24-NBPF6 | ● | ● | ● | - | ● | - | |
| ATXN7L3-FAM171A2 | ● | ● | ● | - | ● | - | |
| C16orf62-IQCK | ● | ● | ● | - | ● | ● | |
| TBL1XR1-RGS17 | ● | - | - | - | - | ||
| BCAS4-BCAS3 | ● | ● | ● | ● | - | ● | |
| RPS6KB1-TMEM49 | ● | ● | ● | - | ○ | ● | |
| ABCA5-PPP4R1L | - | ● | - | - | - | ||
| C16orf45-ABCC1 | - | - | - | - | - | - | |
| TP/FP | 17/14 | 21/83 | 18/132 | 8/11 | 17/126 | 15/26 | |
| Precision | 0.55 | 0.18 | 0.12 | 0.42 | 0.12 | 0.37 | |
| Recall | 0.74 | 0.91 | 0.78 | 0.35 | 0.74 | 0.65 | |
| Overall | Precision | 0.60 | 0.20 | 0.12 | 0.46 | 0.13 | 0.43 |
| Recall | 0.79 | 0.93 | 0.82 | 0.39 | 0.75 | 0.71 | |
| F1 score | 0.68 | 0.33 | 0.21 | 0.42 | 0.22 | 0.53 | |
| Mapping program | SSAHA2 | SOAP2 | GMAP | Bowtie | GSNAP | Bowtie | |
| BWA | |||||||
| Transcriptome | RefGene | Ensembl | Ensembl | RefGene | RefGene | Ensembl | |
‘●’ and ‘○’ indicate that the case was predicted successfully, with direction reversed in ‘○’. Precision = TP/(TP + FP), Recall = TP/(TP + FN), F1 score = 2 × precision×recall/(precision + recall).
FS, FusionScan; SF, SOAPfuse; dF, defuse; FH, FusionHunter; FM, FusionMap; THF, TopHat-Fusion; TP, true-positive; FP, false-positive; FN, false-negative.
Fig. 3.Venn diagram of fusion predictions. We show the total number of predicted fusion genes for all three cell lines from five different programs. Numbers in the parenthesis indicate the number of true positive cases.
Fig. 4.Precision and recall curves for performance evaluation using simulation data sets. (A) Precision rates at the read length of 100 bp, 75 bp, and 50 bp. (B) Recall rates at the read length of 100 bp, 75 bp, 50 bp. 10×, 30×, and 50× indicate the sequencing depth of the simulation data.
Fig. 5.Comparison of CPU time (A) and memory usage (B). CPU time and memory usage are shown in hours and GB, respectively. RNA sequencing data for K562 cell line from the Encyclopedia of DNA Elements (ENCODE) project (SRR521464) was analyzed on a 64-bit machine AMD Opteron Processor 6176 (2.3 GHz, 8 core) with 32 GB RAM.