| Literature DB >> 18826589 |
Shaolin Wang1, Zhenxia Sha, Tad S Sonstegard, Hong Liu, Peng Xu, Benjaporn Somridhivej, Eric Peatman, Huseyin Kucuktas, Zhanjiang Liu.
Abstract
BACKGROUND: SNPs are abundant, codominantly inherited, and sequence-tagged markers. They are highly adaptable to large-scale automated genotyping, and therefore, are most suitable for association studies and applicable to comparative genome analysis. However, discovery of SNPs requires genome sequencing efforts through whole genome sequencing or deep sequencing of reduced representation libraries. Such genome resources are not yet available for many species including catfish. A large resource of ESTs is to become available in catfish allowing identification of large number of SNPs, but reliability of EST-derived SNPs are relatively low because of sequencing errors. This project was designed to answer some of the questions relevant to quality assessment of EST-derived SNPs.Entities:
Mesh:
Year: 2008 PMID: 18826589 PMCID: PMC2570692 DOI: 10.1186/1471-2164-9-450
Source DB: PubMed Journal: BMC Genomics ISSN: 1471-2164 Impact factor: 3.969
Summary of the EST Assembly
| Number of sequences for assembly | 54,960 |
| blue catfish | 10,523 |
| channel catfish | 44,437 |
| Number of contigs | 5,670 |
| Number of singletons | 23,598 |
| Number of putative transcripts | 29,268 |
| Average contig size | 5.5 |
| Average contig length (bp) | 1,001 |
| No of contig with: | |
| 2 ESTs | 3,003 |
| 3 ESTs | 980 |
| 4 ESTs | 468 |
| 5 ESTs | 263 |
| 6–10 ESTs | 469 |
| 11–20 ESTs | 246 |
| 21–30 ESTs | 95 |
| 31–50 ESTs | 72 |
| > 50 ESTs | 74 |
Initial identification of SNPs as detected by AutoSNP software
| 2 | 2,488 | 15,220 | 2,253,452 | 0.68 |
| 3 | 928 | 9,314 | 914,950 | 1.02 |
| 4 | 458 | 6,423 | 506,023 | 1.27 |
| 5 | 98 | 361 | 104,164 | 0.35 |
| 6–10 | 168 | 538 | 179,846 | 0.30 |
| 11–20 | 69 | 246 | 72,058 | 0.34 |
| 21–30 | 49 | 220 | 56,804 | 0.39 |
| 31–50 | 58 | 317 | 69,615 | 0.46 |
| > 50 | 71 | 955 | 93,065 | 1.03 |
| Total | 4,387 | 33,594 | 4,249,977 | 0.79* |
*Average SNP frequency per 100 bp.
Figure 1Distribution of minor allele frequency in domestic and wild channel catfish strains. The name of the populations is labeled on the top of each panel. MAF: minor allele frequency.
Overall summary of the EST-derived SNP genotyping using the Illumina Bead Array technology
| Successful genotype calling | 266 | 0.87 |
| Polymorphic SNPs | 156 | 0.87 |
| Non-polymorphic SNPs | 110 | 0.87 |
| Failed SNPs | 118 | 0.90 |
| Total number of loci tested | 384 | 0.88 |
SNP polymorphic rates as a function of contig size and minor sequence allele frequency
| # of sequences in the contig | # Successful Loci | Sequence ratio* | Minimal Minor Sequence Frequency | Polymorphic rate (%) |
| 2 | 24 | 1:1 | 50% | 33.3 |
| 3 | 37 | 1:2 | 33.3% | 45.9 |
| 4 | 26 | 1:3 | 25% | 15.4 |
| Subtotal | 87 | 33.3* | ||
| 4 | 44 | 2:2 | 50% | 70.5 |
| 5–6 | 60 | 2:3 & 2:4 & 3:3 | 33.% | 60.0 |
| 7–8 | 17 | 3:4 & 3:5 & 4:4 | 37.5% | 64.7 |
| 9–12 | 21 | 4:5 & 4:6 & 4:7 & 4:8 & 5:5 & 5:6 & 5:7 & 6:6 | 33.3% | 76.2 |
| >12 | 37 | 5:7 & 6:6 & 5:8 & 6:7......... & 12:57 | 17.4% | 89.2 |
| Subtotal | 179 | 70.9* | ||
| Total | 266 | 58.6* | ||
*Average polymorphic rate in respective categories.
Effect of low sequence quality (as defined by the presence of hot spots of SNP occurrence) and the presence of predicted intron on success rate of SNP genotyping
| Number of loci with SNP located in regions containing low quality sequences | 14 | 7 | 50% |
| Number of loci with known introns | 5 | 5 | 100% |
| Number of failed loci without gene information | 99 | ||
| With Significant Blast hits | 92 | 92.9% | |
| SNP positions can be located by similarity comparisons with zebrafish genome | 50 | 54.3% | |
| Number of Loci with SNP predicted to be positioned at exon-intron border | 32 | 64% | |
| Total number of loci potentially with SNP positioned at exon-intron border | 37 | 67.3% |
Figure 2SNP quality assessment based on EST contig size and sequence frequency of the alleles. Arrows indicate the trend of SNP quality, with the black arrows indicating trend of heterozygosity within a subset of contigs with the same number of the minor allele sequence, and the red arrow indicating overall SNP quality trend.
Figure 3Schematic illustration of the effect of introns involved in SNP genotyping. In the first case, all the genotyping primers are located in the same exon nearby, leading to successful genotyping (+); in the second case (middle), one of the genotyping primers (P3 as shown) was located at the exon-intron border, causing non-base pairing that lead to failure of genotyping (-); and in the third case, even though all primers were located in exon regions. However, an intron was involved that demands PCR extension to across the intron. Apparently, the Bead array technology provide very limited extension capability, leading to genotyping failure (-) as well.