| Literature DB >> 27347456 |
Richard G J Hodel1, M Claudia Segovia-Salcedo2, Jacob B Landis1, Andrew A Crowl1, Miao Sun3, Xiaoxian Liu1, Matthew A Gitzendanner4, Norman A Douglas4, Charlotte C Germain-Aubrey3, Shichao Chen5, Douglas E Soltis6, Pamela S Soltis7.
Abstract
Microsatellites, or simple sequence repeats (SSRs), have long played a major role in genetic studies due to their typically high polymorphism. They have diverse applications, including genome mapping, forensics, ascertaining parentage, population and conservation genetics, identification of the parentage of polyploids, and phylogeography. We compare SSRs and newer methods, such as genotyping by sequencing (GBS) and restriction site associated DNA sequencing (RAD-Seq), and offer recommendations for researchers considering which genetic markers to use. We also review the variety of techniques currently used for identifying microsatellite loci and developing primers, with a particular focus on those that make use of next-generation sequencing (NGS). Additionally, we review software for microsatellite development and report on an experiment to assess the utility of currently available software for SSR development. Finally, we discuss the future of microsatellites and make recommendations for researchers preparing to use microsatellites. We argue that microsatellites still have an important place in the genomic age as they remain effective and cost-efficient markers.Entities:
Keywords: genotyping by sequencing (GBS); microsatellite development; next-generation sequencing (NGS); restriction site associated DNA sequencing (RAD-Seq); simple sequence repeats (SSR); transcriptomes
Year: 2016 PMID: 27347456 PMCID: PMC4915923 DOI: 10.3732/apps.1600025
Source DB: PubMed Journal: Appl Plant Sci ISSN: 2168-0450 Impact factor: 1.936
The number of loci found in an SSR search and the number of loci found per mega base pair sequence for each software package in each of two data sets used to highlight the vast potential resources available for researchers who cannot generate their own sequence data to search for SSRs.
| Software package | Total no. of loci | Loci/Mbp sequence | Total no. of loci | Loci/Mbp sequence |
| GMATo | 448,569 | 39.3 | 181,223 | 671.2 |
| MISA | 448,746 | 39.4 | 181,449 | 672.0 |
| MSATCOMMANDER | 372,436 | 32.7 | 151,455 | 560.9 |
| PAL_FINDER | NA | NA | 140,463 | 520.2 |
| Phobos (Geneious, STAMP) | 450,948 | 39.6 | 181,616 | 672.7 |
| SSR Locator | 447,185 | 39.2 | 180,763 | 669.5 |
Note: NA = not applicable.
114 million single-end reads (1 × 100); 11.4-Gbp sequence; FASTA file: 12.5 GB.
900,000 paired-end reads (2 × 150; 1.8 million total reads); 270-Mbp sequence; FASTA file: 275 MB.
The number and percentage of each repeat motif type using each software package found in the SSR search for each test data set.
| Software package | ||
| GMATo | ||
| No. of dinucleotides (%) | 335,835 (74.9) | 103,688 (57.2) |
| No. of trinucleotides (%) | 93,812 (20.9) | 74,129 (40.9) |
| No. of tetranucleotides (%) | 15,997 (3.6) | 610 (0.3) |
| No. of pentanucleotides (%) | 2233 (0.5) | 307 (0.2) |
| No. of hexanucleotides (%) | 692 (0.2) | 2489 (1.4) |
| MISA | ||
| No. of dinucleotides (%) | 335,883 (74.8) | 103,768 (57.2) |
| No. of trinucleotides (%) | 93,933 (20.9) | 74,208 (40.9) |
| No. of tetranucleotides (%) | 16,005 (3.6) | 612 (0.3) |
| No. of pentanucleotides (%) | 2233 (0.5) | 322 (0.2) |
| No. of hexanucleotides (%) | 692 (0.2) | 2539 (1.4) |
| MSATCOMMANDER | ||
| No. of dinucleotides (%) | 284,725 (76.4) | 83,377 (55.1) |
| No. of trinucleotides (%) | 74,305 (20.0) | 65,374 (43.2) |
| No. of tetranucleotides (%) | 11,409 (3.1) | 473 (0.3) |
| No. of pentanucleotides (%) | 1631 (0.4) | 197 (0.1) |
| No. of hexanucleotides (%) | 366 (0.1) | 2034 (1.3) |
| PAL_FINDER | ||
| No. of dinucleotides (%) | NA | 83,421 (59.4) |
| No. of trinucleotides (%) | NA | 54,787 (39.0) |
| No. of tetranucleotides (%) | NA | 589 (0.4) |
| No. of pentanucleotides (%) | NA | 313 (0.2) |
| No. of hexanucleotides (%) | NA | 2053 (1.5) |
| Phobos (Geneious, STAMP) | ||
| No. of dinucleotides (%) | 336,595 (74.6) | 103,807 (57.2) |
| No. of trinucleotides (%) | 95,423 (21.2) | 74,311 (40.9) |
| No. of tetranucleotides (%) | 16,005 (3.5) | 613 (0.3) |
| No. of pentanucleotides (%) | 2233 (0.5) | 322 (0.2) |
| No. of hexanucleotides (%) | 692 (0.2) | 2563 (1.4) |
| SSR Locator | ||
| No. of dinucleotides (%) | 334,836 (74.9) | 103,599 (57.3) |
| No. of trinucleotides (%) | 93,427 (20.9) | 73,938 (40.9) |
| No. of tetranucleotides (%) | 15,998 (3.6) | 604 (0.3) |
| No. of pentanucleotides (%) | 2232 (0.5) | 290 (0.2) |
| No. of hexanucleotides (%) | 692 (0.2) | 2332 (1.3) |
Note: NA = not applicable.
114 million single-end reads (1 × 100); 11.4-Gbp sequence; FASTA file: 12.5 GB.
900,000 paired-end reads (2 × 150; 1.8 million total reads); 270-Mbp sequence; FASTA file: 275 MB.
Description of software packages used in this study, including operating systems, important features, URL where software can be obtained, number of citations, authors, and brief comments describing the ease of use.
| Software | Operating system | Features | URL | Citations (Web of Science/Google Scholar) | Reference | Comments |
| Geneious | Linux, Mac OSX, Windows | Integrates multiple functions with plugins, user-friendly interface | 395/633 | Very user friendly, but requires a paid license to run. The microsatellite development plugin (Phobos, | ||
| GMATo | Linux, Mac OSX, Windows | Both GUI and command line interface; SSR mining and statistics at genome level | 0/8 | Runs quickly and has clear output, but it is hard for the user to change important parameter settings. | ||
| HighSSR | Linux, Mac OSX, Windows | A Java program is designed for NGS data and capable of microsatellite detection, elimination of redundancy and primer development, and interacting with PostgreSQL, MUSCLE, and Primer3. | 7/12 | Cannot open large files (>2 GB); unsuitable for most NGS data. | ||
| MISA | Linux, Mac OSX, Windows | Preprocessing sequences, motif search, and interacts with Primer3 for primer designs | 669/1150 | Fast, easy to configure, generates primers. | ||
| MSATCOMMANDER | Linux, Mac OSX, Windows | Motif search, interacts with Primer3 for primer design, and primer auto-tag | 428/509 | Output is difficult to view; requires lots of filtering to find basic statistics. | ||
| PAL_FINDER | Linux, Mac OSX, Windows | Identifies and characterizes microsatellite repeat loci from shotgun genomic sampling by 454 or Illumina paired-end reads, and designs PCR primers by interacting with Primer3 | 87/115 | Slow, struggles with large files; compatibility issues with many FASTQ formats; would not complete with the largest data set. | ||
| QDD3 | Windows and Linux | A computer program to select microsatellite markers from raw sequence reads obtained from 454 or Illumina and design primers from large sequences at genomic level, dealing with the essential bioinformatics and equipped with both command line and a user-friendly graphical interface on the Galaxy server. | 3/9 | Relatively long running time; user cannot easily change parameter settings. | ||
| SSR Locator | Windows | Integrates SSR searches, frequency of occurrence of motifs and arrangements, primer design, PCR simulation, global alignments, and identity and homology searches; eliminates overlaps between adjacent sequences; interacts with Primer3 | http://www.hindawi.com/journals/ijpg/2008/412696/ | 0/98 | Requires tedious file reformatting; only available for Windows. | |
| SSR_pipeline | Linux, Mac OSX, Windows | Identifies simple sequence repeats, e.g., microsatellites from paired-end high-throughput Illumina DNA sequencing data | http://pubs.usgs.gov/ds/778/ | 3/3 | We could not get it to run after ∼24 h of effort. | |
| STAMP | Linux, Mac OSX, Windows | Based on STADEN package, with comprehensive integration of a set of extension modules to facilitate the processing of microsatellite markers, like Phobos, TROLL, Primer3, SQLite module. | http://www.awi.de/en/research/scientific_computing/bioinformatics/software/ | 11/15 | Powerful, but not user friendly. Other modules associated with it (Phobos) quickly and conveniently locate potential loci. |
Software packages, the number of loci they find in an SSR search, and the number of loci they find per mega base pair sequence in each of the four test data sets for four sequencing platforms (MiSeq, HiSeq1, HiSeq2, PacBio).
| MiSeq (ERR365834) | HiSeq1 (ERR368422) | HiSeq2 (ERR965681) | PacBio (SRR1284764) | |||||
| Software package | Total no. of loci | Loci/Mbp sequence | Total no. of loci | Loci/Mbp sequence | Total no. of loci | Loci/Mbp sequence | Total no. of loci | Loci/Mbp sequence |
| GMATo | 482,084 | 146.1 | 171,016 | 114.0 | 722,636 | 83.1 | 104,630 | 219.8 |
| MISA | 482,336 | 146.2 | 171,095 | 114.1 | 723,062 | 83.1 | 104,778 | 220.1 |
| MSATCOMMANDER | 388,663 | 117.8 | 135,168 | 90.1 | 543,610 | 62.5 | 82,588 | 173.5 |
| PAL_FINDER | 310,495 | 94.1 | 158,163 | 105.4 | 591,617 | 68.0 | 48,831 | 102.6 |
| Phobos (Geneious, STAMP) | 483,037 | 146.4 | 172,309 | 114.9 | 723,917 | 83.2 | 104,896 | 220.4 |
| SSR Locator | 481,863 | 146.0 | 170,934 | 114.0 | 722,580 | 83.1 | 104,120 | 218.7 |
6.6 million paired-end reads (2 × 250; 13.2 million total reads); 3.3-Gbp sequence; FASTA file: 3.9 GB.
10.9 million single-end reads (1 × 100); 1.5-Gbp sequence; FASTA file: 2.2 GB.
48.5 million paired-end reads (2 × 100; 97 million total reads); 8.7-Gbp sequence; FASTA file: 5.7 GB.
163,500 reads; 476-Mbp sequence; FASTA file: 445 MB.
The number and percentage of each repeat motif type found in the SSR search in each of the four test data sets for four sequencing platforms (MiSeq, HiSeq1, HiSeq2, PacBio).
| Software package | MiSeq (ERR365834) | HiSeq1 (ERR368422) | HiSeq2 (ERR965681) | PacBio (SRR1284764) |
| GMATo | ||||
| No. of dinucleotides (%) | 395,657 (82.1) | 123,902 (72.5) | 565,192 (78.2) | 95,584 (91.4) |
| No. of trinucleotides (%) | 82,874 (17.2) | 42,764 (25.0) | 151,596 (21.0) | 8366 (8.0) |
| No. of tetranucleotides (%) | 2333 (0.5) | 2290 (1.3) | 3390 (0.5) | 556 (0.5) |
| No. of pentanucleotides (%) | 525 (0.1) | 803 (0.5) | 895 (0.1) | 99 (0.1) |
| No. of hexanucleotides (%) | 695 (0.1) | 1257 (0.7) | 1563 (0.2) | 25 (0.0) |
| MISA | ||||
| No. of dinucleotides (%) | 395,740 (82.0) | 123,918 (72.4) | 565,328 (78.2) | 95,634 (91.3) |
| No. of trinucleotides (%) | 83,016 (17.2) | 42,817 (25.0) | 151,850 (21.0) | 8454 (8.1) |
| No. of tetranucleotides (%) | 2357 (0.5) | 2294 (1.3) | 3406 (0.5) | 564 (0.5) |
| No. of pentanucleotides (%) | 525 (0.1) | 806 (0.5) | 905 (0.1) | 99 (0.1) |
| No. of hexanucleotides (%) | 698 (0.1) | 1260 (0.7) | 1573 (0.2) | 27 (0.0) |
| MSATCOMMANDER | ||||
| No. of dinucleotides (%) | 325,676 (83.8) | 99,465 (73.6) | 432,335 (79.5) | 77,096 (93.4) |
| No. of trinucleotides (%) | 60,629 (15.6) | 32,118 (23.8) | 107,818 (19.8) | 5148 (6.2) |
| No. of tetranucleotides (%) | 1613 (0.4) | 1824 (1.3) | 1925 (0.4) | 286 (0.3) |
| No. of pentanucleotides (%) | 313 (0.1) | 650 (0.5) | 619 (0.1) | 45 (0.1) |
| No. of hexanucleotides (%) | 432 (0.1) | 1111 (0.8) | 913 (0.2) | 13 (0.0) |
| PAL_FINDER | ||||
| No. of dinucleotides (%) | 251,678 (81.1) | 114,219 (72.2) | 460,072 (77.8) | 41,581 (85.2) |
| No. of trinucleotides (%) | 56,389 (18.2) | 40,088 (25.3) | 126,509 (21.4) | 6595 (13.5) |
| No. of tetranucleotides (%) | 1570 (0.5) | 2042 (1.3) | 2909 (0.5) | 531 (1.1) |
| No. of pentanucleotides (%) | 359 (0.1) | 717 (0.5) | 774 (0.1) | 98 (0.2) |
| No. of hexanucleotides (%) | 499 (0.2) | 1097 (0.7) | 1353 (0.2) | 26 (0.1) |
| Phobos (Geneious, STAMP) | ||||
| No. of dinucleotides (%) | 396,367 (82.1) | 124,755 (72.4) | 566,081 (78.2) | 95,743 (91.3) |
| No. of trinucleotides (%) | 83,088 (17.2) | 43,156 (25.0) | 151,949 (21.0) | 8462 (8.1) |
| No. of tetranucleotides (%) | 2359 (0.5) | 2314 (1.3) | 3409 (0.5) | 565 (0.5) |
| No. of pentanucleotides (%) | 525 (0.1) | 810 (0.5) | 905 (0.1) | 99 (0.1) |
| No. of hexanucleotides (%) | 698 (0.1) | 1274 (0.7) | 1573 (0.2) | 27 (0.0) |
| SSR Locator | ||||
| No. of dinucleotides (%) | 395,436 (82.1) | 123,818 (72.4) | 565,033 (78.2) | 95,062 (91.3) |
| No. of trinucleotides (%) | 82,881 (17.2) | 42,773 (25.0) | 151,690 (21.0) | 8373 (8.0) |
| No. of tetranucleotides (%) | 2335 (0.5) | 2288 (1.3) | 3384 (0.5) | 561 (0.5) |
| No. of pentanucleotides (%) | 516 (0.1) | 800 (0.5) | 904 (0.1) | 97 (0.1) |
| No. of hexanucleotides (%) | 695 (0.1) | 1255 (0.7) | 1569 (0.2) | 27 (0.0) |
6.6 million paired-end reads (2 × 250; 13.2 million total reads); 3.3-Gbp sequence; FASTA file: 3.9 GB.
10.9 million single-end reads (1 × 100); 1.5-Gbp sequence; FASTA file: 2.2 GB.
48.5 million paired-end reads (2 × 100; 97 million total reads); 8.7-Gbp sequence; FASTA file: 5.7 GB.
163,500 reads; 476-Mbp sequence; FASTA file: 445 MB.
Costs associated with microsatellite development for 12–15 loci and RAD/GBS costs to generate thousands of loci. Both budgets assume 96 individuals are included in the study.
| Item | Base cost | Quantity | Total cost |
| Microsatellites | |||
| QIAGEN PCR multiplex kit | $270 | 4 | $1080 |
| Unlabeled primer pairs | $6 | 48 | $288 |
| Labeled primer (single) | $80 | 24 | $1920 |
| Genotyping one plate | $100 | 4 | $400 |
| Total | |||
| RAD/GBS (high estimate) | |||
| Double digest | $1261 | 1 | $1261 |
| Digest optimization | $539 | 1 | $539 |
| QC Bioanalyzer chips | $101 | 10 | $1010 |
| Qubit quantification | $400 | 1 | $400 |
| Illumina HiSeq lane | $1225 | 1 | $1225 |
| dsDNA Fluorophore quantification | $69 | 1 | $69 |
| Reagents | $500 | 1 | $500 |
| Total | |||
| RAD/GBS (low estimate) | |||
| Per sample cost | $31.92 | 96 | $3064.32 |
| Fixed cost | $340 | 1 | $340 |
| Total |