| Literature DB >> 19900278 |
Silvia T Rodríguez-Ramilo1, Miguel A Toro, Jesús Fernández.
Abstract
BACKGROUND: The inference of the hidden structure of a population is an essential issue in population genetics. Recently, several methods have been proposed to infer population structure in population genetics.Entities:
Mesh:
Year: 2009 PMID: 19900278 PMCID: PMC2776585 DOI: 10.1186/1297-9686-41-49
Source DB: PubMed Journal: Genet Sel Evol ISSN: 0999-193X Impact factor: 4.297
Figure 1Genetic distance (a) and Δ. Example of ten replicates of a single scenario (K = 5).
Parameter set, genetic variability values and Wright F statistics considered in each evaluated scenario
| Microsatellite loci | ||||
|---|---|---|---|---|
| Scenario | 1 | 2 | 3 | 4 |
| Generations | 10000 | 10000 | 20 | 20 |
| Subpopulation size | 100 | 100 | 20 | 20 |
| Number of markers | 10 | 50 | 10 | 50 |
| Number of alleles | 10 | 10 | 10 | 10 |
| Genetic variability: | ||||
| 7.72 ± 0.14 | 7.78 ± 0.05 | 8.79 ± 0.11 | 8.66 ± 0.05 | |
| 0.55 ± 0.02 | 0.56 ± 0.01 | 0.59 ± 0.01 | 0.60 ± 0.01 | |
| 0.56 ± 0.02 | 0.56 ± 0.01 | 0.59 ± 0.01 | 0.60 ± 0.01 | |
| 0.64 ± 0.02 | 0.65 ± 0.01 | 0.82 ± 0.00 | 0.83 ± 0.00 | |
| Wright | ||||
| 0.01 ± 0.00 | 0.00 ± 0.00 | 0.01 ± 0.01 | 0.00 ± 0.01 | |
| 0.13 ± 0.00 | 0.13 ± 0.00 | 0.27 ± 0.01 | 0.27 ± 0.01 | |
| 0.13 ± 0.01 | 0.12 ± 0.01 | 0.28 ± 0.01 | 0.28 ± 0.01 | |
| Generations | 1000 | 1000 | 20 | 20 |
| Subpopulation size | 100 | 100 | 20 | 20 |
| Number of markers | 60 | 300 | 60 | 300 |
| Number of alleles | 2 | 2 | 2 | 2 |
| Genetic variability: | ||||
| 1.53 ± 0.10 | 1.60 ± 0.01 | 2.00 ± 0.00 | 2.00 ± 0.00 | |
| 0.19 ± 0.01 | 0.18 ± 0.00 | 0.33 ± 0.01 | 0.33 ± 0.00 | |
| 0.19 ± 0.01 | 0.18 ± 0.00 | 0.33 ± 0.00 | 0.33 ± 0.00 | |
| 0.22 ± 0.01 | 0.21 ± 0.00 | 0.46 ± 0.00 | 0.46 ± 0.00 | |
| Wright | ||||
| 0.00 ± 0.00 | 0.00 ± 0.00 | 0.01 ± 0.01 | 0.00 ± 0.00 | |
| 0.14 ± 0.01 | 0.14 ± 0.00 | 0.27 ± 0.01 | 0.27 ± 0.01 | |
| 0.14 ± 0.01 | 0.14 ± 0.00 | 0.28 ± 0.01 | 0.27 ± 0.01 | |
The following parameters were fixed in all data sets: diploidy, hermaphroditic, random mating, finite island model, five subpopulations, equal number of individuals in all subpopulations, constant population size, migration rate m = 0.01, KAM mutation model, equal frequencies for all allelic states in the initial population, free recombination between loci, mutation rate: 10-3 for microsatellite loci and 5 × 10-7 for SNP loci. n: number of alleles; H: observed heterozygosity; H: mean subpopulation gene diversity; H: mean total gene diversity
Genetic variability and Wright statistics with different migrations, K = 10, HIM, HWD and LD
| Scenario 2 | HIM | |||||
|---|---|---|---|---|---|---|
| 50 markers | 300 markers | |||||
| Genetic variability: | ||||||
| 7.68 ± 0.04 | 7.75 ± 0.08 | 7.73 ± 0.05 | 8.03 ± 0.04 | 9.58 ± 0.03 | 1.14 ± 0.02 | |
| 0.60 ± 0.01 | 0.61 ± 0.01 | 0.62 ± 0.01 | 0.49 ± 0.00 | 0.50 ± 0.00 | 0.02 ± 0.00 | |
| 0.60 ± 0.01 | 0.62 ± 0.01 | 0.63 ± 0.01 | 0.50 ± 0.00 | 0.51 ± 0.00 | 0.02 ± 0.00 | |
| 0.62 ± 0.01 | 0.63 ± 0.01 | 0.63 ± 0.01 | 0.67 ± 0.00 | 0.79 ± 0.00 | 0.05 ± 0.00 | |
| Wright | ||||||
| 0.00 ± 0.00 | 0.00 ± 0.00 | 0.01 ± 0.00 | 0.01 ± 0.00 | 0.01 ± 0.00 | 0.01 ± 0.00 | |
| 0.03 ± 0.00 | 0.02 ± 0.00 | 0.01 ± 0.00 | 0.26 ± 0.01 | 0.35 ± 0.00 | 0.50 ± 0.01 | |
| 0.03 ± 0.00 | 0.02 ± 0.00 | 0.02 ± 0.00 | 0.27 ± 0.01 | 0.36 ± 0.00 | 0.50 ± 0.01 | |
| Genetic variability: | ||||||
| 8.45 ± 0.17 | 8.21 ± 0.08 | 7.78 ± 0.19 | 7.18 ± 0.19 | 1.95 ± 0.00 | 1.94 ± 0.00 | |
| 0.51 ± 0.01 | 0.39 ± 0.01 | 0.23 ± 0.01 | 0.09 ± 0.01 | 0.08 ± 0.01 | 0.36 ± 0.00 | |
| 0.60 ± 0.01 | 0.54 ± 0.01 | 0.48 ± 0.02 | 0.46 ± 0.02 | 0.09 ± 0.01 | 0.37 ± 0.00 | |
| 0.82 ± 0.00 | 0.81 ± 0.00 | 0.81 ± 0.01 | 0.79 ± 0.01 | 0.40 ± 0.00 | 0.40 ± 0.00 | |
| Wright | ||||||
| 0.15 ± 0.01 | 0.29 ± 0.01 | 0.52 ± 0.01 | 0.81 ± 0.02 | 0.12 ± 0.02 | 0.02 ± 0.00 | |
| 0.27 ± 0.01 | 0.33 ± 0.01 | 0.40 ± 0.02 | 0.42 ± 0.02 | 0.76 ± 0.02 | 0.07 ± 0.00 | |
| 0.38 ± 0.01 | 0.52 ± 0.01 | 0.71 ± 0.01 | 0.89 ± 0.01 | 0.79 ± 0.02 | 0.09 ± 0.01 | |
Scenario 2 simulated with different migration rates (m) and a higher number of subpopulations (K = 10); hierarchical island model (HIM) with 50 microsatellites and 300 SNP; scenario 3 simulated with selfing (0.3, 0.5, 0.7 and 0.9) to generate Hardy-Weinberg disequilibrium (HWD); scenario 6 with linked loci (recombination rate = 0.06) and 1000 generations with no migration between subpopulations and 10 generations where m = 0.01 or m = 0.1 to generate linkage disequilibrium (LD); see Table 1 for abbreviations and for the explanation of scenarios
Modal value and fraction of replicates where the estimated number of clusters (K) was 5
| Microsatellite loci | SNP loci | |||||||
|---|---|---|---|---|---|---|---|---|
| Scenario | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
| Modal value: | ||||||||
| 5 | 5 | 5 | 5 | 5 | 5 | 5 | 5 | |
| 10 | 5 | 6 | 5 | 14 | 5 | 6 | 5 | |
| 5 | 5 | 5 | 5 | 5 | 5 | 5 | 5 | |
| Replicates | ||||||||
| 0.7 | 1.0 | 0.9 | 0.4 | 0.6 | 1.0 | 0.8 | 0.8 | |
| 0.0 | 1.0 | 0.3 | 0.6 | 0.0 | 0.9 | 0.0 | 0.4 | |
| 0.8 | 1.0 | 1.0 | 1.0 | 0.9 | 1.0 | 1.0 | 1.0 | |
See Table 1 for the explanation of scenarios
Figure 2Mean proportion of correct groupings over replicates in each scenario and method. Bars represent standard errors; see Table 1 for the explanation of the scenarios.
Modal value and fraction of replicates where K = 5 (10) in the remaining scenarios
| Scenario 2 | HIM | |||||
|---|---|---|---|---|---|---|
| 50 markers | 300 markers | |||||
| Modal value: | ||||||
| 5 | 5 | 5 | 9 | 5 | 5 | |
| 5 | 5 | 3 | 10 | 21 | 18 | |
| 5 | 5 | 3 | 10 | 5 | 5 | |
| Replicates | ||||||
| 1.0 | 1.0 | 0.9 | (0.2) | 0.6 | 0.7 | |
| 1.0 | 0.9 | 0.3 | (0.6) | 0.0 | 0.0 | |
| 0.9 | 0.5 | 0.2 | (1.0) | 1.0 | 1.0 | |
| Modal value: | ||||||
| 5 | 5 | 5 | 3 | 5 | 4 | |
| 11 | 10 | 15 | 15 | 9 | 6 | |
| 5 | 5 | 5 | 3 | 5 | 5 | |
| Replicates | ||||||
| 0.8 | 0.5 | 0.7 | 0.3 | 0.8 | 0.1 | |
| 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 0.8 | 0.8 | 0.7 | 0.1 | 1.0 | 0.5 | |
See Table 2 for the explanation of scenarios
Figure 3Proportion of correct groupings with different migration rates, . Mean proportion of correct groupings over replicates for each simulated migration rate (m) and a higher number of subpopulations (K = 10) in scenario 2; hierarchical island model (HIM) with 50 microsatellites and 300 SNP; scenario 3 with selfing (0.3, 0.5, 0.7 and 0.9) to generate Hardy-Weinberg disequilibrium (HWD) and in scenario 6 with linked loci (recombination rate = 0.06) and 1000 generations with no migration between subpopulations and 10 generations where m = 0.01 or m = 0.1 to generate linkage disequilibrium (LD); bars represent standard errors; see Table 1 for the explanation of the scenarios.
Figure 4Schematic representation of the population structure and the relationship with geographic regions in humans. STRUCTURE results taken from Rosenberg et al. [34] and BAPS results from Corander et al. [21]; MGD: maximisation of the genetic distance method, K: number of inferred clusters, N: population size; each box corresponds to a geographical region and the width of the boxes indicates graphically the number of genotyped individuals; Af: Africa (N = 119), E: Europe (N = 161), ME: Middle East (N = 178), CSA: Central-South Asia (N = 210), EA: East Asia (N = 241), O: Oceania (N = 39), Am: America (N = 108), Kal: Kalash (N = 25), Kar: Karitiana (N = 24), S: Surui (N = 21); black lines separate regional affiliations (on the top of the figure) of the individuals; for each analysed K the partition obtained with each methodology is represented with K different colours.