| Literature DB >> 30582651 |
Virginie Merot-L'anthoene1, Rémi Tournebize2, Olivier Darracq1, Vimel Rattina1, Maud Lepelley1, Laurence Bellanger1, Christine Tranchant-Dubreuil2, Manon Coulée2, Marie Pégard2, Sylviane Metairon3, Coralie Fournier3, Piet Stoffelen4, Steven B Janssens4, Catherine Kiwuka5,6, Pascal Musoli5, Ucu Sumirat7, Hyacinthe Legnaté8, Jean-Léon Kambale9, João Ferreira da Costa Neto10, Clara Revel11, Alexandre de Kochko2, Patrick Descombes3, Dominique Crouzillat1, Valérie Poncet2.
Abstract
Coffee species such as Coffea canephora P. (Robusta) and C. arabica L. (Arabica) are important cash crops in tropical regions around the world. C. arabica is an allotetraploid (2n = 4x = 44) originating from a hybridization event of the two diploid species C. canephora and C. eugenioides (2n = 2x = 22). Interestingly, these progenitor species harbour a greater level of genetic variability and are an important source of genes to broaden the narrow Arabica genetic base. Here, we describe the development, evaluation and use of a single-nucleotide polymorphism (SNP) array for coffee trees. A total of 8580 unique and informative SNPs were selected from C. canephora and C. arabica sequencing data, with 40% of the SNP located in annotated genes. In particular, this array contains 227 markers associated to 149 genes and traits of agronomic importance. Among these, 7065 SNPs (~82.3%) were scorable and evenly distributed over the genome with a mean distance of 54.4 Kb between markers. With this array, we improved the Robusta high-density genetic map by adding 1307 SNP markers, whereas 945 SNPs were found segregating in the Arabica mapping progeny. A panel of C. canephora accessions was successfully discriminated and over 70% of the SNP markers were transferable across the three species. Furthermore, the canephora-derived subgenome of C. arabica was shown to be more closely related to C. canephora accessions from northern Uganda than to other current populations. These validated SNP markers and high-density genetic maps will be useful to molecular genetics and for innovative approaches in coffee breeding.Entities:
Keywords: zzm321990C. canephorazzm321990; zzm321990C. eugenioideszzm321990; Coffea arabica origin; SNP array; genetic map; single-nucleotide polymorphism
Mesh:
Substances:
Year: 2019 PMID: 30582651 PMCID: PMC6576098 DOI: 10.1111/pbi.13066
Source DB: PubMed Journal: Plant Biotechnol J ISSN: 1467-7644 Impact factor: 9.803
(a) Utilization and efficiency of the Coffee8.5K array. Evaluation of the Coffee8.5K array for application in the three Coffea species (C. canephora, C. arabica and C. eugenioides), genetic diversity assessment in the C. canephora panel and (b) genetic mapping in the segregating populations of C. canephora and C. arabica
| (a) | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| SNP source | Genomic region | Synthesized loci |
| ( |
| ( |
| ( | |
| Scorable (% | Polymorphic (% | Scorable (%) | Polymorphic in | Polymorphic (%) | Scorable (%) | Polymorphic (%) | |||
|
|
|
|
|
|
|
|
|
| |
| Coding | 208 | 168 | 48 | 154 | 66 | 46 | 138 | 16 | |
| Non‐coding | 2842 | 1810 | 417 | 1636 | 317 | 659 | 1303 | 94 | |
|
|
|
|
|
|
|
|
|
| |
| Coding | 3220 | 3046 | 3000 | 3028 | 801 | 14 | 2912 | 95 | |
| Non‐coding | 2310 | 2041 | 1977 | 2006 | 469 | 5 | 1830 | 50 | |
| Total |
|
|
|
|
|
|
|
| |
| Coding | 3428 | 3214 | 3048 | 3182 | 867 | 60 | 3050 | 111 | |
| Non‐coding | 5152 | 3851 | 2394 | 3642 | 786 | 664 | 3133 | 144 | |
(a) *See Table S2; †The percentage of scorable/used loci to successfully synthesized loci; ‡The percentage of polymorphic loci to scorable loci in the species.
(b) *See Table S2.
Bold values represent cumulative values of coding + non‐coding statistics.
Figure 1Design and development workflow of the Coffee8.5K array. SNPs markers were identified, filtered and validated from (a) the C. arabica or (b) the C. canephora Discovery panels. (c) Summary of variant source origin of the Coffee8.5k SNPs. Filters criteria are mentioned together with related tools and programming languages (in blue).
Figure 2Genetic and genomic distribution of the Coffee8.5K array SNPs. (a) Genome distribution of the 8580 single‐nucleotide polymorphisms (SNPs) synthesized for the array along the 11 pseudo‐chromosomes and the virtual pseudo‐chromosome 0 of unanchored sequences. For each pseudo‐chromosome, we determined: the number of SNPs markers according to their source (C. arabica, Ara or C. canephora, Can), the mean distance between markers and their density related to the estimated size of the pseudo‐chromosome. (b) C. canephora high‐density genetic map of BP409xQ121 progeny, with 11 linkage groups. SNP markers (1307) obtained from the Coffee8.5K array are indicated in red.
Number of SNP markers added to the Robusta genetic map and coverage in cM of each linkage group in the Robusta linkage maps. The Robusta genetic map based on a F1 cross between BP409 (Congolese hybrid) and Q121 (Conilon‐type‐derived accession) comprising 93 individuals. The previous high‐density genetic map has been published by Denoeud et al. (2014) with various types of markers (e.g. SSR, RADseq, RFLP)
| LG | Total Number of markers | Number of SNPs | % | Coverage (cM) | Mean distance between markers (cM) |
|---|---|---|---|---|---|
| A | 287 | 104 | 36 | 112 | 2.6 |
| B | 528 | 220 | 42 | 238 | 2.2 |
| C | 194 | 76 | 39 | 129 | 1.5 |
| D | 266 | 125 | 47 | 109 | 2.4 |
| E | 257 | 125 | 49 | 105 | 2.4 |
| F | 341 | 149 | 44 | 155 | 2.2 |
| G | 284 | 119 | 42 | 105 | 2.7 |
| H | 223 | 108 | 48 | 128 | 1.7 |
| I | 158 | 61 | 39 | 90 | 1.8 |
| J | 263 | 111 | 42 | 104 | 2.5 |
|
| 238 | 109 | 46 | 95 | 2.5 |
| TOTAL | 3039 | 1307 | 43 | 1370 | 2.2 |
The ratio of SNPs from Coffee8.5K array to total number of loci mapped.
Figure 3The population structure of the Coffea canephora diversity panel (27 accessions). Note that the same colour code is used in all graphs (a) Global distribution of the genetic groups across the C. canephora distribution range. Map data from US Dept of State Geographer ©2018 Google Image Landsat/Copernicus Data SIO, NOAA, U.S. Navy, NGA, GEBCO. (b) Population structure analysis using sNMF with three or eight numbers of clusters (K). Each colour represents a single cluster, each vertical column represents one accession. (c) A neighbour‐joining tree based on Euclidean distances between the C. canephora accessions, together with C. arabica and C. eugenioides accessions. Individuals from the same group are represented by the same colour, whereas admixed individuals (AG* and EB *, the two parents of the Robusta progeny, and BE, OE) do not have colour.
Figure 4Coffea arabica and its progenitor species. (a) Range distribution of the three related species. Dotted lines represent their schematic distribution limit, whereas names with colour labels correspond to sampled sites (C. canephora in blue, C. eugenioides in gold, and C. arabica in red), Map data ©2018 Google, ORION‐ME; (b) Origin of the genomes of the allotetraploid species Coffea arabica and of the Dihaploid Et39; (c, d) Haploid Identity‐by‐state distances (IBS) distances between C. arabica and accessions of the C. canephora with colour code as in Figure 3; and (c) C. eugenioides species. The average IBS distances and their standard deviations were calculated over the 17 C. arabica individuals.