Literature DB >> 35961988

Nine out of ten samples were mistakenly switched by The Orang-utan Genome Consortium.

Graham L Banes1,2,3, Emily D Fountain4,5, Alyssa Karklus6,5, Robert S Fulton7, Lucinda Antonacci-Fulton7, Joanne O Nelson7.   

Abstract

The Sumatran orang-utan (Pongo abelii) reference genome was first published in 2011, in conjunction with ten re-sequenced genomes from unrelated wild-caught individuals. Together, these published data have been utilized in almost all great ape genomic studies, plus in much broader comparative genomic research. Here, we report that the original sequencing Consortium inadvertently switched nine of the ten samples and/or resulting re-sequenced genomes, erroneously attributing eight of these to the wrong source individuals. Among them is a genome from the recently identified Tapanuli (P. tapanuliensis) species: thus, this genome was sequenced and published a full six years prior to the species' description. Sex was wrongly assigned to five known individuals; the numbers in one sample identifier were swapped; and the identifier for another sample most closely resembles that of a sample from another individual entirely. These errors have been reproduced in countless subsequent manuscripts, with noted implications for studies reliant on data from known individuals.
© 2022. The Author(s).

Entities:  

Year:  2022        PMID: 35961988      PMCID: PMC9374732          DOI: 10.1038/s41597-022-01602-0

Source DB:  PubMed          Journal:  Sci Data        ISSN: 2052-4463            Impact factor:   8.501


Introduction

Alongside their publication of a Sumatran orang-utan (Pongo abelii) draft genome assembly in 2011, The Orang-utan Genome Consortium re-sequenced the genomes of ten additional unrelated wild-caught individuals – ostensibly five Sumatran and five Bornean (P. pygmaeus) orang-utans – using short-read Illumina sequencing[1]. Their manuscript, and its accompanying 297 Gb of sequence data, has since been cited more than 500 times. During the course of our own studies, however, we noted several inconsistencies between the data made available in the NCBI Sequence Read Archive and their accompanying metadata and descriptors in the paper. We found no record of a sample with the identifier “KB5543”, for example, in the Frozen Zoo repository, the reported source of a sample attributed to the orang-utan, Louis. The closest match in their database to this ID was for another sample, “15543”, which derived from a different individual. We also observed that the identifier “KB9528”, as reported for the orang-utan Baldy in the manuscript’s Tables S4-1, was catalogued as a sample from an “African pig” – though, in a supplemental file, it was correctly denoted as KB9258, which derived from another orang-utan. The sample identifier “SB550”, as reported for the orang-utan Doris, appeared to reference a studbook number (i.e. “SB”) that belonged to another sequenced orang-utan, Sibu. The sex reported for five individuals also contradicted their known sexes, as had been recorded in contemporary studbook records[2], plus differed from the sexes assigned to each sample in Locke et al.’s supplementary data. Thus, we were driven to reconsider the identities of each genome’s source individual, through re-analysis of the published data combined with new molecular studies. Herein, we report that nine of the ten samples and/or published genomes were erroneously labelled in the original Nature publication. We present the corrected data and discuss the implications for other published works.

Methods

We first mapped the re-sequencing reads of all 10 Locke et al. whole genomes, plus those previously published from 27 conspecifics[3,4], to the latest iteration of the (female) orang-utan reference genome (ponAbe3[5]). To this, we had concatenated a recent orang-utan Y chromosome assembly[6]. Using the idxtools function in samtools 1.14[7], we inferred sex by comparing the ratios to which sequence reads were mapped against the X and Y chromosomes. Following two rounds of bootstrapped base recalibration, we then jointly called genotypes with GATK 4.1.8.0[8], all as previously described[9]. We randomly sampled 1,000,000 biallelic autosomal SNPs with no missing genotypes and ≥5% minor allele frequency (MAF), pruned linked loci in PLINK[10] (–indep-pairwise 50 10 0.1), and assigned populations in ADMIXTURE 1.39[11] as supervised with provenance data reported for the conspecifics[3,4] (K = 3). Additionally, we sampled and assayed eight orang-utans known to be first, second or third-degree relatives of seven of those purportedly sequenced by Locke et al., using the Illumina iScan Multi-Ethnic Global Array, also as previously described[12]. The reproduction of those seven, and thus these known relationships, had been contemporaneously recorded[2] (Fig. 1). To convert the microarray intensity data to variant calls, we mapped the probe flank sequences to ponAbe3 (using --fasta-flank) and exported genotypes (--sam-flank) with the bcftoools[7] plugin gtc2vcf (https://github.com/freeseek/gtc2vcf), subject to the following filter parameters: meanR_AB < 0.2, meanR_AA < 0.2, meanR_BB < 0.2, Cluster_Sep < 0.35, meanTHETA_AA > 0.3, meanTHETA_BB < 0.7, meanTHETA_AB < 0.3 and > 0.7, devTHETA_AA > 0.025, devTHETA_AB ≥ 0.07, devTHETA_BB > 0.025 and GenTrain_Score < 0.7. We then re-genotyped all 37 whole genomes at each of the resulting loci, as previously described[9]; merged these with the microarray genotype VCF, and LD-pruned and MAF-filtered biallelic SNPs precisely as aforementioned. With a view to avoiding the spurious kinship associations that typify highly structured data, we then bootstrapped ADMIXTURE’s cross-validation procedure to infer the most suitable K (trialling 1 through 10) before estimating kinship coefficients (Φij) in REAP[13].
Fig. 1

Genogram depicting relationships between orang-utans sampled by Locke et al. and known relatives assayed in this study. Orang-utans used in the kinship analysis (Table 1) are circled. The species affiliations of uncircled orang-utans reflect those purported by contemporary studbook records.

Genogram depicting relationships between orang-utans sampled by Locke et al. and known relatives assayed in this study. Orang-utans used in the kinship analysis (Table 1) are circled. The species affiliations of uncircled orang-utans reflect those purported by contemporary studbook records.
Table 1

Corrected identities and metadata for the samples sequenced and published by Locke et al.[1], as deposited in the NCBI BioSample database.

BioSample IDReported identities and metadataCorrected/validated identities and metadataKnown relative and inferred kinship
Lab IDISBNameSp.SexLab IDISBNameSp.SexX:YISBδ0δ1δ2Exp. ΦijΦij
SAMN00007164KB5404590BillyBFKB5404356DinahBF4.556
SAMN00007165KB4204364DollyBMKB4204590BillyBM0.50635720.6550.2260.1190.1250.116
SAMN00007166KB5406356DinahBFKB5406364DollyBF5.01513870.4320.3760.1920.2500.190
SAMN00007167KB5405360DennisBMKB5405360DennisBM0.54713870.4110.5300.0600.2500.162
SAMN00007168KB5543990LouisBM990LouisBM0.46436190.8230.0250.1520.1250.082
SAMN00007169KB5883550SibuSMKB58831600LikoeSM0.44433510.9710.0000.2210.0630.063
SAMN00007171KB4661695BubblesSMKB4661732BaldySM0.432
SAMN00007172KB43611600LikoeSFKB436153DorisSF4.625
SAMN00007173SB55053DorisSF550SibuSF4.26420690.3130.3900.2980.2500.246
SAMN00007170KB9528732BaldySMKB9258695BubblesTF4.17019800.2000.6320.1680.2500.242
17730.1860.6960.1180.2500.233
34500.7710.2290.0000.1250.036

Originally reported data are reproduced from Locke et al.’s Tables S4-1. “ISB” denotes International Studbook Number; for species, “B” indicates Bornean (P. pygmaeus), “S” indicates Sumatran (P. abelii), and “T” indicates Tapanuli (P. tapanuliensis). The X:Y ratios noted for sex are those inferred, as detailed, from each mapped BAM file. “Lab ID” is the internal identifier used by the sequencing facility, as variously recorded by Locke et al. Relatedness is reported as the probability that each sequenced orang-utan shares 0, 1 and 2 alleles identical by descent with a known relative (i.e. δ0, δ1, and δ2, respectively), plus the expected/theoretical (Exp.) and computed kinship coefficient (Φij).

We adopted a tri-fold method to confirm each sample’s identity. Identities were first inferred with an exclusionary approach, from computed (versus known and reported) sex and species. Each was then confirmed, where available, when observed kinship coefficients resembled those expected from known relationships. Third, we reviewed the historical biomaterial records retained by the Frozen Zoo, the original source of the samples, plus notes from the Laboratory Information System (LIMS) retained at Washington University in Saint Louis, where the samples were originally sequenced. Identity was assigned to a given sample when all these factors concorded.

Results

We observed X:Y sequence ratios in known males to range from 0.369–0.569 (mean 0.476) and in females from 4.114 to 5.827 (mean 4.973). From this, we interpreted that sex had been incorrectly assigned by Locke et al. to the sample SAMN00007170. This sample was inferred to be female (4.170) and thus cannot have derived from Baldy as purported. The species of each sample was correctly reported, though we inferred that the sample SAMN00007170 derived from a Tapanuli orang-utan (Fig. 2). This species was not formally described until 2017[4].
Fig. 2

Ancestry proportions of the Locke et al. orang-utans, as supervised with provenance data from 27 conspecifics sampled across the natural range of the genus[3,4]. Sample SAMN000007170 derives from a Tapanuli orang-utan.

Ancestry proportions of the Locke et al. orang-utans, as supervised with provenance data from 27 conspecifics sampled across the natural range of the genus[3,4]. Sample SAMN000007170 derives from a Tapanuli orang-utan. Kinship analyses linked one or more known relatives to seven of the ten samples sequenced by Locke et al. Specifically, we linked Billy to his granddaughter, Batari (observed/expected kinship coefficients: 0.116/0.125); Dumplin to her parents, Dennis (0.162/0.25) and Dolly (0.19/0.25); Sibu to her son, Tengku (0.246/0.25); Louis to his grandson, Max (0.082/0.125); Likoe to his great granddaughter, Menari (0.063/0.063); and Bubbles to her son, Oliver (0.233/0.25), daughter, Bella (0.242/0.25) and granddaughter, Nairi/Nadira (0.036/0.125). Admixture (K = 3) and kinship were inferred from a total of 1,132,210 biallelic SNPs. We assigned identity to the remaining three samples as molecular sex, sample and LIMS records were all concordant, and as the relatedness data had excluded other possible candidates. Table 1 corrects the record as originally presented in Nature. Corrected identities and metadata for the samples sequenced and published by Locke et al.[1], as deposited in the NCBI BioSample database. Originally reported data are reproduced from Locke et al.’s Tables S4-1. “ISB” denotes International Studbook Number; for species, “B” indicates Bornean (P. pygmaeus), “S” indicates Sumatran (P. abelii), and “T” indicates Tapanuli (P. tapanuliensis). The X:Y ratios noted for sex are those inferred, as detailed, from each mapped BAM file. “Lab ID” is the internal identifier used by the sequencing facility, as variously recorded by Locke et al. Relatedness is reported as the probability that each sequenced orang-utan shares 0, 1 and 2 alleles identical by descent with a known relative (i.e. δ0, δ1, and δ2, respectively), plus the expected/theoretical (Exp.) and computed kinship coefficient (Φij).

Discussion

Because Locke et al. focused solely on genome content, their discrepancies have no bearing on the accuracy of their data or their manuscript’s published findings. These errors have had considerable impact on other studies that utilized the published data, however, particularly those dependent on using data from known individuals. Three of our co-authors (GLB, EDF, AK) write from first-hand experience: reliant on tables and metadata from the original Nature publication, we came perilously close to incorrectly reporting that Baldy, a male orang-utan who lived at the Sacramento Zoo, was the first of the recently described Tapanuli species to be captured and exported from a wild population – a full five decades before his species’ formal description. On the contrary, this dubious honour belongs to Bubbles, a female orang-utan who lived at the San Diego Zoo. Though beyond the scope of our manuscript, the implications of this switch have not escaped our attention: principally, that Bubbles produced eight Sumatran x Tapanuli hybrid descendants, who were previously thought to be Sumatran. The genetic integrity of the captive population is therefore unexpectedly compromised, as we present in detail in a manuscript that is currently under review. Though we eventually caught these errors, others did not. Mattle-Greminger et al. (2018) reproduced eight erroneous sample identities in their paper, meaning each of the genomes they analysed were from different animals than reported[14]. Sudmant et al. (2013) reproduced seven such errors, thus also misattributing samples[15]. Neither Ma et al. (2013) nor Beeravolu et al. (2018) recognized that the sample identities were wrong, though as they reported only sample IDs (versus animal identities), no corrections to their manuscripts are warranted[16,17]. As sample identities are normally only reported in supplemental data – which is not always indexed by search engines – we cannot easily ascertain the full extent to which papers citing Locke et al. have reproduced these errors. Given these findings and implications, we have corrected the samples’ identities in the NCBI BioSample database. Tables detailing the revisions made are included in Supplementary File 1. We respectfully ask that those utilizing these updated identities cite this article in Scientific Data, in addition to the Correction concurrently published in Nature. Supplementary File 1
  15 in total

1.  PLINK: a tool set for whole-genome association and population-based linkage analyses.

Authors:  Shaun Purcell; Benjamin Neale; Kathe Todd-Brown; Lori Thomas; Manuel A R Ferreira; David Bender; Julian Maller; Pamela Sklar; Paul I W de Bakker; Mark J Daly; Pak C Sham
Journal:  Am J Hum Genet       Date:  2007-07-25       Impact factor: 11.025

2.  Fast model-based estimation of ancestry in unrelated individuals.

Authors:  David H Alexander; John Novembre; Kenneth Lange
Journal:  Genome Res       Date:  2009-07-31       Impact factor: 9.043

3.  Morphometric, Behavioral, and Genomic Evidence for a New Orangutan Species.

Authors:  Alexander Nater; Maja P Mattle-Greminger; Anton Nurcahyo; Matthew G Nowak; Marc de Manuel; Tariq Desai; Colin Groves; Marc Pybus; Tugce Bilgin Sonay; Christian Roos; Adriano R Lameira; Serge A Wich; James Askew; Marina Davila-Ross; Gabriella Fredriksson; Guillem de Valles; Ferran Casals; Javier Prado-Martinez; Benoit Goossens; Ernst J Verschoor; Kristin S Warren; Ian Singleton; David A Marques; Joko Pamungkas; Dyah Perwitasari-Farajallah; Puji Rianti; Augustine Tuuga; Ivo G Gut; Marta Gut; Pablo Orozco-terWengel; Carel P van Schaik; Jaume Bertranpetit; Maria Anisimova; Aylwyn Scally; Tomas Marques-Bonet; Erik Meijaard; Michael Krützen
Journal:  Curr Biol       Date:  2017-11-02       Impact factor: 10.834

4.  High-resolution comparative analysis of great ape genomes.

Authors:  Zev N Kronenberg; Ian T Fiddes; David Gordon; Shwetha Murali; Stuart Cantsilieris; Olivia S Meyerson; Jason G Underwood; Bradley J Nelson; Mark J P Chaisson; Max L Dougherty; Katherine M Munson; Alex R Hastie; Mark Diekhans; Fereydoun Hormozdiari; Nicola Lorusso; Kendra Hoekzema; Ruolan Qiu; Karen Clark; Archana Raja; AnneMarie E Welch; Melanie Sorensen; Carl Baker; Robert S Fulton; Joel Armstrong; Tina A Graves-Lindsay; Ahmet M Denli; Emma R Hoppe; PingHsun Hsieh; Christopher M Hill; Andy Wing Chun Pang; Joyce Lee; Ernest T Lam; Susan K Dutcher; Fred H Gage; Wesley C Warren; Jay Shendure; David Haussler; Valerie A Schneider; Han Cao; Mario Ventura; Richard K Wilson; Benedict Paten; Alex Pollen; Evan E Eichler
Journal:  Science       Date:  2018-06-08       Impact factor: 47.728

5.  Comparative and demographic analysis of orang-utan genomes.

Authors:  Devin P Locke; LaDeana W Hillier; Wesley C Warren; Kim C Worley; Lynne V Nazareth; Donna M Muzny; Shiaw-Pyng Yang; Zhengyuan Wang; Asif T Chinwalla; Pat Minx; Makedonka Mitreva; Lisa Cook; Kim D Delehaunty; Catrina Fronick; Heather Schmidt; Lucinda A Fulton; Robert S Fulton; Joanne O Nelson; Vincent Magrini; Craig Pohl; Tina A Graves; Chris Markovic; Andy Cree; Huyen H Dinh; Jennifer Hume; Christie L Kovar; Gerald R Fowler; Gerton Lunter; Stephen Meader; Andreas Heger; Chris P Ponting; Tomas Marques-Bonet; Can Alkan; Lin Chen; Ze Cheng; Jeffrey M Kidd; Evan E Eichler; Simon White; Stephen Searle; Albert J Vilella; Yuan Chen; Paul Flicek; Jian Ma; Brian Raney; Bernard Suh; Richard Burhans; Javier Herrero; David Haussler; Rui Faria; Olga Fernando; Fleur Darré; Domènec Farré; Elodie Gazave; Meritxell Oliva; Arcadi Navarro; Roberta Roberto; Oronzo Capozzi; Nicoletta Archidiacono; Giuliano Della Valle; Stefania Purgato; Mariano Rocchi; Miriam K Konkel; Jerilyn A Walker; Brygg Ullmer; Mark A Batzer; Arian F A Smit; Robert Hubley; Claudio Casola; Daniel R Schrider; Matthew W Hahn; Victor Quesada; Xose S Puente; Gonzalo R Ordoñez; Carlos López-Otín; Tomas Vinar; Brona Brejova; Aakrosh Ratan; Robert S Harris; Webb Miller; Carolin Kosiol; Heather A Lawson; Vikas Taliwal; André L Martins; Adam Siepel; Arindam Roychoudhury; Xin Ma; Jeremiah Degenhardt; Carlos D Bustamante; Ryan N Gutenkunst; Thomas Mailund; Julien Y Dutheil; Asger Hobolth; Mikkel H Schierup; Oliver A Ryder; Yuko Yoshinaga; Pieter J de Jong; George M Weinstock; Jeffrey Rogers; Elaine R Mardis; Richard A Gibbs; Richard K Wilson
Journal:  Nature       Date:  2011-01-27       Impact factor: 69.504

6.  Great ape genetic diversity and population history.

Authors:  Javier Prado-Martinez; Peter H Sudmant; Jeffrey M Kidd; Heng Li; Joanna L Kelley; Belen Lorente-Galdos; Krishna R Veeramah; August E Woerner; Timothy D O'Connor; Gabriel Santpere; Alexander Cagan; Christoph Theunert; Ferran Casals; Hafid Laayouni; Kasper Munch; Asger Hobolth; Anders E Halager; Maika Malig; Jessica Hernandez-Rodriguez; Irene Hernando-Herraez; Kay Prüfer; Marc Pybus; Laurel Johnstone; Michael Lachmann; Can Alkan; Dorina Twigg; Natalia Petit; Carl Baker; Fereydoun Hormozdiari; Marcos Fernandez-Callejo; Marc Dabad; Michael L Wilson; Laurie Stevison; Cristina Camprubí; Tiago Carvalho; Aurora Ruiz-Herrera; Laura Vives; Marta Mele; Teresa Abello; Ivanela Kondova; Ronald E Bontrop; Anne Pusey; Felix Lankester; John A Kiyang; Richard A Bergl; Elizabeth Lonsdorf; Simon Myers; Mario Ventura; Pascal Gagneux; David Comas; Hans Siegismund; Julie Blanc; Lidia Agueda-Calpena; Marta Gut; Lucinda Fulton; Sarah A Tishkoff; James C Mullikin; Richard K Wilson; Ivo G Gut; Mary Katherine Gonder; Oliver A Ryder; Beatrice H Hahn; Arcadi Navarro; Joshua M Akey; Jaume Bertranpetit; David Reich; Thomas Mailund; Mikkel H Schierup; Christina Hvilsom; Aida M Andrés; Jeffrey D Wall; Carlos D Bustamante; Michael F Hammer; Evan E Eichler; Tomas Marques-Bonet
Journal:  Nature       Date:  2013-07-03       Impact factor: 49.962

7.  Twelve years of SAMtools and BCFtools.

Authors:  Petr Danecek; James K Bonfield; Jennifer Liddle; John Marshall; Valeriu Ohan; Martin O Pollard; Andrew Whitwham; Thomas Keane; Shane A McCarthy; Robert M Davies; Heng Li
Journal:  Gigascience       Date:  2021-02-16       Impact factor: 6.524

8.  Genomic targets for high-resolution inference of kinship, ancestry and disease susceptibility in orang-utans (genus: Pongo).

Authors:  Graham L Banes; Emily D Fountain; Alyssa Karklus; Hao-Ming Huang; Nian-Hong Jang-Liaw; Daniel L Burgess; Jennifer Wendt; Cynthia Moehlenkamp; George F Mayhew
Journal:  BMC Genomics       Date:  2020-12-07       Impact factor: 3.969

9.  ABLE: blockwise site frequency spectra for inferring complex population histories and recombination.

Authors:  Champak R Beeravolu; Michael J Hickerson; Laurent A F Frantz; Konrad Lohse
Journal:  Genome Biol       Date:  2018-09-25       Impact factor: 13.583

10.  Dynamic evolution of great ape Y chromosomes.

Authors:  Monika Cechova; Rahulsimham Vegesna; Marta Tomaszkiewicz; Robert S Harris; Di Chen; Samarth Rangavittal; Paul Medvedev; Kateryna D Makova
Journal:  Proc Natl Acad Sci U S A       Date:  2020-10-05       Impact factor: 11.205

View more
  2 in total

1.  Orangutan genome mix-up muddies conservation efforts.

Authors:  Freda Kreier
Journal:  Nature       Date:  2022-10-19       Impact factor: 69.504

2.  Nine out of ten samples were mistakenly switched by The Orang-utan Genome Consortium.

Authors:  Graham L Banes; Emily D Fountain; Alyssa Karklus; Robert S Fulton; Lucinda Antonacci-Fulton; Joanne O Nelson
Journal:  Sci Data       Date:  2022-08-12       Impact factor: 8.501

  2 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.