Literature DB >> 36140757

Detection of Indiscriminate Genetic Manipulation in Thoroughbred Racehorses by Targeted Resequencing for Gene-Doping Control.

Teruaki Tozaki1, Aoi Ohnuma1, Kotono Nakamura2, Kazuki Hano2, Masaki Takasu2, Yuji Takahashi3, Norihisa Tamura3, Fumio Sato3, Kyo Shimizu4, Mio Kikuchi1, Taichiro Ishige1, Hironaga Kakoi1, Kei-Ichi Hirota1, Natasha A Hamilton5, Shun-Ichi Nagata1.   

Abstract

The creation of genetically modified horses is prohibited in horse racing as it falls under the banner of gene doping. In this study, we developed a test to detect gene editing based on amplicon sequencing using next-generation sequencing (NGS). We designed 1012 amplicons to target 52 genes (481 exons) and 147 single-nucleotide variants (SNVs). NGS analyses showed that 97.7% of the targeted exons were sequenced to sufficient coverage (depth > 50) for calling variants. The targets of artificial editing were defined as homozygous alternative (HomoALT) and compound heterozygous alternative (ALT1/ALT2) insertion/deletion (INDEL) mutations in this study. Four models of gene editing (three homoALT with 1-bp insertions, one REF/ALT with 77-bp deletion) were constructed by editing the myostatin gene in horse fibroblasts using CRISPR/Cas9. The edited cells and 101 samples from thoroughbred horses were screened using the developed test, which was capable of identifying the three homoALT cells containing 1-bp insertions. Furthermore, 147 SNVs were investigated for their utility in confirming biological parentage. Of these, 120 SNVs were amenable to consistent and accurate genotyping. Surrogate (nonbiological) dams were excluded by 9.8 SNVs on average, indicating that the 120 SNV could be used to detect foals that have been produced by somatic cloning or embryo transfer, two practices that are prohibited in thoroughbred racing and breeding. These results indicate that gene-editing tests that include variant calling and SNV genotyping are useful to identify genetically modified racehorses.

Entities:  

Keywords:  amplicon sequencing; gene doping; gene editing; horse; thoroughbred

Mesh:

Substances:

Year:  2022        PMID: 36140757      PMCID: PMC9498419          DOI: 10.3390/genes13091589

Source DB:  PubMed          Journal:  Genes (Basel)        ISSN: 2073-4425            Impact factor:   4.141


1. Introduction

Gene doping is a prohibited practice in both human and horse sports to maintain integrity [1]. The global leader of thoroughbred racing, the International Federation of Horseracing Authorities (IFHA, https://www.ifhaonline.org/, accessed on 1 September 2022), prohibits the administration of genetic materials including transgenes, therapeutic oligonucleotides, and genetically modified cells to horses; further, the IFHA and the International Stud Book Committee (ISBC, https://www.internationalstudbook.com/, accessed on 1 September 2022) prohibit the creation of genetically engineered racehorses. Genetically engineered animals (including embryos) have been recently created in many species, including horses [2,3]. The methods for creating genetically modified animals include the introduction of an exogenous gene into the host genome using transposons [4], retroviruses (including lentiviruses) [5], cell-mediated transgenesis [6], and gene editing (homologous recombination [HR] of CRISPR/Cas9) [7,8]. Gene insertion is used in gene therapy to supplement the functions of defective genes [9]. Alternatively, CRISPR/Cas9 can be used to specifically introduce short insertions or deletions (INDELs) to targeted regions via nonhomologous end joining (NHEJ) [10,11]. This facilitates gene knockdown or knockout by introducing a frameshift into the targeted gene. Pathological models of animals using gene-editing techniques to study diseases and livestock with modified genes related to economic traits have also been developed [12,13]. Gene-editing techniques make it theoretically possible to easily perform illegal gene doping to produce genetically modified racehorses, which pose a threat to the integrity of both the horse racing industry and equestrian sports. Many methods to detect inserted transgenes using quantitative PCR have been developed for gene-doping control [14,15,16,17,18,19,20,21,22,23]. By designing hydrolysis probes targeting the exon/exon junction, transgenes can be detected in a target-specific manner. In recent years, next-generation sequencing (NGS) technology has enabled collection of large amounts of massive parallel sequence (MPS) information. NGS technology has also enabled whole-genome resequencing of species with reference genome sequences [24,25,26]. Using NGS, transgenes can be detected as deletions of introns [27,28,29], enabling screening for nontargeted transgenes by comparing the detected intron deletions to annotated gene information (i.e., the intron location in reference sequences). Amplicon sequencing to detect targeted genes using NGS has also been developed for transgene detection [30,31]. Currently, there is no method to detect genome-edited sequences. NGS technology has enabled large-scale whole-genome resequencing (WGR), whole-exome resequencing (WER), and targeted gene or exon resequencing of species with reference genome sequences. WGR and WER are generally used to search for mutations causative of hereditary diseases [32,33], and targeted resequencing is often used to detect mutations in genes associated with cancer [34]. In this study, we developed and validated a method to identify genome-edited sequences using targeted resequencing and examined its applicability by screening gene-edited equine cells. We also evaluated the effectiveness of a single-nucleotide variant (SNV) panel to detect foals that have been produced using embryo transfer and somatic cell nuclear transfer, two artificial breeding techniques that are also prohibited in the thoroughbred racing and breeding industry.

2. Materials and Methods

2.1. Animal Ethics and Sample Collection

Blood and hair sample collection was approved by the Animal Care Committee of the Laboratory of Racing Chemistry (approval number 20-4, 13 February 2020). Samples were collected from horses at the Hidaka Training and Research Center (HTRC) of the Japan Racing Association (JRA), Japan Bloodhorse Breeders’ Association (JBBA), and donated by individual horse owners by the Japan Association for International Racing and Stud Book (JAIRS). Hair roots (n = 46) and whole blood (n = 50) were collected from thoroughbred horses (1–15 years old) as a validation set of samples for the developed gene-editing test. Hair roots from 120 thoroughbred horses (<1 years old) were collected as casework examples to test the utility of the parent verification SNV panel. Blood was collected in BD Vacutainer® spray-coated K2EDTA tubes (Becton, Dickinson and Company, Franklin Lakes, NJ, USA). Hair with roots were pulled from the mane. Blood and hair were stored at −30 °C and 4 °C, respectively, until use.

2.2. DNA Extraction

Genomic DNA was extracted from whole blood (200 μL) and hair (5 and 10 roots) using a DNeasy Blood & Tissue Kit (Qiagen, Hilden, Germany). Extracted DNA was quantified using a Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific, Waltham, MA, USA). The extracts were diluted to 4 ng/μL with Milli-Q.

2.3. Construction of Genome Edited Cells

Ear tissues of necropsied horses were provided by the Equine Research Institute, Japan Racing Association (ERI2021-711, 12 October 2021). Fibroblasts were isolated from ear tissue using an explant culture method. The tissue was finely diced (5–10 mm-thickness) using two scalpels and then cultured at 37 °C in Dulbecco’s modified Eagle’s medium (Sigma-Aldrich, St. Louis, MO, USA) containing 10% fetal bovine serum (Gibco, Grand Island, NY, USA) and gentamicin (20 μg/mL). After cell migration was confirmed, the tissues were removed from the culture medium. Primary cells were passaged once and then harvested for transfection. The Guide-it CRISPR/Cas9 System (Takara Bio Inc., Kusatsu, Shiga, Japan) was used to edit the myostatin (MSTN) gene. gRNAs (Guide RNAs) were designed to target locations 66609921-66609940 (5′-TTTCCAGGCGCAGTTTACTG-3′) on exon 1, 66607737-66607756 (5′-TACCTTGTACCGTCTTTCAT-3′) on exon 2, and 66605475-66605494 (5′-AAATCTCTTCTGGATCGTTT-3′) on exon 3 in chromosome 18 of EquCab3.0 (GCA_002863925.1). DNA oligonucleotides of each gRNA sequence were artificially synthesized (Fasmac Co., Ltd., Atsugi, Kanagawa, Japan) and cloned into the pGuide-it plasmid vector, which contained the U6 promoter and sgRNA scaffold to transcribe gRNA, genes encoding Cas9, and green fluorescent protein (GFP). The constructed pGuide-it vector was electroporated into cultured horse fibroblasts using the Neon® Transfection System (Thermo Fisher Scientific). Transgenic cells expressing GFP were selected and a single transgenic cell line representing each edit was isolated in a 24-well culture plate (Corning, Inc., Corning, NY, USA) using the TransferMan® Nk-2 Micromanipulation System (Eppendorf AG, Hamburg, Germany). As some genome edits did not develop into sufficient numbers of cells, whole-genome amplification (WGA) was performed using the REPLI Single Cell Kit (Qiagen) as required. Finally, genomic DNA was extracted from genome-edited cells or their WGA products using a DNeasy Blood and Tissue Kit (Qiagen). The extracted DNA was quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific). The extracts were diluted to approximately 4 ng/μL using Milli-Q.

2.4. Assay Design for Targeted Genes and SNVs

Fifty-two equine genes (Table 1) selected based on their known functions in several species (primarily horse, human, and mouse) and 147 single-nucleotide variants (SNVs) typically used for DNA typing were targeted in this study [35,36]. Primers for PCR amplification of targeted resequencing were designed by Illumina Concierge using DesignStudio Sequencing Assay Designer (Illumina Inc., San Diego, CA, USA) based on the horse reference genome, EquCab3.0 (GenBank: GCA_002863925.1). Each amplicon ranged from 125 base pairs (bp) to 275 bp in length.
Table 1

Gene name, location, and analyzed exon numbers for the gene-editing test.

Gene NameGene SymbolChromosomeExon Number
angiotensin I converting enzyme ACE 1125
actinin α 2 ACTN2 121
actinin α 3 ACTN3 1221
activin A receptor type 1B ACVR1B 611
adrenoceptor β 2 ADRB2 141
angiotensinogen AGT 14
AKT serine/threonine kinase 1 AKT1 2413
aldehyde dehydrogenase 2 family member ALDH2 813
adenosine monophosphate deaminase 1 AMPD1 516
apolipoprotein A2 APOA2 53
apolipoprotein A5 APOA5 71
bradykinin receptor B2 BDKRB2 242
creatine kinase, M-type CKM 108
clock circadian regulator CLOCK 320
ciliary neurotrophic factor CNTF 122
erythropoietin EPO 135
fibroblast growth factor 1 FGF1 143
fibroblast growth factor 2 FGF2 23
fibroblast growth factor 7 FGF7 13
follistatin FST 216
FTO α-ketoglutarate dependent dioxygenase FTO 310
growth hormone 1 GH1 114
hypoxia inducible factor 1 subunit α HIF1A 2415
insulin-like growth factor 1 IGF1 284
insulin-like growth factor 2 IGF2 123
Insulin-like growth factor binding protein 3 IGFBP3 45
interleukin 15 receptor subunit α IL15RA 298
interleukin 1 receptor antagonist IL1RN 154
interleukin 6 IL6 45
interleukin 6 receptor IL6R 510
lactase LCT 1817
lipoprotein lipase LPL 210
melanocortin 4 receptor MC4R 81
myostatin MSTN 183
methylenetetrahydrofolate reductase MTHFR 212
5-methyltetrahydrofolate-homocysteine methyltransferase MTR 133
5-methyltetrahydrofolate-homocysteine methyltransferase reductase MTRR 2116
neuromedin B NMB 13
nitric oxide synthase 3 NOS3 425
phosphoenolpyruvate carboxykinase 1 PCK1 229
pyruvate dehydrogenase kinase 4 PDK4 412
peroxisome proliferator activated receptor α PPARA 286
peroxisome proliferator activated receptor delta PPARD 206
peroxisome proliferator activated receptor γ PPARG 166
PPARG coactivator 1 α PPARGC1A 313
solute carrier family 16 member 1 SLC16A1 54
solute carrier family 30 member 8 SLC30A8 910
uncoupling protein 2 UCP2 78
uncoupling protein 3 UCP3 76
vitamin D receptor VDR 68
vascular endothelial growth factor A VEGFA 207
zinc finger and AT-hook domain containing ZFAT 917

2.5. Library Preparation and Sequencing

Libraries were prepared using AmpliSeq Library PLUS for Illumina (Illumina Inc.) based on the manufacturer’s recommendations. AmpliSeq CD Indexes Set A-D for Illumina (384 Indexes, Illumina Inc.) were used to index the samples. Sequencing was performed on the NextSeq 500 sequencing platform (Illumina Inc.) using the NextSeq 500/550 Mid Output Kit v2.5 (300 Cycles, Illumina Inc.).

2.6. Data Analysis, Variant Calling, and SNV Genotyping

Variant detection was performed using the RESEQ pipeline (Amelieff Co., Minato, Tokyo, Japan), which was constructed based on QCleaner (Amelieff Co.), Burrows-Wheeler aligner (BWA, version 0.7.17) [37], Genome Analysis Toolkit (GATK, version 4.0.8.1, Broad Institute, Cambridge, MA, USA) (https://software.broadinstitute.org/gatk/best-practices/, accessed on 5 August 2022), and SnpEff (version v4_0) (http://pcingola.github.io/SnpEff/, accessed on 5 August 2022) [38]. While the original pipeline had a step for excluding duplicated reads using Picard (version 2.13.2) (https://broadinstitute.github.io/picard/, accessed on 5 August 2022), this step was not included in this study. Briefly, quality control was performed on raw reads using QCleaner, and low-quality bases (<20 phred scores) were removed. Additionally, reads were removed if 80% of their nucleotides had a quality value <20, if they had sequences of over five unknown nucleotides, if they had <32 base length sequences, or if they did not have a mate-pair. Finally, only high-quality sequences were selected. BWA was used to align the reads using default parameters to the horse reference genome sequence EquCab3.0 assembly from GenBank (GCA_002863925.1). Alignments were converted from the sequence alignment/map (SAM) format to sorted and indexed binary alignment/map (BAM) files using SAMtools (version 1.8) (http://www.htslib.org/, accessed on 5 August 2022). GATK was used to detect SNVs and INDELs using the default parameters, with the exception of: −minIndelFrac = 2.0. SNVs, and INDELs were annotated using SnpEff. In this analysis, the GATK VariantFiltration program was not used. The obtained variant data and their depth information in the 52 targeted gene regions were extracted from each VCF file and BAM file and then combined as one file using GENOMINEE (Amelieff Co.). Mapped reads and detected variants were visualized using Integrative Genomics Viewer (IGV, version 2.3.97, Broad Institute). SNVs were genotyped by collecting variant information (REF-allele, ALT-allele, and their depth) using a BED file in which SNV locations were listed. As HomoREF (REF/REF) was not listed in the VCF file, the depth information of HomoREF was collected from the BAM file using GENOMINEE (Amelieff Co.). REF/REF is HomoREF with sufficient depth information (depth > 50).

3. Results and Discussion

3.1. Extraction of Genomic DNA from Hair Roots and Blood Samples

All thoroughbreds are parent-verified as part of the registration process. This would also be a suitable time to perform the gene-editing test, as a condition of registration is an unmodified genome. Currently, hair roots or blood are the routine sample types collected for parentage testing of thoroughbreds. Amplicon sequencing using NGS generally requires purified genomic DNA (e.g., 4 ng/μL × 5 μL) as the template DNA, and DNA extracted from blood is known to satisfy library preparation quality and quantity requirements for NGS (Table 2). However, the method for extracting purified genomic DNA required for amplicon sequencing from hair roots has not been properly evaluated. Here, we detail a DNA extraction method from hair roots sufficient for amplicon sequencing.
Table 2

Mean concentration of genomic DNA extracted from hair roots or blood in 11 horses that were younger than one year or older than two years.

SampleAgeNumber/VolumeMean Concentration (ng/μL)Standard Deviation
Hair root0-year-old5 roots1.250.31
Hair root0-year-old10 roots1.630.53
Hair root2-years-old5 roots1.670.48
Hair root2-years-old10 roots3.030.91
Blood2-years-old200 μL47.213.5
Genomic DNA was extracted from five and ten hair roots of 11 horses that were less than one-year-old or over two-years-old. The mean concentrations of DNA extracted are shown in Table 2. DNA extracts from either five or ten hair roots did not reach 4 ng/μL, the required quantity for library preparation for NGS. The genomic DNA was concentrated under reduced pressure to obtain amounts of DNA suitable for NGS analysis. This indicates that a gene-editing test could be performed using the same sample collected for the standard parentage test. In a previous study, an average of 8.5 ng/μL (200 μL) of genomic DNA was extracted from 15 hair roots of horses older than one-year-old [39]. However, to streamline the testing of several samples, it may be better to use a small number of hair roots (i.e., five hair roots) to minimize the time spent cutting samples.

3.2. Construction of Genome Edited Cells

Four different genome-edits targeting the MSTN gene were introduced into horse fibroblasts, as shown in Table 3. By designing a gRNA complimentary to sequence in exon 3, a heterozygous edited cell lacking 77 bp at the end of intron 2 and start of exon 3 was constructed (Figure 1A, Edited cell_1). Similarly, gRNAs were designed to target two areas in exon 2, and one in exon 1, resulting in homozygous edited cells with a 1-bp insertion in either exon 2 or 1 (Figure 1B–D, Edited cell_2, 3, and 4). The edited sites were all within the designed gRNA sequences.
Table 3

Summary of gene-edited cells constructed using CRISPR/Cas9.

NameLocationExon/IntronGenotypeReference-AlleleEdited-Allele
Edited cell_1chr18: 66605476-66605554Exon3-Intron2Hetero77-bp insertion77-bp deletion
Edited cell_2chr18: 66607750Exon2HomoCCT
Edited cell_3chr18: 66607753Exon2HomoTTA
Edited cell_4chr18: 66609936Exon1HomoTTA
Figure 1

Sanger sequencing detection of variants introduced using CRISPR/Cas9 into the horse myostatin gene. (A) Edited cell_1 (chr18: 66605476-66605554, exon3-intron2, 77-bp deletion, heterozygote); (B) Edited cell_2 (chr18: 66607750, exon2, 1-bp insertion, homozygote); (C) Edited cell_3 (chr18: 66607753, exon2, 1-bp insertion, homozygote); (D) Edited cell_4 (chr18: 66609936, exon1, 1-bp insertion, homozygote).

One of the simplest types of edits to achieve using CRISPR/Cas9 is insertion of a single nucleotide on both DNA strands, creating a novel homozygous alternative version of a gene (homoALT) [40]. For this reason, three of the four edited cells were constructed as such in this study. The MSTN gene is a negative regulator of muscle growth, and animals (including humans) lacking functional myostatin protein have above-average muscle mass [41,42]. Consequently, disruption of this gene has been the target of genome editing in many animals, including the horse [2,12,43]. In both racing whippets and thoroughbred racehorses, naturally occurring variants of the MSTN gene affect racing performance [44,45]. In thoroughbred horses, where this phenomenon is well-characterized, the variant alters the amount and composition of muscle fiber types [46,47]. Thus, the MSTN gene is an obvious target for gene editing in racehorses, and it was important to show that our screening method would be able to detect MSTN gene editing.

3.3. Genes Targeted for the Gene-Editing Test and Sequencing Summary

In this study, we targeted 52 genes containing 481 exons (as shown in Table 1) to screen as likely targets for the gene-editing detection test. In total, 1012 regions, including the SNVs incorporated for parent verification described below, were amplified using standard Illumina library preparation steps. The designed primer sequences are not disclosed as this test will be used for gene-doping control. To validate the method, 120 samples, including 27 duplicate samples, were used (Table 4). Most of the genome-edited cell lines constructed in this study did not grow well, resulting in only a small number of edited cells being developed. To address this, the genomic DNA obtained from cultured edited cells was subjected to whole-genome amplification (WGA) to obtain enough DNA to achieve successful library preparation. Only the Edited Cell_1 line grew enough cells to allow for sufficient genomic DNA to be extracted for library preparation without WGA. It is not known whether this is because the edit made to Edited Cell_1 resulted in a heterozygous change rather than a homozygous change.
Table 4

DNA extraction conditions for 120 samples used for validation of the gene-editing test.

SampleConditionNumber of Samples Used
Hair rootsExtracted/Purified46
Hair rootsWGA, Extracted/Purified *5 ***
BloodExtracted/Purified50 ****
Edited cell_1Extracted/Purified1
Edited cell_1WGA, Extracted/Purified *2
Edited cell_1WGA **1
Edited cell_2WGA, Extracted/Purified *4
Edited cell_2WGA **2
Edited cell_3WGA, Extracted/Purified *2
Edited cell_3WGA **1
Edited cell_4WGA, Extracted/Purified *4
Edited cell_4WGA **2

*: Amplified genomic DNA amplified was purified using the DNeasy Blood and Tissue Kit. **: Amplified genomic DNA was directly used for library preparation. ***: The five samples were duplicates of samples included in the group of 46. ****: Four of the fifty blood samples were taken from horses that were also typed using 46 hair root samples.

Details of the NGS output for the 120 samples are described in Table 5. On average, 1299 SNVs and 94 INDELs were detected in each sample when mapping to all chromosomes without filtering for quality, although these numbers varied widely across the group. A sufficient amount of FASTQ data and mapped reads were obtained from hair roots, blood (with the exception of one sample), and genome-edited cells both with and without WGA and purification. The sample with the smallest amount of FASTQ data was derived from blood that should be most suitable for NGS. No high-quality variant calls were made in the subsequent analysis, and it was assumed an error occurred in the library preparation of this sample.
Table 5

Summary of sequencing and variant calling in the gene-editing test.

MeanStandard DeviationMaximumMinimum
FASTQ data size218,009,34437,895,950397,302,67417,404,419
Mapped reads *2,075,806362,1313,795,577116,739
SNV **12998545345388
INDEL **945631930

*: The number of reads mapped after quality check. **: No filtrations.

Of the 481 targeted exons, 470 (97.7%) had greater than 50× read coverage, 478 exons (99.4%) were covered by more than 20 reads, and all but one exon (99.8%) averaged more than 10× coverage. The variation in coverage was attributed to variable PCR amplification efficiency of each primer (Figure 2). Whilst the coverage obtained in this study was sufficient to be fit for purpose, future iterations of this test could address this by altering the design of primers targeted at exons that showed low coverage in this instance.
Figure 2

Depth of sequencing coverage of 481 exons in 52 genes. While the coverage varied from exon to exon, over 97% (470 exons) were covered by more than 50 sequence-reads.

3.4. Detection of Artificial Mutations in the MSTN Gene Region

Using NGS and the analysis pipeline on the 120 samples, the 1-bp insertions (homoALT) introduced into Edited cells_2, 3, and 4 were easily detected (Figure 3B–D). However, the heterozygous 77-bp deletion (REF/ALT) introduced into Edited cell_1 was not detected (Figure 3A). The 77-bp deletion incorporated parts of exon 3 and intron 2 and was located 22-bp from the end of the PCR amplicon. The ends of the amplicon were excluded from variable calling by GATK as this program performs a soft clip at the ends of reads, assuming this region is index or library sequence. Thus, large deletions near the ends of PCR amplicons will be difficult to detect using the analysis pipeline developed in this study.
Figure 3

Next-generation sequencing detection of variants introduced using CRISPR/Cas9 into the horse myostatin gene. Each variant was visualized using the Integrative Genomics Viewer. (A) Edited cell_1 (chr18: 66605476-66605554, exon3-intron2, 77-bp deletion, heterozygote, sequence-reads for the allele without the 77-bp deletion were detected but are out of the frame in the figure); (B) Edited cell_2 (chr18: 66607750, exon2, 1-bp insertion, homozygote); (C) Edited cell_3 (chr18: 66607753, exon2, 1-bp insertion, homozygote); (D) Edited cell_4 (chr18: 66609936, exon1, 1-bp insertion, homozygote).

However, all three homozygous, alternative, single-base-insertion alleles generated in the edited cells were detected by the proposed NGS sequencing and pipeline. CRISPR/Cas9 is particularly efficient at inserting single nucleotides on both DNA strands, creating a novel, homozygous, alternative version of a gene (homoALT) [40]. A single base insertion within a coding region will create a frameshift mutation that (depending on the position of the insertion within the coding sequence) is very likely to be destructive to the function of the mature protein. It is thus important that any gene-editing detection method will be able to detect these small but disruptive edits of gene sequence and function. The use of CRISPR/Cas9 is not always 100% efficient and may introduce new alleles in the form of compound heterozygotes (ALT1/ALT2) [48]. Based on the results of this study, these should be detected by comparison with the reference sequence using the methods described. Therefore, artificially introduced homozygous alternative and compound heterozygous sequence changes should be identifiable when screening for gene-editing tests.

3.5. Establishment of Criteria for Detecting Artificially Edited Variants

An important part of screening for prohibited gene editing is to establish criteria to identify the most likely candidate genome-edited variants (i.e., homoALT or ALT1/ALT2). These should differentiate between naturally occurring variants and those that are more likely to have been artificially introduced. In this study, a total of 575 variable loci (492 SNVs and 83 INDELs) were detected in the targeted genes in the validation sample set of 120 samples, which included the gene-edited cells. Homozygous INDELs are likely to be preferred by someone wanting to simply knock out the function of a gene. Of the coding INDELs identified, ten were homoALTs and one was a compound heterozygous ALT1/ALT2. The remaining INDELs were heterozygous. Among the INDELs, four homoALT mutations were excluded as candidates for gene editing because the particular variants had previously been reported in whole-genome sequence data from 101 thoroughbred horses [49]. One compound heterozygous INDEL (ALT1/ALT2) had also been previously observed in the thoroughbred population. Of the remaining six loci, three were excluded as gene-editing candidates because they were covered by extremely low sequence depth in the detected individuals. The remaining three loci were the specific homoALT single-nucleotide insertions introduced into the genome-edited cells (Edited cell_2, 3, and 4), which showed sufficient sequence depth. Based on the above data, we proposed the following filtering criteria for identifying artificial editing candidates: Selection of homoALT or ALT1/ALT2 caused by INDELs as artificial editing candidates; Exclude variants that already exist in the thoroughbred population; Exclude variants in regions with low sequence depth.

3.6. Casework Example

The gene-editing test and detection criteria established in this study were applied to 120 new samples (the casework sample set) collected for the standard parent verification process, which forms an important part of the registration of thoroughbreds in their Stud Book. A total of 508 variants were detected in this group of 120 samples. Those most likely to be gene-editing candidates (homoALT or ALT1/ALT2) constituted seven loci, of which six were already documented as present in the current thoroughbred population. The remaining loci had low coverage. Thus, it is assumed no gene editing was observed in the 120 samples using the developed gene-editing screening test under the described conditions.

3.7. Construction of a SNV Panel to Detect Parentage in Thoroughbred Horses

To create genetically modified animals, it is necessary to transplant edited eggs into surrogate mothers to carry to term [50,51]. Therefore, a simple way to identify a genetically engineered foal is to compare the foal’s DNA with that of the mare that carried it. Therefore, we developed a SNV panel to confirm mother–child inheritance. Primers were designed for amplicon sequencing of 147 loci used in the 2021 SNV comparison test run by the International Society for Animal Genetics (ISAG) [34,35]. These primers were included in the panel for the gene-editing test, and library preparation and sequencing by NGS were performed simultaneously. Three loci failed (chr2:134944, chr9:20379456, and chr18:34581176), leaving 144 SNVs genotyped on the 120 samples used for validating the gene-editing test. Firstly, 29 trios (sire, dam and foal) were genotyped, and the inheritance of the SNVs by the foals investigated. Of the 144 SNV, 24 SNVs showed mis-inheritances among 27 trios (Table 6). The causes of these mis-inheritances were 1) insufficient coverage (i.e., depth < 50) in all or part of the analyzed individuals and 2) different depths of coverage across different alleles. The former is attributed to the difficulty of multiplexing PCR (1012 amplicons), whereas the latter may involve PCR amplification bias. Trio analysis of the remaining 120 SNV confirmed no mis-inheritances (Table 6). Therefore, these 120 SNVs were selected as the SNV parentage panel.
Table 6

Confirmation of inheritance for 144 and 120 SNVs.

144 SNVs120 SNVs
ConsistentInconsistentConsistentInconsistent
Trio_114221200
Trio_213951200
Trio_314221200
Trio_414311200
Trio_514311200
Trio_614311200
Trio_714131200
Trio_814131200
Trio_914311200
Trio_1014041200
Trio_1114041200
Trio_1214131200
Trio_1314131200
Trio_1414131200
Trio_1514311200
Trio_1614131200
Trio_1714401200
Trio_1814041200
Trio_1914131200
Trio_2014311200
Trio_2114401200
Trio_2214221200
Trio_2314221200
Trio_2414221200
Trio_2514221200
Trio_2614221200
Trio_2714131200
Trio_2814131200
Trio_2914131200
Pseudo-Trio_1108369129
Pseudo-Trio_2110348832
Pseudo-Trio_3112329129
Pseudo-Trio_4111339327
Pseudo-Trio_5108369030
Half-pseudo-Trio_11232110218
Half-pseudo-Trio_2110349525
Half-pseudo-Trio_3119259921
Half-pseudo-Trio_4116289723
Half-pseudo-Trio_51261810515
Parent-Foal_114311200
Parent-Foal_214401200
Parent-Foal_314221200
Parent-Foal_414311200
Parent-Foal_514401200
Pseudo-parent-Foal_11321210911
Pseudo-parent-Foal_21281610812
Pseudo-parent-Foal_31311311010
Pseudo-parent-Foal_41321211010
Pseudo-parent-Foal_513681146
Next, we evaluated the accuracy of the SNVs for detecting mis-inheritance using 1) five pseudo-trios (pseudo-sire, pseudo-dam, foal), 2) five half-pseudo-trios (sire, pseudo-dam, and foal; or pseudo-sire, dam, and foal), and 3) five pseudo parents (sire or dam)-foal (Table 6). Pseudo-trios, half-pseudo-trios, and pseudo-parent-foal were detected by 29.5, 20.4, and 9.8 SNVs on average using the 120 SNV panel. Table 7 shows the allele frequency, heterozygosity (He), and paternity exclusion (PE) values of the 120 SNVs. The combined PE1 and PE2 values with 120 SNVs were 0.9999999998 and 0.999997, respectively. Five new single parent–foal combinations were also evaluated and parentage confirmed using the selected SNV panel.
Table 7

Statistical summary of 120 SNVs in the gene-editing test.

SNV IDChromosomeLocationREF-alleleALT-alleleREF frequencyHePE1PE2
BIEC2-11336124061861CT0.5820.4870.1840.118
BIEC2-23891158467370TC0.8470.2590.1130.034
BIEC2-34560180153039TC0.5200.4990.1870.125
BIEC2-34987181012651TG0.7240.3990.1600.080
BIEC421181102064366CT0.2550.3800.1540.072
BIEC2-601861138352227AC0.7550.3700.1510.068
BIEC2-651451153834277CT0.6350.4630.1780.107
BIEC2-785231166558525GA0.5920.4830.1830.117
BIEC2-459311217808463TC0.7400.3850.1560.074
BIEC2-476920246942336CA0.6330.4650.1780.108
BIEC2-491394278433836AG0.5710.4900.1850.120
BIEC2-502451298905614AG0.6730.4400.1720.097
BIEC4208942120778620TC0.6530.4530.1750.103
BIEC2-777914338956274TC0.6220.4700.1800.110
BIEC672139386457199TC0.6940.4250.1670.090
BIEC2-798927388082655AG0.6330.4650.1780.108
BIEC2-799664390095202CA0.6460.4570.1760.105
BIEC2-800511391766696TC0.2350.3590.1470.065
BIEC2-8067713101540587GA0.8160.3000.1270.045
BIEC6819893104518656CT0.5920.4830.1830.117
BIEC2-8118863119599447GA0.3270.4400.1720.097
BIEC7078984178566TC0.7290.3950.1580.078
BIEC2-853347421336785GA0.6020.4790.1820.115
BIEC2-908630545880924CT0.5310.4980.1870.124
BIEC2-910827553722932TG0.6040.4780.1820.114
BIEC2-914714563272427TC0.3780.4700.1800.110
BIEC2-919835573509875GT0.7240.3990.1600.080
BIEC2-946446632343233AG0.7040.4170.1650.087
BIEC798010645478480AG0.3980.4790.1820.115
BIEC811791676708313GT0.6220.4700.1800.110
BIEC81938575215592AC0.4800.4990.1870.125
BIEC2-1007607780202186CT0.5200.4990.1870.125
BIEC870244817383321GC0.6530.4530.1750.103
BIEC871916820028234AT0.5510.4950.1860.122
BIEC2-1052417857040628CT0.7140.4080.1620.083
BIEC2-1066033894179121GT0.5200.4990.1870.125
BIEC2-1066179895184881GA0.7350.3900.1570.076
BIEC2-1080866925656223TC0.5510.4950.1860.122
BIEC915102927114849CT0.5000.5000.1880.125
BIEC8628110922571CT0.4390.4930.1860.121
BIEC2-95522104751573AG0.5310.4980.1870.124
BIEC2-97679107701310GA0.6430.4590.1770.105
BIEC2-1196401044453287TC0.6730.4400.1720.097
BIEC2-1211021048342330AG0.7350.3900.1570.076
BIEC2-1230021052573273TC0.8670.2300.1020.026
BIEC2-1267321060049178AG0.4080.4830.1830.117
BIEC1191581065321527AT0.5820.4870.1840.118
BIEC2-136591117641341TC0.5100.5000.1870.125
BIEC1368211129398313AG0.5920.4830.1830.117
BIEC2-1622451157232386CT0.6530.4530.1750.103
BIEC154171123878382GT0.3370.4470.1730.100
BIEC155175125186277CT0.6980.4220.1660.089
BIEC2-1896661222012560GA0.5510.4950.1860.122
BIEC2-2143461319184556GA0.4900.5000.1870.125
BIEC1801221335676903AG0.6630.4470.1730.100
BIEC183067144666366AG0.5200.4990.1870.125
BIEC2-2459231414775426CT0.7040.4170.1650.087
BIEC2040221464551633AG0.8130.3050.1290.046
BIEC2-2636161469988182CA0.4800.4990.1870.125
BIEC2-2707951483088387CT0.4180.4870.1840.118
BIEC2-278532151201270AG0.7760.3480.1440.061
BIEC2472841548661214GA0.6330.4650.1780.108
BIEC2-3102691552664101TC0.4060.4820.1830.116
BIEC2-326637161496069AG0.6220.4700.1800.110
BIEC266761163860430CT0.5710.4900.1850.120
BIEC2-328954168598011AG0.3980.4790.1820.115
BIEC2-3362691627983909GA0.2350.3590.1470.065
BIEC2-3405951636915439AG0.4080.4830.1830.117
BIEC2-3428811642323485CT0.6020.4790.1820.115
BIEC2-3580611671914513CT0.2960.4170.1650.087
BIEC2-3647411683054140GA0.6840.4330.1690.094
BIEC2-3649451684257390CT0.5410.4970.1870.123
BIEC2-3696991717649634GA0.5100.5000.1870.125
BIEC2-3745711727036258AG0.6530.4530.1750.103
BIEC3227341778660174AG0.4800.4990.1870.125
BIEC2-390920182190684AC0.7450.3800.1540.072
BIEC3245301811431365CT0.6330.4650.1780.108
BIEC2-4072061817621415AG0.4290.4900.1850.120
BIEC2-4153151856944269GA0.5100.5000.1870.125
BIEC2-4305771919611080AG0.2960.4170.1650.087
BIEC2-4390601942595105CT0.6880.4300.1690.092
BIEC2-4428921948488636TC0.6220.4700.1800.110
BIEC4314452016383556TC0.2810.4040.1610.082
BIEC2-5307882033600979TC0.4800.4990.1870.125
BIEC2-5321062040481673CT0.5710.4900.1850.120
BIEC2-5357662049037728TC0.5510.4950.1860.122
BIEC2-575436222011040CT0.5820.4870.1840.118
BIEC2-5849112219064667AG0.3980.4790.1820.115
BIEC4938792230797481CT0.2960.4170.1650.087
BIEC4998602240978756TG0.5510.4950.1860.122
BIEC2-606469234277770GA0.6120.4750.1810.113
BIEC5084102316413178CG0.5420.4970.1870.123
BIEC2-6182842319540553GT0.7500.3750.1520.070
BIEC5234832412615461TC0.2290.3530.1450.062
BIEC2-6369872417444264CT0.6940.4250.1670.090
BIEC5312752427556162CT0.2960.4170.1650.087
BIEC5320602428656489AG0.6120.4750.1810.113
BIEC5416932445741983CT0.6940.4250.1670.090
BIEC5557372512340280GA0.6350.4630.1780.107
BIEC5548132514083078CG0.3470.4530.1750.103
BIEC581695261937434GA0.4690.4980.1870.124
BIEC5717052619267701TC0.5610.4930.1860.121
BIEC2-6925432630370457GA0.6840.4330.1690.094
BIEC5624652642126216AG0.2650.3900.1570.076
BIEC2-7080182718560786CT0.7240.3990.1600.080
BIEC5909862719761701TA0.6630.4470.1730.100
BIEC2-7210612738093323TC0.3670.4650.1780.108
BIEC6044332739476969GT0.1730.2870.1230.041
BIEC2-725322285600796CT0.6040.4780.1820.114
BIEC607490285913256GT0.5000.5000.1880.125
BIEC2-7341032821956151CT0.4390.4930.1860.121
BIEC6170702823842936TC0.7040.4170.1650.087
BIEC2-7377042830865784GT0.4180.4870.1840.118
BIEC6331812913672303GC0.3270.4400.1720.097
BIEC6339792915765096TC0.3570.4590.1770.105
BIEC2-7618512929285909TG0.8670.2300.1020.026
BIEC2-816793307757606TG0.5310.4980.1870.124
BIEC6868003012032090TG0.1530.2590.1130.034
BIEC696480311930525TC0.4900.5000.1870.125
BIEC2-8386303116325150CT0.5710.4900.1850.120
Based on these results, the panel of 120 SNVs is capable of clearly identifying the nonbiological dams of thoroughbred foals. A similar method has been developed using maternally inherited mitochondrial DNA sequences [52]. Such approaches can indirectly screen for genetically modified animals; however, it is not possible to directly identify the edited sequences. These approaches may also be useful to identify horses that have been produced by somatic cell cloning, another practice that is prohibited in thoroughbreds. The development of this DNA profile panel is also important to verify the identity of the tested horses. Any horse that is found to be gene-edited will be required to be DNA profiled to ensure the sample was collected from the correct horse. Having the ability to run the identity analysis simultaneously with the gene-editing test enhances the integrity of the sample analysis. Currently, microsatellite analysis is used for thoroughbred parent verification. However, SNVs are being evaluated for this purpose so that once a final panel of SNVs is determined for standard parent verification across all laboratories this should be incorporated into the gene-editing screening process to allow for simultaneous screening and identification. One of the reasons to use targeted resequencing to identify genome-edited horses instead of whole-genome resequencing is to reduce the cost of such testing. The gene-editing screening test combined with parentage verification using SNVs cost considerably less than WGR, but more than the currently used parentage verification using microsatellites. The two tests combined cost an amount similar to that of parentage verification using SNVs alone. Thus, once SNVs are confirmed for parentage verification, screening thoroughbreds for gene editing at the same time as parentage verification will be cost-efficient. This will benefit all stakeholders including the breeders, owners, racing authorities, and Stud Books in the racing and breeding industries and will also benefit horse welfare by discouraging prohibited and experimental gene editing.

4. Conclusions

The purpose of this study was to construct a screening tool using amplicon sequencing to detect genetically modified thoroughbred horses produced by the use of nonhomologous end-joining with CRISPR/Cas9 in embryos. When 52 gene candidates related to racing performance were targeted and sequenced, 97.7% of the target sequences were covered with sufficient sequence reads to use this method. As a proof of principle, four genetically modified cell lines were screened alongside normal samples. The screening test easily identified the cell lines containing homozygous single-nucleotide insertions, the most efficient product of CRISPR/Cas9 editing. Concurrently, the samples were also DNA profiled, which would allow confirmation of identity if evidence of gene editing were uncovered. It is likely in a real-life gene-editing detection scenario that Sanger sequencing may be necessary to confirm the mutation at that site. The sequencing of PCR amplified amplicons illustrated that it is extremely difficult to uniformly amplify every targeted genomic region, and some regions may not be covered. Further, it will be difficult to differentiate between artificial and natural heterozygous variants. In this situation, further evidence of gene editing being carried out (for example, access to a laboratory with consumables for gene editing) may be required for successful prosecution. As such, identification of artificially introduced heterozygous variants would not be possible with the gene-editing test developed in this study. Nevertheless, this study is still an important step towards being able to identify genetically engineered thoroughbred racehorses.
  52 in total

1.  Molecular definition of an allelic series of mutations disrupting the myostatin function and causing double-muscling in cattle.

Authors:  L Grobet; D Poncelet; L J Royo; B Brouwers; D Pirottin; C Michaux; F Ménissier; M Zanotti; S Dunner; M Georges
Journal:  Mamm Genome       Date:  1998-03       Impact factor: 2.957

2.  Developmental competence of equine oocytes and embryos obtained by in vitro procedures ranging from in vitro maturation and ICSI to embryo culture, cryopreservation and somatic cell nuclear transfer.

Authors:  C Galli; S Colleoni; R Duchi; I Lagutina; G Lazzari
Journal:  Anim Reprod Sci       Date:  2006-10-17       Impact factor: 2.145

3.  Identification of processed pseudogenes in the genome of Thoroughbred horses: Possibility of gene-doping detection considering the presence of pseudogenes.

Authors:  Teruaki Tozaki; Aoi Ohnuma; Mio Kikuchi; Taichiro Ishige; Hironaga Kakoi; Kei-Ichi Hirota; Kanichi Kusano; Shun-Ichi Nagata
Journal:  Anim Genet       Date:  2022-01-25       Impact factor: 3.169

4.  A quantitative PCR screening method for adeno-associated viral vector 2-mediated gene doping.

Authors:  Zibin Jiang; Joanne Haughan; Kaitlyn L Moss; Darko Stefanovski; Kyla F Ortved; Mary A Robinson
Journal:  Drug Test Anal       Date:  2021-09-03       Impact factor: 3.345

5.  Generating Transgenic Mice by Lentiviral Transduction of Spermatozoa Followed by In Vitro Fertilization and Embryo Transfer.

Authors:  Anil Chandrashekran; Colin Casimir; Nick Dibb; Carol Readhead; Robert Winston
Journal:  Methods Mol Biol       Date:  2016

6.  Low-copy transgene detection using nested digital polymerase chain reaction for gene-doping control.

Authors:  Teruaki Tozaki; Aoi Ohnuma; Natasha A Hamilton; Mio Kikuchi; Taichiro Ishige; Hironaga Kakoi; Kei-Ichi Hirota; Kanichi Kusano; Shun-Ichi Nagata
Journal:  Drug Test Anal       Date:  2021-10-18       Impact factor: 3.345

7.  Fast and accurate short read alignment with Burrows-Wheeler transform.

Authors:  Heng Li; Richard Durbin
Journal:  Bioinformatics       Date:  2009-05-18       Impact factor: 6.937

8.  Whole-genome sequencing identifies missense mutation in GRM6 as the likely cause of congenital stationary night blindness in a Tennessee Walking Horse.

Authors:  Yael L Hack; Elizabeth E Crabtree; Felipe Avila; Roger B Sutton; Robert Grahn; Annie Oh; Brian Gilger; Rebecca R Bellone
Journal:  Equine Vet J       Date:  2020-08-03       Impact factor: 2.888

9.  Efficient production of large deletion and gene fragment knock-in mice mediated by genome editing with Cas9-mouse Cdt1 in mouse zygotes.

Authors:  Saori Mizuno-Iijima; Shinya Ayabe; Kanako Kato; Shogo Matoba; Yoshihisa Ikeda; Tra Thi Huong Dinh; Hoai Thu Le; Hayate Suzuki; Kenichi Nakashima; Yoshikazu Hasegawa; Yuko Hamada; Yoko Tanimoto; Yoko Daitoku; Natsumi Iki; Miyuki Ishida; Elzeftawy Abdelaziz Elsayed Ibrahim; Toshiaki Nakashiba; Michito Hamada; Kazuya Murata; Yoshihiro Miwa; Miki Okada-Iwabu; Masato Iwabu; Ken-Ichi Yagami; Atsuo Ogura; Yuichi Obata; Satoru Takahashi; Seiya Mizuno; Atsushi Yoshiki; Fumihiro Sugiyama
Journal:  Methods       Date:  2020-04-22       Impact factor: 3.608

10.  Genome edited sheep and cattle.

Authors:  Chris Proudfoot; Daniel F Carlson; Rachel Huddart; Charles R Long; Jane H Pryor; Tim J King; Simon G Lillico; Alan J Mileham; David G McLaren; C Bruce A Whitelaw; Scott C Fahrenkrug
Journal:  Transgenic Res       Date:  2014-09-10       Impact factor: 2.788

View more

北京卡尤迪生物科技股份有限公司 © 2022-2023.