Zai-Zhen Wu1, Jian-Hua Tang1, Bin Zhang1, Li-Ping Guo2, Hong-Ping Xie1, Bing-Ren Gu3. 1. College of Pharmaceutical Sciences, Soochow University, Suzhou 215123, China. 2. College of Chemistry and Chemical Engineering, Chongqing University of Science and Technology, Chongqing 401331, China. 3. Suzhou Institute for Drug Control, Suzhou 215002, China.
Abstract
This paper studied the expert system of genotype discrimination for the STR locus D5S818 based on near-infrared spectroscopy-principal discriminant variate (PDV). Six genotypes, i.e. genotypes 10-10, 10-11, 11-11, 11-12, 11-13 and 13-13, were selected as research subjects. Based on the optimum polymerase chain reaction (PCR) conditions, about 54 measuring samples for each genotype were obtained; these samples were tested by near-infrared spectroscopy directly. With differences between homozygote genotypes and heterozygote ones, and differences of the total number of core repeat units between the six genotypes, two types of genotyping-tree structure were constructed and their respective PDV models were studied using the near-infrared spectra of the samples as recognition variables. Finally, based on the classification ability of these two genotyping-tree structures, an optimum expert system of genotype discrimination was built using the PDV models. The result demonstrated that the built expert system had good discriminability and robustness; without any preprocessing for PCR products, the six genotypes studied could be discriminated rapidly and correctly. It provided a methodological support for establishing an expert system of genotype discrimination for all genotypes of locus D5S818 and other STR loci.
This paper studied the expert system of genotype discrimination for the STR locus D5S818 based on near-infrared spectroscopy-principal discriminant variate (PDV). Six genotypes, i.e. genotypes 10-10, 10-11, 11-11, 11-12, 11-13 and 13-13, were selected as research subjects. Based on the optimum polymerase chain reaction (PCR) conditions, about 54 measuring samples for each genotype were obtained; these samples were tested by near-infrared spectroscopy directly. With differences between homozygote genotypes and heterozygote ones, and differences of the total number of core repeat units between the six genotypes, two types of genotyping-tree structure were constructed and their respective PDV models were studied using the near-infrared spectra of the samples as recognition variables. Finally, based on the classification ability of these two genotyping-tree structures, an optimum expert system of genotype discrimination was built using the PDV models. The result demonstrated that the built expert system had good discriminability and robustness; without any preprocessing for PCR products, the six genotypes studied could be discriminated rapidly and correctly. It provided a methodological support for establishing an expert system of genotype discrimination for all genotypes of locus D5S818 and other STR loci.
Entities:
Keywords:
Expert system; Genotyping-tree structure; Near-infrared spectroscopy; Principal discriminant variate; Short tandem repeat
Short tandem repeats (STRs), also known as micro-satellites or simple sequence repeats, are polymorphic sites containing core repeat units of between two and seven nucleotides in length that are tandemly repeated from approximately a half dozen to several dozen times [1]. Because of high polymorphism and genetic stability, they have been widely used in the construction of genomic maps [2], paternity tests in forensic medicine [3], [4] and gene diagnosis of diseases [5], [6]. At present, many methods have been reported for detection of STRs, including capillary electrophoresis [7], polyacrylamide gel electrophoresis [8], microdevice electrophoresis [9] and mass spectrometry [10]. For high throughput analysis, the capillary array electrophoresis chip has also been developed [11]. For most of the methods, polymerase chain reaction (PCR), which is used to amplify information from small amounts of available biological materials, is needed. For analysis of PCR products, fluorescent dye markers are needed in the electrophoresis-based methods, and purification of PCR products is necessary in mass spectrometry analysis. Our laboratory has investigated the feasibility of near-infrared spectroscopy–principal discriminant variate (PDV) method in the STR genotyping from the methodology [12]. The results showed that this method has advantages of good fitting, stability and strong prediction. Compared with the afore-mentioned methods, not only fluorescent dye markers but also purification of PCR products is not necessary. At the same time, it is simple, rapid and low-cost. In reference [12], only three genotypes of D16S539 locus with middle degree of difference were selected for methodology research, but in real usage each STR locus has more than three genotypes, which could not be discriminated by only one PDV model. Therefore, this paper would study the expert system of genotype discrimination for many genotypes of one STR locus in methodology based on the method in reference [12].The national database in US has recommended the 13 STR loci for paternity test, including D3S1358, THO1, D21S11, D18S51, VWA, D8S1179, TPOX, FGA, D5S818, D13S317, D7S820, D16S539 and CSF1PO. Since alleles of D5S818 locus have the smallest difference of only one core repeat unit (i.e. the alleles 7, 8, 9, 10, 11, 12, 13, 14, 15 and 16), the requirement for the classification ability of PDV models would be higher than others. Therefore, the genotypes 10–10, 10–11, 11–11, 11–12, 11–13 and 13–13 of the D5S818 locus with middle number of core repeat units and high frequency (i.e. the alleles 10 for 0.195, 11 for 0.33, 12 for 0.23 and 13 for 0.18) were selected as the research subjects.
Theory
PDV method is often used to establish the spectra-based chemical pattern recognition model for multivariate classification because it can effectively handle the collinear problem, which often appears in multi-variable spectra.Supposing there are N objects (n=1, 2, … , N) from K classes C, (k=1, 2, … , K) with the k-th class containing N objects. The object of the PDV method is to find a direction, called the principal discriminant variate , which maximizes the following principal discriminant criterion:where I is the identity (unitary) matrix of the same size as the covariance matrix T. , with a value varied between 0 and 1, is a weight controlling the balance between Fisher linear discriminant analysis (FLDA) and principal component analysis (PCA). That is, it can find a balance between the separability and the stability. By setting λ=1, the principal discrimination criterion (1) becomesand Eq. (2) is a criterion of FLDA. If λ=0, the principal discrimination criterion (1) becomeswhich turns out to be the variance criterion maximized in PCA.In Eq. (1), B is the between-class covariance matrix, which is defined asT is the total covariance matrix given byIn Eqs. (4), (5), and are the mean vectors of all objects and those from C, respectively, as given byAfter determining the first PDV 1, one can find successively other directions by Eq. (1) under the orthogonality constraint that T=0 (i>j). Therefore, the PDV method can find a sequence of discriminant variates until additional discriminant variates do not provide discriminatory information.
Experimental
Genomic DNA extraction
EDTA blood samples available for research purposes at the Department of Forensic Medicine of Soochow University were selected for this research. Genomic DNA samples were extracted according to the Chelex-100 method [12].
PCR amplification
PCR was performed in a final volume of 25 μL reaction mixture containing 1 μL of genomic DNA template, 0.25 mM of each primer (GenScript Inc., Nanjing in China. Primer A: 5′–GAA TGA TTT TCC TCT TTG GT-3′ and primer B: 5′–TGA TTC CAA TCA TAG CCA CA-3′ for D5S818), 0.625U of Taq DNA polymerase (Fermentas Inc., USA), 1×Taq buffer with 1.5 mM of MgCl2, 0.2 mM of each dNTP (GenScript Inc., Nanjing in China) and 16.625 mL of redistilled water.PCR was carried out using the PTC-200 Thermal cycler (BIO-RAD, USA) under the following conditions: held for 3 min at 95 °C, followed by 30 cycles of 30 s at 95 °C, 30 s at 58 °C and 30 s at 72 °C, and then held for 7 min at 72 °C.
Electrophoresis
Agarose gel electrophoresis was performed using the POWER BC6003En electrophoresis apparatus (Shanghai Shenergy Biocolor Bioscience and Technology Company, China). 3 μL of each PCR product was electrophoresed on the 2% agarose gel containing 0.5 μg/mL EtBr for 7 min at 200 V in 1×TAE buffer. Image of the gel was taken using GeneGenius BioImaging Systems (SynGene, England).
Near-infrared spectral analysis
20 μL of each PCR product was diluted with water to 450 μL as measuring sample, and then placed in the quartz cell (path length 10 mm, inside width 2 mm). With air as the background, near-infrared spectrum (NIRS) was measured using the Fourier transformation NEXUS infrared spectrometer (Thermo Electron Company, USA). Spectrometer parameters were as follows: the wave number range of 4000–9400 cm−1, the number of scans that are averaged 32 and the resolution of 8 cm−1.
Results and discussion
Optimization of PCR conditions
Major factors influence on the specificity and efficiency of PCR amplification containing concentrations of dNTP, Mg2+, primers and Taq DNA polymerase as well as annealing temperature. In this study, agarose gel electrophoresis was used to detect specificity and efficiency. We deemed that the PCR was non-efficient when there was no specific band for an amplified product, and then the product should be removed. Certainly, for an efficient amplification, the more was the brightness ratio of specific band to respective nonspecific one, the higher was the amplification specificity. Based on the specificity and efficiency, PCR conditions were optimized. Because of very low genomic DNA quantities available for the six studied samples, a whole genome amplification procedure was performed primarily to obtain enough DNA samples. These DNA samples were diluted to 1000 times with water, which then was used as the DNA templates of the second PCR amplification. The gel electrophoresis image of the second PCR products is shown in Fig. 1.
Figure 1
Agarose gel electrophoretograms of the samples of the six different genotypes of the locus D5S818 (Lanes 1, 2, 3, 4, 5 and 6: Genotypes 10–10, 10–11, 11–11, 11–12, 11–13 and 13–13, respectively).
Agarose gel electrophoretograms of the samples of the six different genotypes of the locus D5S818 (Lanes 1, 2, 3, 4, 5 and 6: Genotypes 10–10, 10–11, 11–11, 11–12, 11–13 and 13–13, respectively).
The preprocessing of spectral data
Using the method in Section 3.4, the NIRS-s of all the measuring samples were obtained. The spectra of all the six genotypes were very similar in shape (as shown in Fig. 2 for example the measuring samples of the genotype 10–10). Obviously, in the no signal range or poor one of the spectra in Fig. 2, it can be found that the NIRS-s of the measuring samples had the shift. In order to ensure that the spectral variation should be only related to genotype differences, the spectral drift needed to be eliminated with the help of the base-corrected method according to the information of no signal range. To eliminate the influence of concentration on the spectra, this paper used a normalization method for the base-corrected spectra.
Figure 2
Measured spectra for the 57 samples of the genotype 10–10.
Measured spectra for the 57 samples of the genotype 10–10.
Establishing PDV discriminant model
For genotypes 10–10, 10–11, 11–11, 11–12, 11–13 and 13–13, the spectral differences resulted from the differences of their genotypes were very small. So the parallel property of PCR amplifications should be very important. Three times of parallel PCR amplification were carried out for each genotype. The small differences of parallel PCR amplifications would be contained in all the measuring samples, which would ensure the robustness of discriminant models. For all the measuring samples of each genotype, two-thirds of sample were selected randomly as calibration sample set and the rest as prediction one (shown in Table 1).
Table 1
Data sets of six different genotypes of STR samples.
Genotype
Calibration set
Prediction set
Prediction accuracy (%)
10–10
38
19
100
10–11
36
18
100
11–11
38
19
100
11–12
37
19
100
11–13
36
19
100
13–13
37
19
100
Data sets of six different genotypes of STR samples.In Ref. [12], both PDV and support vector machine (SVM) methods were used to establish discriminant models. SVM can better classify a small number of calibration samples, and PDV can effectively handle the collinear problems that often appear in multi-variable spectra and spectra-characterized samples with only small difference in composition. In this research, the spectral differences between the six genotypes were very small, which could result in high degree of the collinearity. At the same time, considering the number of the measuring samples were enough large, we selected PDV method to establish discriminant models. For example, the two genotypes 10–10 and 10–11 with minimal difference were discriminated successfully using the optimum PDV model with the weigh λ=10−6 (Fig. 3).
Figure 3
Optimum PDV discriminant model between the genotypes 10–10 and 10–11. (“▴” the calibration sample and “△” the prediction sample for the 10–10 genotype; “•” the calibration sample and “○” the prediction sample for the 10–11 genotype).
Optimum PDV discriminant model between the genotypes 10–10 and 10–11. (“▴” the calibration sample and “△” the prediction sample for the 10–10 genotype; “•” the calibration sample and “○” the prediction sample for the 10–11 genotype).
The expert system of genotype discrimination
For genotyping of the six genotypes studied, two kinds of genotyping-tree structure were studied. One was constructed according to the difference between homozygote genotypes and heterozygote ones, which is shown in Fig. 4. The first layer of this genotyping-tree structure contained one class of homozygous genotypes 10–10, 11–11 and 13–13, and the other class of heterozygous ones 10–11, 11–12 and 11–13. The PDV discriminant model of these two classes was optimized by adjusting the weigh parameter λ. The optimum one is shown in Fig. 5, with the weigh λ=10−6. It could be found that the model did not have good discriminability, which indicated that genotypes could not be classified based on the difference between homozygote genotypes and heterozygote ones.
Figure 4
Built genotyping-tree structure of the six studied genotypes based on the difference of homozygote and heterozygote.
Figure 5
Optimum PDV discriminant model between the class of 10–10, 11–11 and 13–13 and the one of 10–11, 11–12 and 11–13. (“▴” the calibration sample and “△” the prediction sample for the genotypes 10–10; 11–11 and 13–13; “•” the calibration sample and “○” the prediction sample for the ones 10–11, 11–12 and 11–13).
Built genotyping-tree structure of the six studied genotypes based on the difference of homozygote and heterozygote.Optimum PDV discriminant model between the class of 10–10, 11–11 and 13–13 and the one of 10–11, 11–12 and 11–13. (“▴” the calibration sample and “△” the prediction sample for the genotypes 10–10; 11–11 and 13–13; “•” the calibration sample and “○” the prediction sample for the ones 10–11, 11–12 and 11–13).The other genotyping-tree structure is shown in Fig. 6, which was established with the different total number of core repeat units of the six genotypes. In the first layer of this genotyping-tree structure, one class had a range 20–23 of total number of the core repeat units and the other was 24–26. The PDV discriminant models for these two classes were established with the decrease of λ (not shown). When the weight λ=10−6 (Fig. 7), there were a large between-class distance and small within-class distances, and the predictive samples were in the range of the calibration ones. So this model was the best one. All other optimum PDV models (Fig. 3, Figure 8, Figure 9, Figure 10) in this tree-shape structure had the similar properties with Fig. 7, and the weight parameters λ of Figure 8, Figure 9, Figure 10 were 10−5, 10−6 and 10−5, respectively. Based on the Figure 3, Figure 9, Figure 10, the discriminant accuracies of the six genotypes were calculated. In the three models, although some predictive samples were out of the range of the calibration ones, it would not contribute to the prediction accuracy because of the good discriminant ability of the PDV1/PDV2. The prediction accuracies are shown in the last column of Table 1. In this genotyping-tree structure, six genotypes were discriminated successfully, which indicated that the established expert system of genotype discrimination based on the PDV models in this genotyping-tree could be used to genotyping of STR.
Figure 6
Built genotyping-tree structure of the six studied genotypes based on the different total numbers of the core repeat units of the genotypes.
Figure 7
Optimum PDV discriminant model between the class of 10–10, 10–11, 11–11, 11–12 and the one of 11–13, 13–13. (“▴” the calibration sample and “△” the prediction sample for the genotypes 10–10, 10–11, 11–11 and 11–12; “•” the calibration sample and “○” the prediction sample for the ones 11–13 and 13–13).
Figure 8
Optimum PDV discriminant model between the class of 10–10, 10–11 and the one of 11–11, 11–12. (“▴” the calibration sample and “△” the prediction sample for the genotypes 10–10 and 10–11; “•” the calibration sample and “○” the prediction sample for the ones 11–11 and 11–12).
Figure 9
Optimum PDV discriminant model between the genotypes 11–13 and 13–13. (“▴” the calibration sample and “△” the prediction sample for the 11–13 genotype; “•” the calibration sample and “○” the prediction sample for the 13–13 genotype).
Figure 10
Optimum PDV discriminant model between the genotypes 11–11 and 11–12. (“▴” the calibration sample and “△” the prediction sample for the 11–11 genotype; “•” the calibration sample and “○” the prediction sample for the 11–12 genotype).
Built genotyping-tree structure of the six studied genotypes based on the different total numbers of the core repeat units of the genotypes.Optimum PDV discriminant model between the class of 10–10, 10–11, 11–11, 11–12 and the one of 11–13, 13–13. (“▴” the calibration sample and “△” the prediction sample for the genotypes 10–10, 10–11, 11–11 and 11–12; “•” the calibration sample and “○” the prediction sample for the ones 11–13 and 13–13).Optimum PDV discriminant model between the class of 10–10, 10–11 and the one of 11–11, 11–12. (“▴” the calibration sample and “△” the prediction sample for the genotypes 10–10 and 10–11; “•” the calibration sample and “○” the prediction sample for the ones 11–11 and 11–12).Optimum PDV discriminant model between the genotypes 11–13 and 13–13. (“▴” the calibration sample and “△” the prediction sample for the 11–13 genotype; “•” the calibration sample and “○” the prediction sample for the 13–13 genotype).Optimum PDV discriminant model between the genotypes 11–11 and 11–12. (“▴” the calibration sample and “△” the prediction sample for the 11–11 genotype; “•” the calibration sample and “○” the prediction sample for the 11–12 genotype).
Conclusions
This paper studied the expert system of genotype discrimination for the six genotypes of STR locus D5S818 in methodology based on near-infrared spectroscopy–PDV method. The expert system of genotype discrimination in the genotyping-tree structure established with differences of total number of core repeat units had a good discriminability and robustness. Without any preprocessing for PCR products, the six genotypes of D5S818 locus could be classified rapidly and successfully, which provided a methodological support for establishing the expert system of genotype discrimination for other STR loci.
Authors: Nils Goedecke; Brian McKenna; Sameh El-Difrawy; Elizabeth Gismondi; Abigail Swenson; Loucinda Carey; Paul Matsudaira; Daniel J Ehrlich Journal: J Chromatogr A Date: 2005-06-22 Impact factor: 4.759
Authors: John M Butler; Richard Schoske; Peter M Vallone; Margaret C Kline; Alan J Redd; Michael F Hammer Journal: Forensic Sci Int Date: 2002-09-10 Impact factor: 2.395