| Literature DB >> 15647503 |
Bin Tian1, Jun Hu, Haibo Zhang, Carol S Lutz.
Abstract
mRNA polyadenylation is a critical cellular process in eukaryotes. It involves 3' end cleavage of nascent mRNAs and addition of theEntities:
Mesh:
Substances:
Year: 2005 PMID: 15647503 PMCID: PMC546146 DOI: 10.1093/nar/gki158
Source DB: PubMed Journal: Nucleic Acids Res ISSN: 0305-1048 Impact factor: 16.971
Figure 1(A) Schematic representation of a poly(A) site and polyadenylation configuration. In this study, a poly(A) site is a region containing cleavage site(s) (arrowed lines). The 5′-most cleavage site is the reference point (position 0) for the poly(A) site. Thus, the genomic location of a poly(A) site is represented by the location of the 5′-most cleavage site it contains. The sequence −300 to +300 is defined as a terminal sequence. The sites for CPSF and CstF are also depicted. (B) Three types of polyadenylation configuration. A type 1 gene has a single poly(A) site; a type II gene has alternative poly(A) sites all located in the 3′-most exon; and a type III gene has alternative poly(A) sites located in different exons. Types of poly(A) sites are also marked. 1S, a single poly(A) site; 2F, the 5′-most poly(A) site in a type II gene; 2L, the 3′-most poly(A) site in a type II gene; 2M, a middle poly(A) site between 2F and 2L in a type II gene; 3U, a poly(A) site located upstream of the 3′-most exon; and 3S, a single site in the 3′-most exon of a type III gene. Not shown in the graph are 3F, 3M and 3L, which are similar to 2F, 2M and 2L, respectively, except that the formers are located in the 3′-most exon of a type III gene. Exons are represented as boxes; pA, poly(A) site.
Poly(A) sites identified in human and mouse genes
| Homo sapiens | Mus musculus | |
|---|---|---|
| cDNA/EST used | 3 619 860 | 2 676 296 |
| cDNA/EST with poly(A) tail | 396 908 | 108 691 |
| Cleavage sites | 67 440 | 31 179 |
| Poly(A) sites | 29 283 | 16 282 |
| LocusLink entries | 13 942 | 11 155 |
| Poly(A) sites per gene | 2.10 | 1.46 |
aNumber of cDNA/EST sequences in the UniGene database.
bAfter sequence clean-up using approaches detailed in Materials and Methods. cDNA/EST sequences listed in the UniGene database were downloaded from GenBank and dbEST. Sequences were aligned to human and mouse genomes, and poly(A) cleavage sites were identified and clustered (see Materials and Methods for details). cDNA/ESTs were mapped to LocusLink IDs according to the UniGene database.
Figure 2Poly(A) sites of human genes. (A) Histogram of the genomic distance between adjacent poly(A) sites in a gene. (B) Histogram of the distance between adjacent poly(A) sites, both located in the 3′-most exon of a gene (median = 288 nt). (C) Histogram of the distance between the stop codon and its closest downstream poly(A) site (median = 324 nt). The x-axes in all graphs are in base-2 logarithmic scale. For each histogram, a Gaussian smoothing kernel method was used to generate a density line.
Top detected PAS hexamers
| Frequency (%) | Rank | |||||
|---|---|---|---|---|---|---|
| Hs | Mm | Hs.B | Hs | Mm | Hs.B | |
| AAUAAA | 53.18 | 59.16 | 58.2 | 1 | 1 | 1 |
| AUUAAA | 16.78 | 16.11 | 14.9 | 2 | 2 | 2 |
| UAUAAA | 4.37 | 3.79 | 3.2 | 3 | 3 | 3 |
| AGUAAA | 3.72 | 3.28 | 2.7 | 4 | 4 | 4 |
| AAGAAA | 2.99 | 2.15 | 1.1 | 5 | 5 | 10 |
| AAUAUA | 2.13 | 1.71 | 1.7 | 6 | 7 | 5 |
| AAUACA | 2.03 | 1.65 | 1.2 | 7 | 8 | 8 |
| CAUAAA | 1.92 | 1.80 | 1.3 | 8 | 6 | 6 |
| GAUAAA | 1.75 | 1.16 | 1.3 | 9 | 9 | 7 |
| AAUGAA | 1.56 | 0.90 | 0.8 | 10 | 11 | 11 |
| UUUAAA | 1.20 | 1.08 | 1.2 | 11 | 10 | 9 |
| ACUAAA | 0.93 | 0.64 | 0.6 | 12 | 12 | 13 |
| AAUAGA | 0.60 | 0.36 | 0.7 | 13 | 15 | 12 |
Human and mouse genomic sequences located −40 to −1 nt upstream of poly(A) sites were used to detect hexamers that may function as polyadenylation signals. Hs, human sequences; Mm, mouse sequences; and Hs.B, human results reported by Beaudoing et al. (16).
Figure 3Multiple cleavage sites in a poly(A) site. (A) Histogram of the genomic distance between adjacent cleavage sites in genes. (B) Histogram of the distance between the 5′-most cleavage site and other downstream cleavage sites when multiple cleavage sites are present in a poly(A) site (mean = 7.9 nt, median = 5 nt). (C) The relationship between the number of PAS hexamers (AAUAAA and other 11 variants) associated with a poly(A) site and the number of cleavage sites in the poly(A) site. Error bars are standard error of the mean (SEM). (D) Correlation between the number of cleavage sites and the number of supporting cDNA/EST sequences for poly(A) sites (Pearson correlation coefficient R = 0.83). (E) Histogram of the distance between a poly(A) site and the associated PAS when only one PAS is present. Only human poly(A) sites are used in (A–E).
Classification of genes according to the configuration of mRNA polyadenylation
| H.sapiens | M.musculus | |
|---|---|---|
| Number of genes | 13 942 | 11 155 |
| Type I genes | 6418 | 7576 |
| Type II genes | 4416 | 2681 |
| Type III genes | 3108 | 898 |
| Constitutive poly(A) sites | 46.03% | 67.92% |
| Alternative poly(A) sites | 53.97% | 32.08% |
Genes are classified according to the configuration of mRNA polyadenylation depicted in Figure 1B.
Figure 4Conservation of polyadenylation configuration between human and mouse orthologs. (A) Conservation of polyadenylation configuration between human (rows) and mouse (columns) orthologs (χ2-test, P-value = 2.0 × 10−132). Expected values, based on the null hypothesis that there is no correlation, are shown in parentheses. Observed values in (A) are plotted in (B), with the closed bars corresponding to conserved configurations, i.e. human type I versus mouse type I, etc.
Gene Ontology terms disproportionately associated with different types of polyadenylation configuration
| GO terms | Type | H.sapiens | M.musculus |
|---|---|---|---|
| Biological process | |||
| GO:0007166 (cell surface receptor linked site transduction) | I | 1.15E−04 (171, 37, 3) | 9.38E−06 (154, 30, 1) |
| GO:0046907 (intracellular transport) | II and III | 1.29E−07 (60, 59, 7) | 1.57E−08 (64, 68, 5) |
| GO:0015031 (protein transport) | 1.41E−07 (49, 53, 5) | 1.53E−07 (54, 59, 3) | |
| GO:0006886 (intracellular protein transport) | 7.25E−08 (46, 52, 5) | 5.95E−07 (51, 55, 3) | |
| Cellular component | |||
| GO:0005576 (extracellular) | I | 9.98E−05 (168, 32, 8) | 2.54E−10 (518, 115, 24) |
| GO:0005622 (intracellular) | II and III | 6.32E−08 (819, 337, 109) | 1.62E−04 (826, 335, 92) |
| GO:0005524 (nucleus) | III | 7.84E−05 (393, 162, 66) | 2.50E−04 (376, 143, 53) |
| Molecular function | |||
| GO:0004871 (site transducer activity) | I | 4.94E−06 (344, 77, 18) | 3.81E−06 (315, 71, 13) |
| GO:0008565 (protein transporter activity) | II and III | 4.52E−07 (30, 42, 1) | 3.41E−07 (25, 40, 0) |
| GO:0003723 (RNA binding) | III | 1.17E−05 (46, 32, 19) | 1.12E−04 (40, 22, 14) |
aFor each GO term, its GO ID and annotation (in parentheses) are given.
bType is the type of polyadenylation configuration (shown in Figure 1B).
cGO:0006886 (intracellular protein transport) is associated with both GO:0046907 (intracellular transport) and GO:0015031 (protein transport) through an ‘is a’ relationship (for details see Materials and Methods). Three categories of GO (Biological Process, Cellular Component and Molecular Function) were studied to find correlation with polyadenylation configuration. For each GO term in one species, a P-value from Fisher's exact test is provided, which indicates the significance of the association between this GO term and polyadenylation configuration, i.e. the lower the P-value the more significant the association is. Multiple testing adjustment using the Benjamanini and Hochberg method was applied to the selection of significant GO terms. The numbers of genes in three polyadenylation configurations, i.e. type I, II and III, are listed in parentheses. For example, (171, 37, 3) means that 171 type I genes, 37 type II genes and 3 type III genes.
Figure 5Characteristics of different types of poly(A) sites. (A) Association of various PAS hexamers with different types of poly(A) sites [for detailed definition of nine types of poly(A) sites see Figure 1B and Results]. (B) Cluster analysis of PAS hexamers and poly(A) types. The grayscale heat map represents the percentages of usage of PAS hexamers in different poly(A) types, with the sum of all values for each poly(A) type set to 100%. The shade of a cell indicates its value, with darker ones corresponding to higher values. Two-way hierarchical clustering was conducted using Euclidean distance as the metric. (C) Percentage of the number of supporting cDNA/EST sequences for different types of poly(A) sites. The total number of supporting cDNA/EST sequences for a gene is set to 100%. (D) Distribution of the number of cleavage sites per poly(A) site for different types of poly(A) site. Different shades are used to represent the number of cleavage sites per poly(A) site.
Figure 6Nucleotide composition of human terminal sequences. Human terminal sequences containing nine types of poly(A) sites are plotted. The poly(A) site type is marked in each graph, and the number of sequences used for each graph is shown in parentheses. The y-axis for each graph is the percentage of a nucleotide (%) and the x-axis is the genomic location (nt) relative to the poly(A) site. See Figure 1B and Results for detailed definitions of nine poly(A) site types.