| Literature DB >> 31800035 |
Jérôme Bourret1, Samuel Alizon1, Ignacio G Bravo1.
Abstract
Codon Usage Preferences (CUPrefs) describe the unequal usage of synonymous codons at the gene, chromosome, or genome levels. Numerous indices have been developed to evaluate CUPrefs, either in absolute terms or with respect to a reference. We introduce the normalized index COUSIN (for COdon Usage Similarity INdex), that compares the CUPrefs of a query against those of a reference and normalizes the output over a Null Hypothesis of random codon usage. The added value of COUSIN is to be easily interpreted, both quantitatively and qualitatively. An eponymous software written in Python3 is available for local or online use (http://cousin.ird.fr). This software allows for an easy and complete analysis of CUPrefs via COUSIN, includes seven other indices, and provides additional features such as statistical analyses, clustering, and CUPrefs optimization for gene expression. We illustrate the flexibility of COUSIN and highlight its advantages by analyzing the complete coding sequences of eight divergent genomes. Strikingly, COUSIN captures a bimodal distribution in the CUPrefs of human and chicken genes hitherto unreported with such precision. COUSIN opens new perspectives to uncover CUPrefs specificities in genomes in a practical, informative, and user-friendly way.Entities:
Keywords: amino acid composition; bioinformatics; codon adaptation index; codon usage bias; mutational bias; mutation–selection; nucleotide composition; translational selection
Mesh:
Year: 2019 PMID: 31800035 PMCID: PMC6934141 DOI: 10.1093/gbe/evz262
Source DB: PubMed Journal: Genome Biol Evol ISSN: 1759-6653 Impact factor: 3.416
Notations Used to Define COUSIN and CAI Indices
| Symbol | Description |
|---|---|
|
| Codon |
|
| Amino acid |
|
| Frequency |
| ref | Reference |
| que | Query |
|
| Null hypothesis |
|
| Query length |
|
| Set of synonymous codons coding for amino acid |
|
| Amino acids present in both query and reference |
|
| Number of amino acids present in both query and reference |
. 1.—COUSIN (blue curve) and CAI (red curve) scores (y-axis) for a set of hypothetical queries with different frequency for the AAC and AAU codons encoding the asparagine amino acid (x-axis). Values are calculated for a reference set using (A) a strong usage bias of AAC:AAU 80:20 and (B) a slight usage bias of AAC:AAU 60:40. Vertical dashed lines indicate the composition for the Null Hypothesis of equal usage of both codons (gray line) and for the corresponding reference (black line). Horizontal dashed lines show the COUSIN key values that correspond to the Null Hypothesis (gray line) and to the reference (black line). The yellow area indicates queries with CUPrefs opposite to those in the reference, white one queries with similar but weaker CUPrefs than the reference and pink one queries with similar and stronger CUPrefs than the reference. Notice that, by design, the COUSIN values are always 0 and 1 respectively for the H0 and for the reference, independently of the CUPrefs in the reference. By definition, CAI is bounded by 0 and 1. In this example, COUSIN scores below –3 and above 4 are omitted to facilitate results visualization and reading.
. 2.—Density curves for (A) and CAI (B) indices for the complete CDSs of the eight organisms studied (see color legend). For each CDS the values for and CAI were calculated against the average codon usage reference table of the corresponding genome. The normalization renders curves centered around 1 allowing for rapid identification of differential dispersion in the leptokurtic curves for organisms with strong nucleotide compositional biases (e.g., Streptomyces coelicolor, in green) compared with those more platykurtic for organisms with weaker compositional biases (e.g., Escherichia coli in blue). Notice the bimodal distributions for Homo sapiens (black) and Gallus gallus (light blue) in panel A.