Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Absence of classical heat shock response in the citrus pathogen Xylella fastidiosa.

The fastidious bacterium Xylella fastidiosa is associated with important crop diseases worldwide. We have recently shown that X. fastidiosa is a peculiar organism having unusually low values of gene codon bias throughout its genome and, unexpectedly, in the group of the most abundant proteins. Here, we hypothesized that the lack of codon usage optimization in X. fastidiosa would incapacitate this organism to undergo quick and massive changes in protein expression as occurs in a classical stress response. Proteomic analysis of the response to heat stress in X. fastidiosa revealed that no changes in protein expression can be detected. Moreover, stress-inducible proteins identified in the closely related citrus pathogen Xanthomonas axonopodis pv citri were found to be constitutively expressed in X. fastidiosa. These proteins have extremely high codon bias values in the X. citri and other well-studied organisms, but low values in X. fastidiosa. Because biased codon usage is well known to correlate to the rate of protein synthesis, we speculate that the peculiar codon bias distribution in X. fastidiosa is related to the absence of a classical stress response, and, probably, alternative strategies for survival of X. fastidiosa under stressfull conditions.

Bacterial Proteins↗

Organization and expression of algal (Chlamydomonas reinhardtii) mitochondrial DNA.

The mitochondrial genome of Chlamydomonas reinhardtii, a unicellular green alga, is a linear 15.8 kilobase pair (kbp) molecule. In gene arrangement and mode of expression, as well as in size, it differs radically from the large (200-2400 kbp) mitochondrial genomes of higher plants. Heterologous hybridization experiments and nucleotide sequence analysis have revealed that C. reinhardtii mitochondrial DNA (mtDNA) is a compactly organized genome specifying at least eight proteins, a minimum of three transfer RNAs, and large subunit (LS) and small subunit (SS) ribosomal RNAs. Both strands of the mtDNA encode genetic information, with genes organized into perhaps a single transcriptional unit on each strand. Stable transcripts have been identified by Northern hybridization analysis, and transcript termini have been mapped by primer extension and S1 nuclease protection experiments. The results suggest that mature RNAs, which virtually saturate the genome, are generated by precise endonucleolytic cleavage of long precursors, with specific motifs (both primary sequence and secondary structure) implicated as processing signals. Codon usage in C. reinhardtii mitochondria is highly biased, with eight codons entirely absent from all protein-coding genes; however, even though codon usage is restricted, it appears that C. reinhardtii mtDNA cannot encode the minimum number of tRNAs needed to support mitochondrial protein synthesis. The most striking feature of C. reinhardtii mtDNA is the division of SS and LS rRNA genes into a number of separate subgenic coding segments ('modules') that are interspersed with one another and with protein-coding and tRNA genes. We have identified abundant small RNAs, transcribed from these modules, that approximate to the latter in size. This indicates that splicing of rRNA 'pieces' does not occur in this system. Rather, the mature rRNAs apparently exist and function as non-covalent complexes of small RNAs (four in SS rRNA, at least eight in LS rRNA), held together by intermolecular base pairing. These complexes contain all the conserved elements of the minimal secondary structures that define the functional core of conventional LS and SS rRNAs.

Base Sequence↗

Analysis of nucleotide sequences of two ligninase cDNAs from a white-rot filamentous fungus, Phanerochaete chrysosporium.

An analysis of nucleotide sequences of two types of ligninase cDNAs isolated from the basidiomycete Phanerochaete chrysosporium, designated CLG4 and CLG5, are presented here. The amino acid sequences of the corresponding ligninase proteins, designated LG4 and LG5, respectively, have been deduced from the cDNA sequences. Mature ligninases LG4 and LG5 are preceded by leader sequences containing 28 and 27 amino acids (aa), respectively, and each contains 344 aa residues. The estimated Mrs of mature LG4 and LG5 are 36,540 and 36,607, respectively. Potential N-glycosylation site(s) with the general sequence Asn-X-Thr/Ser are found in both LG4 and LG5. Nucleotide sequence homology between the coding region of CLG4 and CLG5 is 71.5%, whereas the amino acid sequence homology between the two ligninases is 68.5%. The codon usage of ligninases is extremely biased in favor of codons rich in cytosine and guanine. Amino acid sequences of two tryptic peptides of ligninase H8 have exactly matching sequences in ligninase LG5. Also, the sequences of the oligodeoxynucleotide probes, which correspond to the sequences in the tryptic peptides of ligninase H8 and which were used in isolating the ligninase clones from the cDNA library, have exactly matching sequences in CLG5. The experimentally determined N-terminal sequence of purified ligninase H8 is found in the deduced N-terminal amino acid sequence of LG5. These results suggest that CLG5 encodes ligninase H8 and that CLG4 represents a related but different ligninase gene.

Amino Acid Sequence↗

Sequences of the E. coli uvrC gene and protein.

We have determined the sequence of a 2400 bp region of E. coli chromosomal DNA containing the uvrC gene. The coding region of uvrc is 2267 bp in length, encodes a polypeptide with a calculated molecular weight of 66,038 daltons, and is preceded by a typical E. coli ribosome binding site. By constructing deletion derivatives we have established that a uvrC promoter lies within the 113 bp region preceding the translational start of uvrC. The codon usage in uvrC is strongly biased in favor of codons used infrequently in E. coli, which may contribute to the relatively low intracellular concentration of uvrC protein.

Amino Acid Sequence↗

Comprehensive plastome variation and RNA editing in Mentha: insights into phylogenetic relationships and candidate DNA barcodes.

INTRODUCTION: Mentha is an economically and medicinally important genus in Lamiaceae, but its taxonomy and species delimitation remain challenging because of frequent hybridization, polyploidy, and marked morphological plasticity. METHODS: In this study, we comparatively analyzed 12 plastomes representing major Mentha species, hybrid taxa, and unresolved accessions, including four newly assembled genomes, to characterize plastome structure, repeat composition, sequence divergence, phylogenetic relationships, and plastid RNA editing. The M. arvensis plastome and RNA-seq datasets originated from independent Swiss and Indian accessions, respectively. RESULTS: The plastomes were highly conserved in overall organization, ranging from 151,824 to 152,154 bp and displaying the typical quadripartite structure. Gene content and order were largely stable across taxa, with only minor variation likely associated with annotation differences at IR/SC boundary regions. Codon usage analysis revealed a clear bias toward A/U-ending synonymous codons, and most shared protein-coding genes showed low Ka/Ks ratios, indicating predominant purifying selection. Repeat analyses showed that simple sequence repeats were mainly composed of A/T-rich mononucleotide motifs, whereas long repeats were concentrated in the 30-40 bp size class. Comparative analyses identified six hypervariable regions, namely ccsA-ndhD, ycf1, ndhD, rpl32-trnL-UAG, rbcL-accD, and petA-psbJ, which represent promising candidate plastid markers for species discrimination. Phylogenetic analysis based on complete plastomes provided strong support for relationships among the sampled taxa and recovered a close affinity among M. aquatica, M. arvensis, and M. canadensis. In addition, RNA-seq analysis of M. arvensis identified 17 candidate plastid RNA editing sites, most of which were C-to-U conversions and nonsynonymous events. DISCUSSION: Together, these results expand plastid genomic resources for Mentha and provide a useful framework for phylogenetic inference, species identification, and future germplasm utilization.

RNA editing↗

Gene expression level shapes the amino acid usages in Prochlorococcus marinus MED4.

Prochlorococcus species are the first example of free-living bacteria with reduced genome. Codon and amino acid usages bias of Prochlorococcus marinus MED4 was investigated using all protein coding genes having length greater than or equal to 100 amino acids. Correspondence analysis on relative synonymous codon usage (RSCU) values shows that there is no such influence of translational selection in shaping the codon usage variation among the genes in this organism. However, amino acid usages were markedly different between the highly and lowly expressed genes in this organism and in particular, GC rich amino acids were found to occur significantly higher in highly expressed genes than the lowly expressed genes. Comparative analysis of the homologous genes of Synechococcus sp. WH8102 and Prochlorococcus marinus MED4 shows that amino acids conservation in highly expressed genes is significantly higher than lowly expressed genes. Based on our results we concluded that conservation of GC rich amino acids in the highly expressed genes to its ancestor is the major source of variation in amino acid usages in the organism.

Bacterial Proteins↗

Evolutionary lability of context-dependent codon bias in bacteria.

In bacteria, synonymous codon usage can be considerably affected by base composition at neighboring sites. Such context-dependent biases may be caused by either selection against specific nucleotide motifs or context-dependent mutation biases. Here we consider the evolutionary conservation of context-dependent codon bias across 11 completely sequenced bacterial genomes. In particular, we focus on two contextual biases previously identified in Escherichia coli; the avoidance of out-of-frame stop codons and AGG motifs. By identifying homologues of E. coli genes, we also investigate the effect of gene expression level in Haemophilus influenzae and Mycoplasma genitalium. We find that while context-dependent codon biases are widespread in bacteria, few are conserved across all species considered. Avoidance of out-of-frame stop codons does not apply to all stop codons or amino acids in E. coli, does not hold for different species, does not increase with gene expression level, and is not relaxed in Mycoplasma spp., in which the canonical stop codon, TGA, is recognized as tryptophan. Avoidance of AGG motifs shows some evolutionary conservation and increases with gene expression level in E. coli, suggestive of the action of selection, but the cause of the bias differs between species. These results demonstrate that strong context-dependent forces, both selective and mutational, operate on synonymous codon usage but that these differ considerably between genomes.

Codon↗

Codon usage patterns in Escherichia coli, Bacillus subtilis, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Drosophila melanogaster and Homo sapiens; a review of the considerable within-species diversity.

The genetic code is degenerate, but alternative synonymous codons are generally not used with equal frequency. Since the pioneering work of Grantham's group it has been apparent that genes from one species often share similarities in codon frequency; under the "genome hypothesis" there is a species-specific pattern to codon usage. However, it has become clear that in most species there are also considerable differences among genes. Multivariate analyses have revealed that in each species so far examined there is a single major trend in codon usage among genes, usually from highly biased to more nearly even usage of synonymous codons. Thus, to represent the codon usage pattern of an organism it is not sufficient to sum over all genes as this conceals the underlying heterogeneity. Rather, it is necessary to describe the trend among genes seen in that species. We illustrate these trends for six species where codon usage has been examined in detail, by presenting the pooled codon usage for the 10% of genes at either end of the major trend. Closely-related organisms have similar patterns of codon usage, and so the six species in Table 1 are representative of wider groups. For example, with respect to codon usage, Salmonella typhimurium closely resembles E. coli, while all mammalian species so far examined (principally mouse, rat and cow) largely resemble humans.

Amino Acids↗

Molecular characterization of the glyceraldehyde-3-phosphate dehydrogenase gene of Phaffia rhodozyma.

The glyceraldehyde-3-phosphate dehydrogenase (GPD; EC1.2.1.12)-encoding gene (gpd) was isolated from a genomic library of Phaffia rhodozyma CBS 6938. Unlike some other eukaryotic organisms the gpd gene is represented by a single copy in P. rhodozyma. The complete nucleotide sequence of the coding, as well as the flanking non-coding regions was determined. The nucleotide sequence of gpd predicted six introns and a polypeptide chain of 339 amino acids. The codon usage in the gpd gene of P. rhodozyma was highly biased and was significantly different from the codon usage in other yeasts. Phylogenetic analysis of different yeasts and filamentous asco- and basidiomycetes gpd sequences indicated that the gpd gene of P. rhodozyma forms a cluster with the corresponding genes of filamentous basidiomycetes.

Amino Acid Sequence↗

Combinatorial codons: a computer program to approximate amino acid probabilities with biased nucleotide usage.

Using techniques from optimization theory, we have developed a computer program that approximates a desired probability distribution for amino acids by imposing a probability distribution on the four nucleotides in each of the three codon positions. These base probabilities allow for the generation of biased codons for use in mutational studies and in the design of biologically encoded libraries. The dependencies between codons in the genetic code often makes the exact generation of the desired probability distribution for amino acids impossible. Compromises are often necessary. The program, therefore, not only solves for the "optimal" approximation to the desired distribution (where the definition of "optimal" is influenced by several types of parameters entered by the user), but also solves for a number of "sub-optimal" solutions that are classified into families of similar solutions. A representative of each family is presented to the program user, who can then choose the type of approximation that is best for the intended application. The Combinatorial Codons program is available for use over the web from http://www.wi.mit.edu/kim/computing.html.

Amino Acids↗

Intraspecific DNA variation in nuclear genes of the mosquito Aedes aegypti.

Single nucleotide polymorphisms (SNPs) are an abundant source of genetic variation among individual organisms. To assess the usefulness of SNPs for genome analysis in the yellow fever mosquito, Aedes aegypti, we sequenced 25 nuclear genes in each of three strains and analysed nucleotide diversity. The average frequency of nucleotide variation was 12 SNPs per kilobase, indicating that nucleotide variation in Ae. aegypti is similar to that in other organisms, including Drosophila and the malaria vector Anopheles gambiae. Transition polymorphisms outnumbered transversion polymorphisms, at a ratio of about 2:1. We examined codon usage and confirmed that mutational bias favours G and C ending codons. Codon bias was most pronounced in highly expressed genes. Nucleotide diversity estimates indicated that substitution rates are positively correlated in coding and non-coding regions. Nucleotide diversity varied from one gene to another. The unequal distribution of SNPs among Ae. aegypti nuclear genes suggests that single base variations are non-neutral and are subject to selective constraints. Our analysis showed that ubiquitously expressed genes have lower polymorphism rates and are likely under strong purifying selection, whereas tissue specific genes and genes with a putative role in parasite defence exhibit higher levels of polymorphism that may be associated with diversifying selection.

Aedes↗

A comparison of homologous developmental genes from Drosophila and Tribolium reveals major differences in length and trinucleotide repeat content.

The flour beetle Tribolium castaneum has become an important model organism for comparative studies of insect development. Many developmentally important genes have now been cloned from both Tribolium and Drosophila and their expression characteristics were studied. We analyze here the complete coding sequences of 17 homologous gene pairs from D. melanogaster and T. castaneum, most of which encode transcription factors. We find that the Tribolium genes are on average 30% shorter than their Drosophila homologues. This appears to be due largely to the almost-complete absence of trinucleotide repeats in the coding sequences of Tribolium as well as the generally lower degree of internal repetitiveness. Clusters of polar and other amino acids such as glutamine, proline, and serine, which are often considered to be important for transcriptional activation domains in Drosophila, are almost completely absent in Tribolium. Codon usage is generally less biased in Tribolium, although we find a similar tendency for the preference of G- or C-ending codons and a higher bias in conserved subregions of the proteins as in Drosophila. Most of the aminoacid substitutions in the DNA-binding domains of the transcription factors occur at residues that do not make a specific contact to DNA, suggesting that the recognition sequences are likely to be conserved between the two species.

Amino Acid Sequence↗

Codon usage in Chlamydia trachomatis is the result of strand-specific mutational biases and a complex pattern of selective forces.

The patterns of synonymous codon choices of the completely sequenced genome of the bacterium Chlamydia trachomatis were analysed. We found that the most important source of variation among the genes results from whether the sequence is located on the leading or lagging strand of replication, resulting in an over representation of G or C, respectively. This can be explained by different mutational biases associated to the different enzymes that replicate each strand. Next we found that most highly expressed sequences are located on the leading strand of replication. From this result, replicational-transcriptional selection can be invoked. Then, when the genes located on the leading strand are studied separately, the correspondence analysis detects a principal trend which discriminates between lowly and highly expressed sequences, the latter displaying a different codon usage pattern than the former, suggesting selection for translation, which is reinforced by the fact that Ks values between orthologous sequences from C. trachomatis and Chlamydia pneumoniae are much smaller in highly expressed genes. Finally, synonymous codon choices appear to be influenced by the hydropathy of each encoded protein and by the degree of amino acid conservation. Therefore, synonymous codon usage in C.trachomatis seems to be the result of a very complex balance among different factors, which rises the problem of whether the forces driving codon usage patterns among microorganisms are rather more complex than generally accepted.

Amino Acids↗

Gene expression level influences amino acid usage, but not codon usage, in the tsetse fly endosymbiont Wigglesworthia.

Wigglesworthia glossinidia brevipalpis, the obligate bacterial endosymbiont of the tsetse fly Glossina brevipalpis, is characterized by extreme genome reduction and AT nucleotide composition bias. Here, multivariate statistical analyses are used to test the hypothesis that mutational bias and genetic drift shape synonymous codon usage and amino acid usage of Wigglesworthia. The results show that synonymous codon usage patterns vary little across the genome and do not distinguish genes of putative high and low expression levels, thus indicating a lack of translational selection. Extreme AT composition bias across the genome also drives relative amino acid usage, but predicted high-expression genes (ribosomal proteins and chaperonins) use GC-rich amino acids more frequently than do low-expression genes. The levels and configuration of amino acid differences between Wigglesworthia and Escherichia coli were compared to test the hypothesis that the relatively GC-rich amino acid profiles of high-expression genes reflect greater amino acid conservation at these loci. This hypothesis is supported by reduced levels of protein divergence at predicted high-expression Wigglesworthia genes and similar configurations of amino acid changes across expression categories. Combined, the results suggest that codon and amino acid usage in the Wigglesworthia genome reflect a strong AT mutational bias and elevated levels of genetic drift, consistent with expected effects of an endosymbiotic lifestyle and repeated population bottlenecks. However, these impacts of mutation and drift are apparently attenuated by selection on amino acid composition at high-expression genes.

Amino Acids↗

Cloning and characterization of the white gene from Anopheles gambiae.

A 14 kb region of genomic DNA containing the X-linked Anopheles gambiae eye colour gene, white, was cloned and sequenced. Genomic clones containing distinct white+ alleles were polymorphic for the insertion of a small transposable element in intron 3, and differed at 1% of nucleotide positions compared. Sequence was also determined from a rare 2914 bp cDNA. Comparison of cDNA and genomic sequences established an intron-exon structure distinct from Drosophila white. Despite a common trend in Anopheles and Drosophila of weak codon bias given low levels of gene expression, codon usage by Anopheles gambiae white was strongly biased. Overall amino acid identity between the predicted mosquito and fruitfly proteins was 64%, but dropped to 14% at the amino terminus. To correlate phenotypically white-eyed strains of A. gambiae with structural lesions in white, five available strains were analysed by PCR and Southern blotting. Although these strains carried allelic mutations, independently generated by gamma radiation (three strains) or spontaneous events (two strains), no white lesions were detected. Significantly, another non-allelic X-linked mutation, causing an identical white-eyed phenotype, has been correlated with a structural defect in the cloned white gene (Benedict et al., 1995). Taken together, these observations suggest that the white-eyed mutants analysed in the present study carry mutations in a second eye colour gene and are most likely white+.

ATP-Binding Cassette Transporters↗

INCA: synonymous codon usage analysis and clustering by means of self-organizing map.

UNLABELLED: INteractive Codon usage Analysis (INCA) provides an array of features useful in analysis of synonymous codon usage in whole genomes. In addition to computing codon frequencies and several usage indices, such as 'codon bias', effective Nc and CAI, the primary strength of INCA has numerous options for the interactive graphical display of calculated values, thus allowing visual detection of various trends in codon usage. Finally, INCA includes a specific unsupervised neural network algorithm, the self-organizing map, used for gene clustering according to the preferred utilization of codons. AVAILABILITY: INCA is available for the Win32 platform and is free of charge for academic use. For details, visit the web page http://www.bioinfo-hr.org/inca or contact the author directly. SUPPLEMENTARY INFORMATION: Software is accompanied with a user manual and a short tutorial.

Algorithms↗

What drives codon choices in human genes?

Synonymous codon usage is based and the bias seems to be different in different organisms. Factors with proposed roles in causing codon bias include degree and timing of gene expression, codon-anticodon interactions, transcription and translation rate and fidelity, codon context, and global and local G + C content. We offer a new perspective and new methods for elucidating codon choices applied especially to the human genome. We present data supporting the thesis that codon choices for human genes are largely a consequence of two factors: (1) amino acid constraints, (2) maintaining DNA structures dependent on base-step conformational tendencies consistent with the organism's genome signature that is determined by genome-wide processes of DNA modification, replication and repair. The related codon signature defined as the dinucleotide relative abundances at the distinct codon positions (1,2), (2,3), and (3,4) (4 = 1 of the next codon) accommodates both the global genome signature and amino acid constraints. In human genes, codon positions (2,3) and (3,4) containing the silent site have similar codon signatures reflecting DNA symmetry. Strong CG and TA dinucleotide underrepresentation is observed at all codon positions as well as in non-coding regions. Estimates of synonymous codon usage based on codon signatures are in excellent agreement with the actual codon usage in human and general vertebrate genes. These properties are largely independent of the isochore compartment (G + C content), gene size, and transcriptional and translational constraints. We hypothesize that major influences on codon usage in human genes result from residue preferences and diresidue associations in proteins coupled to biases on the DNA level, related to replication and repair processes and/or DNA structural requirements.

Codon↗

Codon usage in plant genes.

We have examined codon bias in 207 plant gene sequences collected from Genbank and the literature. When this sample was further divided into 53 monocot and 154 dicot genes, the pattern of relative use of synonymous codons was shown to differ between these taxonomic groups, primarily in the use of G + C in the degenerate third base. Maize and soybean codon bias were examined separately and followed the monocot and dicot codon usage patterns respectively. Codon preference in ribulose 1,5 bisphosphate and chlorophyll a/b binding protein, two of the most abundant proteins in leaves was investigated. These highly expressed are more restricted in their codon usage than plant genes in general.

Amino Acid Sequence↗