Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Evolutionary rate variation in eukaryotic lineage specific human intronless proteins.

The present study examines 783 human-mouse orthologous gene pairs for their pattern of sequence evolution, contrasting mammalia, eukaryota, coelomata, and bilateria specific human intronless genes. Such comparisons may be of use in understanding the general evolution of human genome. Evolutionary rate analyses indicate that mammalia specific human intronless genes are evolving faster as compared to other intronless genes specific to eukaryotic lineage, indicating towards their rapid evolution. The observations indicates that the genes conserved in eukaryota, coelomata, and bilateria, that is, proteins that arose earlier in evolution as compared to mammalia specific genes evolve slowly and are subjected to negative selection. The cause underlying rate variations was also explored. Although mutational bias might slightly fasten the nonsynonymous rates in mammalia specific genes, it is unlikely to be major cause of rate difference between the various categories. Furthermore, rate of divergence of mammalia specific intronless genes has been related to functional classification using the protein family annotation. Protein function was found in some cases to have larger impact on the rate of evolution of genes. Also, the codon usage pattern of mammalia specific intronless genes do not seem to differ much from those of other intronless genes conserved solely in eukaryotic lineage.

Animals↗

Regulation of the nuclear genes encoding the cytoplasmic and mitochondrial leucyl-tRNA synthetases of Neurospora crassa.

We show that the nuclear genes for the cytoplasmic and mitochondrial leucyl-tRNA synthetase (LeuRS) of Neurospora crassa are distinct in their encoded proteins, codon usage, mRNA levels, and regulation. The 4.2-kilobase-pair region representing the structural gene for cytoplasmic LeuRS and flanking regions has been sequenced. The positions of the 5' and 3' ends of mRNA and of a single 62-base-pair intron have been mapped. The methionine-initiated open reading frame encoded a protein of 1,123 amino acids and displayed a strong codon bias. Although cytoplasmic LeuRS shares with mitochondrial LeuRS some general features common to most aminoacyl-tRNA synthetases, there is little amino acid sequence similarity between them, mRNA levels for cytoplasmic LeuRS were much higher than those for mitochondrial LeuRS. This observation and the strong codon bias in the cytoplasmic LeuRS gene may contribute to a greater abundance of cytoplasmic LeuRS than mitochondrial LeuRS. The genes for cytoplasmic and mitochondrial LeuRS are regulated independently. The cytoplasmic LeuRS gene is regulated by the cross-pathway control system in N. crassa, which is analogous to general amino acid control in Saccharomyces cerevisiae. The cytoplasmic LeuRS mRNA levels are induced by amino acid starvation resulting from the addition of aminotriazole. Part of this increase is due to utilization of new transcription start sites. In contrast, the mitochondrial LeuRS gene is not induced by amino acid limitation. However, the mitochondrial LeuRS mRNA levels did increase dramatically upon inhibition of mitochondrial protein synthesis by chloramphenicol or ethidium bromide or in the temperature-sensitive strain leu-5 carrying a mutation in the mitochondrial LeuRS structural gene.

Amino Acid Sequence↗

A simple program to calculate codon bias index.

A computer program (PCBI) was developed to quickly calculate codon bias index (CBI). PCBI can analyze a gene containing introns. The 22 preferred codons defined from Saccharomyces cerevisiae were used in PCBI as the standard to measure the CBI values. However, users can modify the preferred codons to suit each organism. The data PCBI provides include DNA sequence of open reading frame without introns, amino acid sequence of gene product, a table of amino acid composition, a table of codon usage and (G + C) content, parameters for calculating CBI, and the value of CBI. PCBI runs on a DOS or Windows environment, but results can be saved in ASCII text format.

Amino Acids↗

Drosophila melanogaster mitochondrial DNA: gene organization and evolutionary considerations.

The sequence of a 8351-nucleotide mitochondrial DNA (mtDNA) fragment has been obtained extending the knowledge of the Drosophila melanogaster mitochondrial genome to 90% of its coding region. The sequence encodes seven polypeptides, 12 tRNAs and the 3' end of the 16S rRNA and CO III genes. The gene organization is strictly conserved with respect to the Drosophila yakuba mitochondrial genome, and different from that found in mammals and Xenopus. The high A + T content of D. melanogaster mitochondrial DNA is reflected in a reiterative codon usage, with more than 90% of the codons ending in T or A, G + C rich codons being practically absent. The average level of homology between the D. melanogaster and D. yakuba sequences is very high (roughly 94%), although insertion and deletions have been detected in protein, tRNA and large ribosomal genes. The analysis of nucleotide changes reveals a similar frequency for transitions and transversions, and reflects a strong bias against G + C on both strands. The predominant type of transition is strand specific.

Amino Acid Sequence↗

Endosymbiotic origin and codon bias of the nuclear gene for chloroplast glyceraldehyde-3-phosphate dehydrogenase from maize.

The nuclei of plant cells harbor genes for two types of glyceraldehyde-3-phosphate dehydrogenases (GAPDH) displaying a sequence divergence corresponding to the prokaryote/eukaryote separation. This strongly supports the endosymbiotic theory of chloroplast evolution and in particular the gene transfer hypothesis suggesting that the gene for the chloroplast enzyme, initially located in the genome of the endosymbiotic chloroplast progenitor, was transferred during the course of evolution into the nuclear genome of the endosymbiotic host. Codon usage in the gene for chloroplast GAPDH of maize is radically different from that employed by present-day chloroplasts and from that of the cytosolic (glycolytic) enzyme from the same cell. This reveals the presence of subcellular selective pressures which appear to be involved in the optimization of gene expression in the economically important graminaceous monocots.

Amino Acid Sequence↗

Hydrophobicity, expressivity and aromaticity are the major trends of amino-acid usage in 999 Escherichia coli chromosome-encoded genes.

Multivariate analysis of the amino-acid compositions of 999 chromosome-encoded proteins from Escherichia coli showed that three main factors influence the variability of amino-acid composition. The first factor was correlated with the global hydrophobicity of proteins, and it discriminated integral membrane proteins from the others. The second factor was correlated with gene expressivity, showing a bias in highly expressed genes towards amino-acids having abundant major tRNAs. Just as highly expressed genes have reduced codon diversity in protein coding sequences, so do they have a reduced diversity of amino-acid choice. This showed that translational constraints are important enough to affect the global amino-acid composition of proteins. The third factor was correlated with the aromaticity of proteins, showing that aromatic amino-acid content is highly variable.

Amino Acids↗

Identification of the nuclear-encoded chloroplast ribosomal protein L12 of the monocotyledonous plant Secale cereale and sequencing of two different cDNAs with strong codon bias.

Two different cDNA clones (SCL12-1 and SCL12-2) encoding precursors of a chloroplast ribosomal protein with homology to L12 from Escherichia coli were isolated from rye leaf cDNA libraries and sequenced. The corresponding polypeptide of rye chloroplast ribosomes was identified. The sequences for the mature proteins of M(r) 13,447 and 13,609 share 85% amino acid identity. The mature polypeptide of clone SCL12-1 has an amino acid identity of 71%, 72% or 44%, respectively, relative to L12 proteins from spinach, tobacco, or E. coli. Codon usage of the rye L12 cDNAs shows a high preference (97% and 82%) for G or C in the third base position.

Amino Acid Sequence↗

Transcriptional and phylogenetic analysis of five complete ambystomatid salamander mitochondrial genomes.

We report on a study that extended mitochondrial transcript information from a recent EST project to obtain complete mitochondrial genome sequence for 5 tiger salamander complex species (Ambystoma mexicanum, A. t. tigrinum, A. andersoni, A. californiense, and A. dumerilii). We describe, for the first time, aspects of mitochondrial transcription in a representative amphibian, and then use complete mitochondrial sequence data to examine salamander phylogeny at both deep and shallow levels of evolutionary divergence. The available mitochondrial ESTs for A. mexicanum (N=2481) and A. t. tigrinum (N=1205) provided 92% and 87% coverage of the mitochondrial genome, respectively. Complete mitochondrial sequences for all species were rapidly obtained by using long distance PCR and DNA sequencing. A number of genome structural characteristics (base pair length, base composition, gene number, gene boundaries, codon usage) were highly similar among all species and to other distantly related salamanders. Overall, mitochondrial transcription in Ambystoma approximated the pattern observed in other vertebrates. We inferred from the mapping of ESTs onto mtDNA that transcription occurs from both heavy and light strand promoters and continues around the entire length of the mtDNA, followed by post-transcriptional processing. However, the observation of many short transcripts corresponding to rRNA genes indicates that transcription may often terminate prematurely to bias transcription of rRNA genes; indeed an rRNA transcription termination signal sequence was observed immediately following the 16S rRNA gene. Phylogenetic analyses of salamander family relationships consistently grouped Ambystomatidae in a clade containing Cryptobranchidae and Hynobiidae, to the exclusion of Salamandridae. This robust result suggests a novel alternative hypothesis because previous studies have consistently identified Ambystomatidae and Salamandridae as closely related taxa. Phylogenetic analyses of tiger salamander complex species also produced robustly supported trees. The D-loop, used in previous molecular phylogenetic studies of the complex, was found to contain a relatively low level of variation and we identified mitochondrial regions with higher rates of molecular evolution that are more useful in resolving relationships among species. Our results show the benefit of using complete genome mitochondrial information in studies of recently and rapidly diverged taxa.

Ambystomatidae↗

Comparative molecular evolution of primary (Buchnera) and secondary symbionts of aphids based on two protein-coding genes.

A+T content, phylogenetic relationships, codon usage, evolutionary rates, and ratio of synonymous versus non-synonymous substitutions have been studied in partial sequences of the atpD and aroQ/pheA genes of primary ( Buchnera) and secondary symbionts of aphids and a set of selected non-symbiotic bacteria, belonging to the five subdivisions of the Proteobacteria. Compared to the homologous genes of the last group, both genes belonging to Buchnera behave in a similar way, showing a higher A+T content, forming a monophyletic group, a loss in codon bias, especially in third base position, an evolutionary acceleration and an increase in the number of non-synonymous substitutions, confirming previous results reported elsewhere for other genes. When available, these properties have been partly observed with the secondary symbionts, but with values that are intermediate between Buchnera and free living Proteobacteria. They show high A+T content, but not as high as Buchnera, a non-solved phylogenetic position between Buchnera, and the other gamma-Proteobacteria, a loss in codon bias, again not as high as in Buchnera and a significant evolutionary acceleration in the case of the three atpD genes, but not when considering aroQ/pheA genes. These results give support to the hypothesis that they are symbionts at different stages of the symbiotic accommodation to the host.

AT Rich Sequence↗

Dietary arginine drives codon-dependent MHC class I translation and improves immunity in colon tumorigenesis and respiratory viral infection.

Amino acid levels fluctuate across diverse pathological conditions. Whether such amino acid modulations directly shape pathophysiology by regulating host gene expression remains unknown. We found that extracellular arginine restriction, observed in cancer and infection, represses specific arginine tRNAs-directly suppressing translation of major histocompatibility complex I (MHC class I) and antigen presentation. Arginine regulation of MHC class I was codon-usage dependent, as synonymous codon mutations prevented MHC class I modulation. Dietary arginine restriction impaired anti-viral immunity against influenza and SARS-CoV-2 and increased colon tumorigenesis. Conversely, increasing arginine availability via dietary supplementation or myeloid-specific arginase 1 deletion enhanced MHC class I protein levels, suppressed colon tumorigenesis, and improved viral infection outcomes. These disease-modulating effects were abolished in β2-microglobulin (B2m)-deficient mice. Thus, dietary modulation of a single amino acid critically influences codon-biased translation and MHC class I-mediated immunity to respiratory viral infections and cancer, revealing an unexpected mechanism and disease hazard for arginine deficiency and highlighting potential for amino acid-based translation modulation therapy.

Animals↗

Subtype-specific patterns in HIV Type 1 reverse transcriptase and protease in Oyo State, Nigeria: implications for drug resistance and host response.

As the use of antiretroviral therapy becomes more widespread across Africa, it is imperative to characterize baseline molecular variability and subtype-specific peculiarities of drug targets in non-subtype B HIV-1 infection. We sequenced and analyzed 35 reverse transcriptase (RT) and 43 protease (PR) sequences from 50 therapy-naive HIV-1-infected Nigerians. Phylogenetic analyses of RT revealed that the predominant viruses were CRF02_AG (57%), subtype G (26%), and CRF06_cpx (11%). Six of 35 (17%) individuals harbored primary mutations for RT inhibitors, including M41L, V118I, Y188H, P236L, and Y318F, and curiously three of the six were infected with CRF06_cpx. Therefore, CRF06_cpx drug-naive individuals had significantly more drug resistance mutations than the other subtypes (p = 0.011). By combining data on quasisynonymous codon bias with the influence of the differential genetic cost of mutations, we were able to predict some mutations, which are likely to predominate by subtype, under drug pressure. Some subtype-specific polymorphisms occurred within epitopes for HLA B7 and B35 in the RT, and HLA A2 and A*6802 in PR, at positions implicated in immune evasion. Balanced polymorphism was also observed at predicted serine-threonine phosphorylation sites in the RT of subtype G viruses. The subtype-specific codon usage and polymorphisms observed suggest the involvement of differential pathways for drug resistance and host-driven viral evolution in HIV-1 CRF02_AG, subtype G, and CRF06_cpx, compared to subtype B. Subtype-specific responses to HIV therapy may have significant consequences for efforts to provide effective therapy to the populations infected with these HIV-1 subtypes.

Amino Acid Sequence↗

Preferred amino acids and thermostability.

Most organisms grow at temperatures from 20 to 50 degrees C, but some prokaryotes, including Archaea and Bacteria, are capable of withstanding higher temperatures, from 60 to >100 degrees C. Their biomolecules, especially proteins, must be sufficiently stable to function under these extreme conditions; however, the basis for thermostability remains elusive. We investigated the preferential usage of certain groupings of amino acids and codons in thermally adapted organisms, by comparative proteome analysis, using 28 complete genomes from 18 mesophiles (M), 4 thermophiles (T), and 6 hyperthermophiles (HT). Whenever the percent of glutamate (E) and lysine (K) increased in the HT proteomes, the percent of glutamine (Q) and histidine (H) decreased, so that the E + K/Q + H ratio was >4.5; it was <2.5 in the M proteomes, and 3.2 to 4.6 in T. The E + K/Q + H ratios for chaperonins, potentially thermostable proteins, were higher than their proteome ratios, whereas for DNA ligases, which are not necessarily thermostable, they followed the proteome ratios. Analysis of codon usage revealed that HT had more AGR codons for Arg than they did CGN codons, which were more common in mesophiles. The E + K/Q + H ratio may provide a useful marker for distinguishing HT, T and M prokaryotes, and the high percentage of the amino acid couple E + K, consistently associated with a low percentage of the pair Q + H, could contribute to protein thermostability. The preponderance of AGR codons for Arg is a signature of all HT so far analyzed. The E + K/Q + H ratio and the codon bias for Arg are apparently not related to phylogeny. HT members of the Bacteria show the same values as the HT members of the Archaea; the values for T organisms are related to their lifestyle (intermediate temperature) and not to their domain (Archaea) and the values for M are similar in Eukarya, Bacteria and Archaea.

Adaptation, Biological↗

Codon usage and base composition in Rickettsia prowazekii.

Codon usage and base composition in sequences from the A + T-rich genome of Rickettsia prowazekii, a member of the alpha Proteobacteria, have been investigated. Synonymous codon usage patterns are roughly similar among genes, even though the data set includes genes expected to be expressed at very different levels, indicating that translational selection has been ineffective in this species. However, multivariate statistical analysis differentiates genes according to their G + C contents at the first two codon positions. To study this variation, we have compared the amino acid composition patterns of 21 R. prowazekii proteins with that of a homologous set of proteins from Escherichia coli. The analysis shows that individual genes have been affected by biased mutation rates to very different extents: genes encoding proteins highly conserved among other species being the least affected. Overall, protein coding and intergenic spacer regions have G + C content values of 32.5% and 21.4%, respectively. Extrapolation from these values suggests that R. prowazekii has around 800 genes and that 60-70% of the genome may be coding.

Base Composition↗

The VH gene repertoire of splenic B cells and somatic hypermutation in systemic lupus erythematosus.

In systemic lupus erythematosus (SLE) it has been hypothesized that self-reactive B cells arise from virgin B cells that express low-affinity, nonpathogenic germline V genes that are cross-reactive for self and microbial antigens, which convert to high-affinity autoantibodies via somatic hypermutation. The aim of the present study was to determine whether the VH family repertoire and pattern of somatic hypermutation in germinal centre (GC) B cells deviates from normal in SLE. Rearranged immunoglobulin VH genes were cloned and sequenced from GCs of a SLE patient's spleen. From these data the GC V gene repertoire and the pattern of somatic mutation during the proliferation of B-cell clones were determined. The results highlighted a bias in VH5 gene family usage, previously unreported in SLE, and under-representation of the VH1 family, which is expressed in 20-30% of IgM+ B cells of healthy adults and confirmed a defect in negative selection. This is the first study of the splenic GC response in human SLE.

B-Lymphocytes↗

Phylogenetic analyses of penicillia based on partial calmodulin gene sequences.

Partial sequences (about 600 nucleotides) of the calmodulin gene were used for the phylogenetic studies on Eupenicillium, Talaromyces and Penicillium. This region is from the 3rd base of the codon for the 9th amino acid Gln to the 3rd base of the codon for the 122th amino acid Val, flanking parts of the 2nd and 5th exons with complete sequences of two exons and three introns. Seventy-six isolates of 56 taxa of penicillia were involved. The nucleotide sequences with and without introns were analyzed respectively using the neighbor-joining (NJ) and maximum parsimony (MP) methods. The cluster analysis on relative synonymous codon usage (RSCU) of each sequence was also carried out. The fact that species of penicillia belong to the two subfamilies of the Trichocomaceae proposed by Malloch based on traditional methods is supported by our molecular data, whereas, the development of asci and patterns of penicilli show little phylogenetic information. Nine groups in the lineage of Eupenicillium and two in that of Talaromyces were recognized in our studies. In addition to the teleomorph-holomorph-anamorph evolutionary model of penicillia suggested by LoBuglio et al., and Pitt, we proposed that a mutation bias of holomorphs/anamorphs with or without selection is another evolutionary path of these organisms.

Base Sequence↗

Base usage and dinucleotide frequency of infectious bursal disease virus.

Base usage and dinucleotide frequency have been extensively studied in many eukaryotic organisms and bacteria, but not for viruses. In this paper, a comprehensive analysis of these aspects for infectious bursal disease virus (IBDV) was presented. The analysis of base usage indicated that all of the IBDV genes possess equivalent overall nucleotide distributions. However when the base usage at each codon positions was analysed by using cluster analysis, the VP5 open reading frame (ORF) formed a different cluster isolated from the other genes. The unusual base usage of VP5 ORF may indicate that the gene was originated by the virus "overprinting strategy", a strategy in which virus may create novel gene by utilizing the unused reading frames of its existing genes. Meanwhile, the GC content of the IBDV genes and the chicken's coding sequences was comparable; suggesting the virus imitation of the host to increase its translational efficiency. The analysis of dinucleotide frequency indicated that IBDV genome had dinucleotide bias: the frequencies of CpG and TpA were lower and the TpG was higher than the expected. Classical methylation pathway, a process where CpG converted to TpG, may explain the significant correlation between the CpG deficiency and TpG abundance. "Principal component analysis of the dinucleotide frequencies" (DF-PCA) was used to analyse the overall dinucleotide frequencies of IBDV genome. DF-PCA on the hypervariable region and polyprotein (VPX-VP4-VP3) gene showed that the very virulent IBDV (vvIBDV) was segregated from other strains; which meant vvIBDV had a unique dinucleotide pattern. In summary, the study of base usage and dinucleotide frequency had unravelled many overlooked genomic properties of the virus.

Animals↗

Diagrammatization of codon usage in 339 human immunodeficiency virus proteins and its biological implication.

The occurrence frequencies of bases A (adenine), C (cytosine, G (guanine), and T (thymine) occurring in the 1st, 2nd, and 3rd codon positions in the codon usage table of viral genes for the 339 human immunodeficiency virus (HIV) proteins compiled recently have been calculated and diagrammatized. For comparison, the corresponding diagrammatic representations for the 2681 human proteins from the codon usage table for primate genes are also presented. The analyzed results based on these characteristic diagrams indicate that considerably similar features have been found between HIV and human proteins for the 1st and 2nd codon positions; i.e., they are all occupied predominantly by purine, especially base A. However, a significant difference in the 3rd codon position between HIV and human proteins has been observed; i.e., human proteins are of high C + G content and low A + G content in the 3rd codon position, whereas the case is just the opposite for HIV proteins. The biological implication of such a duality on the codon bias of HIV against human proteins is discussed. It is suggested that the 1st and 2nd codon positions can be termed as the structure-determining position, and the 3rd codon position termed as the species-determining position. The diagrammatic representation and analysis method described here possess a great potential for the study of molecular evolution from the viewpoint of the genetic code for which data have been accumulated rapidly and will continue to grow at a much faster pace.

Base Composition↗

A codon-based model designed to describe lentiviral evolution.

A codon-based model designed to describe lentiviral evolution is developed. The model incorporates unequal base compositions in the three codon positions and selection against the CpG dinucleotide within codons to account for a deficit of this dinucleotide exhibited by lentiviral genes. The model is, to a large extent, able to account for the pattern of codon usage exhibited by the HIV1 genes gag, pol, and env, in spite of its parameter paucity. The model is extended to a similar model which operates on pentets (codons and their neighboring bases). The results obtained by the pentet model establish the importance of depression of CpGs across codon boundaries as well as within codons. The goodness of fit of the CpG depression model to the observed evolution in pairwise alignments of HIV1 sequences is assessed. The model provides a significantly better description of the observed evolution than the simpler models examined. The parameter estimates indicate that part of the unusually large biases in nucleotide frequencies observed in HIV1 genes is caused by selection against CpGs. We find that the estimates of expected numbers of substitutions, of transitions to transversions, and of synonymous to nonsynonymous substitution rates are robust to CpG depression, whereas the ratio of CpG-generating substitutions to other substitutions is strongly influenced by the choice of model.

Base Composition↗