Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Occurrence of unmodified adenine and uracil at the first position of anticodon in threonine tRNAs in Mycoplasma capricolum.

Codon usage pattern in the threonine four-codon (ACN) box in Mycoplasma capricolum is strongly biased towards adenine and uracil for the third base of codons. Codons ending in uracil or adenine, especially ACU, predominate over ACC and ACG. This bacterium contains two isoacceptor threonine tRNAs having anticodon sequences AGU and UGU, both with unmodified first nucleotides. It would thus appear that ACN codons are translated in an unusual way; tRNA(Thr)(AGU) would translate the most abundantly used codon ACU exclusively, because adenine at the first anticodon position can, according to the wobble rule, pair only with uracil of the third codon position. The tRNA(Thr)(UGU) would mainly be responsible for translation of three other codons, ACA, ACG, and ACC. Anticodon UGU would also be used for reading codon ACU as a redundancy of tRNA(Thr)-(AGU), as deduced from the mitochondrial code where unmodified uracil at the first anticodon position can pair with adenine, cytosine, guanine, and uracil by four-way wobble. The tRNA(Thr)(AGU) has much higher sequence homology to tRNA(Thr)(UGU) from M. capricolum (88%), Bacillus subtilis (77%) and Escherichia coli (86%) than to tRNA(Thr)(GGU) from B. subtilis (66%) and E. coli (63%), suggesting that tRNA(Thr)-(AGU) has been derived from tRNA(Thr)(UGU), but not from tRNA(Thr)(GGU).

Adenine↗

Codon usage in the Mycobacterium tuberculosis complex.

The usage of alternative synonymous codons in Mycobacterium tuberculosis (and M. bovis) genes has been investigated. This species is a member of the high-G+C Gram-positive bacteria, with a genomic G+C content around 65 mol%. This G+C-richness is reflected in a strong bias towards C- and G-ending codons for every amino acid: overall, the G+C content at the third positions of codons is 83%. However, there is significant variation in codon usage patterns among genes, which appears to be associated with gene expression level. From the variation among genes, putative optimal codons were identified for 15 amino acids. The degree of bias towards optimal codons in an M. tuberculosis gene is correlated with that in homologues from Escherichia coli and Bacillus subtilis. The set of selectively favoured codons seems to be quite highly conserved between M. tuberculosis and another high-G+C Gram-positive bacterium, Corynebacterium glutamicum, even though the genome and overall codon usage of the latter are much less G+C-rich.

Bacillus subtilis↗

Codon usage limitation in the expression of HIV-1 envelope glycoprotein.

BACKGROUND: The expression of both the env and gag gene products of human immunodeficiency virus type 1 (HIV-1) is known to be limited by cis elements in the viral RNA that impede egress from the nucleus and reduce the efficiency of translation. Identifying these elements has proven difficult, as they appear to be disseminated throughout the viral genome. RESULTS: Here, we report that selective codon usage appears to account for a substantial fraction of the inefficiency of viral protein synthesis, independent of any effect on improved nuclear export. The codon usage effect is not specific to transcripts of HIV-1 origin. Re-engineering the coding sequence of a model protein (Thy-1) with the most prevalent HIV-1 codons significantly impairs Thy-1 expression, whereas altering the coding sequence of the jellyfish green fluorescent protein gene to conform to the favored codons of highly expressed human proteins results in a substantial increase in expression efficiency. CONCLUSIONS: Codon-usage effects are a major impediment to the efficient expression of HIV-1 genes. Although mammalian genes do not show as profound a bias as do Escherichia coli genes, other proteins that are poorly expressed in mammalian cells can benefit from codon re-engineering.

Animals↗

Mapping and sequencing of the dihydrofolate reductase gene (DFR1) of Saccharomyces cerevisiae.

The dihydrofolate reductase gene (DFR1) from Saccharomyces cerevisiae has been mapped and sequenced. The gene was isolated on an 8.8-kb BamHI fragment from a yeast genomic library by screening of Escherichia coli transformants for resistance to trimethoprim. A 1.8-kb SalI-BamHI fragment which was able to confer methotrexate resistance in yeast also complemented an E. coli DHFR-deficient (folA) mutant. Nucleotide sequence analysis revealed that the yeast DFR1 gene encoded a polypeptide with a predicted Mr of 24230. The deduced sequence of 211 amino acid residues showed considerable homology with DHFRs from both bacterial and animal sources. The codon bias index of the DFR1 coding region is 0.0083, which indicates a random pattern of codon usage. The upstream region contains two consensus sequences required for binding of the yeast's positive regulatory factor, GCN4, suggesting that the DFR1 gene might be subject to the amino acid general control. Several potential 'TATA' boxes are located in the sequence 5' to the gene. Located in the 3' flanking region are homologies with several canonical sequences thought to be required for efficient transcription termination in yeast. We also mapped the DFR1 gene to a position 1.4 cM proximal to the MET7 locus on chromosome XV.

Amino Acid Sequence↗

Translational effects of differential codon usage among intragenic domains of new genes in Drosophila.

Evolved codon usages often pose a technical challenge over the expressing of eukaryotic genes in microbial systems because of changed translational machinery. In the present study, we investigated the translational effects of intragenic differential codon usage on the expression of the new Drosophila gene, jingwei (jgw), a chimera derived from two unrelated parental genes: Ymp and Adh. We found that jgw possesses a strong intragenic differential usage of synonymous codons, i.e. the Adh-derived C-domain has a significantly higher codon bias than that of the Ymp-derived N-domain (P=0.0023 by t-test). Additional evolutionary analysis revealed the heterogeneous distribution of rare codons, implicating its role in gene regulation and protein translation. The in vitro expression of jgw further demonstrated that the heterogeneous distribution of rare codons has played a role in regulating gene expression, particularly, affecting the quality of protein translation.

Alcohol Dehydrogenase↗

Codon adaptation and synonymous substitution rate in diatom plastid genes.

Diatom plastid genes are examined with respect to codon adaptation and rates of silent substitution (Ks). It is shown that diatom genes follow the same pattern of codon usage as other plastid genes studied previously. Highly expressed diatom genes display codon adaptation, or a bias toward specific major codons, and these major codons are the same as those in red algae, green algae, and land plants. It is also found that there is a strong correlation between Ks and variation in codon adaptation across diatom genes, providing the first evidence for such a relationship in the algae. It is argued that this finding supports the notion that the correlation arises from selective constraints, not from variation in mutation rate among genes. Finally, the diatom genes are examined with respect to variation in Ks among different synonymous groups. Diatom genes with strong codon adaptation do not show the same variation in synonymous substitution rate among codon groups as the flowering plant psbA gene which, previous studies have shown, has strong codon adaptation but unusually high rates of silent change in certain synonymous groups. The lack of a similar finding in diatoms supports the suggestion that the feature is unique to the flowering plant psbA due to recent relaxations in selective pressure in that lineage.

Adaptation, Physiological↗

Intragenic codon bias in a set of mouse and human genes.

To better conceptualize the mechanism underlying the evolution of synonymous codons, we have analysed intragenic codon usage in chosen "regions" of some mouse and human genes. We divided a given gene into two regions: one consisting of a trinucleotide repeat (TNR) and the other consisting of the "rest of the coding region" (RCR). Usually, a TNR is composed of a repetitive single codon, which may reflect its frequency in a gene. In contrast, a non-random frequency of a codon in the RCR versus TNR (or vice versa) of a gene should indicate a bias for that codon within the TNR. We examined this scenario by comparing codon frequency between the RCR and the cognate TNR(s) for a set of human and mouse genes. A TNR length of six amino acids or more was used to identify genes from the Genbank database. Twenty nine human and twenty one mouse genes containing TNRs coding for nine different amino acid runs were identified. The ratio of codon frequency in a TNR versus the corresponding RCR was expressed as "fold change" which was also regarded as a measure of codon bias (defined as preferential use either in TNR or in RCR). Chi-square values were then determined from the distribution of codon frequency in a TNR vs. the cognate RCR. At p<0.001, 22% and 27%, respectively, of human and mouse TNRs showed codon bias. Greater than 40% of the TNRs (29 out of 69 in human, and 18 of 42 in mouse) showed codon bias at p<0.05. In addition, we identify eight single-codon TNRs in mouse and ten in human genes. Thus, our results show intragenic codon bias in both mouse and human genes expressed in diverse tissue types. Since our results are independent of the Codon Adaptation Index (CAI) and starvation CAI, and since the tRNA repertoire in a cell or in a tissue is constant, our data suggest that other constraints besides tRNA abundance played a role in creating intragenic codon bias in these genes.

Amino Acids↗

Isolation and characterization of a complementary DNA clone for an algal pre-apoplastocyanin.

We have isolated a cDNA clone for the Chlamydomonas reinhardtii pre-apoplastocyanin. The sequence contains codons for the complete pre-protein including a two-domain, lumen-targeting transit sequence and the mature apoprotein. The transit sequence (47 amino acids) is the shortest one described for chloroplast lumenal proteins, and like other C. reinhardtii lumen-targeting transit sequences appears to lack an uncharged amino-terminal domain usually present in plant lumen-directing sequences. The mature protein is deduced to be 98 amino acids in length and shows highest primary sequence similarity (74-76% identity) to other unicellular algal plastocyanins. Southern hybridization analysis of C. reinhardtii genomic DNA indicates the presence of a single nuclear gene, as is the case for all other plastocyanin genes characterized to date, although the algal gene might be interrupted. Codon usage in this gene reflects the high GC content of C. reinhardtii nuclear DNA, but is more highly biased than that found in the C. reinhardtii copper-repressible gene for the functionally equivalent pre-apocytochrome c552 (perhaps contributing to the more efficient synthesis in vivo of plastocyanin over cytochrome c552). The deduced physical properties of this plastocyanin are compared to those of the C. reinhardtii plastidic cytochrome c552.

Amino Acid Sequence↗

Correlation between molecular clock ticking, codon usage fidelity of DNA repair, chromosome banding and chromatin compactness in germline cells.

The vertebrate genome is built of long DNA regions, relatively homogeneous in GC content, which likely correspond to bands on stained chromosomes. Large differences in composition have been found among DNA regions belonging to the same genome. They are paralleled by differences in codon usage in genes differently localized. The hypothesis presented here asserts that these differences in composition are caused by different mutational bias of alpha and beta DNA polymerases, these polymerases being involved to different extents in the repair of DNA lesions in compact and relaxed chromatin, respectively, in germline cells.

Animals↗

The effects of Hill-Robertson interference between weakly selected mutations on patterns of molecular evolution and variation.

Associations between selected alleles and the genetic backgrounds on which they are found can reduce the efficacy of selection. We consider the extent to which such interference, known as the Hill-Robertson effect, acting between weakly selected alleles, can restrict molecular adaptation and affect patterns of polymorphism and divergence. In particular, we focus on synonymous-site mutations, considering the fate of novel variants in a two-locus model and the equilibrium effects of interference with multiple loci and reversible mutation. We find that weak selection Hill-Robertson (wsHR) interference can considerably reduce adaptation, e.g., codon bias, and, to a lesser extent, levels of polymorphism, particularly in regions of low recombination. Interference causes the frequency distribution of segregating sites to resemble that expected from more weakly selected mutations and also generates specific patterns of linkage disequilibrium. While the selection coefficients involved are small, the fitness consequences of wsHR interference across the genome can be considerable. We suggest that wsHR interference is an important force in the evolution of nonrecombining genomes and may explain the unexpected constancy of codon bias across species of very different census population sizes, as well as several unusual features of codon usage in Drosophila.

Alleles↗

Comparison of a vitellogenin gene between two distantly related rhabditid nematode species.

Three vitellogenin genes from the free-living nematode Caenorhabditis elegans have previously been characterized at the molecular level. In order to study evolutionary relationships within this poorly understood taxon, we have cloned a vitellogenin gene, CEW1-vit-6, from a distantly related species belonging to the same family as C. elegans. Screening of a genomic library with a probe to total poly(A+) RNA yielded three clones that hybridized more intensely than all others, and all three corresponded to a single gene homologous to C. elegans vit-6. Comparison of CEW1-vit-6 with Ce-vit-6 reveals both strong similarities and surprising differences. Life Ce-vit-6, the gene is about 5 kb long and contains four unusually small introns (38-41 nt), but only one interrupts the gene at the same location as a Ce-vit-6 intron. The promoter region contains five matches to Vitellogenin Promoter Element 1 (VPE1) and no matches to VPE2, both previously shown to be required for vit gene transcription in C. elegans. Codon usage is in general similar to that of the Ce-vit genes, but a few codon biases are quite different. Alignment of the CEW1-vit-6 protein with Ce-vit-6 and Ce-vit-2 products suggests the existence of two domains which have evolved at different rates. Sequence comparison shows that nematode vitellogenins are much more closely related to vertebrate than to insect vitellogenins.

Amino Acid Sequence↗

A synthetic E7 gene of human papillomavirus type 16 that yields enhanced expression of the protein in mammalian cells and is useful for DNA immunization studies.

A synthetic E7 gene of human papillomavirus (HPV) type 16 was generated that consists entirely of preferred human codons. Expression analysis of the synthetic E7 gene in human and animal cells showed levels of E7 protein 20- to 100-fold higher than those obtained with wild-type E7. Enhanced expression of E7 protein resulted from highly efficient translation, as well as increased stability of the E7 mRNA due to its codon optimization. Higher levels of E7 protein in cells transfected with synthetic E7 correlated with significant loss of cell viability in various human cell lines. In contrast, lower E7 protein expression driven by the wild-type gene resulted in a slight induction of cell proliferation. Furthermore, mice inoculated with plasmids expressing the synthetic E7 gene produced significantly higher levels of E7 antibodies than littermates injected with wild-type E7, suggesting that synthetic E7 may be useful for DNA immunization studies and the development of genetic vaccines against HPV-16. In view of these results, we hypothesize that HPVs may have retained a pattern of G + C content and codon usage distinct from that of their host cells in response to selective pressure. Thus, the nonhuman codon bias may have been conserved by HPVs to prevent compromising viability of the host cells by excessive viral early protein expression, as well as to evade the immune system.

Amino Acid Sequence↗

Structure, exon pattern, and chromosome mapping of the gene for cytosolic copper-zinc superoxide dismutase (sod-1) from Neurospora crassa.

A 4.8-kilobase BamHI-HindIII fragment encoding the entire Neurospora crassa CuZn superoxide dismutase gene (herein designated sod-1) was isolated from a genomic library using two 60-base deoxyoligonucleotide probes corresponding to the published N. crassa amino acid sequence. The nucleotide sequence of the gene encodes an amino acid sequence matching the published protein sequence at 152 of 153 positions. Codon preference shows an unusually strong bias such that only 32 of the possible 61 codons are used, with no codons ending in A. Codon usage is that of highly expressed N. crassa genes. The gene contains three introns, none of which corresponds to any of the introns previously identified in the human gene. Analysis of the intron positions provides support for the hypothesis that CuZn superoxide dismutases evolved by gene duplication and fusion followed by the addition of exons encoding an N-terminal beta-hairpin and a zinc-binding subdomain. The N. crassa gene has an intron mapping to amino acid residue 114 in a sequence-conserved region of the protein whereas the human gene has an intron mapping to a similar but not identical position at residue 118. The discordant position of these introns suggests that one of them was inserted relatively recently. The first N. crassa intron contains a sequence that is similar to the transcriptional regulatory site, UAS1, of the yeast CYC1 (iso-1-cytochrome c) gene and to a putative UAS from the yeast manganese superoxide dismutase gene. A 10-nucleotide portion of this region also matches exactly a sequence in intron 2 of the con-10 gene of N. crassa. sod-1 was mapped to the left arm of chromosome I by following the segregation of a restriction fragment length polymorphism in a sexual cross. Although results indicate that there is a single gene for cytosolic CuZn superoxide dismutase, two additional, perhaps distantly related, sequences were identified that hybridized weakly to both oligonucleotide probes.

Amino Acid Sequence↗

Regularities of context-dependent codon bias in eukaryotic genes.

Nucleotides surrounding a codon influence the choice of this particular codon from among the group of possible synonymous codons. The strongest influence on codon usage arises from the nucleotide immediately following the codon and is known as the N1 context. We studied the relative abundance of codons with N1 contexts in genes from four eukaryotes for which the entire genomes have been sequenced: Homo sapiens, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana. For all the studied organisms it was found that 90% of the codons have a statistically significant N1 context-dependent codon bias. The relative abundance of each codon with an N1 context was compared with the relative abundance of the same 4mer oligonucleotide in the whole genome. This comparison showed that in about half of all cases the context-dependent codon bias could not be explained by the sequence composition of the genome. Ranking statistics were applied to compare context-dependent codon biases for codons from different synonymous groups. We found regularities in N1 context-dependent codon bias with respect to the codon nucleotide composition. Codons with the same nucleotides in the second and third positions and the same N1 context have a statistically significant correlation of their relative abundances.

Animals↗

Characterization of the human Ig heavy chain antigen binding complementarity determining region 3 using a newly developed software algorithm, JOINSOLVER.

We analyzed 77 nonproductive and 574 productive human V(H)DJ(H) rearrangements with a newly developed program, JOINSOLVER. In the productive repertoire, the H chain complementarity determining region 3 (CDR3(H)) was significantly shorter (46.7 +/- 0.5 nucleotides) than in the nonproductive repertoire (53.8 +/- 1.9 nucleotides) because of the tendency to select rearrangements with less TdT activity and shorter D segments. Using criteria established by Monte Carlo simulations, D segments could be identified in 71.4% of nonproductive and 64.4% of productive rearrangements, with a mean of 17.6 +/- 0.7 and 14.6 +/- 0.2 retained germline nucleotides, respectively. Eight of 27 D segments were used more frequently than expected in the nonproductive repertoire, whereas 3 D segments were positively selected and 3 were negatively selected, indicating that both molecular mechanisms and selection biased the D segment usage. There was no bias for D segment reading frame (RF) use in the nonproductive repertoire, whereas negative selection of the RFs encoding stop codons and positive selection of RF2 that frequently encodes hydrophilic amino acids were noted in the productive repertoire. Except for serine, there was no consistent selection or expression of hydrophilic amino acids. A bias toward the pairing of 5' D segments with 3' J(H) segments was observed in the nonproductive but not the productive repertoire, whereas V(H) usage was random. Rearrangements using inverted D segments, DIR family segments, chromosome 15 D segments and multiple D segments were found infrequently. Analysis of the human CDR3(H) with JOINSOLVER has provided comprehensive information on the influences that shape this important Ag binding region of V(H) chains.

Algorithms↗

Codon usage is imposed by the gene location in the transcription unit.

A characteristic profile of the fluctuations of codon usage is observed in bacteriophages and mitochondria. By following the DNA in the direction of transcription, one moves slowly from a region where selective pressure favours codons ending with C to a region where the bias is in favour of codons ending with T; then, abruptly, one again enters a region of codons ending in C. The transcription end point takes place in the area of abrupt change in codon usage. By comparing Drosophila yakuba and mouse mitochondrial genomes, it is possible to show that the strategy of codon usage for a given gene depends on its location along the transcription unit and not on the encoded protein. The choice of codons ending in T or C allows large scale variations of DNA stability which could regulate the speed of propagation of the RNA polymerase.

Animals↗

Characterization of the virB operon from an Agrobacterium tumefaciens Ti plasmid.

The virulence genes of the Agrobacterium tumefaciens Ti plasmid are grouped into six transcription units and direct the transfer of T-DNA into plant cells. We report here the nucleotide sequence of the largest vir operon, virB, from the Ti plasmid pTiA6NC. This operon contains 11 open reading frames, 7 of which show evidence of translational coupling. trpE::virB gene fusions were used to confirm the reading frames of genes virB2, 4, 5, 6, 7, 8, 10, and 11. In addition, the native gene products of virB6 and virB9 were identified using maxicell and in vitro transcription-translation techniques, and the VirB9 protein was found to be proteolytically processed. The codon usage of the predicted virB genes is very similar to the other pTiA6 vir genes and is much less biased than Escherichia coli. Since many of the virB gene products have secretion signals common to exported bacterial proteins, it is likely that they will be membrane-associated. We propose that the VirB proteins are involved in the formation of a transmembrane structure which mediates the passage of the transferred T-DNA molecule through the bacterial and plant cell membranes.

Base Sequence↗