Search PubMedSearch

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Evolution of tropomyosin functional domains: differential splicing and genomic constraints.

We have cloned and determined the nucleotide sequence of a complementary DNA (cDNA) encoded by a newly isolated human tropomyosin gene and expressed in liver. Using the least-square method of Fitch and Margoliash, we investigated the nucleotide divergences of this sequence and those published in the literature, which allowed us to clarify the classification and evolution of the tropomyosin genes expressed in vertebrates. Tropomyosin undergoes alternative splicing on three of its nine exons. Analysis of the exons not involved in differential splicing showed that the four human tropomyosin genes resulted from a duplication that probably occurred early, at the time of the amphibian radiation. The study of the sequences obtained from rat and chicken allowed a classification of these genes as one of the types identified for humans. The divergence of exons 6 and 9 indicates that functional pressure was exerted on these sequences, probably by an interaction with proteins in skeletal muscle and perhaps also in smooth muscle; such a constraint was not detected in the sequences obtained from nonmuscle cells. These results have led us to postulate the existence of a protein in smooth muscle that may be the counterpart of skeletal muscle troponin. We show that different kinds of functional pressure were exerted on a single gene, resulting in different evolutionary rates and different convergences in some regions of the same molecule. Codon usage analysis indicates that there is no strict relationship between tissue types (and hence the tRNA precursor pool) and codon usage. G + C content is characteristic of a gene and does not change significantly during evolution.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

Horizontal transfer of accessory chromosomes in fungi - a regulated process for exchange of genetic material?

Horizontal transfer of entire chromosomes has been reported in several fungal pathogens, often significantly impacting the fitness of the recipient fungus. All documented instances of horizontal chromosome transfers (HCTs) showed a marked propensity for accessory chromosomes, consistently involving the transfer of an accessory chromosome while other chromosomes were seldom, if ever, co-transferred. The mechanisms underlying HCTs, as well as the factors regulating the specificity of HCTs for accessory chromosomes, remain unclear. In this perspective, we provide an overview of the observed propensity in reported cases of horizontal chromosome transfers. We hypothesize the existence of a signal that distinguishes mobile, i.e., horizontally transferred, accessory chromosomes from the rest of the donor genome. Recent findings in Metarhizium robertsii and Magnaporthe oryzae, suggest that a mobile accessory chromosome may contain putative histones and/or histone modifiers, which could generate such a signal. Based on this, we propose that mobile accessory chromosomes may encode the machinery required for their own horizontal transmission, implying that HCT could be a regulated process. Finally, we present evidence of substantial differences in codon usage bias between core and accessory chromosomes in 14 out of 19 analysed fungal species and strains. Such differences in codon usage bias could indicate past horizontal transfers of these accessory chromosomes. Interestingly, HCT was previously unknown for many of these species, suggesting that the horizontal transfer of accessory chromosomes may be more widespread than previously thought, and therefore an important factor in fungal genome evolution.

Gene Transfer, Horizontal

High-level expression of staphylococcal nuclease R gene in Escherichia coli.

Staphylococcal nuclease R, an analogue of nuclease A, was overproduced under the transcriptional control of the bacteriophage lambda PRPL promoters regulated by temperature sensitive repressors. The expression level reached 200-300 mg l-1 and showed little host dependence in different strains. The investigations of the recombinant nuclease R have revealed that the amino terminal formyl methionine residue of the nuclease is precisely processed, the protein consists of 155 amino acid residues. The experiment shows that the pBV221-DH5 alpha is a quite suitable vector-host system for high-level expression and precise processing of heterologous genes in Escherichia coli. The comparative studies between the codons used in the staphylococcal nuclease R gene and the optimal codon usage in E. coli indicate that high level expression of heterologous genes in E. coli may not always require a high degree of codon usage bias.

Amino Acid Sequence

Does the 'non-coding' strand code?

The hypothesis that DNA strands complementary to the coding strand contain in phase coding sequences has been investigated. Statistical analysis of the 50 genes of bacteriophage T7 shows no significant correlation between patterns of codon usage on the coding and non-coding strands. In Bacillus and yeast genes the correlation observed is not different from that expected with random synonymous codon usage, while a high correlation seen in 52 E. coli genes can be explained in terms of an excess of RNY codons. A deficiency of UUA, CUA and UCA codons (complementary to termination) seems to be restricted to the E. coli genes, and may be due to low abundance of the relevant cognate tRNA species. Thus the analysis shows that the non-coding strand has the properties expected of a sequence complementary to a coding strand, with no indications that it encodes, or may have encoded, proteins.

Bacillus

An estimate on the effect of point mutation and natural selection on the rate of amino acid replacement in proteins.

We outline a method for estimating quantitatively the influence of point mutations and selection on the frequencies of codons and amino acids. We show how the mutation rate, i.e., the rate of amino acid replacement due to point mutation, can be affected by the codon usage as well as by the rates of the involved base exchanges. A comparison of the mutation rates calculated from reliable values of codon usage and base exchange probabilities with those that would be expected on the basis of chance reveals a notable suppression of replacements leading to tryptophan, glutamate, lysine, and methionine, and particularly of those leading to the termination codons. If selection constraints are neglected and only mutations are taken into account, the best agreement between expected and observed frequencies of both codons and amino acids is obtained for alpha = 1.13-1.15, where (Formula: see text). The "selection values" of codons and amino acids derived by our method show a pattern that partially deviates from others in the literature. For example, the selection pressure on methionine and cysteine turns out to be much more pronounced than expected if only the discrepancies between their observed and expected occurrences in proteins are considered. To estimate to what extent randomly occurring amino acid replacements are accepted by selection, we constructed an "acceptability matrix" from the well-established matrix of accepted point mutations. On the basis of this matrix "acceptability values" of the amino acids can be defined that correlate with their selection values. We also examine the significance of mutations and selection of amino acids with respect to their physicochemical properties and functions in proteins. The conservatism of amino acid replacements with respect to certain properties such as polarity can be brought about by the mutational process alone, whereas the conservatism with respect to other relevant properties--among them all measures of bulkiness--obviously is the result of additional selectional constraints on the evolution of protein structures.

Amino Acid Sequence

Characterization of In0 of Pseudomonas aeruginosa plasmid pVS1, an ancestor of integrons of multiresistance plasmids and transposons of gram-negative bacteria.

Many multiresistance plasmids and transposons of gram-negative bacteria carry related DNA elements that appear to have evolved from a common ancestor by site-specific integration of discrete cassettes containing antibiotic resistance genes or sequences of unknown function. The site of integration is flanked by conserved segments coding for an integraselike protein and for sulfonamide resistance, respectively. These segments, together with the antibiotic resistance genes between them, have been termed integrons (H. W. Stokes and R. M. Hall, Mol. Microbiol. 3:1669-1683, 1989). We report here the characterization of an integron, In0, from Pseudomonas aeruginosa plasmid pVS1, which has an unoccupied integration site and hence may be an ancestor of more complex integrons. Codon usage of the integrase (int) and sulfonamide resistance (sul1) genes carried by this integron suggests a common origin. This contrasts with the codon usage of other antibiotic resistance genes that were presumably integrated later as cassettes during the evolution and spread of these DNA elements. We propose evolutionary schemes for (i) the genesis of the integrons by the site-specific integration of antibiotic resistance genes and (ii) the evolution of the integrons of multiresistance plasmids and transposons, in relation to the evolution of transposons related to Tn21.

Amino Acid Sequence

Overproduction from a cellulase gene with a high guanosine-plus-cytosine content in Escherichia coli.

A recombinant exoglucanase was expressed in Escherichia coli to a level that exceeded 20% of total cellular protein. To obtain this level of overproduction, the exoglucanase gene coding sequence was fused to a synthetic ribosome-binding site, an initiating ATG, and placed under the control of the leftward promoter of bacteriophage lambda contained on the runaway replication plasmid vector pCP3 (E. Remaut, H. Tsao, and W. Fiers, Gene 22:103-113, 1983). With the exception of an inserted asparagine adjacent to the initiating ATG, the highly expressed exoglucanase is identical to the native exoglucanase. The overproduced exoglucanase can be isolated easily in an enriched form as insoluble aggregates, and exoglucanase activity can be recovered by solubilization of the aggregates in 6 M urea or 5 M guanidine hydrochloride. Since the codon usage of the exoglucanase gene is so markedly different from that of E. coli genes, the overproduction of the exoglucanase in E. coli indicates that codon usage may not be a major barrier to heterospecific gene expression in this organism.

Actinomycetales

Fine structural features of the chloroplast genome: comparison of the sequenced chloroplast genomes.

The entire nucleotide sequences of the rice, tobacco and liverwort chloroplast genomes have been determined. We compared all the chloroplast genes, open reading frames and spacer regions in the plastid genomes of these three species in order to elucidate general structural features of the chloroplast genome. Analyses of homology, GC content and codon usage of the genes enabled us to classify them into two groups: photosynthesis genes and genetic system genes. Based on comparisons of homology, GC content and codon usage, unidentified ORFs can also be assigned to each of these groups such that it is possible to speculate about the functions of products which may be produced by these ORFs. The spacer regions and intron sequences were compared and found to have no obvious homology between rice and liverwort or between tobacco and liverwort.

Base Composition

Variation in G + C-content and codon choice: differences among synonymous codon groups in vertebrate genes.

The relationship between G + C-content and codon usage in genes of human, mus, rat, bovine and chicken nuclear genomes was investigated. Correlation and lineal regression analyses were carried out on plots that related the frequency of each codon within each synonymous codon group to the G + C-content of the coding sequence as a whole. Under GC pressure, in most of the quartet codon groups there is a preferential choice of the C-ending codon, except in leucine and valine codon groups where the choice of the G-ending codon is preferred. Among ducts, the choice of codons specifying phenylalanine and glutamate shows the strongest dependence on G + C-content. The relationship found between G + C-content and codon usage in these genomes correlate with taxonomic distance.

Animals

Switches in species-specific codon preferences: the influence of mutation biases.

A model of synonymous codon usage is developed in which the most frequent codons are selectively advantageous because of their coadaptation with tRNA abundances. Random drift opposes the progress of this coevolution by pushing codon frequencies in the direction of the frequency that would result from mutation in the absence of selection. It is predicted that, within a certain range, an increased mutation bias away from an advantageous codon has little influence on its usage in highly expressed genes. However, a subsequent small increase in mutation bias over a critical range leads to a large reduction in the frequency of the codon. The switch in preference from one synonym to another is a sharp transition, with no stable intermediate state in which neither codon is advantageous. Codon usage patterns were compared among three related bacterial species of differing genomic G & C contents, Escherichia coli, Serratia marcescens, and Proteus vulgaris. It was found that although changes in mutation biases do not always result in switches in codon preferences, some switches have occurred in the direction of species-specific mutation biases. Fluctuating mutation biases may therefore be the main cause of differences between species in their codon preferences.

Amino Acids

Characterization of a highly expressed lignin peroxidase-encoding gene from the basidiomycete Phanerochaete chrysosporium.

The genomic clone, LG2, encoding LiP2, the major lignin peroxidase (LiP) isozyme from Phanerochaete chrysosporium strain OGC101, was isolated and characterized. The 5'-untranslated region of LG2 contains sequences similar to CRE and XRE promoter elements. Comparison with its transcript indicates that eight introns, each less than 59 bp, interrupt the coding sequence. Comparison with genes encoding other LiP isozymes shows five related patterns of intron location, whose incidence coincides with described LiP structural subfamilies. Codon bias indices calculated for all known P. chrysosporium genes, including trpC and genes encoding LiP, MnP, and exo-cellobiohydrolase I, demonstrate that LG2 has the most biased codon usage. We conclude that subdivisions of the LiP family may be based on intron location in the encoding genes, and that ranking of isozyme production levels can be estimated by the extent of bias in codon usage in the cognate gene.

Amino Acid Sequence

The nucleotide sequence of an Escherichia coli operon containing genes for the tRNA(m1G)methyltransferase, the ribosomal proteins S16 and L19 and a 21-K polypeptide.

The nucleotide sequence of a 4.6-kb SalI-EcoRI DNA fragment including the trmD operon, located at min 56 on the Escherichia coli K-12 chromosome, has been determined. The trmD operon encodes four polypeptides: ribosomal protein S16 (rpsP), 21-K polypeptide (unknown function), tRNA-(m1G)methyltransferase (trmD) and ribosomal protein L19 (rplS), in that order. In addition, the 4.6-kb DNA fragment encodes a 48-K and a 16-K polypeptide of unknown functions which are not part of the trmD operon. The mol. wt. of tRNA(m1G)methyltransferase determined from the DNA sequence is 28 424. The probable locations of promoter and terminator of the trmD operon are suggested. The translational start of the trmD gene was deduced from the known NH2-terminal amino acid sequence of the purified enzyme. The intercistronic regions in the operon vary from 9 to 40 nucleotides, supporting the earlier conclusion that the four genes are co-transcribed, starting at the major promoter in front of the rpsP gene. Since it is known that ribosomal proteins are present at 8000 molecules/genome and the tRNA-(m1G)methyltransferase at only approximately 80 molecules/genome in a glucose minimal culture, some powerful regulatory device must exist in this operon to maintain this non-coordinate expression. The codon usage of the two ribosomal protein genes is similar to that of other ribosomal protein genes, i.e., high preference for the most abundant tRNA isoaccepting species. The trmD gene has a codon usage typical for a protein made in low amount in accordance with the low number of tRNA-(m1G)methyltransferase molecules found in the cell.

Bacterial Proteins

The cytochrome b region in the mitochondrial DNA of the ant Tetraponera rufoniger: sequence divergence in Hymenoptera may be associated with nucleotide content.

Polymerase chain reaction (PCR) followed by sequencing of single-stranded DNA yielded sequence information from the cytochrome b (cyt b) region in mitochondrial DNA from the ant Tetraponera rufoniger. Compared with the cyt b genes from Apis mellifera, Drosophila melanogaster, and D. yakuba, the overall A+T content (A+T%) of that of T. rufoniger is lower (69.9% vs 80.7%, 74.2%, and 73.9%, respectively) than those of the other three. The codon usage in the cyt b gene of T. rufoniger is biased although not as much as in A. mellifera, D. melanogaster, and D. yakuba; T. rufoniger has eight unused codons whereas D. melanogaster, D. yakuba, and A. mellifera have 21, 20, and 23, respectively. The inferred cyt b polypeptide chain (PPC) of T. rufoniger has diverged at least as much from a common ancestor with D. yakuba as has that of A. mellifera (approximately 3.5 vs approximately 2.9). Despite the lower A+T%, the relative frequencies of amino acids in the cyt b PPC of T. rufoniger are significantly (P < 0.05) associated with the content of adenine and thymine (A+T%) and size of codon families. The mitochondrially located cytochrome oxidase subunit II genes (CO-II) of endopterygote insects have significantly higher average A+T% (approximately 75%) than those of exopterygous (approximately 69%) and paleopterous (approximately 69%) insects. The increase in A+T% of endopterygote insects occurred in Upper Carboniferous and coincided with a significant acceleration of PPC divergence. However, acceleration of PPC divergence is not significantly correlated with the increase of the A+T% (P > 0.1). The high A+T%, the biased codon usage, and the increased PPC divergence of Hymenoptera can in that respect most easily be explained by directional mutation pressure which began in the Upper Carboniferous and still occurs in most members of the order. Given the roughly identical A+T% of the cyt b and CO-II genes from the other insects whose DNA sequences are known (A. mellifera, D. melanogaster, and D. yakuba), it seems most likely that the A+T% of T. rufoniger declined secondarily within the last 100 Myr as a result of a reduced directional mutation pressure.

Amino Acid Sequence

Molecular Evolution and Expression Analysis of the ADH Gene Family in Apple Bud Mutants.

Alcohol dehydrogenase (ADH) catalyzes the reduction of aldehydes to alcohols, key precursor substrates for volatile ester biosynthesis, which determines the characteristic aroma of apple fruit. However, a comprehensive genome-wide investigation of the ADH gene family in apple has been lacking. In this study, we systematically identified ADH genes in the apple genome using integrated bioinformatics approaches, including phylogenetic analysis, synteny evaluation, promoter cis-element prediction, codon usage bias assessment, and protein interaction network modeling. Expression patterns were examined through transcriptomic data and validated by RT-qPCR analysis across different organs and among 'Red Delicious' and its four bud mutant lines. We identified 44 ADH genes, with 12 forming a prominent cluster on chromosome 1. RT-qPCR analysis revealed that MdADH20 was dramatically upregulated in the 'Red Chief' mutant (relative expression of 59.38), suggesting its pivotal role. Phylogenetic analysis revealed a close evolutionary relationship with wild strawberry. The encoded proteins were generally stable and predominantly localized to the cytoplasm. Promoter analysis showed enrichment of growth/development-related and ARE elements, while codon usage analysis identified AGA, GCU, GUU, and CUU as preferred codons. Protein interaction prediction suggested MdADH19 and MdADH20 as hub proteins. Expression profiling and RT-qPCR further identified MdADH20 as a core candidate gene, characterized by its stable and high expression, particularly in the 'Red Delicious' mutant. Its central position in the predicted protein-protein interaction network suggests a potential regulatory role in the aroma biosynthesis pathway of apple fruit. This study provides the first systematic genome-wide characterization of the apple ADH gene family, establishing a theoretical groundwork for deciphering aroma biosynthesis mechanisms and offering potential target genes for flavor improvement through bud mutation breeding strategies.

ADH gene family

Chloramphenicol resistance in Campylobacter coli: nucleotide sequence, expression, and cloning vector construction.

A chloramphenicol-resistance determinant (CmR), originally cloned from Campylobacter coli plasmid pNR9589 in Japan, was isolated and the nucleotide sequence determined, which contained an open reading frame of 621 bp. The gene product was identified as Cm acetyltransferase (CAT), which had a putative amino acid sequence that showed 43% to 57% identity with other CAT proteins of both Gram+ and Gram- origin. Although expression of the cat gene was constitutive in both C. coli and Escherichia coli, results of primer extension experiments indicated that transcription was initiated at different sites in these two species. A kanamycin-resistance determinant, identified as the aphA-3 gene, was located downstream from the cat gene. The codon usage of the cat gene is very different from that used in E. coli, however, the CAT polypeptide was synthesized in large amounts in E. coli maxicells. Therefore, the codon usage bias is not one of the obstacles which affects Campylobacter spp. gene expression in E. coli. New Campylobacter cloning vectors were constructed in this study.

Amino Acid Sequence

Theory of degenerate coding and informational parameters of protein coding genes.

The theory of degenerate coding is presented in a way enabling further application to molecular biology. There are two kinds of redundancy of a degenerate code. The first is due to the excess in codon length and the second to the code degeneracy. If the code is asymmetrically degenerate, the second kind of redundancy can be profitable for control of error rate. This control can be performed just by selective synonymous codon usage. Utilisation of the genetic code is partially influenced by this theoretical possibility. In particular the degree of error protectivity is well correlated with deviation from equiprobability in synonymous codon usage. The biological significance of this fact is discussed.

Animals

Codon catalog usage is a genome strategy modulated for gene expressivity.

The nucleic acid sequence bank now contains 161 mRNAs, 43 new genes are added. One sequence, that of B. mori fibroin, is dropped due to uncertainty on the starting point for translation. Frequencies of all codons are given for each gene added and for each genome type in the total bank. A new series of correspondence analyses on codon use is presented, substantiating the genome hypothesis. Internal regulation of mRNA expression by different third base choices between quartet and duet codons is proposed for bacterial genes.

Amino Acid Sequence

The targeting of somatic hypermutation.

Somatic hypermutation does not occur randomly within immunoglobulin V genes but, rather, is preferentially targeted to certain nucleotide positions (hot spots) and away from others (cold spots). Cold spots often coincide with residues essential for V gene folding. Hotspots, which appear to be strategically located to favour affinity maturation, are most frequently located in the CDRs (particularly CDR1) though conserved hotspots are also found at the base of FR3. Hotspots are in part created by local DNA sequence and the strong biases of codon usage in V genes indicate that the genes have evolved such that somatic hypermutation is targeted to those parts of the V where it is likely to prove most useful. These features of mutational hotspots and biased codon usage are also evident in V genes of lower animals suggesting that diversification by strategic targeting of non-templated mutation may have evolved early in antigen receptor evolution.

Animals