Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Hypermutation generating the sheep immunoglobulin repertoire is an antigen-independent process.

Somatic hypermutation of light chain V genes during development of B cells in sheep ileal Peyer's patches was studied in three experimental conditions: in sterile fragments of the ileum surgically isolated from the gut during fetal life, in germ-free sheep, and in animals thymectomized during early fetal life. The somatic mutation pattern was found identical to control tissues in all three experiments. The same age-dependent amount of mutations, a higher than theoretical R/S ratio in complementarity-determining regions (CDRs), and a similar clustering of mutations in CDRs were observed. The mechanism, as estimated from the silent mutation pattern, appears to target mutations to CDRs; moreover, the major V lambda genes have a specific codon usage with a high purine content at the first two bases of the codons and a low content at the third position, which, together with a specific targeting of mutations to purines, favors replacement mutations in CDRs.

Animals↗

Synonymous substitution rates in enterobacteria.

It has been shown previously that the synonymous substitution rate between Escherichia coli and Salmonella typhimurium is lower in highly than in weakly expressed genes, and it has been suggested that this is due to stronger selection for translational efficiency in highly expressed genes as reflected in their greater codon usage bias. This hypothesis is tested here by comparing the substitution rate in codon families with different patterns of synonymous codon use. It is shown that the decline in the substitution rate across expression levels is as great for codon families that do not appear to be subject to selection for translational efficiency as for those that are. This implies that selection on translational efficiency is not responsible for the decline in the substitution rate across genes. It is argued that the most likely explanation for this decline is a decrease in the mutation rate. It is also shown that a simple evolutionary model in which synonymous codon use is determined by a balance between mutation, selection for an optimal codon, and genetic drift predicts that selection should have little effect on the substitution rate in the present case.

Codon↗

Human metallothionein genes: molecular cloning and sequence analysis of the mRNA.

From a cDNA clone bank prepared from cadmium-treated HeLa cells, we isolated clones representing mRNAs whose concentration is increased after cadmium induction. Several metallothionein cDNA clones were isolated by cross-hybridization to mouse metallothionein-I cDNA. The nucleotide sequence of one of these clones, containing a nearly full-length cDNA copy of human metallothionein-II mRNA, was determined. The homology between the human and mouse metallothionein sequences is strictly limited to the coding region of the mRNA. Codon usage in metallothionein mRNA is not random. Seventy-nine percent of the codons have G or C residues at the third position, resulting in a GC-rich sequence.

Amino Acid Sequence↗

Complete sequence of the amphioxus (Branchiostoma lanceolatum) mitochondrial genome: relations to vertebrates.

The complete nucleotide sequence of the mitochondrial DNA of the amphioxus Branchiostoma lanceolatum has been determined. This mitochondrial genome is small (15 076 bp) because of the short size of the two rRNA genes and the tRNA genes. In addition, this genome contains a very short non-coding region (57 bp) with no sequence reminiscent of a control region. The organisation of the coding genes, as well as of the two rRNA genes, is identical to that of the sea lamprey. Some differences in the repartition of the tRNA genes occur when compared to the lamprey. The mitochondrial codon usage of the amphioxus is reminiscent of that of urochordates since the AGA codon is read as a glycine and not as a stop codon as in vertebrates. Moreover, the base composition at the wobble positions of the codon is strongly biased toward guanine. Altogether, these data clearly emphasise the close relationships between amphioxus and vertebrates, and reinforce the notion that prochordates may be viewed as the brother group of vertebrates.

Animals↗

Adenylylsulphate reductase from the sulphate-reducing archaeon Archaeoglobus fulgidus: cloning and characterization of the genes and comparison of the enzyme with other iron-sulphur flavoproteins.

Adenylylsulphate (adenosine-5'-phosphosulphate, APS) reductase from the extremely thermophilic sulphate-reducing archaeon Archaeoglobus fulgidus is an iron-sulphur flavoprotein containing one non-covalently bound flavin group, eight non-haem iron and six labile sulphide atoms per molecule. Reevaluation of the enzyme structure revealed the presence of two different subunits with molecular masses of 80 and 18.5 kDa. The subunits are arranged in an alpha 2 beta subunit structure. We have cloned and sequenced a 2.7 kb segment of DNA containing the genes for the alpha and beta subunits, which we designate aprA and aprB, respectively. The two genes are separated by 17 bp and localized in the order aprBA. While a putative promoter could not be identified in the vicinity of aprBA a probable termination signal was found just downstream of the translation stop codon of aprA. The codon usage for aprBA shows strong preferences for G and C in the third codon position. aprA encodes a 73.3 kDa polypeptide, which shows significant overall similarities with the flavoprotein subunits of the succinate dehydrogenases from Escherichia coli and Bacillus subtilis and the corresponding flavoprotein of E. coli fumarate reductase. Part of the homologous peptide stretches could be assigned to domains that are involved in the binding of the substrate or of the FAD prosthetic group. aprB encodes a 17.1 kDa polypeptide representing an iron-sulphur protein, seven cysteine residues of which are arranged in two clusters typical of ligands of the iron-sulphur centres in ([Fe3S4][Fe4S4]) 7-Fe ferredoxins.

Amino Acid Sequence↗

A unique set of 11,008 onion expressed sequence tags reveals expressed sequence and genomic differences between the monocot orders Asparagales and Poales.

Enormous genomic resources have been developed for plants in the monocot order Poales; however, it is not clear how representative the Poales are for the monocots as a whole. The Asparagales are a monophyletic order sister to the lineage carrying the Poales and possess economically important plants such as asparagus, garlic, and onion. To assess the genomic differences between the Asparagales and Poales, we generated 11,008 unique ESTs from a normalized cDNA library of onion. Sequence analyses of these ESTs revealed microsatellite markers, single nucleotide polymorphisms, and homologs of transposable elements. Mean nucleotide similarity between rice and the Asparagales was 78% across coding regions. Expressed sequence and genomic comparisons revealed strong differences between the Asparagales and Poales for codon usage and mean GC content, GC distribution, and relative GC content at each codon position, indicating that genomic characteristics are not uniform across the monocots. The Asparagales were more similar to eudicots than to the Poales for these genomic characteristics.

Cytosine↗

Regulation of protein synthesis in Tetrahymena. RNA sequence sets of growing and starved cells.

The complexity of messenger RNA in growing or starved Tetrahymena thermophila is similar and unusually high (approximately 4.5 X 10(7) nucleotides). The complexity of nuclear RNA in growing cells (approximately 7.8 X 10(7) nucleotides) is only about 1.7 times that of mRNA. The concentration of complex class (rare) messages (approximately 53 copies/growing cell and approximately 11 copies/starved cell) is low in comparison to the size of the cell. The concentration of complex nuclear transcripts is also very low (approximately 0.7 copies/growing cell nucleus and approximately 2.6 copies/starved cell nucleus) considering that the macronucleus contains 45 to 90 copies of each single copy sequence. The complex sequence sets found on polysomes of growing and starved cells overlap about 80% and about 60% of the complex nuclear transcripts appear to be held in common. About 60% of macronuclear single copy DNA is transcribed in one or both physiological states. Although growing and starved cells have extremely different fractions of their messages loaded onto polysomes, within each cell type the complex messages in polysomal and nonpolysomal cytoplasmic fractions are indistinguishable, suggesting that exchange may occur between loaded and unloaded messages. Although T. thermophila DNA has an unusually low G + C content (23%), sequences coding for complex RNAs have base ratios similar to those of total DNA. Therefore, codon usage in Tetrahymena must be extremely biased towards adenine- and uridine-rich codons.

Animals↗

Evolution of the GC content of the histone 3 gene in seven Drosophila species.

The molecular evolution of the histone multigene family was studied by cloning and determining the nucleotide sequences of the histone 3 genes in seven Drosophila species, D. takahashii, D. lutescens, D. ficusphila, D. persimilis, D.pseudoobscura, D. americana and D. immigrans. CT repeats, a TATA box and an AGTG motif in the 5' region, and a hairpin loop and purine-rich motifs (CAA(T/G)GAGA) in the 3' region were conserved even in distantly related species. In D. hydei and D.americana, the GC content at the third codon position in the protein coding region was relatively low (49% and 45%), while in D. takahashii and D. lutescens it was relatively high (64% and 65%). The non- significant correlation between the GC contents in the 3' region and at the third codon position as well as the evidence of less constraint in the 3' region suggested that mutational bias may not be the major mechanism responsible for the biased nucleotide change at the third codon position or for codon usage bias.

Animals↗

Ribosome-mediated translational pause and protein domain organization.

Because regions on the messenger ribonucleic acid differ in the rate at which they are translated by the ribosome and because proteins can fold cotranslationally on the ribosome, a question arises as to whether the kinetics of translation influence the folding events in the growing nascent polypeptide chain. Translationally slow regions were identified on mRNAs for a set of 37 multidomain proteins from Escherichia coli with known three-dimensional structures. The frequencies of individual codons in mRNAs of highly expressed genes from E. coli were taken as a measure of codon translation speed. Analysis of codon usage in slow regions showed a consistency with the experimentally determined translation rates of codons; abundant codons that are translated with faster speeds compared with their synonymous codons were found to be avoided; rare codons that are translated at an unexpectedly higher rate were also found to be avoided in slow regions. The statistical significance of the occurrence of such slow regions on mRNA spans corresponding to the oligopeptide domain termini and linking regions on the encoded proteins was assessed. The amino acid type and the solvent accessibility of the residues coded by such slow regions were also examined. The results indicated that protein domain boundaries that mark higher-order structural organization are largely coded by translationally slow regions on the RNA and are composed of such amino acids that are stickier to the ribosome channel through which the synthesized polypeptide chain emerges into the cytoplasm. The translationally slow nucleotide regions on mRNA possess the potential to form hairpin secondary structures and such structures could further slow the movement of ribosome. The results point to an intriguing correlation between protein synthesis machinery and in vivo protein folding. Examination of available mutagenic data indicated that the effects of some of the reported mutations were consistent with our hypothesis.

Bacterial Proteins↗

Mammalian mitochondrial DNA evolution: a comparison of the cytochrome b and cytochrome c oxidase II genes.

The evolution of two mitochondrial genes, cytochrome b and cytochrome c oxidase subunit II, was examined in several eutherian mammal orders, with special emphasis on the orders Artiodactyla and Rodentia. When analyzed using both maximum parsimony, with either equal or unequal character weighting, and neighbor joining, neither gene performed with a high degree of consistency in terms of the phylogenetic hypotheses supported. The phylogenetic inconsistencies observed for both these genes may be the result of several factors including differences in the rate of nucleotide substitution among particular lineages (especially between orders), base composition bias, transition/transversion bias, differences in codon usage, and different constraints and levels of homoplasy associated with first, second, and third codon positions. We discuss the implications of these findings for the molecular systematics of mammals, especially as they relate to recent hypotheses concerning the polyphyly of the order Rodentia, relationships among the Artiodactyla, and various interordinal relationships.

Animals↗

A new measure to study phylogenetic relations in the brown algal order Ectocarpales: the "codon impact parameter".

We analyse forty-seven chloroplast genes of the large subunit of RuBisCO, from the algal order Ectocarpales, sourced from GenBank. Codon-usage weighted by the nucleotide base-bias defines our score called the codon-impact-parameter. This score is used to obtain phylogenetic relations amongst the 47 Ectocarpales. We compare our classification with the ones done earlier.

Base Composition↗

In-phase implies large likelihood for independent codon model: distinguishing coding from non-coding sequences.

It is proven that under the independent codon model, the likelihood of a DNA coding sequence read according to the correct frame is asymptotically larger than that read with an incorrect frame. Based on this proposition, a single set of probabilities of the codon usage is enough for discriminating the six frames of coding sequences under the independent codon model. The direct coding sequence of Escherichia coli genome is taken as an example to examine the codon independency by using the mutual information and chi2 analysis. The contrast between the coding frame and the two offset frames is evident. A self-learning approach for generating training set is proposed to estimate probability parameters.

Codon↗

Molecular characterization of cDNA encoding for adenylate kinase of rice (Oryza sativa L.).

Two types of genes (Adk-a, and Adk-b) encoding for adenylate kinase (AK, EC 2.7.4.3.) were isolated from the cDNA library constructed from poly(A)+ RNA of rice (Oryza sativa L.). Two cDNAs were heterogeneous at 5' and 3' ends of non-coding sequences and had possible polyadenylation signals. One of the genes, Adk-a, had 1154 bp sequences encoding 241 amino acid residues, while the other type, Adk-b, contained 1085 bp sequences encoding for 243 amino acid residues. Homology between Adk-a and Adk-b was 73.7% in nucleotide sequences, and 90.8% in amino acid level. Two genes showed about 53% homology to bovine mitochondrial adenylate kinase (AK2) at nucleotide and amino acid levels. Concerning the codon usage of rice AK genes, T was abundant at the third position of a codon in the reading frames. In order to examine the enzyme activity of the protein encoded by the rice cDNA, Adk-a was cloned into an expression vector, pUC119, which was introduced into Escherichia coli strain CV2, a temperature-sensitive mutant of adenylate kinase. We found that the transformant carrying the rice Adk-a gene in the sense orientation recovered cell growth at non-permissive high temperature (42 degrees C) and expressed enzyme activities higher than the untransformed CV2 and the transformant possessing Adk-a cDNA in the antisense orientation. These observations suggest that rice Adk-a codes a biologically active enzyme. Furthermore, sucrose was found to regulate the transcription of AK genes in rice cell cultures. Organ related accumulation of mRNA in whole plants was also found.

Adenylate Kinase↗

Molecular evolution of ependymin and the phylogenetic resolution of early divergences among euteleost fishes.

The rate and pattern of DNA evolution of ependymin, a single-copy gene coding for a highly expressed glycoprotein in the brain matrix of teleost fishes, is characterized and its phylogenetic utility for fish systematics is assessed. DNA sequences were determined from catfish, electric fish, and characiforms and compared with published ependymin sequences from cyprinids, salmon, pike, and herring. Among these groups, ependymin amino acid sequences were highly divergent (up to 60% sequence difference), but had surprisingly similar hydropathy profiles and invariant glycosylation sites, suggesting that functional properties of the proteins are conserved. Comparison of base composition at third codon positions and introns revealed AT-rich introns and GC-rich third codon positions, suggesting that the biased codon usage observed might not be due to mutational bias. Phylogenetic information content of third codon positions was surprisingly high and sufficient to recover the most basal nodes of the tree, in spite of the observation that pairwise distances (at third codon positions) were well above the presumed saturation level. This finding can be explained by the high proportion of phylogenetically informative nonsynonymous changes at third codon positions among these highly divergent proteins. Ependymin DNA sequences have established the first molecular evidence for the monophyly of a group containing salmonids and esociforms. In addition, ependymin suggests a sister group relationship of electric fish (Gymnotiformes) and Characiformes, constituting a significant departure from currently accepted classifications. However, relationships among characiform lineages were not completely resolved by ependymin sequences in spite of seemingly appropriate levels of variation among taxa and considerably low levels of homoplasy in the data (consistency index = 0.7). If the diversification of Characiformes took place in an "explosive" manner, over a relatively short period of time this pattern should also be observed using other phylogenetic markers. Poor conservation of ependymin's primary structure hinders the design of efficient primers for PCR that could be used in wide-ranging fish systematic studies. However, alternative methods like PCR amplification from cDNA used here should provide promising comparative sequence data for the resolution of phylogenetic relationships among other basal lineages of teleost fishes.

Amino Acid Sequence↗

Comparison and cross-species expression of the acetyl-CoA synthetase genes of the Ascomycete fungi, Aspergillus nidulans and Neurospora crassa.

The genes encoding the acetate-inducible enzyme acetyl-coenzyme A synthetase from Neurospora crassa and Aspergillus nidulans (acu-5 and facA, respectively) have been cloned and their sequences compared. The predicted amino acid sequence of the Aspergillus enzyme has 670 amino acid residues and that of the Neurospora enzyme either 626 or 606 residues, depending upon which of the two possible initiation codons is used. The amino acid sequences following the second alternative AUG show 86% homology between the two species; the extended N-terminal sequences show no homology. The Neurospora protein is characterized by the appearance of the S(T)PXX sequence motif where the amino acid homologies break down. The codon usage is biased in both genes, with a marked deficiency, especially in Neurospora, of codons with A in the third position. The facA transcribed sequence contains six introns: one in the long leader sequence, one in the 5' coding sequence not homologous with acu-5, and four within the sequence that is largely similar to that of acu-5. Only one intron, corresponding in size and position to the furthest downstream of the facA introns, is found in acu-5. The evolution of introns during the divergence of these two Ascomycete fungi is discussed. Each of the two genes has been transferred by transformation into the other species. Each species is evidently able to splice out the other's introns. Most transformants have normal acetate-induction of acetyl-CoA synthetase, implying that the two genes respond to transcriptional control signals common to both species, in spite of the striking divergence of their 5' ends.

Acetate-CoA Ligase↗

Cryptic plasmid of Neisseria gonorrhoeae: complete nucleotide sequence and genetic organization.

The naturally occurring cryptic plasmid pJD1 of Neisseria gonorrhoeae is 4,207 base pairs long and is found in about 96% of gonococcal strains. The total probable coding capacity of pJD1 was determined from the complete nucleotide sequence by using computational probes to identify open reading frames with similar codon usage and by screening for the presence of ribosomal binding sites before the start codons. Candidates for promoters and terminators were also found in the sequence. Based on these findings, we propose a model for the genetic organization of the plasmid. The model predicts two transcriptional units, each composed of five compactly spaced genes. A promoter of one of the transcripts was shown to function in Escherichia coli, and the products of three of the five genes in this operon were identified in minicell expression experiments. Of these, the cppA gene encoded a 9-kilodalton protein, and the cppB and cppC genes both coded for 24-kilodalton proteins. No expression of the other transcriptional unit was detected, but two genes in this operon were expressed in minicells when transcribed from an E. coli promoter. The experimental data were consistent with the model.

Amino Acid Sequence↗

Identification, sequence analysis, and expression of a Corynebacterium glutamicum gene cluster encoding the three glycolytic enzymes glyceraldehyde-3-phosphate dehydrogenase, 3-phosphoglycerate kinase, and triosephosphate isomerase.

To investigate a possible chromosomal clustering of glycolytic enzyme genes in Corynebacterium glutamicum, a 6.4-kb DNA fragment located 5' adjacent to the structural phosphoenolpyruvate carboxylase (PEPCx) gene ppc was isolated. Sequence analysis of the ppc-proximal part of this fragment identified a cluster of three glycolytic genes, namely, the glyceraldehyde-3-phosphate dehydrogenase (GAPDH) gene gap, the 3-phosphoglycerate kinase (PGK) gene pgk, and the triosephosphate isomerase (TPI) gene tpi. The four genes are organized in the order gap-pgk-tpi-ppc and are separated by 215 bp (gap and pgk), 78 bp (pgk and tpi), and 185 bp (tpi and ppc). The predicted gene product of gap consists of 336 amino acids (M(r) of 36,204), that of pgk consists of 403 amino acids (M(r) of 42,654), and that of tpi consists of 259 amino acids (M(r) of 27,198). The amino acid sequences of the three enzymes show up to 62% (GAPDH), 48% (PGK), and 44% (TPI) identity in comparison with respective enzymes from other organisms. The gap, pgk, tpi, and ppc genes were cloned into the C. glutamicum-Escherichia coli shuttle vector pEK0 and introduced into C. glutamicum. Relative to the wild type, the recombinant strains showed up to 20-fold-higher specific activities of the respective enzymes. On the basis of codon usage analysis of gap, pgk, tpi, and previously sequenced genes from C. glutamicum, a codon preference profile for this organism which differs significantly from those of E. coli and Bacillus subtilis is presented.

Amino Acid Sequence↗

The two beta-tubulin genes of Chlamydomonas reinhardtii code for identical proteins.

The two beta-tubulin genes of the unicellular green alga Chlamydomonas reinhardtii are expressed coordinately after deflagellation and produce two transcripts of 2.1 and 2.0 kilobases. Full-length cDNA clones corresponding to the transcript of each gene were isolated. DNA sequences were obtained from the cDNA clones and from cloned tubulin gene fragments. Both genes contained 1,332 base pairs of coding sequence, with only 19 nucleotide differences between the genes. Because all the differences occurred at the third base position of a codon and did not change the predicted amino acid sequence, we concluded that both beta-tubulin genes code for the same protein of 443 amino acids. The predicted amino acid sequence is 89 and 72% homologous with beta-tubulins from chicken and yeast cells, respectively. Each gene had three intervening sequences, which occurred at identical positions. Although the first two intervening sequences were not conserved between the two genes, the nucleotide sequence of the third intervening sequence was 89% conserved between the genes. The codon usage in the tubulin genes of C. reinhardtii was very biased: only 37 different codons were used. Striking differences occurred between the codons used in these nuclear genes and C. reinhardtii chloroplast genes.

Amino Acid Sequence↗