Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Isochore evolution in mammals: a human-like ancestral structure.

Codon usage in mammals is mainly determined by the spatial arrangement of genomic G + C-content, i.e., the isochore structure. Ancestral G + C-content at third codon positions of 27 nuclear protein-coding genes of eutherian mammals was estimated by maximum-likelihood analysis on the basis of a nonhomogeneous DNA substitution model, accounting for variable base compositions among present-day sequences. Data consistently supported a human-like ancestral pattern, i.e., highly variable G + C-content among genes. The mouse genomic structure-more narrow G + C-content distribution-would be a derived state. The circumstances of isochore evolution are discussed with respect to this result. A possible relationship between G + C-content homogenization in murid genomes and high mutation rate is proposed, consistent with the negative selection hypothesis for isochore maintenance in mammals.

Animals↗

Analysis of the primary structure and promoter function of a pyruvate decarboxylase gene (PDC1) from Saccharomyces cerevisiae.

The PDC1 gene of Saccharomyces cerevisiae, encoding pyruvate decarboxylase was sequenced. The gene contains an open reading frame of 1647 base pairs. The codon usage shows the same strong bias as found for some other glycolytic enzymes. Transcription starts mainly at -30 and terminates 100 base pairs downstream of the termination codon. In some strains a second termination site, 46 base pairs upstream of the stop codon was observed. The function of the promoter region was analyzed by fusion to the bacterial structural gene encoding beta-lactamase (bla). On multicopy plasmid or integrated in the genome, the expression of the bla gene showed the regulation of the authentic PDC1 gene.

Amino Acid Sequence↗

Codon frequencies in 119 individual genes confirm consistent choices of degenerate bases according to genome type.

The poor printing of our previous Figure 2 (1) is corrected. Codon usage in mRNA sequences just published is also given. A new correspondence analysis is done, based on simultaneous comparison in all mRNA of use of the 61 codons. This analysis reinforces our claim that most genes in a genome, or genome type, have the same coding strategy; that is, they show similar choices among synonymous codons, or among degenerate bases (2). Like analysis on frequency variation in the amino acids coded reveals an entirely different pattern.

Animals↗

A codon-based model of nucleotide substitution for protein-coding DNA sequences.

A codon-based model for the evolution of protein-coding DNA sequences is presented for use in phylogenetic estimation. A Markov process is used to describe substitutions between codons. Transition/transversion rate bias and codon usage bias are allowed in the model, and selective restraints at the protein level are accommodated using physicochemical distances between the amino acids coded for by the codons. Analyses of two data sets suggest that the new codon-based model can provide a better fit to data than can nucleotide-based models and can produce more reliable estimates of certain biologically important measures such as the transition/transversion rate ratio and the synonymous/nonsynonymous substitution rate ratio.

Animals↗

Origin of transfer of IncF plasmids and nucleotide sequences of the type II oriT, traM, and traY alleles from ColB4-K98 and the type IV traY allele from R100-1.

The complete nucleotide sequences of the ColB4-K98 (ColB4) plasmid transfer genes oriT, traM, and traY as well as the traY gene of R100-1 are presented and compared with the corresponding regions from the conjugative plasmids F, R1, and R100. The sequence encoding the oriT nick sites and surrounding inverted repeats identified in F was conserved in ColB4. The adenine-thymine-rich sequence following these nick sites was conserved in R1 and ColB4 but differed in F and R100, indicating that this region may serve as the recognition site for the traY protein. A series of direct repeats unique to the ColB4 plasmid was found in the region of dyad symmetry following this AT-rich region. This area also encodes 21-base-pair direct repeats which are homologous to those in F and R100. The traM gene product may bind in this region. Overlapping and following these repeats is the promoter(s) for the traM protein. The traM protein from ColB4 is similar to the equivalent products from F, R1, and R100. The traY protein from ColB4 is highly homologous to the R1 traY gene product, while the predicted R100-1 traY product differs at several positions. These differences presumably define the different alleles of traM and traY previously identified for IncF plasmids by genetic criteria. The translational start codons of the ColB4 and R100-1 traY genes are GUG and UUG, respectively, two examples of rare initiator codon usage.

Alleles↗

Low diversity and divergence in the fil1 gene family of Antirrhinum (Scrophulariaceae).

Detailed nucleotide diversity studies revealed that the fil1 gene of Antirrhinum, which has been reported to be single copy, is a member of a gene family composed of at least five genes. In four Antirrhinum majus populations with different mating systems and one A. graniticum population, diversity within populations is very low. Divergence among Antirrhinum species and between Antirrhinum and Digitalis is also low. For three of these genes we also obtained sequences from a more divergent member of the Scrophulariaceae, Verbascum nigrum. Compared with Antirrhinum, little divergence is again observed. These results, together with similar data obtained previously for five cycloidea genes, suggest either that these gene families (or the Antirrhinum genome) are unusually constrained or that there is a low rate of substitution in these lineages. Using a sample of 52 genes, based on two measures of codon usage (ENC and GC3 content), we show that cyc and fil1 are among the least biased Antirrhinum genes, so that their low diversity is not due to extreme codon bias.

Arabidopsis Proteins↗

The mitochondrial genome of Strongyloides stercoralis (Nematoda) - idiosyncratic gene order and evolutionary implications.

The complete mitochondrial genome sequence of the parasitic nematode Strongyloides stercoralis was determined, and its organisation and structure compared with other nematodes for which complete mitochondrial sequence data were available. The mitochondrial genome of S. stercoralis is 13,758 bp in size and contains 36 genes (all transcribed in the clockwise direction) but lacks the atp8 gene. This genome has a high T content (55.9%) and a low C content (8.3%). Corresponding to this T content, there are 16 (poly-T) tracts of >/=12 Ts distributed across the genome. In protein-coding genes, the T bias is greatest (76.4%) at the third codon position compared with the first and second codon positions. Also, the C content is higher at the first (9.3%) and second (13.4%) codon positions than at the third (2%) position. These nucleotide biases have a significant effect on predicted codon usage patterns and, hence, on amino acid compositions of the mitochondrial proteins. Interestingly, six of the 12 protein-coding genes are predicted to employ a unique initiation codon (TTT), which has not yet been reported for any other animal mitochondrial genome. The secondary structures predicted for the 22 transfer RNA (trn) genes and the two ribosomal RNA (rrn) genes are similar to those of other nematodes. In contrast, the gene arrangement in the mitochondrial genome of S. stercoralis is different from all other nematodes studied to date, revealing only a limited number of shared gene boundaries (atp6-nad2 and cox2-rrnL). Evolutionary analyses of mitochondrial nucleotide and amino acid sequence data sets for S. stercoralis and seven other nematodes demonstrate that the mitochondrial genome provides a rich source of phylogenetically informative characters. In conclusion, the S. stercoralis mitochondrial genome, with its unique gene order and characteristics, should provide a resource for comparative mitochondrial genomics and systematics studies of parasitic nematodes.

Amino Acid Sequence↗

Similarity in oligonucleotide usage in introns and intergenic regions contributes to long-range correlation in the Caenorhabditis elegans genome.

A method is presented which allows detection of a sequence correlation effect not related to patchiness in base composition or to preferences in codon usage. Recurrence plots providing local views of oligonucleotide recurrence regimen show that introns and intergenic regions are often characterised by a highly recurrent use of oligonucleotides. By window analysis it is possible to score a long sequence for the recurrence of a given subset of oligos while filtering away the effects of short-range correlations. Long-range exploration of chromosome III from Caenorhabditis elegans reveals that consistent use of recurrent oligonucleotides in introns and intergenic regions generates a correlation effect that extends over several megabases.

Animals↗

A codon-based model designed to describe lentiviral evolution.

A codon-based model designed to describe lentiviral evolution is developed. The model incorporates unequal base compositions in the three codon positions and selection against the CpG dinucleotide within codons to account for a deficit of this dinucleotide exhibited by lentiviral genes. The model is, to a large extent, able to account for the pattern of codon usage exhibited by the HIV1 genes gag, pol, and env, in spite of its parameter paucity. The model is extended to a similar model which operates on pentets (codons and their neighboring bases). The results obtained by the pentet model establish the importance of depression of CpGs across codon boundaries as well as within codons. The goodness of fit of the CpG depression model to the observed evolution in pairwise alignments of HIV1 sequences is assessed. The model provides a significantly better description of the observed evolution than the simpler models examined. The parameter estimates indicate that part of the unusually large biases in nucleotide frequencies observed in HIV1 genes is caused by selection against CpGs. We find that the estimates of expected numbers of substitutions, of transitions to transversions, and of synonymous to nonsynonymous substitution rates are robust to CpG depression, whereas the ratio of CpG-generating substitutions to other substitutions is strongly influenced by the choice of model.

Base Composition↗

Cloning of the gene for inorganic pyrophosphatase from a thermoacidophilic archaeon, Sulfolobus sp. strain 7, and overproduction of the enzyme by coexpression of tRNA for arginine rare codon.

The gene encoding an extremely stable inorganic pyrophosphatase from Sulfolobus sp. strain 7, a thermoacidophilic archaeon, was cloned and sequenced. An open reading frame consisted of 516 base pairs coding for a protein of 172-amino acid residues. The deduced sequence was supported by partial amino acid sequence analyses. All the catalytically important residues were conserved. A unique 17-base-pair sequence motif was found to be repeated four times in frame in the gene, encoding a cluster of acidic amino acids essential for the function. Although the codon usage of the gene was quite different from that of Escherichia coli, the gene was effectively expressed in E. coli. Coexpression of tRNA(Arg), cognate for the rare codon AGA in E. coli, however, further improved the production of the enzyme, which occupied more than 85% of the soluble proteins obtained after removal of heat denatured E. coli proteins.

Amino Acid Sequence↗

Comparative studies on the structural gene for the ribosomal protein S1 in ten bacterial species.

By applying the Southern blot technique we compared the structural gene rpsA for ribosomal protein S1 and its preceding sequence from Escherichia coli with nine other bacterial species. We found high homology among the structural genes of E. coli and other gram-negative but not gram-positive bacteria. In contrast, the regulatory sequence preceding the structural gene was not highly conserved among the organisms studied. Cloning and DNA sequence analysis of a 1.2 kb fragment coding for most of the structural gene for S1 from Providencia localized some strongly conserved parts of the DNA sequence, despite the fact that the codon usage showed considerable divergence from that of E. coli.

Amino Acid Sequence↗

Endothelin-1 mRNA expression in the rat kidney.

Cultured pig and bovine endothelial cells are capable of synthesizing endothelin-1 (ET-1). Thus the observation that the kidney contains a large number of binding sites for ET distributed in close proximity to endothelial cells suggests that ET-1 may be released from the endothelium to act locally on these receptors. In support of this hypothesis, using the technique of reverse transcription with specific amplification of cDNA, we report here that ET-1 mRNA is expressed in the rat kidney. The partial sequence of the amplified rat ET-1 cDNA confirms that the mature rat peptide is identical to that of the mouse, man and pig, but with some differences in codon usage.

Animals↗

Long-term evolution and functional diversification in the members of the nucleophosmin/nucleoplasmin family of nuclear chaperones.

The proper assembly of basic proteins with nucleic acids is a reaction that must be facilitated to prevent protein aggregation and formation of nonspecific nucleoprotein complexes. The proteins that mediate this orderly protein assembly are generally termed molecular (or nuclear) chaperones. The nucleophosmin/nucleoplasmin (NPM) family of molecular chaperones encompasses members ubiquitously expressed in many somatic tissues (NPM1 and -3) or specific to oocytes and eggs (NPM2). The study of this family of molecular chaperones has experienced a renewed interest in the past few years. However, there is a lack of information regarding the molecular evolution of these proteins. This work represents the first attempt to characterize the long-term evolution followed by the members of this family. Our analysis shows that there is extensive silent divergence at the nucleotide level suggesting that this family has been subject to strong purifying selection at the protein level. In contrast to NPM1 and NPM-like proteins in invertebrates, NPM2 and NPM3 have a polyphyletic origin. Furthermore, the presence of selection for high frequencies of acidic residues as well as the existence of higher levels of codon bias was detected at the C-terminal ends, which can be ascribed to the critical role played by these residues in constituting the acidic tracts and to the preferred codon usage for phosphorylatable amino acids at these regions.

Animals↗

HIV1 integrase expressed in Escherichia coli from a synthetic gene.

Human immunodeficiency virus type 1 (HIV1) integrase is cleaved from the gag-pol precursor by the HIV1 protease. The resulting 32-kDa protein is used by the infecting virus to insert a linear, double-stranded DNA copy of its genome, prepared by reverse transcription of viral RNA, into the host cell's chromosomal DNA. In order to achieve high levels of expression, to minimize an internal initiation problem and to facilitate mutagenesis, we have designed and synthesized a gene encoding the integrase from the infectious molecular clone, pNL4-3. Codon usage was optimized for expression in Escherichia coli and unique restriction sites were incorporated throughout the gene. A 905-bp cassette containing a ribosome-binding site, a start codon and the integrase-coding sequence, sandwiched between EcoRI and HindIII sites, was synthesized by overlap extension of nine long synthetic oligodeoxyribonucleotides [90-120 nucleotides (nt)] and subsequent amplification using two primers (28-30 nt). The cassette was subcloned into the vector pKK223-3 for expression under control of a tac promoter. The protein produced from this highly expressed gene has the expected N-terminal sequence and molecular mass, and displays the DNA processing, DNA joining and disintegration activities expected from recombinant integrase. These studies have demonstrated the utility of codon optimization, and lay the groundwork for structure-function studies of HIV1 integrase.

Amino Acid Sequence↗

Generation of 10,154 expressed sequence tags from a leafy gametophyte of a marine red alga, Porphyra yezoensis.

A total of 10,154 5'-end expressed sequence tags (EST) were established from the normalized and size-selected cDNA libraries of a marine red alga, Porphyra yezoensis. Among the ESTs, 2140 were unique species, and the remaining 8014 were grouped into 1127 species. Database search of the 3267 non-redundant ESTs by BLAST algorithm showed that the sequences of 1080 species (33.1%) have similarity to those of registered genes from various organisms including higher plants, mammals, yeasts, and cyanobacteria, while 2187 (66.9%) are novel. Codon usage analysis in the coding regions of 101 non-redundant EST groups showing significant similarity to known genes indicated the higher GC contents at the third position of codons (79.4%) than the first (62.2%) and the second position (45.0%), suggesting that the genome has been exposed to high GC pressure during evolution. The sequence data of individual ESTs are available at the web site http://www.kazusa.or.jp/en/plant/porphyra/EST/.

Algorithms↗

Horizontal transfer of a phosphatase gene as evidence for mosaic structure of the Salmonella genome.

The genomes of Escherichia coli and Salmonella typhimurium are similar with respect to base composition, chromosome size, and the order, orientation and spacing of genes, but differ with respect to some 29 'loops', regions unique to one species. To evaluate the genetic basis for the structure and organization of the enteric bacterial genomes, we examined the gene encoding a non-specific acid phosphatase (phoN) which maps to a loop at 96 min on the S.typhimurium chromosome. We detected atypical base composition, codon usage pattern and trinucleotide frequencies. The 1.4 kb region containing phoN had an overall base composition of 43% G+C, while the G+C content at the third positions of codons in the phoN reading frame is only 39%, much lower than the Salmonella chromosome which averages 52%. Non-specific acid phosphatase activity, assayed in 14 Gram-negative species, was detected only in Morganella morganii and Providencia stuartii, organisms with low genomic G+C contents. Upstream of the phoN gene in Salmonella is a sequence with high similarity to the oriT region of incFII plasmids, suggesting that the phoN gene, and perhaps the entire loop structure, was acquired by lateral transmission in a plasmid-mediated event.

Acid Phosphatase↗

Codon context.

The analysis of coding sequences reveals nonrandomness in the context of both sense and stop codons. Part of this is related to nucleotide doublet preference, seen also in non-coding sequences and thought to arise from the dependence of mutational events on surrounding sequence. Another nonrandom context element, relating the wobble nucleotides of successive codons, is observed even when doublet preference, codon usage and bias in amino acid doublets are all allowed for. Several phenomena related to protein synthesis have been shown in vivo to be affected by the nucleotide sequence around codons. Thus, nonsense and missense suppression, elongation rate, precision of tRNA selection and polypeptide chain termination are all affected by codon context. At present, it remains unclear how these phenomena may influence the evolution of nonrandomness in the context of codons in natural sequences.

Codon↗

RNA folding is unaffected by the nonrandom degenerate codon choice.

The frequent suggestion that the nonrandom codon usage is explained by its forming more stable mRNAs is tested in 22 genes. Only the histones, globins, and the rat preproinsulin gene show a correlation between the preferred degenerate codons and the stability of the secondary structure of the their mRNAs. However, the examined members from the histone and globin gene families, both among the oldest, in evolutionary sense, eukaryotic genes, have a high GC content (approx. 56% compared to an average of 42% in all eukaryotes) which is reflected in their degenerate codon choice and thus in their more stable folding.

Animals↗