Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Regulation of the nuclear genes encoding the cytoplasmic and mitochondrial leucyl-tRNA synthetases of Neurospora crassa.

We show that the nuclear genes for the cytoplasmic and mitochondrial leucyl-tRNA synthetase (LeuRS) of Neurospora crassa are distinct in their encoded proteins, codon usage, mRNA levels, and regulation. The 4.2-kilobase-pair region representing the structural gene for cytoplasmic LeuRS and flanking regions has been sequenced. The positions of the 5' and 3' ends of mRNA and of a single 62-base-pair intron have been mapped. The methionine-initiated open reading frame encoded a protein of 1,123 amino acids and displayed a strong codon bias. Although cytoplasmic LeuRS shares with mitochondrial LeuRS some general features common to most aminoacyl-tRNA synthetases, there is little amino acid sequence similarity between them, mRNA levels for cytoplasmic LeuRS were much higher than those for mitochondrial LeuRS. This observation and the strong codon bias in the cytoplasmic LeuRS gene may contribute to a greater abundance of cytoplasmic LeuRS than mitochondrial LeuRS. The genes for cytoplasmic and mitochondrial LeuRS are regulated independently. The cytoplasmic LeuRS gene is regulated by the cross-pathway control system in N. crassa, which is analogous to general amino acid control in Saccharomyces cerevisiae. The cytoplasmic LeuRS mRNA levels are induced by amino acid starvation resulting from the addition of aminotriazole. Part of this increase is due to utilization of new transcription start sites. In contrast, the mitochondrial LeuRS gene is not induced by amino acid limitation. However, the mitochondrial LeuRS mRNA levels did increase dramatically upon inhibition of mitochondrial protein synthesis by chloramphenicol or ethidium bromide or in the temperature-sensitive strain leu-5 carrying a mutation in the mitochondrial LeuRS structural gene.

Amino Acid Sequence↗

Distribution of isoaccepting tRNAs and codons for proline and glycine in collagenous and noncollagenous chicken tissues.

The relation between codon usage and tRNA content for proline and glycine, the major constituents of collagen, was studied in two tissues: the magnum of laying hen oviduct and the leg tendons of chick embryo where collagen is produced. Although the relative contents of tRNA(GCCGly) and tRNA(IGGPro) in tendons, as compared to magnum indicate a specialization of the tRNA population for collagen synthesis, the distribution of the preponderant codons in collagen mRNA is correlated but at a lesser extent to that of their cognate tRNAs.

Animals↗

Nucleotide sequence of the Thiobacillus ferrooxidans chromosomal gene encoding mercuric reductase.

The nucleotide sequence of the Thiobacillus ferrooxidans chromosomal mercuric-reductase-encoding gene (merA) has been determined. The merA gene contains 1635 bp, and shares 78.2% and 76.6% sequence homology with the transposon, Tn501, and plasmid R100 merA genes, respectively. From the sequence, a 545-amino acid (aa) polypeptide was deduced, and comparison with those of Tn501 and R100 revealed 80.6% and 80.0% homology, respectively, at the aa sequence level. Divergence among the three merA aa sequences was clustered within a specific region (aa positions 41-87). By analysis of codon usage frequency, it is speculated that the T. ferrooxidans merA gene originated from Tn501, R100, or a common ancestral gene, but not from T. ferrooxidans itself.

Acidithiobacillus thiooxidans↗

Cloning and sequencing of the genes encoding the alpha and beta subunits of C-phycocyanin from the cyanobacterium Agmenellum quadruplicatum.

Synthetic oligonucleotide probes were used to identify a cloned DNA fragment from the cyanobacterium Agmenellum quadruplicatum that contains the genes for the alpha and beta subunits of C-phycocyanin. The coding region for the alpha-subunit gene begins 108 base pairs downstream from the 3' end of the beta-subunit structural gene. The sequences of the coding regions for both genes have been determined as well as 379 base pairs of 5' flanking region, 204 base pairs of 3' flanking region, and the 108 base pairs between the two genes. The site of transcriptional initiation is located approximately equal to 325 base pairs upstream from the beta-subunit gene, and an open reading frame 114 base pairs long is found within this region. The significance of this additional open reading frame is not yet known. The derived amino acid sequences for both C-phycocyanin subunits were compared with other known C-phycocyanin sequences for homology. Homologies between the A. quadruplicatum alpha subunit and alpha subunits from other species were approximately equal to 70%, as were homologies between the A. quadruplicatum beta subunit and other beta subunits. Homologies between the various alpha and beta subunits were 21%-27%. Codon usage for both the C-phycocyanin alpha- and beta-subunit genes shows asymmetries for many amino acids that correspond closely to those seen in highly expressed Escherichia coli genes.

Amino Acid Sequence↗

Human proprotein convertase 2 homologue from a plant nematode: cloning, characterization, and comparison with other species.

Proprotein convertases (PCs) are evolutionarily conserved enzymes responsible for processing the precursors of many bioactive peptides in mammals. The invertebrate homologues of PC2 play important roles during development that makes the enzyme a good target for practical applications in pest management. Screening of a plant nematode Heterodera glycines cDNA library resulted in isolation of a full-length clone encoding a PC2-like precursor. The deduced protein (74.2 kD) exhibits strong amino acid homology to all known PC2s, including human, and shares the main structural characteristics: signal peptide; prosegment; catalytic domain, with D/H/S catalytic triad, PC2-specific residues, and 7B2 binding sites; P domain (with RRGDT pentapeptide); and carboxyl terminus. Comparative analysis of PC2s from 15 species discloses the presence of an insert in the catalytic domain unique to nematodes. Expression of PC2-like mRNA found in eggs and juveniles was undetectable in adult stages of H. glycines. Nucleotide analysis reveals distinctive differences in base composition and codon usage between H. glycines and Caenorhabditis elegans PC2s. The H. glycines cDNA clone encoding PC2 is the first one isolated from plant-parasitic nematodes.

Animals↗

Aminoacyl-tRNA formation in the extreme thermophile Thermus thermophilus.

Thermophilic organisms must be capable of accurate translation at temperatures in which the individual components of the translation machinery and also specific amino acids are particularly sensitive. Thermus thermophilus is a good model organism for studies of thermophilic translation because many of the components in this process have undergone structural and biochemical characterization. We have focused on the pathways of aminoacyl-tRNA synthesis for glutamine, asparagine, proline, and cysteine. We show that the T. thermophilus prolyl-tRNA synthetase (ProRS) exhibits cysteinyl-tRNA synthetase (CysRS) activity although the organism also encodes a canonical CysRS. The ProRS requires tRNA for cysteine activation, as is known for the characterized archaeal prolyl-cysteinyl-tRNA synthetase (ProCysRS) enzymes. The heterotrimeric T. thermophilus aspartyl-tRNA(Asn) amidotransferase can form Gln-tRNA in addition to Asn-tRNA: however, a 13-amino-acid C-terminal truncation of the holoenzyme A subunit is deficient in both activities when assayed with homologous substrates. A survey of codon usage in completed prokaryotic genomes identified a higher Glu:Gln ratio in proteins of thermophiles compared to mesophiles.

Amino Acyl-tRNA Synthetases↗

Another putative heat-shock gene and aminoacyl-tRNA synthetase gene are located upstream from the grpE-like and dnaK-like genes in Chlamydia trachomatis.

The 4.1-kb sequence of genomic DNA located upstream from the Chlamydia trachomatis grpE-like and dnaK-like heat shock (HS) genes was determined. Another putative HS gene was located just 5' to grpE along with an inverted repeat (IR) sequence proposed to be involved in HS regulation. The overall organization of this locus in Chlamydia resembles that of Bacillus subtilis, rather than Escherichia coli. Two other open reading frames (ORFs) were found in the sequence, one of which has homology to aminoacyl-tRNA synthetases. The other ORF has no significant homology to reported genes. We also examined the codon usage bias for these newly identified chlamydial ORFs and for previously reported chlamydial genes, and found them to be different from E. coli.

Amino Acid Sequence↗

A comprehensive software suite for the analysis of cDNAs.

We have developed a comprehensive software suite for bioinformatics research of cDNAs; it is aimed at rapid characterization of the features of genes and the proteins they code. Methods implemented include the detection of translation initiation and termination signals, statistical analysis of codon usage, comparative study of amino acid composition, comparative modeling of the structures of product proteins, prediction of alternative splice forms, and metabolic pathway reconstruction.

Alternative Splicing↗

Adaptation of standard spreadsheet software for the analysis of DNA sequences.

This paper presents use of spreadsheet software to derive various statistics on nucleotide and codon distribution and frequency from gene sequence data, including chi 2 test and their derivatives, and proportion of nucleotides in the third position. The basic principles can be easily extended to other more complex functions. In addition, it can be used to translate a nucleic acid sequence into a protein sequence or to produce a codon usage table. This adaptation permits sequence analysis without expensive dedicated software and easy modification of the analysis according to the user's requirements. It is also ideal for introducing biological science students to the programming potential of spreadsheets.

Base Sequence↗

Codon optimization of Caenorhabditis elegans GluCl ion channel genes for mammalian cells dramatically improves expression levels.

Organisms use synonymous codons in a highly non-random fashion. These codon usage biases sometimes frustrate attempts to express high levels of exogenous genes in hosts of widely divergent species. The Caenorhabditis elegans GluClalpha1 and GluClbeta genes form a functional glutamate and ivermectin-gated chloride channel when expressed in Xenopus oocytes, but expression is weak in mammalian cells. We have constructed synthetic genes that retain the amino acid sequence of the wild-type GluCl channel proteins, but use codons that are optimal for mammalian cell expression. We have tagged the native and codon-optimized GluCl cDNAs with enhanced yellow fluorescent protein (EYFP, GluClalpha1 subunit) and enhanced cyan fluorescent protein (EFCP, GluClbeta subunit), expressed the channels in E18 rat hippocampal neurons and measured the relative expression levels of the two genes with fluorescence microscopy as well as with electrophysiology. Codon optimization provides a 6- to 9-fold increase in expression, allowing the conclusions that the ivermectin-gated channel has an EC(50) of 1.2 nM and a Hill coefficient of 1.9. We also confirm that the Y182F mutation in the codon-optimized beta subunit results in a heteromeric channel that retains the response to ivermectin while reducing the response to 100 microM glutamate by 7-fold. The engineered GluCl channel is the first codon-optimized membrane protein expressed in mammalian cells and may be useful for selectively silencing specific neuronal populations in vivo.

Animals↗

The Chlorella H+/hexose cotransporter gene.

The complete genomic sequence of the inducible Chlorella kessleri H+/hexose cotransporter (HUP1) has been obtained from two overlapping clones isolated from a lambda gt10 library. The HUP1 gene is interrupted by 14 introns with the first intron being located in the 5'-untranslated part of the gene. The average intron length is 220 bp, yielding a very regular intron/exon pattern in the gene. The codon usage in this gene is strongly biased with a clear preference for C and a strong suppression of A. A consensus sequence for a putative algal polyadenylation sequence is shown and compared with other algal cDNA sequences.

Base Sequence↗

Two different macronuclear EF-1 alpha-encoding genes of the ciliate Euplotes crassus are very dissimilar in their sequences, copy numbers and transcriptional activities.

Genes (EFA) encoding the translation elongation factor EF-1 alpha (EFA) or the prokaryotic homolog EFTu frequently occur in multiple copies in the same organism. This has been interpreted either in terms of a potential of differential gene expression during different phases of development, or as gene dosage adaptation to the need of high-level production of the gene products. Since ciliates can differentially amplify their genes, the latter argument would lead to the expectation of only one EFA gene in the macronucleus. However, we have found two such genes which strongly differ in both copy number and codon usage. Both transcripts are detectable at very different levels. The expression of the genes takes place both in the vegetative and sexual phases, i.e.,during conjugation.

Amino Acid Sequence↗

Stochastic traits of molecular evolution--acceptance of point mutations in native actin genes.

A stochastic matrix of nucleotide mutation probabilities is derived by counting differences and identities in alignments of native actin genes, with the aim of obtaining a more reliable data base for regular modes of molecular evolution. The evolution of DNA sequences is thereby considered as a Markov process consisting of events (point mutations) characterized by a stochastic matrix for codon-codon interchanges. The genetic distance is set to 1 PAM (percentage of accepted point mutations). The results can be reproduced by Monte Carlo simulations which are subjected to selective constraints. The latter are observed as nonrandom codon usage and ratios of silent to recognizable point mutations. Specific patterns within the matrix of mutation probabilities attest to preferences of natural selection in the evolution of a specific protein.

Actins↗

Effect of tandem rare codon substitution and vector-host combinations on the expression of the EBV gp110 C-terminal domain in Escherichia coli.

Gp110 of Epstein-Barr virus (EBV) is a glycoprotein that functions exclusively during the assembly of EBV nucleocapsid and the release of infectious EBV. Its C-terminal tail domain (gp110 CTD) is essential for gp110's function and may provide signals that are responsible for the assembly and release of EBV. In the present study, to get large amounts of gp110 CTD for structural analysis, the effects of vector system, codon usage, and host strain on expression levels of gp110 CTD in Escherichia coli have been investigated. The coding region of gp110 CTD (11 kDa) was subcloned into the expression vectors pSE 280, pET-15b, pET-29a, pMAL-c2x, and pGEX-4T-1. Except the pMAL-c2x construct, all the others failed to express detectable amounts of recombinant gp110 CTD. Substituting a tandem rare AGA (Arg) codon with a synonymous CGC (Arg) codon facilitated expression of the recombinant protein, while a protease-deficient host E. coli strain helped in the accumulation of a soluble form of gp110 CTD fusion. The secondary structures of the obtained recombinant gp110 CTD purified from soluble extracts and inclusion bodies were compared using circular dichroism analysis. In aqueous solutions, both samples equally adopt a mixed alpha-helix and beta-sheet conformation as well as a partly unordered structure. Notably, in the membrane-mimicking environments the helical propensity of gp110 CTD increased up to the previously predicted level based on its sequence, suggesting that gp110 CTD may fold into a more stable conformation through interactions with the cell membrane.

Circular Dichroism↗

Evolutionary analysis of the two-component systems in Pseudomonas aeruginosa PAO1.

Gene organization and functional motif analyses of the 123 two-component system (2CS) genes in Pseudomonas aeruginosa PAO1 were carried out. In addition, NJ and ML trees for the sensor kinases and the response regulators were constructed, and the distances measured and comparatively analyzed. It was apparent that more than half of the sensor-regulator gene pairs, especially the 2CSs with OmpR-like regulators, are derivatives of a common ancestor and have most likely co-evolved through gene pair duplication. Several of the 2CS pairs, especially those with NarL-like regulators, however, appeared to be relatively divergent. This is supportive of the recruitment model, in which a sensor gene and regulator gene with different phylogenetic history are assembled to form a 2CS. Correlation of the classification of sensor kinases and response regulators provides further support for these models. Upon comparison of the phylogenetic trees comprised of sensors and regulators, we have identified six congruent clades, which represent the group of the most recently duplicated 2CS gene pairs. Analyses of the congruent 2CS pairs of each of the clades revealed that certain paralogous 2CS pairs may carry a redundant function even after a gene duplication event. Nevertheless, comparative analysis of the putative promoter regions of the paralogs suggested that functional redundancy could be prevented by a differential control. Both codon usage and G+C content of these 2CS genes were found to be comparable with those of the P. aeruginosa genome, suggesting that they are not newly acquired genes.

Base Composition↗

Sequence analysis of the lysin gene region of the prolate lactococcal bacteriophage c2.

Approximately 80% of the genome of the prolate-headed lactococcal bacteriophage c2 was cloned into shuttle vectors pSA3 and pFX3 in Escherichia coli and transferred to Lactococcus lactis. A 1.67-kilobase EcoRV fragment containing the gene for the phage lysin was identified and the position and orientation of the phage lysin gene in the physical map of the phage were determined. The phage lysin was expressed in E. coli and its sequence was determined and compared with the sequences of other bacteriophage lytic genes. The sequence was similar, but not identical, to that of the related lactococcal phage m13, having a number of silent substitutions and an apparent deletion that altered the carboxy terminus of the protein. Possible alternative translation initiation codons for the lysin gene and two possible alternative mechanisms for access of the lysin enzyme to the cell wall are discussed. An open reading frame upstream of the putative lysin gene was found to be 177 base pairs longer than that reported for phage m13. A codon usage table for the lysin genes of several phages as well as for reported gene sequences from L. lactis and lactococcal bacteriophages is presented.

Amino Acid Sequence↗

The nucleotide sequence of human rhinovirus 1B: molecular relationships within the rhinovirus genus.

We have determined the complete nucleotide sequence of human rhinovirus 1B and made comparisons with other rhinoviruses. Extensive homology was found with serotypes 2 and 89 but the similarity to serotype 14 was considerably less. Rhinovirus-specific characteristics have been noted, in particular the length of the 5' non-coding region and the pattern of codon usage, and these may be sufficient to define the rhinoviruses as a distinct genus rather than being considered as members of the enteroviruses as has been suggested previously.

Amino Acid Sequence↗

Nucleotide sequence of the tcml gene (ribosomal protein L3) of Saccharomyces cerevisiae.

The yeast tcml gene, which codes for ribosomal protein L3, has been isolated by using recombinant DNA and genetic complementation. The DNA fragment carrying this gene has been subcloned and we have determined its DNA sequence. The 20 amino acid residues at the amino terminus as inferred from the nucleotide sequence agreed exactly with the amino acid sequence data. The amino acid composition of the encoded protein agreed with that determined for purified ribosomal protein L3. Codon usage in the tcml gene was strongly biased in the direction found for several other abundant Saccharomyces cerevisiae proteins. The tcml gene has no introns, which appears to be atypical of ribosomal protein structural genes.

Amino Acid Sequence↗