Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

Post-transcriptional regulation of the steady-state levels of mitochondrial tRNAs in HeLa cells.

In human mitochondrial DNA (mtDNA), the tRNA genes are located in three different transcription units that are transcribed at three different rates. To analyze the regulation of tRNA formation by the three transcription units, we have examined the steady-state levels and metabolic properties of the tRNAs of HeLa cell mitochondria. DNA excess hybridization experiments utilizing separated strands of mtDNA and purified tRNA samples from exponential cells long term labeled with [32P]orthophosphate have revealed a steady-state level of 6 x 10(5) tRNA molecules/cell, with three-fourths being encoded in the H-strand and one-fourth in the L-strand. Hybridization of the tRNAs with a panel of M13 clones of human mtDNA containing, in most cases, single tRNA genes and a quantitation of two-dimensional electrophoretic fractionations of the tRNAs have shown that the steady-state levels of tRNA(Phe) and tRNA(Val) are two to three times higher than the average level of the other H-strand-encoded tRNAs and three to four times higher than the average level of the L-strand-encoded tRNAs. Similar experiments carried out with tRNAs isolated from cells labeled with very short pulses of [5-3H]uridine have indicated that the rates of formation of the individual tRNA species are proportional to their steady-state amounts. Therefore, the approximately 25-fold higher rate of transcription of the tRNA(Phe) and tRNA(Val) genes relative to the other H-strand tRNA genes and the 10-16-fold higher rate of transcription of the L-strand tRNA genes relative to the H-strand tRNA genes are not reflected in the steady-state levels or the rates of formation of the corresponding tRNAs. A comparison of the steady-state levels of the individual tRNAs with the corresponding codon usage for protein synthesis, as determined from the DNA sequence and the rates of synthesis of the various polypeptides, has not revealed any significant correlation between the two parameters.

Cloning, Molecular↗

A functional significance for codon third bases.

Most amino acids are specified by more than one trinucleotide codon. Here we show that amino acids of differing functional importance may be distinguished by the pattern of synonymous codon usage. GC-rich genes tend to be of a greater transcriptional (p<0.01) and mitogenic (p<0.0001) significance than AT-rich genes, consistent with GC-->AT mutational drift in methylated genomic regions. Third-base GC retention also identifies critical amino acids within individual proteins, as indicated by non-random patterns of codon variation between gene homologs and also by differential sequelae of site-directed mutagenesis. Sequence analysis of human receptor tyrosine kinase genes confirms that functionally important transmembrane hydrophobic amino acids are specified by codons containing GC third bases more often than are transmembrane neutral amino acids (chi(2)=134.2). Amino acids encoded by GC third bases thus appear more tightly linked to cell function and survival than are those encoded by AT third bases.

Amino Acids↗

The A-superfamily of conotoxins: structural and functional divergence.

The generation of functional novelty in proteins encoded by a gene superfamily is seldom well documented. In this report, we define the A-conotoxin superfamily, which is widely expressed in venoms of the predatory cone snails (Conus), and show how gene products that diverge considerably in structure and function have arisen within the same superfamily. A cDNA clone encoding alpha-conotoxin GI, the first conotoxin characterized, provided initial data that identified the A-superfamily. Conotoxin precursors in the A-superfamily were identified from six Conus species: most (11/16) encoded alpha-conotoxins, but some (5/16) belong to a family of excitatory peptides, the kappaA-conotoxins that target voltage-gated ion channels. alpha-Conotoxins are two-disulfide-bridged nicotinic antagonists, 13-19 amino acids in length; kappaA-conotoxins are larger (31-36 amino acids) with three disulfide bridges. Purification and biochemical characterization of one peptide, kappaA-conotoxin MIVA is reported; five of the other predicted conotoxins were previously venom-purified. A comparative analysis of conotoxins purified from venom, and their precursors reveal novel post-translational processing, as well as mutational events leading to polymorphism. Patterns of sequence divergence and Cys codon usage define the major superfamily branches and suggest how these separate branches arose.

Amino Acid Sequence↗

Codon optimization effect on translational efficiency of DNA vaccine in mammalian cells: analysis of plasmid DNA encoding a CTL epitope derived from microorganisms.

Interspecific difference of codon usage is one of the major obstacles for effective induction of specific immune responses against bacteria and protozoa by DNA immunization. Using genes encoding major histocompatibility complex class I-restricted cytotoxic T-lymphocyte (CTL) epitopes, derived from an intracellular bacterium, Listeria monocytogenes and a mouse malaria parasite, Plasmodium yoelii, we report here that the codon optimization level of the genes is not precisely proportional to, but does correlate well with the translational efficiency in mammalian cells, which is concomitantly associated with the induction level of specific CTL response in the mouse. These results suggest that DNA immunization using the gene codon-optimized to mammals through the entire region is very effective.

Animals↗

Preferential codons enhancing the expression level of human beta-defensin-2 in recombinant Escherichia coli.

Human beta-defensin-2 (hBD2) is a small antimicrobial peptide with potential as a therapeutic agent. The effect of codon usage on the expression of hBD2 in Escherichia coli was studied. Two coding sequences encoding the same hBD2 precursor were both expressed as fusion protein with thioredoxin in E. coli BL21 (DE3). One is the wild-type human cDNA and the other is a gene synthesized by a PCR-based method in which rare codons were altered to those frequently used in E. coli. The expression level of recombinant hBD2 was over 50% of the total cellular protein when the synthetic gene with preferential codons was employed which was a 9-fold enhancement over the wild-type cDNA. The result shows the codon bias of the host was a major barrier in high-level expression of recombinant hBD2 and suggests a similar approach may be used in the expression of other defensins in E. coli.

Amino Acid Sequence↗

Structural and thermodynamic properties of DNA uncover different evolutionary histories.

We propose an index of DNA homogeneity (IDH) based on a binary distribution model that quantifies structural and thermodynamic aggregates present in DNA primary structures. Extensive analysis of sequence databases with the IDH uncovers significant constraints on DNA sequence other than those derived from codon usage or protein function. This index clearly distinguishes between organisms of different evolutive origins and places them in disjoint domains of DNA sequence space.

Biological Evolution↗

Strand asymmetry patterns in trypanosomatid parasites.

The genome organization of kinetoplastid parasites is unusual, with chromosomes containing several long regions of polycistronically transcribed genes. The regions where the direction of transcription switches have been hypothesized to contain origins of replication and possibly also centromers and promoters. We report that overall strand asymmetry patterns can be observed in Trypanosoma cruzi and Trypanosoma brucei with optima on strand-switch regions. The base skews of T. cruzi and T. brucei divergent strand-switches show patterns analogous to those for bacterial origins of replication, but they differ from those of Leishmania major. Bias in codon usage and the trypanosomatid unidirectional gene clusters predict most of this skew, but fail to properly explain the same trend in intergenic regions, as does the current knowledge of regulatory sequences.

Animals↗

An experimental validation of orphan genes of Buchnera, a symbiont of aphids.

Although Buchnera sp. APS, an intracellular symbiont of pea aphids, is a close relative of Escherichia coli, its genome has been extensively modified because of its prolonged intracellular life. In our previous studies on the Buchnera genome, computer analysis predicted three "orphan" genes, yba2, yba3, and yba4, which are open reading frames (ORFs) with no homologs in the database. In this paper, we successfully validated all these orphan genes by RT-PCR and Northern hybridization. The present study also revealed that yba3 and yba4 formed an operon, suggesting that they function in concert. Sequences around transcriptional start sites suggests that these genes are under the control of sigma 70. In view of codon usage and AT bias observed in these genes, it is likely that Buchnera have maintained them for an evolutionarily long time.

Amino Acid Sequence↗

Primary structure of Escherichia coli ribosomal protein S1 and of its gene rpsA.

The primary structure of proteins S1, the largest protein component of the Escherichia coli ribosome, has been elucidated by determining the amino acid sequence of the protein (from E. coli MRE600) and the nucleotide sequence of the S1 gene (rpsA, of a K-12 strain). The two methods gave results in perfect agreement except of two positions where possible strain specific differences were found. Protein S1 (MRE600) is composed of 557 amino acid residues (no modified amino acids were detected) and has Mr 61,159. The DNA sequence for protein S1 (K-12) suggests 556 amino acid residues. A computer survey of the sequence revealed three regions in S1 with a high degree of internal homology. The ribosome binding domain of S1 (NH2 terminus) does not show any preponderance of basic amino acids. The two cysteine and the majority of tryptophan residues of S1 as well as two od the three homologous regions were located in its middle region which contains the nucleic acid binding domain. The pattern of degenerate codon usage in the S1 gene is nonrandom and similar to that reported for other ribosomal protein genes.

Amino Acid Sequence↗

Structure of the type II collagen gene.

In summary, the exon/intron structure of the chicken type II collagen gene is identical with that of the chicken alpha 2(I) collagen gene and differs at only one known position from the human and mouse alpha 1(I) genes. However, the chicken type II gene is different from the chicken alpha 2(I) gene in that it is considerably shorter because of a much smaller average intron size and in that the G+C composition of the introns is much higher. The codon usage of the type II genes also shows characteristic differences. There is a single copy of the chick type II gene per haploid genome.

Animals↗

Organization of the agropine synthesis region of the T-DNA of the Ri plasmid from Agrobacterium rhizogenes.

The agropine/mannopine synthesis region of the TR region of the Ri plasmid of Agrobacterium rhizogenes strain A4 was localized on the basis of sequence similarity with probes from Ti plasmids of Agrobacterium tumefaciens and analysis of transposon insertions. The nucleotide sequence of the right part of the TR-DNA of pRiA4, encompassing the three genes involved in mannityl-opine synthesis, was determined and compared to the sequence of the corresponding region of the octopine-type Ti plasmid pTi15955. The organization of this region is strongly conserved between Ri and Ti plasmids, but the similarity is restricted to the coding sequences: no homology was detected in the 5' and 3' flanking sequences. The mas1' and ags proteins are the most conserved, showing more than 68% amino acid conservation, whereas the mas2' proteins are only 59% identical. Significant G/C content and codon usage differences are observed between pTi15955 and pRiA4. An open reading frame strongly similar to that of bacterial repressors is situated immediately to the right of the TR region.

Amino Acid Sequence↗

Amino acids runs and genomic compositional biases in vertebrates.

A compositional analysis of a sample of 50 zebrafish proteins containing at least one alanine run and of their open reading frames (ORFs) has been performed. The sample of poly(Ala) proteins showed a tendency to have runs of other amino acids (His/H, Gln/Q, Ser/S, Pro/P). Their ORFs and the first and second codon positions had higher GC contents than a reference gene set. The "universal" correlation between the GC content of the first+second and third codon positions (GC1+2 vs GC3) does not hold, but I provide an explanation in terms of genomic heterogeneity. Significant correlation between AHQS content and GC3 was obtained, reflecting codon bias favoring G/C at the third codon position of these amino acids. A correspondence analysis (COA) of relative synonymous codon usage showed that the poly(Ala) proteins have a biased distribution according to the second axis of the COA, which correlates with gene expression in zebrafish. A comparison with human is undertaken.

Amino Acids↗

Cloning, nucleotide sequence and expression in Streptomyces lividans and Escherichia coli of pabB from Lactococcus lactis subsp. lactis NCDO 496.

A gene (pabB) encoding the aminase activity of p-aminobenzoate (PABA) synthase in Lactococcus lactis subsp. lactis was cloned in pIJ41 and expressed in Streptomyces lividans strains defective in PABA biosynthesis. Expression of the gene was associated with a 1.2 kb deletion between the aph promoter and the cloning site in pIJ41. Subcloning in pBR322 and expression in Escherichia coli AB3295 of the cloned L. lactis DNA fragment localized the pabB-complementing gene in a 1.9 kb segment. The nucleotide sequence of this segment contained a 1410 bp open reading frame encoding a 470-amino-acid polypeptide of 50937 Da. The deduced amino acid sequence showed substantial similarity to those reported for PabB and TrpE from several organisms. Synonymous codon usage reflected the low G + C content in the genomic DNA of L. lactis subsp. lactis, and therefore differed markedly from the preferred usage in the S. lividans host. The cloned heterologous pabB DNA was expressed in amounts that allowed accumulation of excreted PABA in cultures of S. lividans transformants.

Amino Acid Sequence↗

Cloning and analysis of the Neurospora crassa gene for cytochrome c heme lyase.

The cyt-2-1 mutant of Neurospora crassa is deficient in cytochromes aa3 and c and in cytochrome c heme lyase activity (Mitchell, M.B., Mitchell, H.K., and Tissieres, A. (1953) Proc. Natl. Acad. Sci. U.S.A. 39, 606-613; Nargang, F.E., Drygas, M.E., Kwong, P.L., Nicholson, D.W., and Neupert, W. (1988) J. Biol. Chem. 263, 9388-9394). By rescue of the slow growth character of the cyt-2-1 mutant, we have cloned the cyt-2+ gene from a N. crassa genomic library using sib selection. Analysis of the DNA sequence of the cyt-2+ gene revealed an open reading frame of 346 amino acids that has homology to the yeast cytochrome c heme lyase. The open reading frame is interrupted by two short introns. Codon usage and Northern hybridization analysis suggest that the cyt-2 gene is expressed at low levels. The cyt-2-1 mutant allele was cloned from a partial cyt-2-1 gene bank using the wild-type gene as a probe. Sequence analysis of the mutant gene revealed a 2-base (CT) deletion that alters the reading frame for 21 codons before generating an early stop codon in the protein-coding sequence. It was previously suggested that the cyt-2-1 mutation inactivates one of two regulatory circuits controlling the production of cytochrome aa3. The finding that the cyt-2-1 mutation affects the coding sequence for cytochrome c heme lyase provides a direct explanation for the deficiency of cytochrome c in the mutant and suggests that the lack of cytochrome aa3 is a regulatory response to the deficiency of cytochrome c.

Amino Acid Sequence↗

R1 and R2 retrotransposable elements of Drosophila evolve at rates similar to those of nuclear genes.

The non-long-terminal repeat retrotransposable elements, R1 and R2, insert at unique locations in the 28S ribosomal RNA genes of insects. Based on the nucleotide sequences of these elements in the eight members of the melanogaster species subgroup of the genus Drosophila, they have been maintained by vertical germline transmission for the 17-20 million year history of this subgroup. The stable inheritance of R1 and R2 within these species has enabled a determination of their nucleotide substitution rates. The sequence of the R1 and R2 elements from D. ambigua, a member of the obscura species group, has also been determined to enable an extrapolation of this rate over an estimated 45-60 million years. The mean rate of substitutions at synonymous sites (Ks) was 6.6 and 9.6 times the rate at replacement sites (Ka) in the R1 and R2 elements, respectively. Both elements appear to have been under selective pressure to maintain their open reading frames and thus their ability to retrotranspose for most of their evolution in these lineages. Using the rate of change at synonymous sites (Ks) as the best indicator of the nucleotide substitution rate, the mean Ks values for R1 and R2 were 2.3 and 2.2 times that of the alcohol dehydrogenase (Adh) genes. However, this faster rate is a result of the lower codon usage bias of R1 and R2 compared with that of Adh. When the Ks rates of R1 and R2 were compared with that of a larger number of nuclear genes available from at least two of the nine species under investigation, R1 and R2 were found to evolve in most lineages at rates similar to that of nuclear genes with low codon bias. The ability of R1 and R2 to maintain their presence in this species subgroup by retrotransposition while exhibiting rates of nucleotide evolution similar to nuclear genes suggests these transposition events are rare or not as error prone as that of retroviruses.

Amino Acid Sequence↗

Organization of the mitochondrial genome of Atlantic cod, Gadus morhua.

The mitochondrial DNA (mtDNA) from the Atlantic cod, Gadus morhua, was mapped using 11 different restriction enzymes and cloned into plasmid vectors. Sequence data obtained from more than 10 kilobases of cod mtDNA show that the genome organization, genetic code, and the overall codon usage have been conserved throughout the evolution of vertebrates. Comparison of the derived amino acid sequences of proteins encoded by cod mtDNA to the ones encoded by Xenopus laevis mtDNA revealed that the amino acid identity range from 46% to 93% for the different proteins. ND4L is most divergent while COI is most conserved. GUG was found as the translation initiation codon of the COI gene, indicating a dual coding function for this codon. The sequences of the 997 base pair displacement-loop (D-loop)-containing region and the origin of L-strand replication (oriL), are presented. Only few of the primary and secondary structure features found to be conserved among mammalian mitochondrial D-loops, can be identified in cod. Presence of CSB-2 in the D-loop-containing region and the conserved hairpin structure at oriL, indicates that replication of bony fish mtDNA may follow the same general scheme as described for higher vertebrates.

Amino Acid Sequence↗

A characterization of the H3 and H4 histone genes from the ascidian Styela plicata.

Information about histone sequences and histone gene organization from invertebrate chordates is completely lacking. A genomic clone containing linked H3 and H4 histone genes from the urochordate Styela plicata was analyzed. The nucleotide sequence indicates that the two genes are transcribed in opposite directions and have structural similarities to cell cycle-dependent histones. Unique amino acid replacements occur in H3 at positions 88 (serine) and 98 (arginine). In H4, the methionine at position 84 is replaced by a leucine, a replacement found only in fungi. Codon usage patterns are nonrandom and resemble invertebrate patterns. The H3 and H4 genes are polymorphic and are represented five times per haploid genome.

Amino Acid Sequence↗

Organisation of the biosynthetic gene cluster for rapamycin in Streptomyces hygroscopicus: analysis of genes flanking the polyketide synthase.

Analysis of the gene cluster from Streptomyces hygroscopicus that governs the biosynthesis of the polyketide immuno-suppressant rapamycin (Rp) has revealed that it contains three exceptionally large open reading frames (ORFs) encoding the modular polyketide synthase (PKS). Between two of these lies a fourth gene (rapP) encoding a pipecolate-incorporating enzyme that probably also catalyzes closure of the macrolide ring. On either side of these very large genes are ranged a total of 22 further ORFs before the limits of the cluster are reached, as judged by the identification of genes clearly encoding unrelated activities. Several of these ORFs appear to encode enzymes that would be required for Rp biosynthesis. These include two cytochrome P-450 monooxygenases (P450s), designated RapJ and RapN, an associated ferredoxin (Fd) RapO, and three potential SAM-dependent O-methyltransferases (MTases), RapI, RapM and RapQ. All of these are likely to be involved in 'late' modification of the macrocycle. The cluster also contains a novel gene (rapL) whose product is proposed to catalyze the formation of the Rp precursor, L-pipecolate, through the cyclodeamination of L-lysine. Adjacent genes have putative roles in Rp regulation and export. The codon usage of the PKS biosynthetic genes is markedly different from that of the flanking genes of the cluster.

Amino Acid Sequence↗