Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

Structure of a cluster of mouse histone genes.

The four mouse histone genes (2 H3 genes, an H2b gene and an H2a gene) present in a cloned 12.9 kilobase fragment of DNA have been completely sequenced including both 5' and 3' flanking regions. These genes are expressed in cultured mouse cells and the 3' and 5' ends of the mRNA have been determined by S1 nuclease mapping. These genes code for a minor fraction of the histone mRNAs expressed in cultured mouse cells. They comprise at most 5-8% of the total histone mRNA of each type. The two H3 genes code for H3.2 and H3.1 histone proteins, while the H2b gene codes for an H2b.1 protein with a single amino acid change (val-leu) at position 18. Only the 3' portion of the H2a gene is contained in the clone and there is an amino acid change (alanine-proline) at position 126. Comparison of the 5' and 3' flanking sequences reveals a conserved sequence at the 3' end of the mRNA which forms a hairpin loop structure. The codon usage in the genes is non-random and there has been no discrimination against CG doublets in the coding region of the genes.

Amino Acid Sequence↗

Molecular cloning and nucleotide sequence analysis of the maltose-inducible porin gene of Aeromonas salmonicida.

The gene for the Aeromonas salmonicida maltose-inducible porin (maltoporin) was cloned into phagemid pTZ18R in two restriction fragments, 0.6-kb PstI/KpnI and 1.7-kb SphI, of genomic DNA and their nucleotide sequences were determined. Open reading frames of 1329 and 1335 bp translated into sequences of 443 and 445 amino acids, with a 23 or 25 amino acid signal sequence and a 420 amino acid mature protein of molecular mass 46424 Da. Putative ribosome binding sites, AGGA and GGGAA, occurred 9 bp upstream of two possible ATG initiation codons. The A. salmonicida gene product showed a high degree of similarity with Escherichia coli LamB, and codon usage was very similar to that of another A. salmonicida outer membrane protein but markedly different from those of extracellular proteins.

Aeromonas↗

Effects of an opal termination codon preceding the nsP4 gene sequence in the O'Nyong-Nyong virus genome on Anopheles gambiae infectivity.

The genomic RNA of an alphavirus encodes four different nonstructural proteins, nsP1, nsP2, nsP3, and nsP4. The polyprotein P123 is produced when translation terminates at an opal termination codon between nsP3 and nsP4. The polyprotein P1234 is produced when translational readthrough occurs or when the opal termination codon has been replaced by a sense codon in the alphavirus genome. Evolutionary pressures appear to have maintained genomic sequences encoding both a stop codon (opal) and an open reading frame (arginine) as a general feature of the O'nyong-nyong virus (ONNV) genome, indicating that both are required at some point. Alternate replication of ONNVs in both vertebrate and invertebrate hosts may determine predominance of a particular codon at this locus in the viral quasispecies. However, no systematic study has previously tested this hypothesis in whole animals. We report here the results of the first study to investigate in a natural mosquito host the functional significance of the opal stop codon in an alphavirus genome. We used a full-length cDNA clone of ONNV to construct a series of mutants in which the arginine between nsP3 and nsP4 was replaced with an opal, ochre, or amber stop codon. The presence of an opal stop codon upstream of nsP4 nearly doubled (75.5%) the infectivity of ONNV over that of virus possessing a codon for the amino acid arginine at the corresponding position (39.8%). Although the frequency with which the opal virus disseminated from the mosquito midgut did not differ significantly from that of the arginine virus on days 8 and 10, dissemination did began earlier in mosquitoes infected with the opal virus. Although a clear fitness advantage is provided to ONNV by the presence of an opal codon between nsP3 and nsP4 in Anopheles gambiae, sequence analysis of ONNV RNA extracted from mosquito bodies and heads indicated codon usage at this position corresponded with that of the virus administered in the blood meal. These results suggest that while selection of ONNV variants is occurring, de novo mutation at the position between nsP3 and nsP4 does not readily occur in the mosquito. Taken together, these results suggest that the primary fitness advantage provided to ONNV by the presence of an opal codon between nsP3 and nsP4 is related to mosquito infectivity.

Alphavirus↗

A simple program to calculate codon bias index.

A computer program (PCBI) was developed to quickly calculate codon bias index (CBI). PCBI can analyze a gene containing introns. The 22 preferred codons defined from Saccharomyces cerevisiae were used in PCBI as the standard to measure the CBI values. However, users can modify the preferred codons to suit each organism. The data PCBI provides include DNA sequence of open reading frame without introns, amino acid sequence of gene product, a table of amino acid composition, a table of codon usage and (G + C) content, parameters for calculating CBI, and the value of CBI. PCBI runs on a DOS or Windows environment, but results can be saved in ASCII text format.

Amino Acids↗

Competitive expression of two heterologous genes inserted into one plasmid in Saccharomyces cerevisiae.

Plasmids were constructed which contained two expression units encoding single-chain insulin precursors. Surprisingly, the total amount of insulin precursor produced was similar to that produced from plasmids containing a single expression unit. In this system, therefore, two expression cassettes can be brought to compete for the limited ability of the yeast cell for synthesis and secretion. Using genes encoding B(1-29)-A(1-21) and B(1-29)-Ala-Ala-Lys-A-(1-21), the slightly different precursors could be quantified individually after separation by high-performance liquid chromatography from the culture supernatant. The two-cassette system allowed a sensitive and well controlled comparison of parameters important for optimal expression of a heterologous gene in Saccharomyces cerevisiae. The system was used to compare two promoter constructions and also to evaluate the position of expression cassettes in the plasmid. Finally the codon usage in the gene to be expressed was found to influence its ability to compete for expression.

Amino Acid Sequence↗

Nucleotide sequences of the trpI, trpB, and trpA genes of Pseudomonas syringae: positive control unique to fluorescent pseudomonads.

A 904-bp probe from Pseudomonas aeruginosa was used to identify the trpB, trpA and trpI genes of Pseudomonas syringae. Transcription initiation at the P. syringae trpBA promoter in vitro was activated by the P. aeruginosa TrpI protein in the presence of indoleglycerol phosphate. Thus, trpB and trpA are regulated positively in three species of fluorescent pseudomonads, P. aeruginosa, P. putida, and P. syringae, but in no other eubacteria so far investigated [Crawford, Annu. Rev. Microbiol. 43 (1989) 567-600]. In addition to conservation of protein-coding sequences, there is a high degree of nucleotide sequence identity in the intergenic control region that includes the divergent trpI and trpBA promoters, especially in the binding sites for TrpI protein. Differences in patterns of codon usage distinguish the trpI genes of P. syringae and P. putida from P. aeruginosa trpI and from the trpB and trpA genes of all three species.

Amino Acid Sequence↗

Compositional nonrandomness upstream of start codons in archaebacteria.

Since the regions directing transcriptional and translational initiation in archaebacteria are poorly characterized, the purpose of this study was to characterize them using measurements of nonrandomness upstream of start codons on eight fully sequenced archaebacterial genomes. Two distinctly different regions with conservation were identified. The location of the first corresponded well to the classical Shine-Dalgarno region (phi peak), and the other was located approximately 20-35 nucleotides upstream of start codons (alpha region), but both regions are not present in all strains. The composition of the region around the phi peak showed an overrepresentation of guanine, whereas composition in the alpha region had an overrepresentation of adenine and thymine. It is furthermore shown that the alpha region surprisingly is associated with start codon usage and other characteristics of the genes. The alpha region is likely to correspond to TATA-boxes and thereby indicates use of leaderless (or short-leadered) transcripts.

5' Flanking Region↗

Characterization of the Candida rugosa lipase system and overexpression of the lip1 isoenzyme in a non-conventional yeast.

The fungus C. rugosa produces lipase isoenzymes (CRLs) homologous to the Geotrichum candidum and Yarrowia lipolytica lipases to which they share ca. 40 and 30% sequence identity, with a domain of sequence conservation at the N-terminal half of the protein. CRL proteins have high sequence homology but are not identical in their catalytic activity, therefore calling for the resolution of isoforms via heterologous expression. The non-conventional use of a serine codon in several Candida species frustrates overexpression in the currently available host systems. The LIP1 gene, coding for the major CRL form, was therefore expressed in C. maltosa, a related fungus with the same codon usage as C. rugosa. A recombinant lipase was produced and secreted in an active form in the culture medium upon engineering the 5' and 3' ends of the gene.

Candida↗

Cloning and sequencing of the adenylate kinase gene (adk) of Escherichia coli.

Adenylate kinase, the product of the adk locus in Escherichia coli K12, catalyzes the conversion of AMP and ATP to two molecules of ADP. The gene has been cloned by complementation of an adk temperature sensitive mutation. The DNA sequence of the complete coding region and of 5'- and 3'-untranslated regions were determined. The resulting protein sequence was found to contain several regions of high homology with cytosolic adenylate kinase of pig muscle (AK1), whose three-dimensional structure has been determined. The most significant of the amino acid exchanges is the replacement of histidine 36 with glutamine. This residue is believed to play a role in catalysis through metal ion binding. The codon usage pattern and the determination of adenylate kinase molecules per cell shows that the enzyme is one of the more abundant soluble proteins of the bacterial cells.

Adenylate Kinase↗

Inferring the number of evolutionary events from DNA coding sequence differences.

The estimation of the amount of evolutionary divergence that has taken place between two DNA coding sequences depends strongly on the degree of constraint on amino acid replacements. If amino acid replacements are relatively unconstrained, the individual nucleotide is the appropriate unit of analysis and the method of Tajima and Nei can be used. If amino acid replacements are constrained, however, this method is shown to be inapplicable. For sequences with strong amino acid constraints, a method is outlined analogous to the Tajima and Nei method using codons as the unit of analysis. Only synonymous substitutions are used. Codon usage data can be employed to estimate the necessary parameters of the calculation, or a priori models of substitution may be employed. Sequences with significant but intermediate constraints on amino acid replacements are, in principle, unanalyzable.

Animals↗

Inquiries into the cloning and expression of the tuf gene of Thermus thermophilus HB8.

The tuf gene, which encodes the elongation factor Tu (EF-Tu), of Thermus thermophilus HB8 was cloned and sequenced. The whole G + C content of the tuf gene was 64.9%, and 84.5% of the third base in codon usage was either G or C, in which G was much favorable than C. The tuf gene was expressed in Escherichia coli under the control of the E. coli trp promoter. For the highly efficient expression, the SD sequence of the E. coli trpL was used instead of that of T. thermophilus HB8 and the length between the SD sequence and the initiation codon was controlled to keep 9 nucleotides.

Base Sequence↗

Genetic code deviations in the ciliates: evidence for multiple and independent events.

In several species of ciliates, the universal stop codons UAA and UAG are translated into glutamine, while in the euplotids, the glutamine codon usage is normal, but UGA appears to be translated as cysteine. Because the emerging position of this monophyletic group in the eukaryotic lineage is relatively late, this deviant genetic code represents a derived state of the universal code. The question is therefore raised as to how these changes arose within the evolutionary pathways of the phylum. Here, we have investigated the presence of stop codons in alpha tubulin and/or phosphoglycerate kinase gene coding sequences from diverse species of ciliates scattered over the phylogenetic tree constructed from 28S rRNA sequences. In our data set, when deviations occur they correspond to in frame UAA and UAG coding for glutamine. By combining these new data with those previously reported, we show that (i) utilization of UAA and UAG codons occurs to different extents between, but also within, the different classes of ciliates and (ii) the resulting phylogenetic pattern of deviations from the universal code cannot be accounted for by a scenario involving a single transition to the unusual code. Thus, contrary to expectations, deviations from the universal genetic code have arisen independently several times within the phylum.

Animals↗

The primary structure of the leu1+ gene of Schizosaccharomyces pombe.

A DNA fragment which carries the leu1 gene encoding beta-isopropylmalate dehydrogenase in Schizosaccharomyces pombe has been isolated by complementation of an E. coli leuB mutation. This 1.5 kb DNA fragment complements not only the S. pombe leu1 mutation, but also the S. cerevisiae leu2 mutation. The nucleotide sequence of the essential part of the leu1 gene and its flanking regions was determined. This sequence contains an open reading frame of 371 codons, from which a protein having a Mr = 39,732 can be predicted. The deduced amino acid sequence and its codon usage were compared with those of the S. cerevisiae LEU2 protein. The cloned DNA will be a useful marker when transforming S. pombe.

3-Isopropylmalate Dehydrogenase↗

Cloning and sequence analysis of the fermentative alcohol-dehydrogenase-encoding gene of Escherichia coli.

A 6-kb fragment of DNA, which complemented defects in the alcohol dehydrogenase (ADH)-encoding gene (adhE) of Escherichia coli, was cloned into a multicopy vector. Both ADH and coenzyme-A-linked acetaldehyde dehydrogenase (ACDH) activities were encoded by the plasmid, pHIL8. The adhE gene was identified as an open reading frame of 891 codons encoding an Mr 96,008 protein (minus the initiating methionine). Codon usage analysis indicates that adhE should be highly expressed. This gene shows no significant homology to any previously sequenced ADH-encoding gene.

Alcohol Dehydrogenase↗

DING proteins are from Pseudomonas.

DING proteins have been described as animal and plant proteins with potential biomineralisation, receptor or signalling roles that have been characterised by an N-terminal DINGGG-sequence. However, these sequences have only ever been identified as either N-terminal peptides or partial cDNA sequences, and have yet to be detected in any of the many genomic animal and plant genomes now available. Microbial relatives of the DING proteins have been described, which appear to be periplasmic phosphate-binding proteins. Recently, full-length Pseudomonas aeruginosa UCBPP-PA14 and Hypericum perforatum genes have been sequenced that show high homology to the published DING protein N-terminal sequences, and small peptides previously identified in conjunction with the peptide sequencing of DING proteins can also be mapped to regions across these full-length sequences. Searching with these sequences identifies other plant and animal cDNA fragments in the public nucleotide databases, and, additionally, an unordered rat genomic contig that contains a DING-like sequence on a small fragment. Analysing the codon usage of these DNA sequences identifies all of these sequences as of Pseudomonas origin, suggesting that DING proteins do not exist in eukaryotes, but instead are potentially due to microbial contamination or infection.

Amino Acid Sequence↗

Complete nucleotide sequence of the structural gene for colicin A, a gene translated at non-uniform rate.

The complete nucleotide sequence of the structural gene for colicin A has been established. This sequence consists of 1776 base-pairs. According to the predicted amino acid sequence, the colicin A polypeptide chain comprises 592 amino acids and has a molecular weight of 62,989. The amino-terminal part is rich in proline and glycine and accordingly secondary structure prediction indicates that this region (1 to 185) is beta-structured. The rest of the molecule (residues 186 to 592) is very rich in alpha-helix. An uncharged amino acid sequence of 48 residues is located in the C-terminal part of the molecule, which is involved in the membrane depolarization caused by colicin A. A similar region has been found in colicin E1, which has the same mode of action as colicin A. Three peptides of these bacteriocins were found to be homologous, but a comparison of the bacteriocin genes did not reveal any significant homology out of the corresponding regions. The codon usage of both genes, however, exhibits some similarity and is quite different from that of genes coding for highly or weakly expressed proteins of Escherichia coli.

Amino Acid Sequence↗

Complete DNA sequence of the mitochondrial genome of the black chiton, Katharina tunicata.

The DNA sequence of the 15,532-base pair (bp) mitochondrial DNA (mtDNA) of the chiton Katharina tunicata has been determined. The 37 genes typical of metazoan mtDNA are present: 13 for protein subunits involved in oxidative phosphorylation, 2 for rRNAs and 22 for tRNAs. The gene arrangement resembles those of arthropods much more than that of another mollusc, the bivalve Mytilus edulis. Most genes abut directly or overlap, and abbreviated stop codons are inferred for four genes. Four junctions between adjacent pairs of protein genes lack intervening tRNA genes; however, at each of these junctions there is a sequence immediately adjacent to the start codon of the downstream gene that is capable of forming a stem-and-loop structure. Analysis of the tRNA gene sequences suggests that the D arm is unpaired in tRNA(ser)(AGN), which is typical of metazoan mtDNAs, and also in tRNA(ser)(UCN), a condition found previously only in nematode mtDNAs. There are two additional sequences in Katharina mtDNA that can be folded into structures resembling tRNAs; whether these are functional genes is unknown. All possible codons except the stop codons TAA and TAG are used in the protein-encoding genes, and Katharina mtDNA appears to use the same variation of the mitochondrial genetic code that is used in Drosophila and Mytilus. Translation initiates at the codons ATG, ATA and GTG. A + T richness appears to have affected codon usage patterns and, perhaps, the amino acid composition of the encoded proteins. A 142-bp non-coding region between tRNA(glu) and CO3 contains a 72-bp tract of alternating A and T.

Amino Acid Sequence↗

Complete nucleotide sequence of type 6 M protein of the group A Streptococcus. Repetitive structure and membrane anchor.

The DNA sequence of the gene for type 6 M protein of Streptococcus pyogenes contains two extended tandem repeat regions and one nontandem repeat region. We suggest that the duplication and deletion of these repeats generates the observed diversity in size and sequence among the family of M proteins in the group A streptococci. In addition, the DNA sequence reveals the presence of a 42-amino-acid signal peptide, a region rich in proline that is thought to be located in the cell wall, and a membrane anchor sequence at the carboxyl-terminal end of the protein. Signals similar to the consensus sequences recognized for the initiation of transcription and translation in Gram-positive bacteria have been identified in the DNA sequence. Codon usage is similar to that of other Gram-positive bacteria and significantly different from that of Escherichia coli.

Amino Acid Sequence↗