Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Influence of the codon following the AUG initiation codon on the expression of a modified lacZ gene in Escherichia coli.

In a lacZ expression vector (pMC1403Plac), all 64 codons were introduced immediately 3' from the AUG initiation codon. The expression of the second codon variants was measured by immunoprecipitation of the plasmid-coded fusion proteins. A 15-fold difference in expression was found among the codon variants. No distinct correlation could be made with the level of tRNA corresponding to the codons and large differences were observed between synonymous codons that use the same tRNA. Therefore the effect of the second codon is likely to be due to the influence of its composing nucleotides, presumably on the structure of the ribosomal binding site. An analysis of the known sequences of a large number of Escherichia coli genes shows that the use of codons in the second position deviates strongly from the overall codon usage in E. coli. It is proposed that codon selection at the second position is not based on requirements of the gene product (a protein) but is determined by factors governing gene regulation at the initiation step of translation.

Base Sequence↗

The close proximity of Escherichia coli genes: consequences for stop codon and synonymous codon use.

It is shown that synonymous codon usage is less biased in favor of those codons preferred by highly expressed genes at the end of Escherichia coli genes than in the middle. This appears to be due to the close proximity of many E. coli genes. It is shown that a substantial number of genes overlap either the Shine-Dalgarno sequence or the coding sequence of the next gene on the chromosome and that the codons that overlap have lower synonymous codon bias than those which do not. It is also shown that there is an increase in the frequency of A-ending codons, and a decrease in the frequency of G-ending codons at the end of E. coli genes that lie close to another gene. It is suggested that these trends in composition could be associated with selection against the formation of mRNA secondary structure near the start of the next gene on the chromosome. Stop codon use is also affected by the close proximity of genes; many genes are forced to use TGA and TAG stop codons because they terminate either within the Shine-Dalgarno or coding sequence of the next gene on the chromosome. The implications these results have for the evolution of synonymous codon use are discussed.

Base Sequence↗

Hydropathic characteristics of adenovirus hexons.

The complete nucleotide sequence and the predicted amino acid sequence of the adenovirus type 7 hexon gene were determined. The hydropathy of the hexon proteins from human adenovirus types 2, 3, 4, 5, 7, 12, 16, 40, 41, and 48, bovine adenovirus type 3, murine adenovirus type 1, and avian adenovirus types 1 and 10 was analysed. The presence of purines and pyrimidines in the second position of the codons was correlated to hydrophilicity and hydrophobicity, respectively. Comparison of the hydrophilicity plots of eight hexons showed seven hypervariable regions to be distributed on the surface. A large portion of the hypervariable regions manifests hydrophilicity. The strength of the surface charge accumulated on the hydrophilic and hydrophobic regions correlated to the tissue tropism of the different adenovirus types. Analysis of codon usage for adenovirus hexons showed that among synonymous codons those with cytidine in the third position were preferably used to a great extent. Analysis of the nucleotide and amino acid sequence pair distances and the phylogenetic tree of 14 hexon proteins showed members of subgenera B, D and E to be closely related, especially Ad4 and Ad16, and subgenus A to be closely related to subgenus F.

Adenoviruses, Human↗

Nucleotide sequence of gene pfkB encoding the minor phosphofructokinase of Escherichia coli K-12.

The nucleotide sequence of a 1.3-kb DNA fragment containing the entire pfkB gene which codes for Pfk-2 of Escherichia coli, a minor phosphofructokinase (Pfk) enzyme, is reported. The Pfk-2 protein subunit is encoded by 924 bp, has 308 amino acids and an Mr of 33 000. Like other weakly expressed E. coli genes the codon usage in the pfkB gene is random; there is no strong bias for the usage of major tRNA isoaccepting species, and the codon preference rules of Grosjean and Fiers [Gene, 18 (1982) 199-209] are followed. This is the first report of the complete gene sequence of a phosphofructokinase.

Amino Acid Sequence↗

First complete mitochondrial genome of Uzelothrips scabrosus (Thysanoptera: Uzelothripidae) provides insights into gene rearrangements and phylogenetic position within Terebrantia.

The family Uzelothripidae is represented by a single genus Uzelothrips and can be distinguished from others by the presence of whip-like antennae, a circular ventral sensorium on antennal segment III, a well-developed tentorium, and a membranous ovipositor. Here, we generated the first complete mitochondrial genome of Uzelothrips scabrosus (15,674 bp) using next-generation sequencing to explore the gene rearrangements and phylogenetic relationships. It consists of 13 protein-coding genes, 22 transfer RNAs, two ribosomal RNAs, and two putative control regions. The genome exhibits strong AT bias (71.35%) with negative AT and GC skew. Codon usage analyses indicate a strong bias towards A/U-ending codons and influenced by both natural selection and mutation pressure. All PCGs were under purifying selection, with cox1 being the most conserved and nad4L the most variable. The gene order of the family Uzelothripidae is highly rearranged compared to the ancestral insect gene order. Comparative analysis revealed that gene block B was the most widely conserved, whereas the remaining gene blocks exhibited family or lineage-specific conservation patterns, reflecting extensive mitochondrial gene rearrangements during the evolution of the Thysanoptera. Moreover, 228 synapomorphic and 68 autapomorphic gene boundaries were identified across thysanopteran mitogenomes. Phylogenies indicated that the family Uzelothripidae is in a sister relationship with Stenurothripidae, and the Uzelothripidae + Stenurothripidae clade is sister to Thripidae. This study provides the first mitogenomic insights into Uzelothripidae and highlights the need for broader taxon sampling and nuclear genomic data to resolve deep evolutionary relationships within Thysanoptera.

Comparative analysis↗

A ribosomal protein from Thermus thermophilus is homologous to a general shock protein.

The gene encoding the ribosomal protein from Thermus thermophilus, TL5, which binds to the 5S rRNA, has been cloned and sequenced. The codon usage shows a clear preference for G/C rich codons that is characteristic for many genes in thermophilic bacteria. The deduced amino acid sequence consists of 206 residues. The sequence of TL5 shows a strong similarity to a general shock protein from Bacillus subtilis, named CTC. The protein CTC is homologous in its N-terminal part to the 5S rRNA binding protein, L25, from E coli. An alignment of the TL5, CTC and L25 sequences displays a number of residues that are totally conserved. No clear sequence similarity was found between TL5 and other proteins which are known to bind to 5S rRNA. The evolutionary relationship of a heat shock protein in mesophiles and a ribosomal protein in thermophilic bacteria as well as a possible role of TL5 in the ribosome are discussed.

Amino Acid Sequence↗

The translational signal database, TransTerm, is now a relational database.

TransTerm-97 contains more than 97 500 non-redundant coding-sequence initiation and termination contexts compiled from GenBank, release 101 (15-June-1997). In addition, several coding sequence parameters are available: coding sequence length, Nc, GC3, and, when it is computable, codon adaptation index (CAI). Codon usage tables and summaries of start and stop codon contexts are also included. The information covers more than 325 species and organelles, including seven complete bacterial genomes and one complete eukaryotic genome. To promote research in translational control of protein synthesis, TransTerm has been converted into a relational database to ease the process of making queries. The relational database manager, Postgresql, gives access to the database using SQL (Structured Query Language). A World Wide Web interface using forms is being completed to allow the casual user access to the database. Extensions are planned to include the full 5'-UTR, full coding sequence and 3'-UTR. TransTerm-97 is available on the World Wide Web at:http://biochem. otago.ac.nz:800/Transterm/homepage.html

Animals↗

Spiroplasma virus 4: nucleotide sequence of the viral DNA, regulatory signals, and proposed genome organization.

The replicative form (RF) of spiroplasma virus 4 (SpV4) has been cloned in Escherichia coli, and the cloned RF has been shown to be infectious by transfection (M. C. Pascarel-Devilder, J. Renaudin, and J.-M. Bové, Virology 151:390-393, 1986). The cloned SpV4 RF was randomly subcloned and was fully sequenced by the dideoxy chain termination technique, using the M13 cloning and sequencing system. The nucleotide sequence of the SpV4 genome contains 4,421 nucleotides with a G+C content of 32 mol%. The triplet TGA is not a termination codon but, as in Mycoplasma capricolum (F. Yamao, A. Muto, Y. Kawauchi, M. Iwami, S. Iwagani, Y. Azumi, and S. Osawa, Proc. Natl. Acad. Sci. USA 82:2306-2309, 1985), probably codes for tryptophan. With these assumptions, nine open reading frames (ORFs) were identified. All nine are characterized by an ATG or GTG initiation codon, one or several termination codons, and a Shine-Dalgarno sequence upstream of the initiation codon. The nine ORFs are distributed in all three reading frames. One of the ORFs (ORF1) corresponds to the 60,000-dalton capsid protein gene. Analysis of codon usage showed that T- and A-terminated codons are preferably used, reflecting the low G+C content (32 mol%) of the SpV4 genome. The viral DNA contains two G+C-rich inverted repeat sequences. One could be involved in transcription termination and the other in initiation of cDNA strand synthesis. The SpV4 genome was found to contain at least three promoterlike sequences quasi-identical to those of eubacteria. These results fully support the bacterial origin of spiroplasmas.

Amino Acid Sequence↗

Cloning of two glutamate dehydrogenase cDNAs from Asparagus officinalis: sequence analysis and evolutionary implications.

Two different amplification products, termed c1 and c2, showing a high similarity to glutamate dehydrogenase sequences from plants, were obtained from Asparagus officinalis using two degenerated primers and RT-PCR (reverse transcriptase polymerase chain reaction). The genes corresponding to these cDNA clones were designated aspGDHA and aspGDHB. Screening of a cDNA library resulted in the isolation of cDNA clones for aspGDHB only. Analysis of the deduced amino acid (aa) sequence from the full-length cDNA suggests that the gene product contains all regions associated with metabolic function of NAD glutamate dehydrogenase (NAD-GDH). A first phylogenetic analysis including only GDHs from plants suggested that the two GDH genes of A. officinalis arose by an ancient duplication event, pre-dating the divergence of monocots and dicots. Codon usage analysis showed a bias towards A/T ending codons. This tendency is likely due to the biased nucleotide composition of the asparagus genome, rather than to the translational selection for specific codons. Using principal coordinate analysis, the evolutionary relatedness of plant GDHs with homologous sequences from a large spectrum of organisms was investigated. The results showed a closer affinity of plant GDHs to GDHs of thermophilic archaebacterial and eubacterial species, when compared to those of unicellular eukaryotic fungi. Sequence analysis at specific amino acid signatures, known to affect the thermal stability of GDH, and assays of enzyme activity at non-physiological temperatures, showed a greater adaptation to heat-stress conditions for the asparagus and tobacco enzymes compared with the Saccharomyces cerevisiae enzyme.

Amino Acid Sequence↗

Spiroplasmas: gene structure and expression.

Upon sequencing of the SpV4 genome, eight putative open reading frames (ORFs) including that for the 65-kilodalton (kDa) capsid protein were detected. They involve all three reading frames. Three promoter sequences were found, as well as a transcription terminator and the initiation site for complementary strand synthesis. Ribosome binding sites and regulatory sequences are closely related to those of Eubacteria. Codon usage analysis showed that A and T terminated codons are preferably used. UAA is the major termination codon. Upon cloning of the full-size SpV4 replicative form, the capsid protein gene could not be expressed in Escherichia coli, whereas the spiralin gene cloned in the same bacterium is expressed. These results suggest that in spiroplasmas, as in Mycoplasma capricolum, UGA is not a termination codon, but very probably codes for tryptophan. Spiralin contains no tryptophan. Hence, its gene contains no UGA codons and can thus be expressed in E. coli. On the other hand, the gene for capsid protein has nine UGA codons and cannot be fully expressed in the bacterium. Our results fully support the bacterial origin of spiroplasmas.

Bacteria↗

The araBAD operon of Salmonella typhimurium LT2. I. Nucleotide sequence of araB and primary structure of its product, ribulokinase.

Hybrid plasmids containing the araBAD operon of Salmonella typhimurium LT2 were characterized by Southern blot and genetic analyses. The nucleotide sequence of araB was determined. The araB gene product, ribulokinase (EC 2.7.1.16), was purified and the results of amino acid composition analysis and partial amino acid sequence are in agreement with predictions from the DNA sequence. Ribulokinase is 569 amino acid residues long and has a calculated Mr of 61 793. Ribulokinase shares significant homology with xylulose kinase from Escherichia coli. Codon usage in the araB gene does not favor those codons which have intermediate codon-anticodon binding energy.

Amino Acid Sequence↗

Analysis of nucleotide sequences of two ligninase cDNAs from a white-rot filamentous fungus, Phanerochaete chrysosporium.

An analysis of nucleotide sequences of two types of ligninase cDNAs isolated from the basidiomycete Phanerochaete chrysosporium, designated CLG4 and CLG5, are presented here. The amino acid sequences of the corresponding ligninase proteins, designated LG4 and LG5, respectively, have been deduced from the cDNA sequences. Mature ligninases LG4 and LG5 are preceded by leader sequences containing 28 and 27 amino acids (aa), respectively, and each contains 344 aa residues. The estimated Mrs of mature LG4 and LG5 are 36,540 and 36,607, respectively. Potential N-glycosylation site(s) with the general sequence Asn-X-Thr/Ser are found in both LG4 and LG5. Nucleotide sequence homology between the coding region of CLG4 and CLG5 is 71.5%, whereas the amino acid sequence homology between the two ligninases is 68.5%. The codon usage of ligninases is extremely biased in favor of codons rich in cytosine and guanine. Amino acid sequences of two tryptic peptides of ligninase H8 have exactly matching sequences in ligninase LG5. Also, the sequences of the oligodeoxynucleotide probes, which correspond to the sequences in the tryptic peptides of ligninase H8 and which were used in isolating the ligninase clones from the cDNA library, have exactly matching sequences in CLG5. The experimentally determined N-terminal sequence of purified ligninase H8 is found in the deduced N-terminal amino acid sequence of LG5. These results suggest that CLG5 encodes ligninase H8 and that CLG4 represents a related but different ligninase gene.

Amino Acid Sequence↗

Cloning and molecular characterization of the acetamidase-encoding gene (amdS) from Aspergillus oryzae.

We have isolated an acetamidase-encoding gene (amdS) from Aspergillus oryzae by heterologous hybridization using the corresponding Aspergillus nidulans gene as a probe. The gene is located on a 3.5-kb SacI fragment and its nucleotide (nt) sequence was determined. Compared with the A. nidulans amdS gene, the coding region of A. oryzae gene consists of seven exons interrupted by six introns and encodes 545 amino acid (aa) residues. The deduced aa sequence has a high degree of homology with that of the A. nidulans acetamidase protein. Three introns (IVS-1, IVS-2, and IVS-4) exist at the same positions as those of A. nidulans amdS, whilst three additional introns (IVS-3, IVS-5, and IVS-6) are also present. There is no preference in its codon usage (G + C content in the third position of codons is 51%). Gene disruption experiments demonstrate that the resulting mutants show significantly reduced growth on acetamide-containing medium, indicating that the A. oryzae amdS gene encodes a functional acetamidase that is required for acetamide utilization. Transcriptional analysis by Northern blot reveals a 1.8-kb transcript in RNA extracted from mycelium grown in medium containing acetamide or acetate plus beta-alanine as the sole carbon and nitrogen sources.

Amidohydrolases↗

Gene synthesis, expression in Escherichia coli, purification and characterization of the recombinant bovine acyl-CoA-binding protein.

A synthetic gene encoding the 86 amino acid residues of mature acyl-CoA-binding protein (ACBP), and the initiating methionine was constructed. The synthetic gene was assembled from eight partially overlapping oligonucleotides. Codon usage and nucleotides surrounding the ATG translation-initiation codon were chosen to allow efficient expression in Escherichia coli as well as in yeast. The synthetic gene was inserted into the expression vector pKK223-3 and expressed in E. coli. In maximally induced cultures, recombinant ACBP constitutes 12-15% of total cellular protein. A fraction highly enriched for recombinant ACBP was obtained by extracting induced E. coli cells with 1 M-acetic acid. Recombinant ACBP was purified to homogeneity by successive use of gel-filtration chromatography, ion-exchange chromatography and reverse-phase h.p.l.c. Recombinant ACBP differed from native ACBP by lacking the N-terminal acetyl group. The acyl-CoA-binding characteristics of recombinant ACBP did not differ from those of native ACBP, and the two proteins showed the same ability to induce medium-chain acyl-CoA synthesis by goat mammary-gland fatty acid synthetase. It was concluded that the N-terminal acetyl group is not important for acyl-CoA binding.

Acyl Coenzyme A↗

Syk mutation in Jurkat E6-derived clones results in lack of p72syk expression.

The human leukemic Jurkat cell line is commonly used as a model cellular system to study T lymphocyte signal transduction. Various clonal derivatives of Jurkat T cells exist which display different characteristics with regard to responses to external stimuli. Among these, the E6-1 clone of Jurkat T cells has been used as a parental line from which numerous important somatic mutant clones have been generated. During the course of experiments examining signals initiated by the T cell antigen receptor in an E6-1-derived Jurkat cell clone J.CaM1, we observed that the 72-kilodalton Syk protein tyrosine kinase previously found in other Jurkat cells was not detected. Upon further analysis it was determined that Syk transcripts from the J.CaM1 cells as well as the parental E6-1 cells contain a single guanine nucleotide insertion at position 92. This nucleotide insertion results in a shift in the Syk open reading frame leading to alternate codon usage as well as the generation of a termination codon at position 109. Thus, Syk transcripts in E6-1 cells and E6-1-derived clones are predicted to be capable of encoding only the first 33 amino acids of the 630-amino acid wild type Syk. These findings are incompatible with a recently proposed model of T cell antigen receptor signal transduction based, in part, on experiments conducted using E6-1-derived cells, suggesting that Syk might play a role upstream of Lck and Zap70.

Amino Acid Sequence↗

Primary structure of the ompF gene that codes for a major outer membrane protein of Escherichia coli K-12.

The nucleotide sequence of the ompF gene coding for a major outer membrane protein of Escherichia coli K-12 has been determined and the amino acid sequence of the OmpF protein was deduced from it. The OmpF protein contains 340 amino acid residues, and is produced from a precursor having 22 extra amino acid residues, the signal peptide, at the amino terminus. The expected secondary structure of the OmpF protein had a high beta-sheet content with a low alpha-helix content. The promoter region and the transcription termination region of the ompF gene had a significantly high AT content, while the AT content of the coding region was about the same as the average AT content of the E. coli chromosome. Following the termination codon, a typical rho-independent transcription termination signal was observed. The codon usage in the ompF gene was highly nonrandom; the codons preferably utilized are those recognized by the most abundant species of isoaccepting tRNAs or those, among synonymous codons recognized by the same tRNA, that can interact more properly with the anticodon.

Amino Acid Sequence↗

Sequences of the E. coli uvrC gene and protein.

We have determined the sequence of a 2400 bp region of E. coli chromosomal DNA containing the uvrC gene. The coding region of uvrc is 2267 bp in length, encodes a polypeptide with a calculated molecular weight of 66,038 daltons, and is preceded by a typical E. coli ribosome binding site. By constructing deletion derivatives we have established that a uvrC promoter lies within the 113 bp region preceding the translational start of uvrC. The codon usage in uvrC is strongly biased in favor of codons used infrequently in E. coli, which may contribute to the relatively low intracellular concentration of uvrC protein.

Amino Acid Sequence↗

Nucleotide sequence of the LuxC gene and the upstream DNA from the bioluminescent system of Vibrio harveyi.

The nucleotide sequence of the luxC gene (1431 bp) and the upstream DNA (1049 bp) of the luminescent bacterium Vibrio harveyi has been determined. The luxC gene can be translated into a polypeptide of 55 kDa in excellent agreement with the molecular mass of the reductase polypeptide required for synthesis of the aldehyde substrate for the bioluminescent reaction. Analyses of codon usage showed a high frequency (1.9%) of the isoleucine codon, AUA, in the luxC gene compared to that found in Escherichia coli genes (0.2%) and its absence in the luxA, B and D genes. The low G/C content of the luxC gene and upstream DNA (38-39%) compared to that found in the other lux genes of V. harveyi (45%) was primarily due to a stretch of 500 nucleotides with only a 24% G/C content, extending from 200 bp inside lux C to 300 bp upstream. Moreover, an open reading frame did not extend for more than 48 codons between the luxC gene and 600 bp upstream at which point a gene transcribed in the opposite direction started. As the lux system in the luminescent bacterium, V. fischeri, contains a regulatory gene immediately upstream of luxC transcribed in the same direction, these results show that the organization and regulation of the lux genes have diverged in different luminescent bacteria.

Amino Acid Sequence↗