Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Nucleotide sequence and deduced amino acid sequence of a Plasmodium falciparum actin gene.

The nucleotide sequence of a Plasmodium falciparum actin gene has been established. The gene codes for a protein of 376 amino acids and is not interrupted by introns. The nucleotide sequence reveals an extreme bias in codon usage. Not less than 85% of the codons possess an A or T at the third position. As has been found for the actins in other unicellular eukaryotes, P. falciparum actin is related both to vertebrate cytoplasmic and vertebrate muscle specific actins. However, the malarial actin is one of the most alpha-like actins hitherto found in lower eukaryotes.

Actins↗

Nucleotide sequence of a gene from the Pseudomonas transposon Tn501 encoding mercuric reductase.

We have determined the nucleotide sequence of the merA gene from the mercury-resistance transposon Tn501 and have predicted the structure of the gene product, mercuric reductase. The DNA sequence predicts a polypeptide of Mr 58 660, the primary structure of which shows strong homologies to glutathione reductase and lipoamide dehydrogenase, but mercuric reductase contains as additional N-terminal region that may form a separate domain. The implications of these comparisons for the tertiary structure and mechanism of mercuric reductase are discussed. The DNA sequence presented here has an overall G+C content of 65.1 mol%, typical of the bulk DNA of Pseudomonas aeruginosa from which Tn501 was originally isolated. Analysis of the codon usage in the merA gene shows that codons with C or G at the third position are preferentially utilized.

Amino Acid Sequence↗

Occurrence of unmodified adenine and uracil at the first position of anticodon in threonine tRNAs in Mycoplasma capricolum.

Codon usage pattern in the threonine four-codon (ACN) box in Mycoplasma capricolum is strongly biased towards adenine and uracil for the third base of codons. Codons ending in uracil or adenine, especially ACU, predominate over ACC and ACG. This bacterium contains two isoacceptor threonine tRNAs having anticodon sequences AGU and UGU, both with unmodified first nucleotides. It would thus appear that ACN codons are translated in an unusual way; tRNA(Thr)(AGU) would translate the most abundantly used codon ACU exclusively, because adenine at the first anticodon position can, according to the wobble rule, pair only with uracil of the third codon position. The tRNA(Thr)(UGU) would mainly be responsible for translation of three other codons, ACA, ACG, and ACC. Anticodon UGU would also be used for reading codon ACU as a redundancy of tRNA(Thr)-(AGU), as deduced from the mitochondrial code where unmodified uracil at the first anticodon position can pair with adenine, cytosine, guanine, and uracil by four-way wobble. The tRNA(Thr)(AGU) has much higher sequence homology to tRNA(Thr)(UGU) from M. capricolum (88%), Bacillus subtilis (77%) and Escherichia coli (86%) than to tRNA(Thr)(GGU) from B. subtilis (66%) and E. coli (63%), suggesting that tRNA(Thr)-(AGU) has been derived from tRNA(Thr)(UGU), but not from tRNA(Thr)(GGU).

Adenine↗

Drosophila melanogaster mitochondrial DNA: gene organization and evolutionary considerations.

The sequence of a 8351-nucleotide mitochondrial DNA (mtDNA) fragment has been obtained extending the knowledge of the Drosophila melanogaster mitochondrial genome to 90% of its coding region. The sequence encodes seven polypeptides, 12 tRNAs and the 3' end of the 16S rRNA and CO III genes. The gene organization is strictly conserved with respect to the Drosophila yakuba mitochondrial genome, and different from that found in mammals and Xenopus. The high A + T content of D. melanogaster mitochondrial DNA is reflected in a reiterative codon usage, with more than 90% of the codons ending in T or A, G + C rich codons being practically absent. The average level of homology between the D. melanogaster and D. yakuba sequences is very high (roughly 94%), although insertion and deletions have been detected in protein, tRNA and large ribosomal genes. The analysis of nucleotide changes reveals a similar frequency for transitions and transversions, and reflects a strong bias against G + C on both strands. The predominant type of transition is strand specific.

Amino Acid Sequence↗

Ureaplasma urealyticum urease genes; use of a UGA tryptophan codon.

Nucleotide sequence analysis of a Ureaplasma urealyticum DNA fragment, homologous to cloned urease genes of other prokaryotes, revealed three consecutive open reading frames. The molecular weights of the three deduced polypeptides are 11.2 kD, 13.6 kD and 66.6 kD. These values are consistent with the size of the three subunits previously reported for purified native urease. A significant sequence homology was found between the three polypeptides of the ureaplasmal urease and the single polypeptide of jack bean (Canavalia ensiformis) urease. Codon usage indicates that UGA is a tryptophan codon in this mollicute. Use of polymerase chain reactions has disclosed the existence of genetic polymorphism among the urease genes of different serotypes of U. urealyticum.

Amino Acid Sequence↗

Identification, genetic analysis and DNA sequence of a 7.8-kb virulence region of the Salmonella typhimurium virulence plasmid.

The 90-kilobase (kb) virulence plasmid of Salmonella typhimurium is responsible for invasion from the intestines to mesenteric lymph nodes and spleens of orally inoculated mice. We used Tn5 and aminoglycoside phosphotransferase (aph) gene insertion mutagenesis and deletion mutagenesis of a previously identified 14-kb virulence region to reduce this virulence region to 7.8kb. The 7.8-kb virulence region subcloned into a low copy-number vector conferred a wild-type level of splenic infection to virulence plasmid-cured S. typhimurium and conferred essentially a wild-type oral LD50. Insertion mutagenesis identified five loci essential for virulence, and DNA sequence analysis of the virulence region identified six open reading frames. Expected protein products were identified from four of the six genes, with three of the proteins identified as doublet bands in Escherichia coli minicells. Three of the five mutated genes were able to be complemented by clones containing only the corresponding wild-type gene. Only one of the five deduced amino acid sequences, that of the positive regulatory element, SpvR, possessed significant homology to other proteins. The codon usage for the virulence genes showed no codon bias, which is consistent with the low levels of expression observed for the corresponding proteins. Consensus promoters for several different sigma factors were identified upstream of several of the genes, whereas only consensus Rho-dependent termination sequences were observed between certain of the genes. The operon structure of this virulence region therefore appears to be complex. The construction of the cloned 7.8-kb virulence region and the determination of the DNA sequence will aid in the further genetic analysis of the five plasmid-encoded virulence genes of S. typhimurium.

Amino Acid Sequence↗

Molecular cloning of DNA complementary to bovine growth hormone mRNA.

We have cloned DNA complementary to mRNA coding for bovine growth hormone (bGH). Double-stranded DNA complementary to bovine pituitary mRNA was inserted into the Pst I site of plasmid pBR322 by the dC x dG tailing technique and amplified in E. coli x 1776. A recombinant plasmid containing bGH cDNA ws identified by hybridization to cloned rat growth hormone cDNA. It contains the entire coding and 3'-untranslated regions and 31 bases in the 5'-untranslated region. Nucleotide sequence analysis determined the sequence of the 26-amino acid signal peptide and confirmed the published amino acid sequence of the secreted hormone at all but 2 residues. Codon usage is nonrandom, with 81.7% of the codons ending in G or C. The nucleotide sequence of bGH mRNA is 83.9% homologous with rat GH mRNA and 76.5% homologous with human GH mRNA, while the respective amino acid sequence homologies are 83.5% and 66.8%.

Amino Acid Sequence↗

Molecular cloning of the chicken melanocortin 2 (ACTH)-receptor gene.

The chicken melanocortin 2-receptor (MC2-R) gene was isolated. It is found to be a single copy gene encoding a 357 amino acid protein, sharing 65.8-68.7% identity with mammalian counterparts. The chicken MC2-R mRNA is expressed in the adrenal and spleen, suggesting that the receptor mediates both endocrine and immunoregulatory functions of ACTH in the chicken. The amino acid sequence of the chicken MC2-R is collinear with those of other subtypes of MC-R, whereas all cloned mammalian MC2-Rs contain a gap in the third intracellular loop, suggesting that mammalian MC2-R molecules have evolved by lacking a part of the domain which determines the specificity of signal transduction in G-protein coupled receptors. Interestingly, the codon usage differs dramatically between MC1-R and MC2-R in the chicken; the GC-contents at the third codon position in MC1-R and MC2-R are 94.6 and 50.6%, respectively. It may reflect selective constraints on the usage of synonymous codons.

Adrenal Glands↗

The large mitochondrial genome of Syndiclis anlungensis (Lauraceae): Genome structure, comparative analysis, and phylogenetic relationships among Syndiclis species.

The complete mitochondrial genome (mitogenome) of Syndiclis anlungensis, a critically endangered tropical tree, was determined in this study. The mitogenome spans 2,368,454&#xa0;bp across four contigs and harbors 41 protein-coding genes, 22 tRNA genes, and three rRNA genes. Potential mutation regions, including 1317 repeat sequences and 698 simple sequence repeats (SSRs), were accurately located in the S. anlungensis mitogenome. Sixty-five transferred fragments of the repeats were found between its mitochondrial and chloroplast genomes. When compared to three other Laurales mitogenomes, extensive gene order shuffling is evident, leaving only five conserved gene clusters intact. Codon usage analysis reveals a pronounced A/T bias in both mitochondrial and chloroplast genes, and three mitochondrial genes (atp9, rps19, and sdh3) stand out for their high divergence across eleven Syndiclis taxa. Selection analyses indicate strong purifying pressure on rpl2, rpl16, and sdh3 (Ka/Ks&#xa0;<&#xa0;1), with no positive selection detected. Using 41 mitochondrial protein-coding gene sequences from sixteen and three individuals of Syndiclis and Beilschmiedia species, respectively, our phylogenetic tree recovers Syndiclis as monophyletic, with two well-supported clades: one includes S. anlungensis, S. chinensis, S. lotungensis, S. marlipoensis, and a putative new Syndiclis species from Yunnan; the other contains S. furfuracea, S. hongkongensis, S. kwangsiensis, and three putative new Syndiclis species from Guangdong and Vietnam.

Genome, Mitochondrial↗

Influence of the codon following the AUG initiation codon on the expression of a modified lacZ gene in Escherichia coli.

In a lacZ expression vector (pMC1403Plac), all 64 codons were introduced immediately 3' from the AUG initiation codon. The expression of the second codon variants was measured by immunoprecipitation of the plasmid-coded fusion proteins. A 15-fold difference in expression was found among the codon variants. No distinct correlation could be made with the level of tRNA corresponding to the codons and large differences were observed between synonymous codons that use the same tRNA. Therefore the effect of the second codon is likely to be due to the influence of its composing nucleotides, presumably on the structure of the ribosomal binding site. An analysis of the known sequences of a large number of Escherichia coli genes shows that the use of codons in the second position deviates strongly from the overall codon usage in E. coli. It is proposed that codon selection at the second position is not based on requirements of the gene product (a protein) but is determined by factors governing gene regulation at the initiation step of translation.

Base Sequence↗

The close proximity of Escherichia coli genes: consequences for stop codon and synonymous codon use.

It is shown that synonymous codon usage is less biased in favor of those codons preferred by highly expressed genes at the end of Escherichia coli genes than in the middle. This appears to be due to the close proximity of many E. coli genes. It is shown that a substantial number of genes overlap either the Shine-Dalgarno sequence or the coding sequence of the next gene on the chromosome and that the codons that overlap have lower synonymous codon bias than those which do not. It is also shown that there is an increase in the frequency of A-ending codons, and a decrease in the frequency of G-ending codons at the end of E. coli genes that lie close to another gene. It is suggested that these trends in composition could be associated with selection against the formation of mRNA secondary structure near the start of the next gene on the chromosome. Stop codon use is also affected by the close proximity of genes; many genes are forced to use TGA and TAG stop codons because they terminate either within the Shine-Dalgarno or coding sequence of the next gene on the chromosome. The implications these results have for the evolution of synonymous codon use are discussed.

Base Sequence↗

Hydropathic characteristics of adenovirus hexons.

The complete nucleotide sequence and the predicted amino acid sequence of the adenovirus type 7 hexon gene were determined. The hydropathy of the hexon proteins from human adenovirus types 2, 3, 4, 5, 7, 12, 16, 40, 41, and 48, bovine adenovirus type 3, murine adenovirus type 1, and avian adenovirus types 1 and 10 was analysed. The presence of purines and pyrimidines in the second position of the codons was correlated to hydrophilicity and hydrophobicity, respectively. Comparison of the hydrophilicity plots of eight hexons showed seven hypervariable regions to be distributed on the surface. A large portion of the hypervariable regions manifests hydrophilicity. The strength of the surface charge accumulated on the hydrophilic and hydrophobic regions correlated to the tissue tropism of the different adenovirus types. Analysis of codon usage for adenovirus hexons showed that among synonymous codons those with cytidine in the third position were preferably used to a great extent. Analysis of the nucleotide and amino acid sequence pair distances and the phylogenetic tree of 14 hexon proteins showed members of subgenera B, D and E to be closely related, especially Ad4 and Ad16, and subgenus A to be closely related to subgenus F.

Adenoviruses, Human↗

Nucleotide sequence of gene pfkB encoding the minor phosphofructokinase of Escherichia coli K-12.

The nucleotide sequence of a 1.3-kb DNA fragment containing the entire pfkB gene which codes for Pfk-2 of Escherichia coli, a minor phosphofructokinase (Pfk) enzyme, is reported. The Pfk-2 protein subunit is encoded by 924 bp, has 308 amino acids and an Mr of 33 000. Like other weakly expressed E. coli genes the codon usage in the pfkB gene is random; there is no strong bias for the usage of major tRNA isoaccepting species, and the codon preference rules of Grosjean and Fiers [Gene, 18 (1982) 199-209] are followed. This is the first report of the complete gene sequence of a phosphofructokinase.

Amino Acid Sequence↗

First complete mitochondrial genome of Uzelothrips scabrosus (Thysanoptera: Uzelothripidae) provides insights into gene rearrangements and phylogenetic position within Terebrantia.

The family Uzelothripidae is represented by a single genus Uzelothrips and can be distinguished from others by the presence of whip-like antennae, a circular ventral sensorium on antennal segment III, a well-developed tentorium, and a membranous ovipositor. Here, we generated the first complete mitochondrial genome of Uzelothrips scabrosus (15,674&#xa0;bp) using next-generation sequencing to explore the gene rearrangements and phylogenetic relationships. It consists of 13 protein-coding genes, 22 transfer RNAs, two ribosomal RNAs, and two putative control regions. The genome exhibits strong AT bias (71.35%) with negative AT and GC skew. Codon usage analyses indicate a strong bias towards A/U-ending codons and influenced by both natural selection and mutation pressure. All PCGs were under purifying selection, with cox1 being the most conserved and nad4L the most variable. The gene order of the family Uzelothripidae is highly rearranged compared to the ancestral insect gene order. Comparative analysis revealed that gene block B was the most widely conserved, whereas the remaining gene blocks exhibited family or lineage-specific conservation patterns, reflecting extensive mitochondrial gene rearrangements during the evolution of the Thysanoptera. Moreover, 228 synapomorphic and 68 autapomorphic gene boundaries were identified across thysanopteran mitogenomes. Phylogenies indicated that the family Uzelothripidae is in a sister relationship with Stenurothripidae, and the Uzelothripidae&#xa0;+&#xa0;Stenurothripidae clade is sister to Thripidae. This study provides the first mitogenomic insights into Uzelothripidae and highlights the need for broader taxon sampling and nuclear genomic data to resolve deep evolutionary relationships within Thysanoptera.

Comparative analysis↗

A ribosomal protein from Thermus thermophilus is homologous to a general shock protein.

The gene encoding the ribosomal protein from Thermus thermophilus, TL5, which binds to the 5S rRNA, has been cloned and sequenced. The codon usage shows a clear preference for G/C rich codons that is characteristic for many genes in thermophilic bacteria. The deduced amino acid sequence consists of 206 residues. The sequence of TL5 shows a strong similarity to a general shock protein from Bacillus subtilis, named CTC. The protein CTC is homologous in its N-terminal part to the 5S rRNA binding protein, L25, from E coli. An alignment of the TL5, CTC and L25 sequences displays a number of residues that are totally conserved. No clear sequence similarity was found between TL5 and other proteins which are known to bind to 5S rRNA. The evolutionary relationship of a heat shock protein in mesophiles and a ribosomal protein in thermophilic bacteria as well as a possible role of TL5 in the ribosome are discussed.

Amino Acid Sequence↗

The translational signal database, TransTerm, is now a relational database.

TransTerm-97 contains more than 97 500 non-redundant coding-sequence initiation and termination contexts compiled from GenBank, release 101 (15-June-1997). In addition, several coding sequence parameters are available: coding sequence length, Nc, GC3, and, when it is computable, codon adaptation index (CAI). Codon usage tables and summaries of start and stop codon contexts are also included. The information covers more than 325 species and organelles, including seven complete bacterial genomes and one complete eukaryotic genome. To promote research in translational control of protein synthesis, TransTerm has been converted into a relational database to ease the process of making queries. The relational database manager, Postgresql, gives access to the database using SQL (Structured Query Language). A World Wide Web interface using forms is being completed to allow the casual user access to the database. Extensions are planned to include the full 5'-UTR, full coding sequence and 3'-UTR. TransTerm-97 is available on the World Wide Web at:http://biochem. otago.ac.nz:800/Transterm/homepage.html

Animals↗

Regularities of context-dependent codon bias in eukaryotic genes.

Nucleotides surrounding a codon influence the choice of this particular codon from among the group of possible synonymous codons. The strongest influence on codon usage arises from the nucleotide immediately following the codon and is known as the N1 context. We studied the relative abundance of codons with N1 contexts in genes from four eukaryotes for which the entire genomes have been sequenced: Homo sapiens, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana. For all the studied organisms it was found that 90% of the codons have a statistically significant N1 context-dependent codon bias. The relative abundance of each codon with an N1 context was compared with the relative abundance of the same 4mer oligonucleotide in the whole genome. This comparison showed that in about half of all cases the context-dependent codon bias could not be explained by the sequence composition of the genome. Ranking statistics were applied to compare context-dependent codon biases for codons from different synonymous groups. We found regularities in N1 context-dependent codon bias with respect to the codon nucleotide composition. Codons with the same nucleotides in the second and third positions and the same N1 context have a statistically significant correlation of their relative abundances.

Animals↗

Spiroplasma virus 4: nucleotide sequence of the viral DNA, regulatory signals, and proposed genome organization.

The replicative form (RF) of spiroplasma virus 4 (SpV4) has been cloned in Escherichia coli, and the cloned RF has been shown to be infectious by transfection (M. C. Pascarel-Devilder, J. Renaudin, and J.-M. Bové, Virology 151:390-393, 1986). The cloned SpV4 RF was randomly subcloned and was fully sequenced by the dideoxy chain termination technique, using the M13 cloning and sequencing system. The nucleotide sequence of the SpV4 genome contains 4,421 nucleotides with a G+C content of 32 mol%. The triplet TGA is not a termination codon but, as in Mycoplasma capricolum (F. Yamao, A. Muto, Y. Kawauchi, M. Iwami, S. Iwagani, Y. Azumi, and S. Osawa, Proc. Natl. Acad. Sci. USA 82:2306-2309, 1985), probably codes for tryptophan. With these assumptions, nine open reading frames (ORFs) were identified. All nine are characterized by an ATG or GTG initiation codon, one or several termination codons, and a Shine-Dalgarno sequence upstream of the initiation codon. The nine ORFs are distributed in all three reading frames. One of the ORFs (ORF1) corresponds to the 60,000-dalton capsid protein gene. Analysis of codon usage showed that T- and A-terminated codons are preferably used, reflecting the low G+C content (32 mol%) of the SpV4 genome. The viral DNA contains two G+C-rich inverted repeat sequences. One could be involved in transcription termination and the other in initiation of cDNA strand synthesis. The SpV4 genome was found to contain at least three promoterlike sequences quasi-identical to those of eubacteria. These results fully support the bacterial origin of spiroplasmas.

Amino Acid Sequence↗