Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

Characterization and phylogenetic utility of the mammalian protamine p1 gene.

We sequenced the protamine P1 gene (ca. 450 bp) from 20 bats (order Chiroptera) and the flying lemur (order Dermoptera). We compared these sequences with published sequences from 19 other mammals representing seven orders (Artiodactyla, Carnivora, Cetacea, Perissodactyla, Primates, Proboscidea, and Rodentia) to assess structure, base compositional bias, and phylogenetic utility. Approximately 80% of second codon positions were guanine, resulting in protamine proteins containing a high frequency of arginine residues. Our data indicate that codon usage for arginine differs among higher mammalian taxa. Parsimony analysis of 40 species representing nine orders produced a well-resolved tree in which most nodes were supported strongly, except at the lowest taxonomic levels (e.g., within Artiodactyla and Vespertilionidae). These data support monophyly of several taxa proposed by morphologic and molecular studies (all nine orders: Laurasiatheria, Cetartiodactytla, Yangochiroptera, Noctilionoidea, Rhinolophoidea, Vespertilionoidea, Phyllostomidae, Natalidae, and Vespertilionidae) and, in agreement with recent molecular studies, reject monophyly of Archonta, Volitantia, and Microchiroptera. Bats were sister to a clade containing Perissodactyla, Carnivora, and Cetartiodactyla, and, although not unequivocally, rhinolophoid bats (traditional microchiropterans) were sister to megachiropterans. Sequences of the protamine P1 gene are useful for resolving relationships at and above the familial level in bats, and generally within and among mammalian orders, but with some drawbacks. The coding and intervening sequences are small, producing few phylogenetically informative characters, and aligning the intron is difficult, even among closely related families. Given these caveats, the protamine P1 gene may be important to future systematic studies because its functional and evolutionary constraints differ from other genes currently used in systematic studies.

Amino Acid Sequence↗

Nucleotide sequence determination and genetic analysis of the Bacteroides plasmid, pBI143.

The nucleotide sequence and genetic organization of the Bacteroides plasmid pBI143 were determined. The plasmid was 2747 base pairs (bp) and had a G+C content of 41% (GenBank Accession No. U30316). There were two open reading frames greater than 50 codons and these were designated mobA and repA. A 56-bp inverted repeat divided pBI143 into modules with repA and mobA in separate regions. There was a marked difference in the G+C content and codon usage for the two regions; repA had 33% G+C and mobA was 44% G+C. MobA had homology to other Bacteroides mobilization proteins and RepA shared homology to a replication protein from Zymomonas mobilis plasmid pZM2. These two putative replication proteins formed a subgroup of the rolling-circle replication.proteins belonging to the pSN2 family of gram-positive plasmids. Consistent with this finding, single-stranded pBI143 DNA was detected in plasmid containing Bacteroides fragilis cultures. Availability of the pBI143 sequence allowed the elucidation of the complete nucleotide sequence for pFD288 an 8.9-kb Bacteroides shuttle vector (GenBank Accession No. U30830).

Amino Acid Sequence↗

Cloning and characterization of a high-copy-number novel insertion sequence from chemolithotrophic Thiobacillus ferrooxidans.

Two distinct families of repetitive DNA elements (1.4 and 1.2 kb) were identified from S1 nuclease-treated genomic DNA of four strains of Thiobacillus ferrooxidans. The 1.4-kb fragment hybridized with IST2, an insertion sequence of T. ferrooxidans. The 1.2-kb fragment was cloned and sequenced. The sequence (IST445), 1219 bp in length, with features characteristic of an insertion element, has a terminal inverted repeat of 8 bp, which can be further extended to 23 or 48 bp with 9 and 26 mismatches, respectively. It displays 54.4% identity in 967 nucleotides of overlap with ISAE1 of Alcaligenes eutrophus. The IST445 contains three open reading frames which have codon usage almost similar to 56 different coding genes of T. ferrooxidans. In Southern blots of restricted genomic DNAs probed with IST445, each of the several strains of T. ferrooxidans gives a distinctive fingerprint. IST445 is present in the range of 10-20 copies per genome in the four strains studied.

Amino Acid Sequence↗

Comparative sequence analysis of plasmids pME2001 and pME2200 of methanothermobacter marburgensis strains Marburg and ZH3.

Comparison of the updated complete nucleotide sequences of the two related plasmids pME2001 and pME2200 from the thermophilic archaeon Methanothermobacter marburgensis (formerly Methanobacterium thermoautotrophicum) strains Marburg and ZH3, respectively, revealed an almost identical common backbone structure and five plasmid-specific inserted fragments (IFs), four of which are flanked by perfect or nearly perfect direct repeats 25-52 bp in length. A 4354-bp minimal replicon was derived from the alignment of the two plasmids, which encodes one putative antisense RNA related to replication control and five open reading frames (ORFs) organized in two operons. The first operon consists of four ORFs, the third of which, i.e. ORF3, contains a helix-turn-helix motif and a purine NTP-binding motif often found in proteins involved in DNA metabolic processes. The database search results suggest that ORF3 might function as a replication initiator protein. The large putative Rep protein encoded by pME2001 was overexpressed in Escherichia coli as an N-terminal His-tagged version using pET28a and a compatible helper plasmid that coexpresses minor tRNAs, argU and ileX to compensate for codon usage difference. ORFs 1, 2, and 3 are organized in a sequence reminiscent of that described in E. coli plasmids of the R1 family, cop-tap-rep. ORF6 encoded by IF1, one of the pME2200-specific elements, showed significant similarity to ORF6 encoded by archaeal phage psiM2 of M. marburgensis strain Marburg and may confer the apparent immunity of its host strain ZH3 to infection by phage psiM2. Our data indicate that M. marburgensis plasmids may evolve by a series of gene duplication and excision events.

Adenosine Triphosphatases↗

High-level bacterial expression of human glutathione transferase P1-1 encoded by semisynthetic DNA.

A cDNA clone, lambda GTHP1del, encoding glutathione transferase (GST) P1-1, was isolated from a human K562 erythroleukemia cell line cDNA library. The coding sequence was lacking the codons for the N-terminal 34 amino acids. A DNA segment was designed in order to obtain the missing portion and a structure representing the entire protein. The synthetic DNA sequence was constructed to achieve efficient base pairing with Escherichia coli 16S ribosomal RNA, avoidance of internal secondary structure, and optimal codon usage for high-level protein expression in accord with the known preferences in E. coli. The truncated GST P1-1 cDNA sequence and the synthetic segment were ligated into a plasmid to give an inducible expression system. Among the resulting clones a limited number was selected by immunodetection for highest yield of GST P1-1. Maximal expression was obtained from a spontaneously mutated sequence with altered as well as deleted bases as compared to the original construct. This clone, pKXHP1, allowed heterologous expression in E. coli in yields of > 200 mg enzyme per liter culture medium. The physicochemical and catalytic properties of the recombinant protein were indistinguishable from those of the enzyme purified from human placenta.

Base Sequence↗

Optimized heterologous expression of the polymorphic human glutathione transferase M1-1 based on silent mutations in the corresponding cDNA.

An expression clone for large-scale production of the polymorphic human glutathione transferase (GST) M1-1 has been developed. Heterologous expression in Escherichia coli afforded a yield of 100 mg of GST M1-1 per 3 liters of culture medium, corresponding to an approximately 10-fold increased yield compared to the parental expression construct. Overproduction of the enzyme was dependent on the codon usage in the 5' region of the DNA sequence encoding glutathione transferase M1-1. High-level expression clones were generated by a combination of random silent mutations in selected wobble positions in the coding sequence and immunoselection of clones from the library of random mutants. The strategy used is generally applicable for the production of recombinant proteins provided that a suitable selection procedure is available for identifying the desired mutants.

Amino Acid Sequence↗

Heterologous gene expression in a membrane-protein-specific system.

We have constructed an expression system for heterologous proteins which uses the molecular machinery responsible for the high level production of bacteriorhodopsin in Halobacterium salinarum. Cloning vectors were assembled that fused sequences of the bacterio-opsin gene (bop) to coding sequences of heterologous genes and generated DNA fragments with cloning sites that permitted transfer of fused genes into H. salinarum expression vectors. Gene fusions include: (i) carboxyl-terminal-tagged bacterio-opsin; (ii) a carboxyl-terminal fusion with the catalytic subunit of the Escherichia coli aspartate transcarbamylase; (iii) the human muscarinic receptor, subtype M1; (iv) the human serotonin receptor, type 5HT2c; and (v) the yeast alpha mating factor receptor, Ste2. Characterization of the expression of these fusions revealed that the bop gene coding region contains previously undescribed molecular determinants which are critical for high level expression. For example, introduction of immunogenic and purification tag sequences into the C-terminal coding region significantly decreased bop gene mRNA and protein accumulation. The bacteriorhodopsin-aspartate transcarbamylase fusion protein was expressed at 7 mg per liter of culture, demonstrating that E. coli codon usage bias did not limit the system's potential for high level expression. The work presented describes initial efforts in the development of a novel heterologous protein expression system, which may have unique advantages for producing multiple milligram quantities of membrane-associated proteins.

Amino Acid Sequence↗

Optimization of the expression of equistatin in Pichia pastoris.

To improve the expression of equistatin, a proteinase inhibitor from the sea anemone Actinia equina, in the yeast Pichia pastoris, we prepared gene variants with yeast-preferred codon usage and lower repetitive AT and GC content. The full gene optimization approximately doubled the level of steady-state mRNA and protein accumulated in the culture medium. The removal of a short stretch of 12 additional nucleotides from the multiple cloning site (MCS) sequence in the vector pPIC9 had an enhancement effect similar to full gene optimization (factor 1.5) at the mRNA level. However, at the protein level, this increase was 4- to 10-fold. The optimized gene without the MCS sequence yielded 1.66 g/L active protein in a bioreactor and was purified by a new two-step procedure with a recovery of activity that was >95%. This production level constitutes an overall improvement of about 20-fold relative to our previously published results. The characteristics of the MCS sequence element are discussed in the light of its apparent ability to act as negative expression regulator.

Amino Acid Sequence↗

Isolation, cloning, and complete nucleotide sequence of a phenotypically distinct Brazilian isolate of human T-lymphotropic virus type II (HTLV-II).

Analysis of human T-lymphotropic virus type II (HTLV-II) isolates from North America and Europe have demonstrated the existence of two molecular subtypes of the virus, HTLV-IIa and HTLV-IIb. Recently, studies on HTLV-II infections in Brazil have revealed isolates that are related phylogenetically to the HTLV-IIa subtype but have a HTLV-IIb phenotype with respect to the transactivating protein, tax. To more clearly define this relationship, HTLV-II was isolated from peripheral blood of an IVDA from Sao Paulo, Brazil (SP-WV), and the complete provirus was cloned and sequenced. Comparison of HTLV-II(SP-WV) nucleotide sequences to other available complete HTLV-II proviral sequences revealed that HTLV-II(SP-WV) is most closely related to HTLV-II(Mo), the prototypic HTLV-IIa subtype sequence. Phylogenetic analysis of LTR, env, and tax regions unequivocally demonstrated that HTLV-II(SP-WV) and all other Brazilian sequences examined are members of the IIa subtype. The predicted amino acid sequences of the major coding regions of HTLV-II(SP-WV) are also most closely related to HTLV-II(Mo), with the important exception of tax. The tax protein encoded by HTLV-II(SP-WV) is 96-99% identical to the tax of IIb isolates and is similar in that it has an additional 25 amino acids at the carboxy-terminus compared to the HTLV-II(Mo) tax with which it shares 91% identity. Analysis of tax stop codon usage of a number of HTLV-IIa isolates from North American, Europe, and Brazil demonstrated that isolates from the last region appear to be unique in their extended tax phenotype. It could be demonstrated that the extended tax proteins in the HTLV-IIb and Brazilian isolates had equivalent ability to transactivate the viral LTR, and studies with deletion mutants indicated that the extended C-terminus is not essential for transactivation. In contrast, the HTLV-IIa tax was found to have a greatly diminished ability to transactivate the viral LTR, which appeared to be a consequence of reduced expression of the protein. The studies show that although the Brazilian strains do not represent an entirely new subtype based on nucleotide sequence analysis they are a phenotypically unique molecular variant within the HTLV-IIa subtype.

Brazil↗

The origin and evolution of species differences in Escherichia coli and Salmonella typhimurium.

Since diverging from a common ancestor some 120 million years, Escherichia coli and Salmonella typhimurium have accumulated numerous phenotypic characteristics which have traditionally been used to distinguish these enteric species. While most of the genetic differences between these species are due to the accumulation of point mutations, the majority of the observed variation in phenotypic characters is attributable to segments of the genome confined to only one of the species. We have analyzed the map positions, G+C contents, nucleotide sequences and functions of regions unique to the Salmonella chromosome in an attempt to determine the ancestry of species-specific sequences. Some of the Salmonella-specific regions had uncharacteristically low base compositions and contained open reading frames of atypical codon usage patterns suggesting that portions of the genome were acquired by horizontal transfer from distantly-related bacterial species. The role of these species-specific sequences was assayed by constructing mutant strains harboring deletions in the corresponding regions of the genome. Several functions were ascribed to these unique portions of the Salmonella chromosome, including one encoding proteins involved in virulence and invasion of host epithelial cells.

Biological Evolution↗

Vicilin-like seed storage proteins in the gymnosperm interior spruce (Picea glauca/engelmanii).

A seed storage protein cDNA was characterized from a library of interior spruce (Picea glauca/engelmanii complex) cotyledonary stage somatic embryos. The deduced amino acid sequence predicts a 448 amino acid (50 kDa) polypeptide with 28-38% identity with angiosperm vicilin-like 7S globulins. XXC/G codon usage is low (47%) relative to monocot angiosperms while pairwise comparisons show that spruce, monocot, and dicot vicilins are approximately equal in amino acid divergence. Although small by comparison, the spruce vicilin contains an N terminal hydrophilic region characteristic of angiosperm 'large' vicilins. Genomic Southern blotting predicts that the cDNA is encoded by a gene family.

Amino Acid Sequence↗

Isolation and characterisation of cDNA clones representing the genes encoding the major tuber storage protein (dioscorin) of yam (Dioscorea cayenensis Lam.).

cDNA clones encoding dioscorins, the major tuber storage proteins (M(r) 32,000) of yam (Dioscorea cayenesis) have been isolated. Two classes of clone (A and B, based on hybrid release translation product sizes and nucleotide sequence differences) which are 84.1% similar in their protein coding regions, were identified. The protein encoded by the open reading frame of the class A cDNA insert is of M(r) 30,015. The difference in observed and calculated molecular mass might be attributed to glycosylation. Nucleotide sequencing and in vitro transcription/translation suggest that the class A dioscorin proteins are synthesised with signal peptides of 18 amino acid residues which are cleaved from the mature peptide. The class A and class B proteins are 69.6% similar with respect to each other, but show no sequence identity with other plant proteins or with the major tuber storage proteins of potato (patatin) or sweet potato (sporamin). Storage protein gene expression was restricted to developing tubers and was not induced by growth conditions known to induce expression of tuber storage protein genes in other plant species. The codon usage of the dioscorin genes suggests that the Dioscoreaceae are more closely related to dicotyledonous than to monocotyledonous plants.

Amino Acid Sequence↗

The reconstruction and expression of a Bacillus thuringiensis cryIIIA gene in protoplasts and potato plants.

A Bacillus thuringiensis (B.t.) cryIIIA delta-endotoxin gene was designed for optimal expression in plants. The modified cry gene has the codon usage pattern of an average dicot gene and does not contain AT-rich nucleotide sequences typical of native B.t. cry genes. We assembled the 1.8 kb cryIIIA gene in nine blocks of three oligonucleotide pairs. For two DNA blocks, the polymerase chain reaction was used to enrich for correctly ligated pairs. We compared modified cryIIIA gene with native gene expression by electroporation of dicot (carrot) and monocot (corn) protoplasts. CryIIIA-specific RNA and protein was detected in carrot and corn protoplasts only after electroporation with the rebuilt gene. Transgenic potato lines were generated containing the redesigned cryIIIA gene under the transcriptional control of a chimeric CaMV 35S/mannopine synthetase (Mac) promoter. Out of 63 transgenic potato lines, 58 controlled first-instar Colorado potato beetle (CPB) larvae in bioassays. Egg masses which produced ca. 250,000 CPB larvae were placed on replicate clones of 56 transgenic potatoes. No CPB larvae developed past the second instar on any of these plants. Plants expressing high levels of delta-endotoxin were identified by their toxicity to more resistant third-instar larvae. We show there was good correlation between insect control and the levels of delta-endotoxin RNA and protein.

Amino Acid Sequence↗

A putative beta-glucanase pseudogene behind the potato GBSS gene.

We identified an open reading frame (ORF) which is located closely behind the gene encoding granule-bound starch synthase (GBSS) of potato (Solanum tuberosum L.). The ORF ends with a perfect 43 bp direct repeat, which carries the stop triplet precisely at the beginning of the second repeat. The deduced protein shows homology with all known isoforms of plant beta-1,3-glucanases and beta-1,3-1,4-glucanases. Although the DNA sequence is unique in potato and tomato (Lycopersicon esculentum L.), no expression of the gene was found in these species. Taken together with the unusual codon usage and length of the predicted protein, this sequence could be a pseudogene.

Amino Acid Sequence↗

Evolution of chorion gene families in lepidoptera: characterization of 15 cDNAs from the gypsy moth.

Fifteen unique chorion protein-encoding cDNAs from gypsy moth have been completely sequenced. These sequences are encoded by a family of genes, based on pairwise similarity values of 78-100% within a 225-nt region. Pairwise comparisons and maximum parsimony analysis strongly support the existence of two clusters of 11 and four sequences each, called noc1 and noc2. While noc2 consists of two subclusters, there is little character support for subclusters within noc1. The highly localized character-state distribution on the parsimony tree in gypsy moth is reminiscent of that in Bombyx mori, specifically for those chorion families that have been shown to undergo gene conversion. Gene conversion thus becomes a reasonable explanation for the homogeneity of noc1 sequences and for their distinctness from noc2. The relationship between the two major clusters of chorion sequences in gypsy moth (noc1, noc2) and Bombyx mori (Bm alpha, Bm beta) has been addressed through mixed-species tree construction. All four groups cluster separately, thus providing no direct evidence of orthologous sequences. However, the occurrence of gene conversion could have eliminated such evidence. The relationship between the chorion gene tree and the species cladogenic event is discussed, as are biases in codon usage, base composition, and nucleotide transformations.

Animals↗

The complete nucleotide sequence and gene organization of carp (Cyprinus carpio) mitochondrial genome.

The complete sequence of the carp mitochondrial genome of 16,575 base pairs has been determined. The carp mitochondrial genome encodes the same set of genes (13 proteins, 2 rRNAs, and 22 tRNAs) as do other vertebrate mitochondrial DNAs. Comparison of this teleostean mitochondrial genome with those of other vertebrates reveals a similar gene order and compact genomic organization. The codon usage of proteins of carp mitochondrial genome is similar to that of other vertebrates. The phylogenetic relationship for mitochondrial protein genes is more apparent than that for the mitochondrial tRNA and rRNA genes.

Amino Acid Sequence↗

Conservation of the mammalian RNA polymerase II largest-subunit C-terminal domain.

We have isolated and sequenced a portion of the gene encoding the carboxy-terminal domain (CTD) of the largest subunit of RNA polymerase II from three mammals. These mammalian sequences include one rodent and two primate CTDs. Comparisons of the new sequences to mouse and Chinese hamster show a high degree of conservation among the mammalian CTDs. Due to synonymous codon usage, the nucleotide differences between hamster, rat, ape, and human result in no amino acid changes. The amino acid sequence for the mouse CTD appears to have one different amino acid when compared to the other four sequences. Therefore, except for the one variation in mouse, all of the known mammalian CTDs have identical amino acid sequences. This is in marked contrast to the situation among more divergent species. The present study suggests that there is a strong evolutionary pressure to maintain the primary structure of the mammalian CTD.

Animals↗

Mo-MuLV nucleotide sequence exhibits three levels of oligomeric repetitions, suggesting a stepwise molecular evolution.

An exhaustive computer-assisted analysis of the Moloney murine leukemia virus nucleotide sequence shows numerous deviations in the oligomeric distribution, suggesting three overlapping levels of a stepwise duplicative evolution. (1) The sequence fits the universal rule of TG/CT excess which has been proposed as the construction principle of all sequences, and maintains some degree of symmetry between the two complementary strands. (2) Oligomeric repeating units share a core consensus regularly scattered throughout the sequence. This consensus is not merely predictable from the doublet frequencies and codon usage, but could correspond to an intermediary stage in a so-called periodic-to-chaotic transition. (3) Probable stepwise local duplications could be accounted for by slippagelike mechanisms. Comparison with the human spumaretrovirus (HSRV) shows similar segments in the overrepresented oligomers of the two sequences. The intermediary stage of transition oligomeric repeating units is not so clearly suggested in HSRV, perhaps because of numerous stepwise local duplications. In any case, a common evolutionary origin for the two viruses is not ruled out.

Base Sequence↗