Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

Transformation of NIH3T3 cells with synthetic c-Ha-ras genes.

Synthetic human c-Ha-ras genes in which amino acid codons were altered to those which are frequently used in highly expressed Escherichia coli genes were ligated to the 3'-end of Rous sarcoma virus long terminal repeat. When NIH3T3 cells were transfected with the plasmids having those genes with valine at codon 12, leucine at codon 61 or arginine at codon 61, transformants were efficiently produced. These results indicated that the synthetic c-Ha-ras genes are expressed in a mammalian system even though their codon usage is altered to correspond with that of E. coli. This expression vector system should be useful for studies on the structure-function relationships of c-Ha-ras, since the synthetic gene can be easily modified to have multiple base alterations, and can also be used simultaneously for the production of large amounts of p21 in E. coli for biochemical and biophysical studies.

Avian Sarcoma Viruses↗

DNA sequence of the D-serine deaminase activator gene dsdC.

We have determined the DNA sequence of dsdC, the gene that encodes the D-serine deaminase activator protein of Escherichia coli K-12. The sequence contains a single open reading frame that terminates in a UGA codon. One the basis of the size of the protein, 33 kilodaltons, and the amino acid sequence encoded by the open reading frame, we identified a likely translation initiation codon 731 base pairs upstream of the translation initiation codon for the divergently transcribed D-serine deaminase gene. There is a broad range of codon usage, not surprising in view of the weak expression of the gene. The N-terminal two-thirds of the activator is arginine-lysine rich and quite polar; the remainder is more neutral. The segment of the protein that seems most likely to have potential to form the helix-turn-helix structure characteristic of DNA-regulatory proteins is located near the end of the polar region. The protein contains a region with significant homology to lambda attB.

Amino Acid Sequence↗

Approaches to stabilization of inter-domain recombination in polyketide synthase gene expression plasmids.

Regions of extremely high sequence identity are recurrent in modular polyketide synthase (PKS) genes. Such sequences are potentially detrimental to the stability of PKS expression plasmids used in the combinatorial biosynthesis of polyketide metabolites. We present two different solutions for circumventing intra-plasmid recombination within the megalomicin PKS genes in Streptomyces coelicolor. In one example, a synthetic gene was used in which the codon usage was reengineered without affecting the primary amino acid sequence. The other approach utilized a heterologous subunit complementation strategy to replace one of the problematic regions. Both methods resulted in PKS complexes capable of 6-deoxyerythronolide B analogue biosynthesis in S. coelicolor CH999, permitting reproducible scale-up to at least 5-l stirred-tank fermentation and a comparison of diketide precursor incorporation efficiencies between the erythromycin and megalomicin PKSs.

Base Sequence↗

Cloning, expression and characterization of phycoerythrin gene from Ceramium boydenn.

Phycobiliproteins function as a major light harvesting protein-pigment complex in the cyanobacteria and the eukaryotic algae. Phycoerythrin (PE) is a kind of phycobiliproteins, widely located in all rhodophytes, some species of cyanobacteria and cryptophytes, and different ecotypes of Prochlorococcus populations. PeBA encoding beta and alpha subunits of PE from Ceramium boydenn was cloned and sequenced in this research. A peBA specific PCR primer was synthesized, based on the peBA gene conserved sequences. The beta subunit encoding gene (peB) contained an open reading frame of 534 bp, while the alpha subunit (peA) was 495 bp. Recombinant expression plasmid pET-peAB was constructed and expressed in Escherichia coli BL21. The molecular weight of expressive product of peB and peA was about 23.3 and 18.2 KD, respectively. Results of codon usage analysis show that G + C content is heterogeneous among different groups of PE and spacers have dramatically lower G + C contents than coding regions. Also there is a high variance in G + C content among sequences at the third position sites. It is also found in this paper that several sequence regions, which might reflect functional or structural requirements of the PE organization, and several residues known for their functional importance are conserved in almost all the sequences.

Amino Acid Sequence↗

Nucleotide sequence of the alpha-amylase-pullulanase gene from Clostridium thermohydrosulfuricum.

The nucleotide sequence of the gene (apu) encoding the thermostable alpha-amylase-pullulanase of Clostridium thermohydrosulfuricum was determined. An open reading frame of 4425 bp was present. The deduced polypeptide (Mr 165,600), including a 31 amino acid putative signal sequence, comprised 1475 amino acids, with no cysteine residues. The structural gene was preceded by the consensus promoter sequence TTGACA TATAAT, a putative regulatory sequence and a putative ribosome-binding sequence AAAGGGGG. The codon usage resembled that of Bacillus genes. The deduced sequence of the mature apu product showed similarities to various amylolytic enzymes, especially the neopullulanase of Bacillus stearothermophilus, whereas the signal sequence showed similarity to those of the alpha-amylases of B. stearothermophilus and B. subtilis. Three regions thought to be highly conserved in the primary structure of alpha-amylases could also be distinguished in the apu product, two being partly 'duplicated' in this alpha-1,4/alpha-1,6-active enzyme.

Amino Acid Sequence↗

RNA polymerase II of Drosophila. Relation of its 140,000 Mr subunit to the beta subunit of Escherichia coli RNA polymerase.

We have determined the nucleotide sequence of the gene coding for the 140,000 Mr subunit of the DNA-dependent RNA polymerase II from Drosophila melanogaster. This analysis revealed features that are typical for a household-function gene. The codon usage is relaxed and there is no apparent TATA box upstream from the transcription start site. A comparison of the deduced amino acid sequence with that of the Escherichia coli RNA polymerase beta subunit shows a total of nine regions of homology. These regions are also conserved in chloroplast DNA of tobacco. This supports the notion that the two large subunits of the eukaryotic RNA polymerases are the structural and functional equivalents of E. coli beta' and beta subunits.

Amino Acid Sequence↗

Isolation and characterization of a complementary DNA clone for an algal pre-apoplastocyanin.

We have isolated a cDNA clone for the Chlamydomonas reinhardtii pre-apoplastocyanin. The sequence contains codons for the complete pre-protein including a two-domain, lumen-targeting transit sequence and the mature apoprotein. The transit sequence (47 amino acids) is the shortest one described for chloroplast lumenal proteins, and like other C. reinhardtii lumen-targeting transit sequences appears to lack an uncharged amino-terminal domain usually present in plant lumen-directing sequences. The mature protein is deduced to be 98 amino acids in length and shows highest primary sequence similarity (74-76% identity) to other unicellular algal plastocyanins. Southern hybridization analysis of C. reinhardtii genomic DNA indicates the presence of a single nuclear gene, as is the case for all other plastocyanin genes characterized to date, although the algal gene might be interrupted. Codon usage in this gene reflects the high GC content of C. reinhardtii nuclear DNA, but is more highly biased than that found in the C. reinhardtii copper-repressible gene for the functionally equivalent pre-apocytochrome c552 (perhaps contributing to the more efficient synthesis in vivo of plastocyanin over cytochrome c552). The deduced physical properties of this plastocyanin are compared to those of the C. reinhardtii plastidic cytochrome c552.

Amino Acid Sequence↗

Structure and molecular evolutionary analysis of a plant cytochrome c gene: surprising implications for Arabidopsis thaliana.

We have isolated a cytochrome c gene from Arabidopsis thaliana (cv. Columbia), which is the first cytochrome c gene to be cloned from a higher plant. Genomic DNA blot analysis indicates that there is only one copy of cytochrome c in Arabidopsis. The gene consists of three exons separated by two introns. Gene features such as regulatory regions, codon usage, and conserved splicing-specific sequences are all present and typical of dicotyledonous plant nuclear genes. We have constructed phenograms and cladograms for cytochrome c amino acid sequences and histone H3, alcohol dehydrogenase, and actin DNA sequences. For both cytochrome c and histone H3, Arabidopsis clusters poorly with other higher plants. Instead, it clusters with Neurospora and/or the yeasts. We suggest that perhaps this observation should be considered when using Arabidopsis as a model system for higher plants.

Actins↗

Phosphonate biosynthesis: molecular cloning of the gene for phosphoenolpyruvate mutase from Tetrahymena pyriformis and overexpression of the gene product in Escherichia coli.

The phosphoenolpyruvate mutase gene from Tetrahymena pyriformis has been cloned and overexpressed in Escherichia coli. To our knowledge, this is the first Tetrahymena gene to be expressed in E. coli, a task made more complicated by the idiosyncratic codon usage by Tetrahymena. The N-terminal amino acid sequence of phosphoenolpyruvate mutase purified from T. pyriformis has been used to generate a precise oligonucleotide probe for the gene, using in vitro amplification from total genomic DNA by the polymerase chain reaction. Use of this precise probe and oligo(T) as primers for in vitro amplification from a T. pyriformis cDNA library has allowed the cloning of the mutase gene. A similar amplification strategy from genomic DNA yielded the genomic sequence, which contains three introns. The sequence of the DNA that encodes 10 amino acids upstream of the N-terminal sequence of the isolated protein was found by oligonucleotide hybridization to a subgenomic library. These 10 N-terminal amino acids are cleanly removed in Tetrahymena in vivo. The full mutase gene sequence codes for a protein of 300 amino acids, and it includes two amber (TAG) codons in the open reading frame. In Tetrahymena, TAG codes for glutamine. When the two amber codons are each changed to a glutamine codon (CAG) that is recognized by E. coli and the gene is placed behind a promoter driven by the T7 RNA polymerase, expression in E. coli is observed. The mutase gene also contains a large number of arginine AGA codons, a codon that is very rarely used by E. coli. Cotransformation with a plasmid carrying the dnaY gene [which encodes tRNA(Arg)(AGA)] results in more than 4-fold higher expression. The mutase then comprises about 25% of the total soluble cell protein in E. coli transformants. The mutase gene bears significant similarity to one other gene in the available data bases, that of carboxyphosphonoenolpyruvate mutase from Streptomyces hygroscopicus, an enzyme that catalyzes a closely related transformation. Due to the large evolutionary distance between Tetrahymena and Streptomyces, this similarity can be interpreted as the first persuasive evidence that the biosynthesis of phosphonates is an ancient metabolic process.

Amino Acid Sequence↗

Nucleotide sequence of korB, a replication control gene of broad host-range plasmid RK2.

The korB gene is a major regulatory element in the replication and maintenance of broad host-range plasmid RK2. It negatively controls the replication gene trfA, the host-lethal determinants kilA and kilB, and the korA-korB operon. Here, we present the nucleotide sequence of an 1167 base-pair region that encodes korB. Using sequence data from korB mutants, we identified the korB structural gene. The predicted polypeptide product is negatively charged and has a molecular weight of 39,015, which is considerably less than that estimated by its electrophoretic mobility in SDS/polyacrylamide gels. Secondary-structure predictions of korB polypeptide revealed three closely spaced helix-turn-helix regions with significant homology to similar structures in known DNA-binding proteins. The korB gene, like all other sequenced RK2 genes, shows a strong preference for codons ending in a G or C residue. This is similar to codon usage by genes of Klebsiella and Pseudomonas, the original hosts for RK2 and some closely related plasmids. We also sequenced the site of transposon Tn76 insertion in the host-range mutant pRP761 and found it to be located immediately upstream from korB in the incC gene. Finally, we report the presence of sequences resembling a replication origin within the korB structural gene: a cluster of four 19 base-pair direct repeats and a nearby potential binding site for Escherichia coli dna A replication protein.

Base Composition↗

Allelic polymorphism of T-cell receptor constant domains is widespread in fishes.

T-cell receptor chains contain membrane-proximal constant domains of the immunoglobulin superfamily that are relatively invariant in mammalian species. In contrast, recent studies in the bicolor damselfish have demonstrated surprising allelic polymorphism in the TCR alpha ( A) and TCR beta ( B) "constant" (C) domain genes. This report extends these initial observations beyond Perciformes to two other orders of teleost fishes. Studies in both the Atlantic cod and zebrafish show high levels of polymorphism in the TCRA constant genes. Levels of 13% and 15% amino acid nonidentity were found within cod and zebrafish, respectively. Evolutionary analysis of codon usage suggests that positive selection maintains the high number of TCRAC alleles in these fish populations. Additionally, investigation of a TCRB constant gene from the Beau Gregory, a sister species of the bicolor damselfish, shows no evidence of transpecies maintenance of constant region alleles. These data argue that the T-cell receptor constant domain is being employed by many vertebrates in a manner inconsistent with our current understanding, and may indicate unheralded complexity in signal transduction through the TCR/CD3 complex.

Alleles↗

Coding properties of Oxytricha trifallax (Sterkiella histriomuscorum) macronuclear chromosomes: analysis of a pilot genome project.

The macronuclear genomes of spirotrichous ciliates are almost entirely polyploid, single-gene chromosomes ("nanochromosomes"). We recently performed a pilot genome project for a member of this group, Oxytricha trifallax ( Sterkiella histriomuscorum), in which approximately 2000 nanochromosomes were cloned at random and end-sequenced. Here we describe the global properties of the coding regions predicted for these molecules, including nucleotide composition, codon usage, and intron properties. In identifying splice donor, acceptor and branch sites, we found that longer introns in Oxytricha have a stronger signal at the donor site than do smaller introns, as has been found for Caenorhabditis elegans and Drosophila, despite the overall small size of the introns. A systematic search for multi-gene chromosomes identified 11 candidate nanochromosomes. We compare the results from this large dataset with those obtained from earlier studies and with statistics recorded from ciliates and other eukaryotes.

Animals↗

Complete sequences of the highly rearranged molluscan mitochondrial genomes of the Scaphopod Graptacme eborea and the bivalve Mytilus edulis.

We have determined the complete sequence of the mitochondrial genome of the scaphopod mollusk Graptacme eborea (14,492 nts) and completed the sequence of the mitochondrial genome of the bivalve mollusk Mytilus edulis (16,740 nts). (The name Graptacme eborea is a revision of the species formerly known as Dentalium eboreum.) G. eborea mtDNA contains the 37 genes that are typically found and has the genes divided about evenly between the two strands, but M. edulis contains an extra trnM and is missing atp8, and it has all genes on the same strand. Each has a highly rearranged gene order relative to each other and to all other studied mtDNAs. G. eborea mtDNA has almost no strand skew, but the coding strand of M. edulis mtDNA is very rich in G and T. This is reflected in differential codon usage patterns and even in amino acid compositions. G. eborea mtDNA has fewer noncoding nucleotides than any other mtDNA studied to date, with the largest noncoding region only 24 nt long. Phylogenetic analysis using 2,420 aligned amino acid positions of concatenated proteins weakly supports an association of the scaphopod with gastropods to the exclusion of Bivalvia, Cephalopoda, and Polyplacophora, but it is generally unable to convincingly resolve the relationships among major groups of the Lophotrochozoa, in contrast to the good resolution seen for several other major metazoan groups.

Animals↗

Four basic symmetry types in the universal 7-cluster structure of microbial genomic sequences.

Coding information is the main source of heterogeneity (non-randomness) in the sequences of microbial genomes. The heterogeneity corresponds to a cluster structure in triplet distributions of relatively short genomic fragments (200-400 bp). We found a universal 7-cluster structure in microbial genomic sequences and explained its properties. We show that codon usage of bacterial genomes is a multi-linear function of their genomic G+C-content with high accuracy. Based on the analysis of 143 completely sequenced bacterial genomes available in Genbank in August 2004, we show that there are four "pure" types of the 7-cluster structure observed. All 143 cluster animated 3D-scatters are collected in a database which is made available on our web-site (http://www.ihes.fr/~zinovyev/7clusters). The findings can be readily introduced into software for gene prediction, sequence alignment or microbial genomes classification.

Codon↗

The nucleotide sequence of the Escherichia coli fus gene, coding for elongation factor G.

We have determined the nucleotide sequence of the Escherichia coli fus gene, which codes for elongation factor G. The protein product of the sequenced gene contains 703 amino acids, with a predicted molecular weight of 77,444. The fus gene shows the nonrandom pattern of codon usage typical of ribosomal proteins and other proteins synthesized at a high level. We have identified several potential promoter sequences within the gene. One of these sequences may correspond to the secondary promoter for expression of the downstream tufA gene (encoding elongation factor Tu) whose activity has been described previously (1,2). A comparison of the nucleotide and amino acid sequences of elongation factors G and Tu reveals a limited but significant homology between the two proteins within the 150 amino acid residues at their amino-terminal ends.

Amino Acid Sequence↗

Analysis of pFQ31, a 8551-bp cryptic plasmid from the symbiotic nitrogen-fixing actinomycete Frankia.

The actinomycete Frankia has never been transformed genetically. To favour the development of Frankia cloning vectors, we have fully sequenced the Frankia alni pFQ31 cryptic plasmid and performed analyses to characterise its coding and non-coding regions. This plasmid is 8551 bp-long and contains 72% G+C. Computer-assisted analyses identified 18 open reading frames (ORFs). These ORFs show a synonymous codon usage different from the one of Frankia chromosomal genes, suggesting an evolutionary bias linked to the nature of the replicon or a horizontal transfer. Three ORFs were found to encode genes likely to be involved in plasmid replication and stability: parFA (partition protein), ptrFA (transcriptional repressor of the GntR family) and repFA (initiation of replication). DNA signatures of a replication origin were identified in the ptrFA-repFA intergenic region. These structural motifs are similar to those observed among origins of iteron-containing plasmids replicating via a θ mode.

Actinomycetales↗

Characterization of the virB operon from an Agrobacterium tumefaciens Ti plasmid.

The virulence genes of the Agrobacterium tumefaciens Ti plasmid are grouped into six transcription units and direct the transfer of T-DNA into plant cells. We report here the nucleotide sequence of the largest vir operon, virB, from the Ti plasmid pTiA6NC. This operon contains 11 open reading frames, 7 of which show evidence of translational coupling. trpE::virB gene fusions were used to confirm the reading frames of genes virB2, 4, 5, 6, 7, 8, 10, and 11. In addition, the native gene products of virB6 and virB9 were identified using maxicell and in vitro transcription-translation techniques, and the VirB9 protein was found to be proteolytically processed. The codon usage of the predicted virB genes is very similar to the other pTiA6 vir genes and is much less biased than Escherichia coli. Since many of the virB gene products have secretion signals common to exported bacterial proteins, it is likely that they will be membrane-associated. We propose that the VirB proteins are involved in the formation of a transmembrane structure which mediates the passage of the transferred T-DNA molecule through the bacterial and plant cell membranes.

Base Sequence↗

Molecular analysis of an essential gene upstream of rpoN in Rhizobium NGR234.

Rhizobium sp. NGR234 is a broad-host range strain. The rpoN gene of this organism encodes a sigma factor which is a primary co-regulator of endosymbiosis. We characterized the locus upstream of rpoN, and identified a contiguous open reading frame, here termed ORF1. DNA sequence analysis of this ORF showed that it encoded a polypeptide highly conserved with a corresponding ORF of Rhizobium meliloti. The gene product contained two ATP/GTP binding pockets. Codon usage in the ORF and the nitrogenase operon nifKDH of NGR234 was similar. Although we used a non-transposable cassette flanked by appropriate sized DNA fragments, we were unable to isolate site-directed mutants in the ORF, whose ATP/GTP binding protein product is thus probably of essential biological function. ORF1 and rpoN exhibited conserved linkage among diverse rhizobia, and in Azotobacter vinelandii. Intragenomic and interspecific homology studies confirmed directly that ORF1 (NGR234) belonged to a large family of ATP-binding protein genes.

Adenosine Triphosphate↗