Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Linkage limits the power of natural selection in Drosophila.

Population genetic theory shows that the efficacy of natural selection is limited by linkage-selection at one site interferes with selection at linked sites. Such interference slows adaptation in asexual genomes and may explain the evolutionary advantage of sex. Here, we test for two signatures of constraint caused by linkage in a sexual genome, by using sequence data from 255 Drosophila melanogaster and Drosophila simulans loci. We find that (i) the rate of protein adaptation is reduced in regions of low recombination, and (ii) evolution at strongly selected amino acid sites interferes with optimal codon usage at weakly selected, tightly linked synonymous sites. Together these findings suggest that linkage limits the rate and degree of adaptation even in recombining genomes.

Adaptation, Physiological↗

Human F1-ATPase: molecular cloning of cDNA for the beta subunit.

F1-ATPase is the major enzyme for ATP synthesis, and its beta subunit is the catalytic site. To date, no full-length cDNA for the eukaryotic F1 gene has been reported. Human F1 was studied because of its importance in medicine and cell biology. Here we report molecular cloning of a full-length cDNA for the human F1 beta subunit and purification of the human F1 beta subunit. The HeLa cell cDNA library constructed in an expression vector gamma gt11 was screened with antiserum against the yeast F1 beta subunit. One of the positive phage DNAs containing the human F1 beta gene and its flanking regions (1.8 kilobase pairs) was sequenced by the dideoxy chain termination method. The open reading frame started from a putative signal presequence, which was rich in both serine and arginine. There was a homologous segment in the signal presequence of human ornithine transcarbamoylase and that of F1 beta. The precursor of F1 beta was expressed in E. coli harboring a plasmid which had been constructed with T5 promotor and the F1 beta cDNA. Both the precursor and mature form of F1 beta were detected in HeLa cells in a pulse-chase experiment. The amino acid sequence of 480 residues (51,568.3 daltons) following the presequence was highly homologous with that of mature beef heart F1 beta (97.5%) and E. coli F1 beta (71.7%), but the codon usage in the human gene was very different from those of reported genes coding for F1 beta of other species.

Amino Acid Sequence↗

Comparison of the lipoprotein gene among the Enterobacteriaceae. DNA sequence of Erwinia amylovora lipoprotein gene.

A DNA sequence of 816 base pairs encompassing the entire Erwinia amylovora lipoprotein gene was determined. Sequence comparison between E. amylovora, Escherichia coli, and Serratia marcescens suggests that the structure of the lipoprotein has been highly conserved under the constraint of efficient gene expression selecting promoter structure, mRNA secondary structure, and codon usage in addition to the polypeptide function. The sequence also suggests that the lpp gene of the three bacteria diverged sequentially in the course of evolution.

Base Sequence↗

Expressed sequence tags of medaka (Oryzias latipes) liver mRNA.

A medaka liver cDNA library was constructed in lambda ZAPII. The number of clones in this library is approximately 5 x 10(6). Three hundred sixty-one clones were randomly selected and, of these, 33 clones with over 1 kb were sequenced. These sequences were compared with GenBank and the dbEST (Re.96.0). Twenty-five clones of these 33 clones encoded 18 different genes and 2 different expressed sequence tags (ESTs). Sequences of 10 of the 18 clones had not previously been reported in fish. The codon usage of medaka genes is similar to that of Xenopus, mouse, and human genes, but not similar to that of yeast and Escherichia coli genes.

Animals↗

Concerted action of multiple cis-acting sequences is required for Rev dependence of late human immunodeficiency virus type 1 gene expression.

Based on the human immunodeficiency virus type 1 (HIV-1) gag gene, subgenomic reporter constructs have been established allowing the contributions of different cis-acting elements to the Rev dependency of late HIV-1 gene products to be determined. Modification of intragenic regulatory elements achieved by adapting the codon usage of the complete gene to highly expressed mammalian genes resulted in constitutive nuclear export allowing high levels of Gag expression independent from the Rev/Rev-responsive element system and irrespective of the absence or presence of the isolated major splice donor. Leptomycin B inhibitor studies revealed that the RNAs derived from the codon-optimized gag gene lacking AU-rich inhibitory elements are directed to a distinct, CRM1-independent, nuclear export pathway.

Base Sequence↗

High-level production of fully active human alpha 1-antitrypsin in Escherichia coli.

The human alpha-1-antitrypsin (A1AT) gene expressed in Escherichia coli as a full-length, non-fusion gene product accumulates to a relatively low level approaching less than or equal to 0.1% of total cellular protein. In contrast, deletion of the first 5, 10 or 15 codons leads to production of truncated A1AT derivatives at levels between 10 and 30% of total cellular protein. The protein with the largest truncation was insoluble and inactive following solubilization by chaotropic agents. In contrast, the two derivatives with the smaller truncations were found to be soluble, and exhibit identical specific activities in both trypsin and elastase inhibition assays to authentic human A1AT. The expression of the full-length A1AT was also optimized by making silent third position mutations within its first 15 codons. These mutations were chosen to optimize codon usage and minimize the possibility of RNA secondary structure formation in this region. Via this approach, expression of full-length, authentic, fully active A1AT was increased at least 20-fold to 2% of total cellular protein. Optimal expression was obtained using as few as three silent mutations in the first five codons, confirming the importance of this 5'-terminal region as had been defined by our deletion mutants. Both the full-length derivatives as well as the small N-terminal deletion derivative can be readily purified from bacterial extracts in fully active form suitable for the examination of their potential therapeutic application.

Base Sequence↗

The organization, localization and nucleotide sequence of the histone genes of the midge Chironomus thummi.

Several histone gene repeating units containing the genes for histones H1, H2A, H2B, H3 and H4 were isolated by screening a genomic DNA library from the midge Chironomus thummi ssp. thummi. The nucleotide sequence of one complete histone gene repeating unit was determined. This repeating unit contains one copy of each of the five histone genes in the order and orientation mean value of H3 H4 mean value of H2A H2B H1 mean value of. The overall length is 6262 bp. The orientation, nucleotide sequence and inferred amino acid sequence as well as the chromosomal arrangement and localization are different from those reported for Drosophila melanogaster. The codon usage also shows marked differences between Chironomus and Drosophila. Thus the histone gene structure reported for Drosophila is not typical of all insects.

Amino Acid Sequence↗

Evidence for horizontal transfer from Streptococcus to Escherichia coli of the kfiD gene encoding the K5-specific UDP-glucose dehydrogenase.

Capsular polysaccharides are important virulence factors both in Gram-positive and Gram-negative bacteria. A similar cluster organization of the genes involved in the synthesis of bacterial exopolysaccharides has been postulated in both cases, suggesting that these clusters evolved by module assembly. Horizontal gene transfer has been postulated to explain the polymorphism found in these cellular polymers. The cap1 K and cap3A genes coding for the pneumococcal type 1 and type 3 UDP-glucose dehydrogenases, respectively, have been compared with other UDP-sugar dehydrogenases. We have observed that the evolutionary distance between Cap1K and Cap3A is approximately equal to that found between Cap1K (or Cap3A) and other UDP-GlcDH of families evolutionarily distant like KfiD, the dehydrogenase from Escherichia coli K5. On the basis of comparisons of G + C content, patterns of synonymous and nonsynonymous substitutions, dinucleotide frequencies, and codon usage bias, we conclude that the kfiD gene has been introduced into E. coli from an exogenous source, probably from a streptococcal species.

Bacterial Capsules↗

Structure of a full-length cDNA clone for the prepro alpha 1(I) chain of human type I procollagen.

A full-length cDNA clone for the human prepro alpha 1(I) chain of type I procollagen was characterized. Nucleotide sequencing of the first 1500 nucleotide residues of the 5'-end of the cDNA clone provided 729 nucleotide residues and the codons for 243 amino acid residues not previously defined from any species. The data made it possible, for the first time, to compare completely codon usage for the human alpha 1(I) and alpha 2(I) chains.

Amino Acid Sequence↗

CpsK of Streptococcus agalactiae exhibits alpha2,3-sialyltransferase activity in Haemophilus ducreyi.

Streptococcus agalactiae (GBS) is a major cause of serious newborn bacterial infections. Crucial to GBS evasion of host immunity is the production of a capsular polysaccharide (CPS) decorated with sialic acid, which inactivates the alternative complement pathway. The CPS operons of serotypes Ia and III GBS have been described, but the CPS sialyltransferase gene was not identified. We identified cpsK, an open reading frame in the CPS operon of most serotypes, which was homologous to the lipooligosaccharide (LOS) sialyltransferase gene, lst, of Haemophilus ducreyi. To determine if cpsK might encode a sialyltransferase, we complemented a H. ducreyi lst mutant with cpsK. CpsK was expressed in H. ducreyi and LOS was isolated and analysed for sialic acid content by SDS-PAGE and high-performance liquid chromatography (HPLC). Sialo-LOS was seen in the wild-type, cpsK- or lst-complemented mutant strains, but not in the mutant without cpsK. Addition of Neu5Ac to the LOS was confirmed by mass spectroscopy. Lectin binding studies detected terminal Neu5Ac(alpha 2-->3)Gal(beta 1- on LOS produced by the wild-type, cpsK or lst-complemented mutant strain LOS, compared with the mutant alone. Our data characterize the first sialyltransferase gene from a Gram- positive bacterium and provide compelling evidence that its product catalyses the alpha2,3 addition of Neu5Ac to H. ducreyi LOS and therefore the terminal side-chain of GBS CPS. Phylogenetic studies further indicated that lst and cpsK are related but distinct from sialyltransferases of most other bacteria and, along with their similar codon usage bias and G + C content, suggests acquisition by lateral transfer from an ancestral low G + C organism.

Amino Acid Sequence↗

Metabolic efficiency and amino acid composition in the proteomes of Escherichia coli and Bacillus subtilis.

Biosynthesis of an Escherichia coli cell, with organic compounds as sources of energy and carbon, requires approximately 20 to 60 billion high-energy phosphate bonds [Stouthamer, A. H. (1973) Antonie van Leeuwenhoek 39, 545-565]. A substantial fraction of this energy budget is devoted to biosynthesis of amino acids, the building blocks of proteins. The fueling reactions of central metabolism provide precursor metabolites for synthesis of the 20 amino acids incorporated into proteins. Thus, synthesis of an amino acid entails a dual cost: energy is lost by diverting chemical intermediates from fueling reactions and additional energy is required to convert precursor metabolites to amino acids. Among amino acids, costs of synthesis vary from 12 to 74 high-energy phosphate bonds per molecule. The energetic advantage to encoding a less costly amino acid in a highly expressed gene can be greater than 0.025% of the total energy budget. Here, we provide evidence that amino acid composition in the proteomes of E. coli and Bacillus subtilis reflects the action of natural selection to enhance metabolic efficiency. We employ synonymous codon usage bias as a measure of translation rates and show increases in the abundance of less energetically costly amino acids in highly expressed proteins.

Amino Acids↗

Expressed sequence tags from immature female sexual organ of a liverwort, Marchantia polymorpha.

A total of 970 expressed sequence tag (EST) clones were generated from immature female sexual organ of a liverwort, Marchantia polymorpha. The 376 ESTs resulted in 123 redundant groups, thus the total number of unique sequences in the EST set was 717. Database search by BLAST algorithm showed that 302 of the unique sequences shared significant similarities to known nucleotide or amino acid sequences. Six unique sequences showed significant similarities to genes that are involved in flower development and sexual reproduction, such as cynarase, fimbriata-associated protein and S-receptor kinase genes. The remaining unique 415 sequences have no significant similarity with any database-registered genes or proteins. The redundant 123 ESTs implied the presence of gene families and abundant transcripts of unknown identity. Analyses of the coding sequences of 61 unique sequences, which contained no ambiguous bases in the predicted coding regions, highly homologous to known sequences at the amino acid level with a similarity score greater than 400, and with stop codons at similar positions as their possible orthologues, indicated the presence of biased codon usage and higher GC content within the coding sequences (50.4%) than that within 3' flanking sequences (41.9%).

Amino Acid Sequence↗

Molecular cloning and characterization of the rfc gene of Pseudomonas aeruginosa (serotype O5).

Previous work from our laboratory has shown that cosmid clone pFV100, containing a 26 kb insert, is able to restore O-antigen synthesis in serotype O5 rough mutants of Pseudomonas aeruginosa. Mobilization of pFV100 into two P. aeruginosa semi-rough (SR) mutants, AK14O1 and rd7513, resulted in O-antigen expression, indicating that pFV100 may contain an O-polymerase (rfc) gene. pFV.TK6, a subclone of pFV100 that contains a 5.6 kb chromosomal insert, was able to complement O-antigen expression in these SR mutants. Mutagenesis of pFV.TK6 using Tn1000 exposed a 1.5 kb region that was essential for complementing O-antigen expression in AK14O1. A 2.0 kb XhoI-HindIII fragment, containing this region, was cloned into vector pUCP26 and the resulting plasmid called pFV.TK8. In Southern analysis of the 20 P. aeruginosa serotypes using a probe generated from the 1.5 kb XhoI fragment of pFV.TK8, the rfc probe hybridized to a common fragment of the cross-reactive O2-O5-O16-O18-O20 serogroup, suggesting that these serotypes may share a common O-polymerase gene. In functional studies of the rfc gene, the PAO1 (serotype O5) chromosomal rfc was mutated using a gene-replacement strategy. These knockout mutants expressed the SR lipopolysaccharide (LPS) phenotype, which indicated that they were no longer producing a functional O-polymerase enzyme. Nucleotide sequence analysis of the insert DNA of pFV.TK8 revealed one open reading frame (ORF), designated ORF48.9, which could code for a 48.9 kDa protein. In comparisons of the P. aeruginosa rfc nucleotide and amino acid sequences with DNA and protein databases, no significant homology was found. However, the deduced structure of the P. aeruginosa Rfc protein indicated that it is very hydrophobic and contains 11 putative membrane-spanning domains. Therefore, the predicted structure is similar to that of other reported Rfc proteins. Furthermore, comparison of the amino acid composition and codon usage of the P. aeruginosa Rfc with other Rfc proteins revealed significant similarity between them.

Amino Acid Sequence↗

Chemical synthesis and in vivo hyperexpression of a modular gene coding for Escherichia coli translational initiation factor IF1.

An artificial gene encoding the Escherichia coli translational initiation factor IF1 was synthesized based on the primary structure (71 amino acid residues) of the protein. Codons for individual amino acids were selected on the basis of the preferred codon usage found in the structural genes for the initiation factor IF2 of E. coli and Bacillus stearothermophilus, both of which can be expressed at high levels in E. coli cells. We gave the IF1 gene a modular structure by introducing specific restriction enzyme sites into the sequence, resulting in units of three to ten codons. This was conceived to facilitate site-directed mutagenesis of the gene and thus to obtain IF1 with specific amino acid alterations at desired positions. The IF1 gene was assembled by shot-gun ligation of 9 synthetic oligodeoxyribonucleotides ranging in size from 31 to 65 nucleotides and cloned into an expression vector to place the gene under the control of an inducible promoter. Upon induction, E. coli cells harbouring the artificial gene were found to produce large amounts (greater than or equal to 60 mg/100 g cells) of a protein indistinguishable from natural IF1 in both chemical and biological properties.

Base Sequence↗

Identification of the nuclear-encoded chloroplast ribosomal protein L12 of the monocotyledonous plant Secale cereale and sequencing of two different cDNAs with strong codon bias.

Two different cDNA clones (SCL12-1 and SCL12-2) encoding precursors of a chloroplast ribosomal protein with homology to L12 from Escherichia coli were isolated from rye leaf cDNA libraries and sequenced. The corresponding polypeptide of rye chloroplast ribosomes was identified. The sequences for the mature proteins of M(r) 13,447 and 13,609 share 85% amino acid identity. The mature polypeptide of clone SCL12-1 has an amino acid identity of 71%, 72% or 44%, respectively, relative to L12 proteins from spinach, tobacco, or E. coli. Codon usage of the rye L12 cDNAs shows a high preference (97% and 82%) for G or C in the third base position.

Amino Acid Sequence↗

Multiplicative versus additive selection in relation to genome evolution: a simulation study.

The evolution of molecular quantitative traits, such as codon usage bias or base frequencies, can be explained as the result of mutational biases alone, or as the result of mutation and selection. Whereas mutation models can be investigated easily, realistic modelling of selection-directed genome evolution is analytically intractable, and numerical calculations require substantial computer resources. We investigated the evolution of optimal codon frequency under additive and multiplicative effects of selected linked codons. We show that additive selective effects of many linked sites cannot be effective in genomes when the number of selected sites is greater than the effective population size, a realistic assumption according to current molecular data. We then discuss the implications of these results for isochore evolution in vertebrates.

Codon↗

Isolation and characterization of the actin gene from the cellulolytic fungus Humicola grisea and analysis of transcription levels of actin and cellulase genes.

An actin gene was isolated from the cellulolytic fungus Humicola grisea. The gene structure, which has 5 introns in the coding region, is similar to those of the so far cloned fungal actin genes. But there are some differences in intron sizes and codon usage. Transcription levels of actin and cellulase genes were also investigated.

5' Untranslated Regions↗