Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,675 records · Page 93Linked to original sources

Monitoring phase-specific gene expression in Histoplasma capsulatum with telomeric GFP fusion plasmids.

Dimorphism is an essential feature of Histoplasma capsulatum pathogenesis, and much attention has been focused on characteristics that are unique to the saprophytic mycelial phase or the parasitic yeast phase. Recently, we identified a secreted calcium-binding protein, CBP, that is produced in large amounts by yeast cells but is undetectable in mycelial cultures. In this study, the green fluorescent protein (GFP) was established as a reporter in H. capsulatum to study regulation of CBP1 expression in cultures and in single cells grown under different conditions and inside macrophages. One GFP version that was optimized for human codon usage yielded highly fluorescent Histoplasma yeast cells. By monitoring GFP fluorescence during the transition from mycelia to yeast, we demonstrated that the CBP1 promoter is only fully active after complete morphological conversion to the yeast form, indicating for the first time that CBP1 is developmentally regulated rather than simply temperature regulated. Continuous activity of the CBP1 promoter during infection of macrophages supports the hypothesis that CBP secretion plays an important role for Histoplasma survival within the phagolysosome. Broth cultures of Histoplasma yeasts carrying a CBP-GFP protein fusion construct were able to secrete a full-length fluorescent fusion protein that remained localized within the phagolysosomes of infected macrophages. Additionally, a comparison of two Histoplasma strains carrying the CBP1 promoter fusion construct either epichromosomally or integrated into the chromosome revealed cell-to-cell variation in plasmid copy number due to uneven plasmid partitioning into daughter cells.

Animals↗

Elevated evolutionary rates in the laboratory strain of Saccharomyces cerevisiae.

By using the maximum likelihood method, we made a genome-wide comparison of the evolutionary rates in the lineages leading to the laboratory strain (S288c) and a wild strain (YJM789) of Saccharomyces cerevisiae and found that genes in the laboratory strain tend to evolve faster than in the wild strain. The pattern of elevated evolution suggests that relaxation of selection intensity is the dominant underlying reason, which is consistent with recurrent bottlenecks in the S. cerevisiae laboratory strain population. Supporting this conclusion are the following observations: (i) the increases in nonsynonymous evolutionary rate occur for genes in all functional categories; (ii) most of the synonymous evolutionary rate increases in S288c occur in genes with strong codon usage bias; (iii) genes under stronger negative selection have a larger increase in nonsynonymous evolutionary rate; and (iv) more genes with adaptive evolution were detected in the laboratory strain, but they do not account for the majority of the increased evolution. The present discoveries suggest that experimental and possible industrial manipulations of the laboratory strain of yeast could have had a strong effect on the genetic makeup of this model organism. Furthermore, they imply an evolution of laboratory model organisms away from their wild counterparts, questioning the relevancy of the models especially when extensive laboratory cultivation has occurred. In addition, these results shed light on the evolution of livestock and crop species that have been under human domestication for years.

Biological Evolution↗

Isolation of anti-CD22 Fv with high affinity by Fv display on human cells.

In vitro antibody affinity maturation has generally been achieved by display of mouse or human antibodies on the surface of microorganisms (phage, bacteria, and yeast). However, problems with protein folding, posttranslational modification, and codon usage still limit the number of improved antibodies that can be obtained. An ideal system would select and improve antibodies in a mammalian cell environment where they are naturally made. Here we show that human embryonic kidney 293T cells that are widely used for transient protein expression can be used for cell surface display of single-chain Fv antibodies for affinity maturation. In a proof-of-concept experiment, cells expressing a rare mutant antibody with higher affinity were enriched 240-fold by a single-pass cell sorting from a large excess of cells expressing WT antibody with a slightly lower affinity. Furthermore, we successfully obtained a highly enriched mutant with increased binding affinity for CD22 after a single selection of a combinatory library randomizing an intrinsic antibody hotspot. Important features are that one display selection cycle requires only 1 week, and transfection of cells in a single 100-mm dish produces 10(7) individual clones so that a repertoire of 10(9) is feasible under current experimental conditions.

Antibody Affinity↗

Genome analysis of the smallest free-living eukaryote Ostreococcus tauri unveils many unique features.

The green lineage is reportedly 1,500 million years old, evolving shortly after the endosymbiosis event that gave rise to early photosynthetic eukaryotes. In this study, we unveil the complete genome sequence of an ancient member of this lineage, the unicellular green alga Ostreococcus tauri (Prasinophyceae). This cosmopolitan marine primary producer is the world's smallest free-living eukaryote known to date. Features likely reflecting optimization of environmentally relevant pathways, including resource acquisition, unusual photosynthesis apparatus, and genes potentially involved in C(4) photosynthesis, were observed, as was downsizing of many gene families. Overall, the 12.56-Mb nuclear genome has an extremely high gene density, in part because of extensive reduction of intergenic regions and other forms of compaction such as gene fusion. However, the genome is structurally complex. It exhibits previously unobserved levels of heterogeneity for a eukaryote. Two chromosomes differ structurally from the other eighteen. Both have a significantly biased G+C content, and, remarkably, they contain the majority of transposable elements. Many chromosome 2 genes also have unique codon usage and splicing, but phylogenetic analysis and composition do not support alien gene origin. In contrast, most chromosome 19 genes show no similarity to green lineage genes and a large number of them are specialized in cell surface processes. Taken together, the complete genome sequence, unusual features, and downsized gene families, make O. tauri an ideal model system for research on eukaryotic genome evolution, including chromosome specialization and green lineage ancestry.

Animals↗

Cloning and nucleotide sequence of DNA coding for bovine preproparathyroid hormone.

We have cloned in Escherichia coli a DNA copy of mRNA coding for bovine preproparathyroid hormone. Double-stranded DNA was inserted into the Pst I site in plasmid pBR322 by using the poly(dG)-poly(dC) homopolymer extension technique to join the DNA molecules. Recombinant plasmids coding for preproparathyroid hormone were identified by the plasmid's ability to arrest specifically the translation of preproparathyroid hormone mRNA. The nucleotide sequence of the largest recombinant was determined by using both chemical and enzymatic techniques. The parathyroid insert contains 470 nucleotides--102 nucleotides from the 5' noncoding region of the mRNA, 345 nucleotides representing the entire coding region, and 23 nucleotides from the 3' noncoding region. The coding sequence clarifies the hormone's amino acid sequence, which has been disputed. Codon usage is discussed.

Amino Acid Sequence↗

Nucleotide sequence of a cyanobacterial nifH gene coding for nitrogenase reductase.

The nucleotide sequence of nifH, the structural gene for nitrogenase reductase (component II or Fe protein of nitrogenase) from the cyanobacterium Anabaena 7120 has been determined. Also reported are 194 bases of the 5'-flanking sequence and 170 bases of the 3'-flanking sequence. The predicted amino acid sequence was compared with that determined for the complete nitrogenase reductase of Clostridium pasteurianum and the cysteine-containing peptides of the protein from Azotobacter vinelandii. Amino acid sequences around five cysteines, located in the NH(2)-terminal two-thirds of the protein, are highly conserved in all three species. Codon usage in the first gene from a cyanobacterium to be sequenced shows striking asymmetries for eight amino acids.

Journal Article↗

Sequence of the gene coding for the beta-subunit of dinitrogenase from the blue-green alga Anabaena.

The nitrogen fixation nif K gene of the blue-green alga Anabaena, which codes for the beta-subunit of dinitrogenase, has been subjected to sequence analysis. The nif K protein is predicted to be 512 amino acids long, to have a M(r) or 57,583, and to contain six cysteine residues. Three of these cysteines are within peptides homologous to FeS cluster-binding cysteinyl peptides from ferredoxins and from a high potential iron protein and, thus, may be ligands to which FeS clusters bind in dinitrogenase. The sequences surrounding the cysteine residues are 70% homologous to the corresponding cysteinyl tryptic peptides of the Azotobacter vinelandii dinitrogenase, although the positions of the cysteine residues are not always conserved between the two proteins. A 15-amino acid coding sequence precedes nif K on its transcript. Amino acid codon usage is highly asymmetric and parallels that found for the Anabaena dinitrogenase reductase gene (nif H). Putative promoter and ribosome binding site sequences were identified for nif K. These regulatory sequences are homologous to sequences preceding nif D; nif D codes for the alpha-subunit of dinitrogenase but is separated from nif K on the chromosome by 11,000 nucleotides. The nif K promoter also is virtually identical to a promoter-like sequence that immediately precedes the start of the transcript for the large subunit of ribulosebisphosphate carboxylase from maize chloroplasts. This homology appears to support the theory that chloroplasts evolved from blue-green algae.

Journal Article↗

Isolation and sequence of the gene for the large subunit of ribulose-1,5-bisphosphate carboxylase from the cyanobacterium Anabaena 7120.

Cloned DNA probes containing genes coding for the large subunit of ribulose-1,5-bisphosphate carboxylase (rbcA) of corn and of Chlamydomonas were used to identify, by heterologous hybridization, DNA fragments from Anabaena 7120 carrying the corresponding gene sequence. The same probes were used to isolate, from a recombinant lambda library, a 17-kilobase-pair EcoRI Anabaena DNA fragment containing the coding sequence for the rbcA gene. The entire coding sequence, as well as 210 base pairs of 5' flanking region and 210 base pairs of 3' flanking region, was determined. Comparison of the nucleotide and amino acid sequences with those of corn, spinach, Chlamydomonas, and Synechococcus rbcA genes revealed homology of 71-77% at the nucleotide level and 80-85% at the amino acid level. Conservation of sequence is lost immediately outside the coding region on either side. Codon usage in the Anabaena rbcA gene is not significantly different from that in the Anabaena genes for nitrogenase reductase and nitrogenase beta subunit.

Journal Article↗

Sequence of the initiation factor IF2 gene: unusual protein features and homologies with elongation factors.

The gene for protein synthesis initiation factor IF2 in Escherichia coli, infB, is located downstream from nusA on the same operon. We sequenced about 3 kilobases of DNA beginning within nusA and including the entire infB structural gene plus another 392 bases downstream. This region contains no obvious strong promoter signals, but a possible transcriptional termination or pausing site occurs downstream from infB. The putative initiator codon for IF2 alpha (97,300 daltons) is AUG; that for IF2 beta (79,700 daltons) is GUG, located 471 bases downstream in the same reading frame. The codon usage for IF2 is typical of other highly expressed proteins in E. coli and suggests that IF2 mRNA is efficiently translated. IF2 alpha contains two adjacent regions (residues 104-155 and 167-214) that are rich in alanine and charged amino acids and that show striking periodicities in their sequences. These regions may alternate between flexible and helical conformations, thereby drawing together the NH2-terminal and COOH-terminal globular domains of the factor as IF2 interacts with ribosomes or tRNA. Certain regions of the DNA and protein sequences of IF2 share strong homologies with elongation factor EF-Tu and lesser homology with EF-G. In particular, a region of EF-Tu implicated in GTP binding contains sequences and secondary structure that are conserved in IF2. The homologies indicate that the genes for IF2 and the elongation factors are derived at least in part from a common ancestor.

Amino Acid Sequence↗

Tobacco chloroplast tRNA(UUU) gene contains a 2.5-kilobase-pair intron: An open reading frame and a conserved boundary sequence in the intron.

The nucleotide sequence of a tRNA(Lys)(UUU) gene on tobacco (Nicotiana tabacum) chloroplast DNA has been determined. This gene is located 215 base pairs upstream from the gene for the 32,000-dalton thylakoid membrane protein on the same DNA strand and has a 2526-base-pair intron in the anticodon loop. The intron boundary sequence does not follow the G-U/A-G rule but is similar to those of tobacco chloroplast split genes for tRNA(Gly)(UCC) and ribosomal proteins L2 and S12. The intron contains one major open reading frame of 509 codons. The codon usage in the open reading frame resembles those observed in the genes for tobacco chloroplast proteins so far analyzed. The primary transcript of this tRNA gene is 2.7 kilobases long.

Journal Article↗

The murine Fc receptor for immunoglobulin: purification, partial amino acid sequence, and isolation of cDNA clones.

The murine Fc receptor for IgG (Fc gamma R) was purified to homogeneity by immunoaffinity chromatography from detergent lysates of the macrophage cell line J774. Microsequencing of intact protein yielded a single amino-terminal sequence, which was confirmed and extended to 20 residues by the isolation of an overlapping peptide. The isolation of additional proteolytic fragments obtained by using Staphylococcus aureus V8 protease, cyanogen bromide, and lysine C proteinase, facilitated sequence analysis of a total of 119 amino acid residues. Codon usage charts were used to construct oligonucleotide probes based on the amino acid sequences of three nonoverlapping peptides. These probes were used to screen a cDNA library derived from the WEHI-3B myelomonocytic cell line, and a single cDNA clone (pFc24) to which all three probes hybridized was isolated. This clone, containing a 1.02-kilobase cDNA insert, has been characterized by restriction mapping and partial DNA sequencing, and it has been shown to encode the Fc gamma R. The sequence at the 5' end of the clone contained the coding information for the amino-terminal sequence of the Fc gamma R as well as a putative 13-amino acid signal sequence. The 3' end of the clone encoded a peptide identified in purified receptor preparations. Thus, the presence of coding information at the 5' and 3' ends of this clone suggests that full-length Fc receptor cDNA spans greater than 1 kilobase.

Amino Acid Sequence↗

Primary structure of the (1-->3,1-->4)-beta-D-glucan 4-glucohydrolase from barley aleurone.

During germination of barley grains, the cell walls of the starchy endosperm are degraded by (1-->3,1-->4)-beta-glucanases (EC 3.2.1.73) secreted from the aleurone and scutellar tissues. The complete sequence of the aleurone (1-->3,1-->4)-beta-glucanase isoenzyme II comprises 306 amino acids and was determined by sequencing nine tryptic peptides (110 residues) and aligning them with the amino acid sequence deduced from a cDNA clone encoding the 291 NH(2)-terminal residues. Although no amino acid sequence homology with a bacterial (1-->3)(1-->4)-beta-glucanase is apparent, close to 50% homology is found with two large regions of a (1-->3)-beta-glucanase from tobacco pith tissue. The gene for barley (1-->3,1-->4)-beta-glucanase isoenzyme II shares with that for the alpha-amylase isoenzyme 1 a strongly preferred use of codons with G and C in the wobble position (94% and 90%, respectively). Both enzymes are secreted from the aleurone cells during germination. Such one-sided codon usage is not characteristic for the gene encoding the (1-->3)-beta-glucanase of tobacco pith tissue or the hor2-4 gene encoding the B(1) hordein storage protein in the endosperm.

Journal Article↗

Molecular cloning of a gene belonging to the carcinoembryonic antigen gene family and discussion of a domain model.

Carcinoembryonic antigen (CEA) is a glycoprotein important as a tumor marker for colonic cancer. Immunological and biochemical studies have shown it to be closely related to a number of other glycoproteins, which together make up a gene family. We have cloned a member of this gene family by using long oligonucleotide probes (42-54 nucleotides) based on our protein sequence data for CEA and NCA (nonspecific cross-reacting antigen) and on human codon usage. The clone obtained (lambda 39.2) hybridizes with six probes and has a 15-kilobase insert. The 5' end of the gene is contained within a 2700-base-pair EcoRI fragment, which hybridizes with five of the six synthetic probes. Sequencing of the 5' end region revealed the location and structure of one exon and two putative intron boundaries. The exon encodes part of the leader sequence and the NH2-terminal 107 amino acids of NCA. Southern blot analysis of human normal and tumor DNA, using as probes two lambda 39.2 fragments that contain coding sequences, suggests the existence of 9-11 genes for the CEA family. One of the restriction fragments described here has been used by Zimmermann et al. [Zimmermann, W., Ortlieb, B., Friedrich, R. & von Kleist, S. (1987) Proc. Natl. Acad. Sci. USA 84, 2960-2964] to isolate partial cDNA clones for CEA. The identity of this clone was verified with our protein sequence data [Paxton, R., Mooser, G., Pande, H., Lee, T.D. & Shively, J.E. (1987) Proc. Natl. Acad. Sci. USA 84, 920-924]. We discuss a domain structure for CEA based on the CEA sequence data and the NCA exon sequence data. It is likely that this gene family evolved from a common ancestor shared with neural cell adhesion molecule and alpha 1 B-glycoprotein and is perhaps a family within the immunoglobulin superfamily.

Amino Acid Sequence↗

Translation of bicistronic viral mRNA in transfected cells: regulation at the level of elongation.

The S1 species of mammalian reovirus mRNA, like a number of other viral but not cellular mRNAs, codes for two dissimilar polypeptides by initiation of translation at two 5'-proximal, out-of-frame AUG codons. To determine if uninfected cells can utilize bicistronic genes, a bovine papilloma virus-based vector system was used to select mouse C127 cell lines containing multiple integrated copies of the reovirus S1 gene. These cell lines produced both reovirus polypeptides from a single mRNA. In addition, studies of COS cells transfected with the S1 gene containing small changes around the first AUG suggest that bicistronic mRNA translation is regulated at the level of elongation. A model is proposed in which ribosomes engaged in translation of one reading frame interfere with movement of ribosomes in the other frame because of differences in codon usage. Expression of bicistronic genes may be similarly regulated in virus-infected cells.

Animals↗

Chemical synthesis of the thymidylate synthase gene.

A 978-base-pair gene that encodes thymidylate synthase (TS; 5,10-methylenetetrahydrofolate:dUMP C-methyltransferase, EC 2.1.1.45) from Lactobacillus casei has been synthesized and inserted into Escherichia coli expression vectors. The DNA sequence contains 35 unique restriction sites that are located an average of 28 base pairs apart throughout the entire length of the gene. A ribosome binding site was included 9 base pairs upstream from the translation start site and codon usage was adjusted to ensure efficient translation in E. coli. The TS gene is flanked by unique EcoRI and HindIII restriction sites that render the gene portable to any of several E. coli expression vectors. Catalytically active TS encoded by the synthetic gene is expressed in large amounts (10-20% of the soluble protein) and is indistinguishable from that isolated from L. casei. The utility of the synthetic gene for mutagenesis is demonstrated by a single experiment in which His-199 was replaced with 14 different amino acids. Analysis of the mutants by genetic complementation indicates that TS can tolerate a number of amino acid substitutions at that position and shows that His-199 is not strictly required for catalytic activity.

Amino Acid Sequence↗

Isolation and structural characterization of a cDNA clone encoding the human DNA repair protein for O6-alkylguanine.

O6-Methylguanine-DNA methyltransferase (MGMT; DNA-O6-methylguanine:protein-L-cysteine S-methyltransferase, EC 2.1.1.63), a unique DNA repair protein present in most organisms, removes the carcinogenic and mutagenic adduct O6-alkylguanine from DNA by stoichiometrically accepting the alkyl group on a cysteine residue in a suicide reaction. The mammalian protein is highly regulated in both somatic and germ-line cells. In addition, the toxicity of certain alkylating drugs in tumor and normal cells is inversely related to the levels of this protein. The cDNA of the human gene, henceforth named MGMT, has been cloned in an expression vector on the basis of its rescue of a methyltransferase-deficient (ada-) Escherichia coli host. A 22-kDa active methyltransferase encoded entirely by the cDNA contains an amino acid sequence of 61 residues that bears 60-65% similarity with segments of E. coli methyltransferase (products of the ada and ogt genes), which encompass the alkyl-acceptor residues. The human cDNA has no sequence similarity with the ada and ogt genes, due in part to differences in codon usage, and shows no detectable homology with E. coli genomic DNA. However, it hybridizes with distinct restriction fragments of human, mouse, and rat DNAs. The lack of methyltransferase observed in many human cell lines is due to the absence of the MGMT gene or to lack of synthesis and/or stability of its 0.95-kilobase poly(A)+ RNA transcript.

Amino Acid Sequence↗

Evolution of Drosophila mitochondrial DNA and the history of the melanogaster subgroup.

The nucleotide sequences of a common region of 15 mitochondrial DNAs (mtDNAs) sampled from the Drosophila melanogaster subgroup were determined. The region is 2527 base pairs long, including most of the NADH dehydrogenase subunit 2 and cytochrome oxidase subunit 1 genes punctuated by three tRNA genes. The comparative study revealed (i) the extremely low saturation level of transitional differences, (ii) recombination or variable substitution rates even within species, (iii) long persistence times of distinct types of mtDNA in Drosophila simulans and Drosophila mauritiana, and (iv) an apparent lack of within-type variations in island species. Also found was a high correlation among the transitional rate, the saturation level, and the G + C content (or codon usage). It appears that D. simulans and D. mauritiana have maintained highly structured populations for more than 1 million years. Such structures are consistent with the origination of Drosophila sechellia from D. simulans. Yet geographic isolation is so weak as to show no evidence for further speciation. Moreover, one type of mtDNA shared by D. simulans and D. mauritiana suggests either recent divergence or ongoing introgression.

Animals↗

Roles of selection and recombination in the evolution of type I restriction-modification systems in enterobacteria.

Restriction-modification systems can protect bacteria against viral infection. Sequences of the hsdM gene, encoding one of the three subunits of type I restriction-modification systems, have been determined for four strains of enterobacteria. Comparison with the known sequences of EcoK and EcoR124 indicates that all are homologous, though they fall into three families (exemplified by EcoK, EcoA, and EcoR124), the first two of which are apparently allelic. The extent of amino acid sequence identity between EcoK and EcoA is so low that the genes encoding them might be better termed pseudoalleles; this almost certainly reflects genetic exchange among highly divergent species. Within the EcoK family the ratio of intra- to interspecific divergence is very high. The extent of divergence between the genes from Escherichia coli K-12 and Salmonella typhimurium LT2 is similar to that for other genes with the same level of codon usage bias. In contrast, intraspecific divergence (between E. coli strains B and K-12) is extremely high and may reflect the action of frequency-dependent selection mediated by bacteriophages. There is also evidence of lateral transfer of a short sequence between E. coli and S. typhimurium.

Amino Acid Sequence↗