Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Nucleotide sequence of Candida pelliculosa beta-glucosidase gene.

The nucleotide sequence of the DNA fragment containing the beta-glucosidase gene of Candida pelliculosa was determined. Analysis of the sequence revealed three open reading frames which could encode 65,825, and 412 amino acid residues. The presence of the second frame was found to be sufficient for the expression of the beta-glucosidase gene in a heterologous host Saccharomyces cerevisiae. Putative protein encoded by this gene had hydrophobic amino acids, resembling a signal peptide, at its N-terminal region and 19 potential glycosylation sites. Codon usage of Candida genes had the similar pattern shown in S.cerevisiae. Codon bias of the beta-glucosidase gene of Candida was relatively low, compared with that of the highly expressed genes of S. cerevisiae.

Amino Acid Sequence↗

Purification, cloning, and sequence of outer membrane protein P1 of Haemophilus influenzae type b.

Outer membrane protein P1 from Haemophilus influenzae type b MinnA was purified and partially characterized. Antiserum was generated against the purified protein and was used to immunologically screen a lamba EMBL3 genomic library prepared from strain MinnA DNA. A 4.2-kilobase-pair EcoRI-BamHI fragment containing the P1 gene was subcloned into pBR322. The recombinant protein was synthesized by Escherichia coli K-12, in which it localized to the outer membrane. The N-terminal sequence of the purified protein was determined and found to correspond to residues 23 through 36. The 22-amino-acid leader peptide had a typical structure, with two lysine residues near the amino terminus, a stretch of hydrophobic residues, and alanine residues at positions 20 and 22. The Mr of the processed protein was 47,752, which is in good agreement with the estimate of 50,000 from sodium dodecyl sulfate-polyacrylamide gel electrophoresis. Putative -35 and -10 promoter sequences were identified upstream from the translational start site. Codon usage was examined and determined to be substantially different than the codon preference in E. coli.

Amino Acid Sequence↗

Insertion of N-linked glycosylation sites in the variable regions of the human immunodeficiency virus type 1 surface glycoprotein through AAT triplet reiteration.

Variable regions with sequence length variation in the human immunodeficiency virus type 1 envelope exhibit an unusual pattern of codon usage with AAT, ACT, and AGT together composing > 70% of all codons used. We postulate that this distribution is caused by insertion of AAT triplets followed by point mutations and selection. Accumulation of the encoded amino acids (asparagine, serine, and threonine) leads to the creation of new N-linked glycosylation sites, which helps the virus to escape from the immune pressure exerted by virus-neutralizing antibodies.

Base Sequence↗

On the evolution of codon volatility.

Volatility of a codon is defined as the probability that a random point mutation in the codon generates a nonsynonymous change. It has been proposed that higher-than-expected mean codon volatility of a gene indicates that positive selection for nonsynonymous changes has acted on the gene in the recent past. I show that strong frequency-dependent selection (minority advantage) in large populations can increase codon volatility slightly, whereas directional positive selection has no effect on volatility. Factors unrelated to positive selection, such as expression-related or GC-content-related codon usage bias, also affect volatility. These and other considerations suggest that codon volatility has only limited utility for detecting positive selection at the DNA sequence level.

Animals↗

Amplification and molecular cloning of the IMP dehydrogenase gene of Leishmania donovani.

A mutant (MPA100) strain of Leishmania donovania was generated from a wild type (D1700) population by virtue of its ability to survive the selective pressure of gradually increasing concentrations of mycophenolic acid (MPA), an inhibitor of IMP dehydrogenase (IMPDH) activity. Comparative growth experiments revealed that the MPA100 strain was 100-fold more resistant to MPA toxicity and cross-resistant to ribavarin, another inhibitor of IMPDH. A direct comparison of IMPDH levels in D1700 and MPA100 cells showed that the latter expressed at least 20-fold higher enzyme activity. In order to evaluate the mechanism by which MPA100 cells overexpressed IMPDH, the leishmanial gene encoding IMPDH was isolated from a genomic library in EMBL3 by cross-hybridization to a mouse IMPDH cDNA, and a 2.3-kilobase EcoRV-PstI fragment was subcloned into a Bluescript vector and sequenced. The EcoRV-PstI fragment contained an open reading frame of 514 amino acids that encompassed the entire leishmanial IMPDH coding sequence. The predicted amino acid sequence showed a 52.5% identity with that of the corresponding human IMPDH. The codon usage of the leishmanial IMPDH gene reflected a strong bias toward codons containing either G or C in the wobble position. The EcoRV-PstI fragment hybridized to a 3.0-kilobase mRNA that was expressed at 10-20-fold greater levels in the MPA100 cells. Using the EcoRV-PstI fragment as a probe, the increased amount of IMPDH activity and IMPDH mRNA in the MPA100 cells could be attributed to an approximately 10-20-fold amplification of the leishmanial IMPDH gene.

Amino Acid Sequence↗

Open reading frames in the antisense strands of genes coding for glycolytic enzymes in Saccharomyces cerevisiae.

Open reading frames longer than 300 bases were observed in the antisense strands of the genes coding for the glycolytic enzymes phosphoglucose isomerase, phosphoglycerate mutase, pyruvate kinase and alcohol dehydrogenase I. The open reading frames on both strands are in codon register. It has been suggested that proteins coded in codon register by complementary DNA strands can bind to each other. Consequently, it was interesting to investigate whether the open reading frames in the antisense strands of glycolytic enzyme genes are functional. We used oligonucleotide-directed mutagenesis of the PGI1 phosphoglucose isomerase gene to introduce pairs of closely spaced base substitutions that resulted in stop codons in one strand and only silent replacements in the other. Introduction of the two stop codons into the PGI1 sense strand caused the same physiological defects as already observed for pgil deletion mutants. No detectable effects were caused by the two stop codons in the antisense strand. A deletion that removed a section from -31 bp to +109 bp of the PGI1 gene but left 83 bases of the 3' region beyond the antisense open reading frame had the same phenotype as a deletion removing both reading frames. A similar pair of deletions of the PYK1 gene and its antisense reading frame showed identical defects. Our own Northern experiments and those reported by other authors using double-stranded probes detected only one transcript for each gene. These observations indicate that the antisense reading frames are not functional. On the other hand, evidence is provided to show that the rather long reading frames in the antisense strands of these glycolytic enzyme genes could arise from the strongly selective codon usage in highly expressed yeast genes, which reduces the frequency of stop codons in the antisense strand.

Base Sequence↗

Molecular evolution of duplicated ray finned fish HoxA clusters: increased synonymous substitution rate and asymmetrical co-divergence of coding and non-coding sequences.

In this study the molecular evolution of duplicated HoxA genes in zebrafish and fugu has been investigated. All 18 duplicated HoxA genes studied have a higher non-synonymous substitution rate than the corresponding genes in either bichir or paddlefish, where these genes are not duplicated. The higher rate of evolution is not due solely to a higher non-synonymous-to-synonymous rate ratio but to an increase in both the non-synonymous as well as the synonymous substitution rate. The synonymous rate increase can be explained by a change in base composition, codon usage, or mutation rate. We found no changes in nucleotide composition or codon bias. Thus, we suggest that the HoxA genes may experience an increased mutation rate following cluster duplication. In the non-Hox nuclear gene RAG1 only an increase in non-synonymous substitutions could be detected, suggesting that the increased mutation rate is specific to duplicated Hox clusters and might be related to the structural instability of Hox clusters following duplication. The divergence among paralog genes tends to be asymmetric, with one paralog diverging faster than the other. In fugu, all b-paralogs diverge faster than the a-paralogs, while in zebrafish Hoxa-13a diverges faster. This asymmetry corresponds to the asymmetry in the divergence rate of conserved non-coding sequences, i.e., putative cis-regulatory elements. These results suggest that the 5' HoxA genes in the same cluster belong to a co-evolutionary unit in which genes have a tendency to diverge together.

Animals↗

Codon distribution in vertebrate genes may be used to predict gene length.

I have analysed the coding regions of 96 eukaryotic genes for their use of iso-coding codons. Specific codons occur more frequently in specific positions in all members of some gene families than would be expected if codon choice was determined solely by the frequency of codon usage. In the absence of evidence a priori for selection for particular codons at particular positions, I term such co-occurring codons "coincident codons". Coincident codons are not confined to particular regions of genes, and their occurrence is not detectably linked with the location of introns in the genomic sequence. Their presence is partly but not completely explained by the exchange of sequence between similar functional genes within a species: homologous genes from different organisms also possess the same codons at some sites with greater than expected frequencies. The relative excess of coincident codons correlates well with the overall length of the genes analysed, but not with the length of mRNA or coding regions, or with qualitative features of gene structure or expression. This, and the unusual sequence environment of coincident codons, suggests that they are a feature of the overall secondary structure of the heterogeneous nuclear RNA. Such considerations suggest approaches for optimizing the expression of exogenous genes in eukaryotic systems, and for predicting the structure of genes for which only partial sequence data is available.

Actins↗

Nucleotide sequence of an actin-encoding gene from Hydra attenuata: structural characteristics and evolutionary implications.

We have determined the complete nucleotide sequence of an actin-encoding gene from Hydra attenuata as well as partial sequences of cDNA clones from two additional actin-encoding genes. The gene from the genomic clone contains a single intron, and has promoter and polyadenylation signals similar to those found in other species. The hydra genome has a very A + T-rich base composition (71%). This is reflected in the codon usage of the actin-encoding genes, which is strongly biased towards codons having A or T in the third position. The hydra actin-encoding gene family consists of three or more transcribed genes, two of which are very closely related to each other and probably arose by a recent gene duplication. Hydra actin, like other invertebrate actins, is more similar to the non-muscle isotypes of vertebrates than to the vertebrate muscle actins. Hydra actin is more similar to animal actins than to those of plants or fungi, which is consistent with the view that all metazoans arose from a single protist ancestor.

Actins↗

Effects of codon-optimization on protein expression by the human herpesvirus 6 and 7 U51 open reading frame.

Codon-optimization refers to the alteration of gene sequences, to make codon usage match the available tRNA pool within the cell/species of interest. Codon-optimization has emerged as a powerful tool to increase protein expression by genes from small RNA and DNA viruses, which commonly contain overlapping reading frames as well as structural elements that are embedded within coding regions; these features are not widespread among large DNA viruses. We therefore examined whether codon-optimization might influence protein expression from a herpesvirus gene. We focused on the U51 gene from human herpesviruses-6 and -7, which was cloned in both native and codon-optimized form, with an N-terminal HA epitope tag to allow protein detection. Codon-optimization was associated with a profound (10-100 fold) increase in U51 expression in human (293A, HSG, K562) or hamster (CHO) cell lines, suggesting this may represent a valuable tool to facilitate functional studies on recalcitrant herpesvirus genes. Finally, it is postulated that the suboptimal expression of native U51 may reflect a regulatory mechanism that controls viral gene expression.

Animals↗

Codon adaptation and synonymous substitution rate in diatom plastid genes.

Diatom plastid genes are examined with respect to codon adaptation and rates of silent substitution (Ks). It is shown that diatom genes follow the same pattern of codon usage as other plastid genes studied previously. Highly expressed diatom genes display codon adaptation, or a bias toward specific major codons, and these major codons are the same as those in red algae, green algae, and land plants. It is also found that there is a strong correlation between Ks and variation in codon adaptation across diatom genes, providing the first evidence for such a relationship in the algae. It is argued that this finding supports the notion that the correlation arises from selective constraints, not from variation in mutation rate among genes. Finally, the diatom genes are examined with respect to variation in Ks among different synonymous groups. Diatom genes with strong codon adaptation do not show the same variation in synonymous substitution rate among codon groups as the flowering plant psbA gene which, previous studies have shown, has strong codon adaptation but unusually high rates of silent change in certain synonymous groups. The lack of a similar finding in diatoms supports the suggestion that the feature is unique to the flowering plant psbA due to recent relaxations in selective pressure in that lineage.

Adaptation, Physiological↗

Organization and nucleotide sequences of ten ribosomal protein genes from the region equivalent to the S10 operon in the archaebacterium, Halobacterium halobium.

A determination was made of the nucleotide sequence of the 7340-bp region of a ribosomal protein gene cluster of Halobacterium halobium, which is equivalent to the S10 operon of Escherichia coli. The sequence was analyzed with the codonpreference program deduced from the halobacterial codon usage table that showed a very high GC content of the third codon position. The sequence was comprised of a string of 13 tightly linked ORFs. Most of the ORFs were homologous with ribosomal protein genes (ORF1-ORF2-rpl3-rpl4-rpl23--rpl2- rps19-rpl22-rps3-rpl29-ORF11-rps17-r pl14). The 13-gene string was preceded by three putative AT-rich promoter sequences. The order of the genes in H. halobium essentially agreed with that of the corresponding genes of E. coli (S10-operon), except for certain deletions or insertions of additional protein genes.

Amino Acid Sequence↗

Inferring weak selection from patterns of polymorphism and divergence at "silent" sites in Drosophila DNA.

Patterns of codon usage and "silent" DNA divergence suggest that natural selection discriminates among synonymous codons in Drosophila. "Preferred" codons are consistently found in higher frequencies within their synonymous families in Drosophila melanogaster genes. This suggests a simple model of silent DNA evolution where natural selection favors mutations from unpreferred to preferred codons (preferred changes). Changes in the opposite direction, from preferred to unpreferred synonymous codons (unpreferred changes), are selected against. Here, selection on synonymous DNA mutations is investigated by comparing the evolutionary dynamics of these two categories of silent DNA changes. Sequences from outgroups are used to determine the direction of synonymous DNA changes within and between D. melanogaster and Drosophila simulans for five genes. Population genetics theory shows that differences in the fitness effect of mutations can be inferred from the comparison of ratios of polymorphism to divergence. Unpreferred changes show a significantly higher ratio of polymorphism to divergence than preferred changes in the D. simulans lineage, confirming the action of selection at silent sites. An excess of unpreferred fixations in 28 genes suggests a relaxation of selection on synonymous mutations in D. melanogaster. Estimates of selection coefficients for synonymous mutations (3.6 < magnitude of Nes < 1.3) in D. simulans are consistent with the reduced efficacy of natural selection (magnitude of Nes < 1) in the three- to sixfold smaller effective population size of D. melanogaster. Synonymous DNA changes appear to be a prevalent class of weakly selected mutations in Drosophila.

Animals↗

Mitochondrial genomes of Galathealinum, Helobdella, and Platynereis: sequence and gene arrangement comparisons indicate that Pogonophora is not a phylum and Annelida and Arthropoda are not sister taxa.

We report a contiguous region of more than half (> 7,500 nt) of the mitochondrial genomes for Platynereis dumerii (Annelida: Polychaeta), Helobdella robusta (Annelida: Hirudinida), and Galathealinum brachiosum (Pogonophora: Perviata). The relative arrangements of all 22 genes identified for Helobdella and Galathealinum are identical to one another and to their arrangements in the mtDNA of the previously studied oligochaete annelid Lumbricus. In contrast, Platynereis differs from these taxa in the positions of several tRNA genes and in having two additional tRNA genes (trnC and trnM) and a large noncoding sequence in this region. Comparisons of relative gene arrangements and of the nucleotide and inferred amino acid sequences among these and other published taxa provide strong support for an annelid-mollusk clade that excludes arthropods, and for the inclusion of pogonophorans within Annelida, rather than giving them separate phylum status. Gene arrangement comparisons include the first use of a recently described method on previously unpublished data. Although a variety of alternative initiation codons are typically used by mitochondrial protein-encoding genes, ATG appears to be the initiator for all but one reported here. The large noncoding region (1,091 nt) identified in Platynereis has no significant sequence similarity to the noncoding region of Lumbricus, although each contains runs of TA dinucleotides and of homopolymers, which could potentially serve as signaling elements. There is strong bias for synonymous codon usage in Helobdella and especially in Galathealinum. In this latter taxon, 5 codons are completely unused, 13 are used three or fewer times, and G appears at third codon positions in only 26 of the 2,236 codons. Nucleotide composition bias appears to influence amino acid composition of the proteins.

Amino Acid Sequence↗

Identification of a novel operon in Lactococcus lactis encoding three enzymes for lactic acid synthesis: phosphofructokinase, pyruvate kinase, and lactate dehydrogenase.

The discovery of a novel multicistronic operon that encodes phosphofructokinase, pyruvate kinase, and lactate dehydrogenase in the lactic acid bacterium Lactococcus lactis is reported. The three genes in the operon, designated pfk, pyk, and ldh, contain 340, 502, and 325 codons, respectively. The intergenic distances are 87 bp between pfk and pyk and 117 bp between pyk and ldh. Plasmids containing pfk and pyk conferred phosphofructokinase and pyruvate kinase activity, respectively, on their host. The identity of ldh was established previously by the same approach (R. M. Llanos, A. J. Hillier, and B. E. Davidson, J. Bacteriol. 174:6956-6964, 1992). Each of the genes is preceded by a potential ribosome binding site. The operon is expressed in a 4.1-kb transcript. The 5' end of the transcript was determined to be a G nucleotide positioned 81 bp upstream from the pfk start codon. The pattern of codon usage within the operon is highly biased, with 11 unused amino acid codons. This degree of bias suggests that the operon is highly expressed. The three proteins encoded on the operon are key enzymes in the Embden-Meyerhoff pathway, the central pathway of energy production and lactic acid synthesis in L. lactis. For this reason, we have called the operon the las (lactic acid synthesis) operon.

Amino Acid Sequence↗

Estimating the "effective number of codons": the Wright way of determining codon homozygosity leads to superior estimates.

In 1990, Frank Wright introduced a method for measuring synonymous codon usage bias in a gene by estimation of the "effective number of codons," N(c). Several attempts have been made recently to improve Wright's estimate of N(c), but the methods that work in cases where a gene encodes a protein not containing all amino acids with degenerate codons have not been tested against each other. In this article I derive five new estimators of N(c) and test them together with the two published estimators, using resampling under rigorous testing conditions. Estimation of codon homozygosity, F, turns out to be a key to the estimation of N(c). F can be estimated in two closely related ways, corresponding to sampling with or without replacement, the latter being what Wright used. The N(c) methods that are based on sampling without replacement showed much better accuracy at short gene lengths than those based on sampling with replacement, indicating that Wright's homozygosity method is superior. Surprisingly, the methods based on sampling with replacement displayed a superior correlation with mRNA levels in Escherichia coli.

Codon↗

Structure of the Escherichia coli K12 regulatory gene tyrR. Nucleotide sequence and sites of initiation of transcription and translation.

The nucleotide sequence of 1964 base pairs of the Escherichia coli K12 chromosome containing the autogenously regulated regulatory gene tyrR has been determined. The site of initiation of transcription of tyrR has been mapped by primer-extension analysis, and the initiation codon has been identified by site-specific deletion mutagenesis. The nucleotide sequence predicts a subunit molecular weight of 53,099 for the TyrR protein. Codon usage in the tyrR structural gene shows a bias toward those synonymic codons which are used rarely in efficiently expressed E. coli genes. The nucleotide sequence of a 22-base pair region adjacent to the promoter and distal to the structural gene exhibits considerable identity with corresponding regions of other genes regulated by tyrR. It is proposed that this is a site for repression by the TyrR protein.

Amino Acid Sequence↗

Isolation and sequence analysis of a cDNA clone encoding the entire catalytic subunit of phosphorylase kinase.

Synthetic oligonucleotides have been used to isolate a 1.85 kb clone containing the full length coding sequence for the catalytic subunit of rabbit skeletal muscle phosphorylase kinase from a cDNA library constructed in lambda gt10. Sequence analysis of the clone predicted an amino acid sequence in agreement with a published primary structure. Inspection of the codon usage revealed a strong preference for G or C nucleotides at the third codon position as found for several other skeletal muscle proteins. This cDNA clone should facilitate identification of functional domains, including the calmodulin-binding site, and investigation of the molecular basis of X-linked phosphorylase kinase deficiencies.

Amino Acid Sequence↗