Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Open reading frames in the antisense strands of genes coding for glycolytic enzymes in Saccharomyces cerevisiae.

Open reading frames longer than 300 bases were observed in the antisense strands of the genes coding for the glycolytic enzymes phosphoglucose isomerase, phosphoglycerate mutase, pyruvate kinase and alcohol dehydrogenase I. The open reading frames on both strands are in codon register. It has been suggested that proteins coded in codon register by complementary DNA strands can bind to each other. Consequently, it was interesting to investigate whether the open reading frames in the antisense strands of glycolytic enzyme genes are functional. We used oligonucleotide-directed mutagenesis of the PGI1 phosphoglucose isomerase gene to introduce pairs of closely spaced base substitutions that resulted in stop codons in one strand and only silent replacements in the other. Introduction of the two stop codons into the PGI1 sense strand caused the same physiological defects as already observed for pgil deletion mutants. No detectable effects were caused by the two stop codons in the antisense strand. A deletion that removed a section from -31 bp to +109 bp of the PGI1 gene but left 83 bases of the 3' region beyond the antisense open reading frame had the same phenotype as a deletion removing both reading frames. A similar pair of deletions of the PYK1 gene and its antisense reading frame showed identical defects. Our own Northern experiments and those reported by other authors using double-stranded probes detected only one transcript for each gene. These observations indicate that the antisense reading frames are not functional. On the other hand, evidence is provided to show that the rather long reading frames in the antisense strands of these glycolytic enzyme genes could arise from the strongly selective codon usage in highly expressed yeast genes, which reduces the frequency of stop codons in the antisense strand.

Base Sequence↗

Codon distribution in vertebrate genes may be used to predict gene length.

I have analysed the coding regions of 96 eukaryotic genes for their use of iso-coding codons. Specific codons occur more frequently in specific positions in all members of some gene families than would be expected if codon choice was determined solely by the frequency of codon usage. In the absence of evidence a priori for selection for particular codons at particular positions, I term such co-occurring codons "coincident codons". Coincident codons are not confined to particular regions of genes, and their occurrence is not detectably linked with the location of introns in the genomic sequence. Their presence is partly but not completely explained by the exchange of sequence between similar functional genes within a species: homologous genes from different organisms also possess the same codons at some sites with greater than expected frequencies. The relative excess of coincident codons correlates well with the overall length of the genes analysed, but not with the length of mRNA or coding regions, or with qualitative features of gene structure or expression. This, and the unusual sequence environment of coincident codons, suggests that they are a feature of the overall secondary structure of the heterogeneous nuclear RNA. Such considerations suggest approaches for optimizing the expression of exogenous genes in eukaryotic systems, and for predicting the structure of genes for which only partial sequence data is available.

Actins↗

Nucleotide sequence of an actin-encoding gene from Hydra attenuata: structural characteristics and evolutionary implications.

We have determined the complete nucleotide sequence of an actin-encoding gene from Hydra attenuata as well as partial sequences of cDNA clones from two additional actin-encoding genes. The gene from the genomic clone contains a single intron, and has promoter and polyadenylation signals similar to those found in other species. The hydra genome has a very A + T-rich base composition (71%). This is reflected in the codon usage of the actin-encoding genes, which is strongly biased towards codons having A or T in the third position. The hydra actin-encoding gene family consists of three or more transcribed genes, two of which are very closely related to each other and probably arose by a recent gene duplication. Hydra actin, like other invertebrate actins, is more similar to the non-muscle isotypes of vertebrates than to the vertebrate muscle actins. Hydra actin is more similar to animal actins than to those of plants or fungi, which is consistent with the view that all metazoans arose from a single protist ancestor.

Actins↗

Effects of codon-optimization on protein expression by the human herpesvirus 6 and 7 U51 open reading frame.

Codon-optimization refers to the alteration of gene sequences, to make codon usage match the available tRNA pool within the cell/species of interest. Codon-optimization has emerged as a powerful tool to increase protein expression by genes from small RNA and DNA viruses, which commonly contain overlapping reading frames as well as structural elements that are embedded within coding regions; these features are not widespread among large DNA viruses. We therefore examined whether codon-optimization might influence protein expression from a herpesvirus gene. We focused on the U51 gene from human herpesviruses-6 and -7, which was cloned in both native and codon-optimized form, with an N-terminal HA epitope tag to allow protein detection. Codon-optimization was associated with a profound (10-100 fold) increase in U51 expression in human (293A, HSG, K562) or hamster (CHO) cell lines, suggesting this may represent a valuable tool to facilitate functional studies on recalcitrant herpesvirus genes. Finally, it is postulated that the suboptimal expression of native U51 may reflect a regulatory mechanism that controls viral gene expression.

Animals↗

Codon adaptation and synonymous substitution rate in diatom plastid genes.

Diatom plastid genes are examined with respect to codon adaptation and rates of silent substitution (Ks). It is shown that diatom genes follow the same pattern of codon usage as other plastid genes studied previously. Highly expressed diatom genes display codon adaptation, or a bias toward specific major codons, and these major codons are the same as those in red algae, green algae, and land plants. It is also found that there is a strong correlation between Ks and variation in codon adaptation across diatom genes, providing the first evidence for such a relationship in the algae. It is argued that this finding supports the notion that the correlation arises from selective constraints, not from variation in mutation rate among genes. Finally, the diatom genes are examined with respect to variation in Ks among different synonymous groups. Diatom genes with strong codon adaptation do not show the same variation in synonymous substitution rate among codon groups as the flowering plant psbA gene which, previous studies have shown, has strong codon adaptation but unusually high rates of silent change in certain synonymous groups. The lack of a similar finding in diatoms supports the suggestion that the feature is unique to the flowering plant psbA due to recent relaxations in selective pressure in that lineage.

Adaptation, Physiological↗

Organization and nucleotide sequences of ten ribosomal protein genes from the region equivalent to the S10 operon in the archaebacterium, Halobacterium halobium.

A determination was made of the nucleotide sequence of the 7340-bp region of a ribosomal protein gene cluster of Halobacterium halobium, which is equivalent to the S10 operon of Escherichia coli. The sequence was analyzed with the codonpreference program deduced from the halobacterial codon usage table that showed a very high GC content of the third codon position. The sequence was comprised of a string of 13 tightly linked ORFs. Most of the ORFs were homologous with ribosomal protein genes (ORF1-ORF2-rpl3-rpl4-rpl23--rpl2- rps19-rpl22-rps3-rpl29-ORF11-rps17-r pl14). The 13-gene string was preceded by three putative AT-rich promoter sequences. The order of the genes in H. halobium essentially agreed with that of the corresponding genes of E. coli (S10-operon), except for certain deletions or insertions of additional protein genes.

Amino Acid Sequence↗

Inferring weak selection from patterns of polymorphism and divergence at "silent" sites in Drosophila DNA.

Patterns of codon usage and "silent" DNA divergence suggest that natural selection discriminates among synonymous codons in Drosophila. "Preferred" codons are consistently found in higher frequencies within their synonymous families in Drosophila melanogaster genes. This suggests a simple model of silent DNA evolution where natural selection favors mutations from unpreferred to preferred codons (preferred changes). Changes in the opposite direction, from preferred to unpreferred synonymous codons (unpreferred changes), are selected against. Here, selection on synonymous DNA mutations is investigated by comparing the evolutionary dynamics of these two categories of silent DNA changes. Sequences from outgroups are used to determine the direction of synonymous DNA changes within and between D. melanogaster and Drosophila simulans for five genes. Population genetics theory shows that differences in the fitness effect of mutations can be inferred from the comparison of ratios of polymorphism to divergence. Unpreferred changes show a significantly higher ratio of polymorphism to divergence than preferred changes in the D. simulans lineage, confirming the action of selection at silent sites. An excess of unpreferred fixations in 28 genes suggests a relaxation of selection on synonymous mutations in D. melanogaster. Estimates of selection coefficients for synonymous mutations (3.6 < magnitude of Nes < 1.3) in D. simulans are consistent with the reduced efficacy of natural selection (magnitude of Nes < 1) in the three- to sixfold smaller effective population size of D. melanogaster. Synonymous DNA changes appear to be a prevalent class of weakly selected mutations in Drosophila.

Animals↗

Mitochondrial genomes of Galathealinum, Helobdella, and Platynereis: sequence and gene arrangement comparisons indicate that Pogonophora is not a phylum and Annelida and Arthropoda are not sister taxa.

We report a contiguous region of more than half (> 7,500 nt) of the mitochondrial genomes for Platynereis dumerii (Annelida: Polychaeta), Helobdella robusta (Annelida: Hirudinida), and Galathealinum brachiosum (Pogonophora: Perviata). The relative arrangements of all 22 genes identified for Helobdella and Galathealinum are identical to one another and to their arrangements in the mtDNA of the previously studied oligochaete annelid Lumbricus. In contrast, Platynereis differs from these taxa in the positions of several tRNA genes and in having two additional tRNA genes (trnC and trnM) and a large noncoding sequence in this region. Comparisons of relative gene arrangements and of the nucleotide and inferred amino acid sequences among these and other published taxa provide strong support for an annelid-mollusk clade that excludes arthropods, and for the inclusion of pogonophorans within Annelida, rather than giving them separate phylum status. Gene arrangement comparisons include the first use of a recently described method on previously unpublished data. Although a variety of alternative initiation codons are typically used by mitochondrial protein-encoding genes, ATG appears to be the initiator for all but one reported here. The large noncoding region (1,091 nt) identified in Platynereis has no significant sequence similarity to the noncoding region of Lumbricus, although each contains runs of TA dinucleotides and of homopolymers, which could potentially serve as signaling elements. There is strong bias for synonymous codon usage in Helobdella and especially in Galathealinum. In this latter taxon, 5 codons are completely unused, 13 are used three or fewer times, and G appears at third codon positions in only 26 of the 2,236 codons. Nucleotide composition bias appears to influence amino acid composition of the proteins.

Amino Acid Sequence↗

Identification of a novel operon in Lactococcus lactis encoding three enzymes for lactic acid synthesis: phosphofructokinase, pyruvate kinase, and lactate dehydrogenase.

The discovery of a novel multicistronic operon that encodes phosphofructokinase, pyruvate kinase, and lactate dehydrogenase in the lactic acid bacterium Lactococcus lactis is reported. The three genes in the operon, designated pfk, pyk, and ldh, contain 340, 502, and 325 codons, respectively. The intergenic distances are 87 bp between pfk and pyk and 117 bp between pyk and ldh. Plasmids containing pfk and pyk conferred phosphofructokinase and pyruvate kinase activity, respectively, on their host. The identity of ldh was established previously by the same approach (R. M. Llanos, A. J. Hillier, and B. E. Davidson, J. Bacteriol. 174:6956-6964, 1992). Each of the genes is preceded by a potential ribosome binding site. The operon is expressed in a 4.1-kb transcript. The 5' end of the transcript was determined to be a G nucleotide positioned 81 bp upstream from the pfk start codon. The pattern of codon usage within the operon is highly biased, with 11 unused amino acid codons. This degree of bias suggests that the operon is highly expressed. The three proteins encoded on the operon are key enzymes in the Embden-Meyerhoff pathway, the central pathway of energy production and lactic acid synthesis in L. lactis. For this reason, we have called the operon the las (lactic acid synthesis) operon.

Amino Acid Sequence↗

Structure of the Escherichia coli K12 regulatory gene tyrR. Nucleotide sequence and sites of initiation of transcription and translation.

The nucleotide sequence of 1964 base pairs of the Escherichia coli K12 chromosome containing the autogenously regulated regulatory gene tyrR has been determined. The site of initiation of transcription of tyrR has been mapped by primer-extension analysis, and the initiation codon has been identified by site-specific deletion mutagenesis. The nucleotide sequence predicts a subunit molecular weight of 53,099 for the TyrR protein. Codon usage in the tyrR structural gene shows a bias toward those synonymic codons which are used rarely in efficiently expressed E. coli genes. The nucleotide sequence of a 22-base pair region adjacent to the promoter and distal to the structural gene exhibits considerable identity with corresponding regions of other genes regulated by tyrR. It is proposed that this is a site for repression by the TyrR protein.

Amino Acid Sequence↗

Isolation and sequence analysis of a cDNA clone encoding the entire catalytic subunit of phosphorylase kinase.

Synthetic oligonucleotides have been used to isolate a 1.85 kb clone containing the full length coding sequence for the catalytic subunit of rabbit skeletal muscle phosphorylase kinase from a cDNA library constructed in lambda gt10. Sequence analysis of the clone predicted an amino acid sequence in agreement with a published primary structure. Inspection of the codon usage revealed a strong preference for G or C nucleotides at the third codon position as found for several other skeletal muscle proteins. This cDNA clone should facilitate identification of functional domains, including the calmodulin-binding site, and investigation of the molecular basis of X-linked phosphorylase kinase deficiencies.

Amino Acid Sequence↗

Hypermutation generating the sheep immunoglobulin repertoire is an antigen-independent process.

Somatic hypermutation of light chain V genes during development of B cells in sheep ileal Peyer's patches was studied in three experimental conditions: in sterile fragments of the ileum surgically isolated from the gut during fetal life, in germ-free sheep, and in animals thymectomized during early fetal life. The somatic mutation pattern was found identical to control tissues in all three experiments. The same age-dependent amount of mutations, a higher than theoretical R/S ratio in complementarity-determining regions (CDRs), and a similar clustering of mutations in CDRs were observed. The mechanism, as estimated from the silent mutation pattern, appears to target mutations to CDRs; moreover, the major V lambda genes have a specific codon usage with a high purine content at the first two bases of the codons and a low content at the third position, which, together with a specific targeting of mutations to purines, favors replacement mutations in CDRs.

Animals↗

Synonymous substitution rates in enterobacteria.

It has been shown previously that the synonymous substitution rate between Escherichia coli and Salmonella typhimurium is lower in highly than in weakly expressed genes, and it has been suggested that this is due to stronger selection for translational efficiency in highly expressed genes as reflected in their greater codon usage bias. This hypothesis is tested here by comparing the substitution rate in codon families with different patterns of synonymous codon use. It is shown that the decline in the substitution rate across expression levels is as great for codon families that do not appear to be subject to selection for translational efficiency as for those that are. This implies that selection on translational efficiency is not responsible for the decline in the substitution rate across genes. It is argued that the most likely explanation for this decline is a decrease in the mutation rate. It is also shown that a simple evolutionary model in which synonymous codon use is determined by a balance between mutation, selection for an optimal codon, and genetic drift predicts that selection should have little effect on the substitution rate in the present case.

Codon↗

Human metallothionein genes: molecular cloning and sequence analysis of the mRNA.

From a cDNA clone bank prepared from cadmium-treated HeLa cells, we isolated clones representing mRNAs whose concentration is increased after cadmium induction. Several metallothionein cDNA clones were isolated by cross-hybridization to mouse metallothionein-I cDNA. The nucleotide sequence of one of these clones, containing a nearly full-length cDNA copy of human metallothionein-II mRNA, was determined. The homology between the human and mouse metallothionein sequences is strictly limited to the coding region of the mRNA. Codon usage in metallothionein mRNA is not random. Seventy-nine percent of the codons have G or C residues at the third position, resulting in a GC-rich sequence.

Amino Acid Sequence↗

Complete sequence of the amphioxus (Branchiostoma lanceolatum) mitochondrial genome: relations to vertebrates.

The complete nucleotide sequence of the mitochondrial DNA of the amphioxus Branchiostoma lanceolatum has been determined. This mitochondrial genome is small (15 076 bp) because of the short size of the two rRNA genes and the tRNA genes. In addition, this genome contains a very short non-coding region (57 bp) with no sequence reminiscent of a control region. The organisation of the coding genes, as well as of the two rRNA genes, is identical to that of the sea lamprey. Some differences in the repartition of the tRNA genes occur when compared to the lamprey. The mitochondrial codon usage of the amphioxus is reminiscent of that of urochordates since the AGA codon is read as a glycine and not as a stop codon as in vertebrates. Moreover, the base composition at the wobble positions of the codon is strongly biased toward guanine. Altogether, these data clearly emphasise the close relationships between amphioxus and vertebrates, and reinforce the notion that prochordates may be viewed as the brother group of vertebrates.

Animals↗

Adenylylsulphate reductase from the sulphate-reducing archaeon Archaeoglobus fulgidus: cloning and characterization of the genes and comparison of the enzyme with other iron-sulphur flavoproteins.

Adenylylsulphate (adenosine-5'-phosphosulphate, APS) reductase from the extremely thermophilic sulphate-reducing archaeon Archaeoglobus fulgidus is an iron-sulphur flavoprotein containing one non-covalently bound flavin group, eight non-haem iron and six labile sulphide atoms per molecule. Reevaluation of the enzyme structure revealed the presence of two different subunits with molecular masses of 80 and 18.5 kDa. The subunits are arranged in an alpha 2 beta subunit structure. We have cloned and sequenced a 2.7 kb segment of DNA containing the genes for the alpha and beta subunits, which we designate aprA and aprB, respectively. The two genes are separated by 17 bp and localized in the order aprBA. While a putative promoter could not be identified in the vicinity of aprBA a probable termination signal was found just downstream of the translation stop codon of aprA. The codon usage for aprBA shows strong preferences for G and C in the third codon position. aprA encodes a 73.3 kDa polypeptide, which shows significant overall similarities with the flavoprotein subunits of the succinate dehydrogenases from Escherichia coli and Bacillus subtilis and the corresponding flavoprotein of E. coli fumarate reductase. Part of the homologous peptide stretches could be assigned to domains that are involved in the binding of the substrate or of the FAD prosthetic group. aprB encodes a 17.1 kDa polypeptide representing an iron-sulphur protein, seven cysteine residues of which are arranged in two clusters typical of ligands of the iron-sulphur centres in ([Fe3S4][Fe4S4]) 7-Fe ferredoxins.

Amino Acid Sequence↗

Regulation of protein synthesis in Tetrahymena. RNA sequence sets of growing and starved cells.

The complexity of messenger RNA in growing or starved Tetrahymena thermophila is similar and unusually high (approximately 4.5 X 10(7) nucleotides). The complexity of nuclear RNA in growing cells (approximately 7.8 X 10(7) nucleotides) is only about 1.7 times that of mRNA. The concentration of complex class (rare) messages (approximately 53 copies/growing cell and approximately 11 copies/starved cell) is low in comparison to the size of the cell. The concentration of complex nuclear transcripts is also very low (approximately 0.7 copies/growing cell nucleus and approximately 2.6 copies/starved cell nucleus) considering that the macronucleus contains 45 to 90 copies of each single copy sequence. The complex sequence sets found on polysomes of growing and starved cells overlap about 80% and about 60% of the complex nuclear transcripts appear to be held in common. About 60% of macronuclear single copy DNA is transcribed in one or both physiological states. Although growing and starved cells have extremely different fractions of their messages loaded onto polysomes, within each cell type the complex messages in polysomal and nonpolysomal cytoplasmic fractions are indistinguishable, suggesting that exchange may occur between loaded and unloaded messages. Although T. thermophila DNA has an unusually low G + C content (23%), sequences coding for complex RNAs have base ratios similar to those of total DNA. Therefore, codon usage in Tetrahymena must be extremely biased towards adenine- and uridine-rich codons.

Animals↗

Ribosome-mediated translational pause and protein domain organization.

Because regions on the messenger ribonucleic acid differ in the rate at which they are translated by the ribosome and because proteins can fold cotranslationally on the ribosome, a question arises as to whether the kinetics of translation influence the folding events in the growing nascent polypeptide chain. Translationally slow regions were identified on mRNAs for a set of 37 multidomain proteins from Escherichia coli with known three-dimensional structures. The frequencies of individual codons in mRNAs of highly expressed genes from E. coli were taken as a measure of codon translation speed. Analysis of codon usage in slow regions showed a consistency with the experimentally determined translation rates of codons; abundant codons that are translated with faster speeds compared with their synonymous codons were found to be avoided; rare codons that are translated at an unexpectedly higher rate were also found to be avoided in slow regions. The statistical significance of the occurrence of such slow regions on mRNA spans corresponding to the oligopeptide domain termini and linking regions on the encoded proteins was assessed. The amino acid type and the solvent accessibility of the residues coded by such slow regions were also examined. The results indicated that protein domain boundaries that mark higher-order structural organization are largely coded by translationally slow regions on the RNA and are composed of such amino acids that are stickier to the ribosome channel through which the synthesized polypeptide chain emerges into the cytoplasm. The translationally slow nucleotide regions on mRNA possess the potential to form hairpin secondary structures and such structures could further slow the movement of ribosome. The results point to an intriguing correlation between protein synthesis machinery and in vivo protein folding. Examination of available mutagenic data indicated that the effects of some of the reported mutations were consistent with our hypothesis.

Bacterial Proteins↗