Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Compositional gradients in Gramineae genes.

In this study, we describe a property of Gramineae genes, and perhaps all monocot genes, that is not observed in eudicot genes. Along the direction of transcription, beginning at the junction of the 5'-UTR and the coding region, there are gradients in GC content, codon usage, and amino-acid usage. The magnitudes of these gradients are large enough to hinder the annotation of the rice genome and to confound the detection of protein homologies across the monocot-eudicot divide.

Amino Acids↗

The priB gene encoding the primosomal replication n protein of Escherichia coli.

The gene encoding protein n of the Escherichia coli primosome has been discovered in the rpsF-rpsR-rplI ribosomal protein operon and designated priB. The low copy number of PriB protein and the distinctive codon usage of its gene argue against its being a ribosomal protein. A strain which overproduces PriB was constructed and has been used to purify the protein to homogeneity. The overproduced protein behaves like that purified from wild-type cells.

Amino Acid Sequence↗

Synthetic neomycin-kanamycin phosphotransferase, type II coding sequence for gene targeting in mammalian cells.

The bacterial neomycin-kanamycin phosphotransferase, type II enzyme is encoded by the neo gene and confers resistance to aminoglycoside drugs such as neomycin and kanamycin-bacterial selection and G418-eukaryotic cell selection. Although widely used in gene targeting in mouse embryonic stem cells, the neo coding sequence contains numerous cryptic splice sites and has a high CpG content. At least the former can cause unwanted effects in cis at the targeted locus. We describe a synthetic sequence, sneo, which encodes the same protein as that encoded by neo. This synthetic sequence has no predicted splice sites in either strand, low CpG content, and increased mammalian codon usage. In mouse embryonic stem cells sneo expressability is similar to neo. The use of sneo in gene targeting experiments should substantially reduce the probability of unwanted effects in cis due to splicing, and perhaps CpG methylation, within the coding sequence of the selectable marker.

Animals↗

Third codon G + C periodicity as a possible signal for an "internal" selective constraint.

Quasi-local analysis methods, such as window Fast Fourier Transform and an information theoretical quantity known as mutual information, have allowed us to gain some further insights on the importance and the contextual dependence of a pattern found in DNA sequences showing a periodicity of three with a G or C base in the third position. We have screened for such a periodicity, in terms of the alternative "strong" (S = C or G) versus "weak" (W = A or T) base, a large sample of DNA coding and non-coding sequences from both prokaryotes and eukaryotes, with the aim of testing whether this pattern could be considered as a significant signal for past or present constraints regarding DNA organization and/or function. This periodicity was indeed found in a number of sequences always associated with open reading frames, generally confined in prokaryotes living in extreme environments or in highly conserved eukaryotic genes. Moreover, codon usage was found to be very similar even in genes coding for very different functions. The data are discussed in view of their possible implications for an adaptive value of such a periodicity, in terms of more accurate translation processing and better overall stability.

Animals↗

Phylogenetic and recombination analysis of coronavirus HKU1, a novel coronavirus from patients with pneumonia.

Phylogenetic trees constructed using predicted amino acid sequences of putative proteins of coronavirus HKU1 (CoV-HKU1) revealed that CoV-HKU1 formed a distinct branch among group 2 coronaviruses. Of the 14 trees from p65 to nsp10, nine showed that CoV-HKU1 was clustered with murine hepatitis virus. From nsp11, the topologies of the trees changed dramatically. For the eight trees from nsp11 to N, seven showed that the CoV-HKU1 branch was the first branch. The codon usage patterns of CoV-HKU1 differed significantly from those in other group 2 coronaviruses. Split decomposition analysis revealed that recombination events had occurred between CoV-HKU1 and other coronaviruses.

Codon↗

Cloning and sequencing of genes encoding the TthHB8I restriction and modification enzymes: comparison with the isoschizomeric TaqI enzymes.

Genes encoding the TthHB8I restriction and modification (R-M) system from Thermus thermophilus HB8 (recognition sequence T decreases CGA) were cloned in Escherichia coli. The genes have the same transcriptional orientation, with the last 13 codons of the methyltransferase (MTase) overlapping the first 13 codons of the endonuclease (ENase). Nucleotide sequence analysis of the TthHB8I ENase revealed a single chain of 263 amino acid (aa) residues that share a 77% identity with the corrected isoschizomeric TaqI ENase. Likewise, the Tth MTase (428 aa) shares a 79% identity with the corrected sequence of the TaqI MTase. This high degree of aa conservation suggests a common origin between the Taq and Tth R-M systems. However, codon usage and G+C content for the R-M genes differed markedly from that of other cloned Thermus genes. This suggests that these R-M genes were only recently introduced into the genus Thermus.

Amino Acid Sequence↗

Macronuclear and micronuclear configurations of a gene encoding the protein synthesis elongation factor EF 1 alpha in Stylonychia lemnae.

The micronuclear and macronuclear configurations of a gene encoding the protein synthesis elongation factor EF 1 alpha in the hypotrich ciliate Stylonychia lemnae were compared. The two sequences are generally colinear. The coding sequence of the micronuclear gene is, however, interrupted by a 64 bp insert flanked by a 2 bp direct repeat in a gene region which is moderately conserved among EF 1 alpha genes of different organisms. The insertion site is distinct from known intron positions in eukaryotic EF 1 alpha genes. The insert sequence shows inverted repeats at its ends and thus exhibits typical features of an internal eliminated sequence (IES). Comparison with other such sequences in the related organism Oyxtricha nova shows that the IES falls into a new group of such elements. The macronuclear gene exhibits a strikingly limited codon usage, which cannot be simply explained by the overall base composition of the DNA but probably also relates to the very high copy number of the macronuclear gene and the putative high amount of the gene product.

Amino Acid Sequence↗

Sequences of four mouse histone H3 genes: implications for evolution of mouse histone genes.

The sequences of four histone H3 genes coding for the replication variant proteins H3.1 and H3.2 have been determined. Three of these genes, two coding for H3.1 proteins and one for an H3.2 protein, are located on chromosome 13 and expressed at low levels. The fourth gene, encoding an H3.2 protein, is located on chromosome 3 and expressed at a high level. The coding regions of the three genes on chromosome 13 are more similar to each other than to the H3 gene on chromosome 3, and equally divergent from it, suggesting that either gene duplication or gene conversion has occurred since the genes were dispersed onto two chromosomes. A 14-base sequence including the CCAAT sequence and located 5' to the genes on chromosome 13 has been conserved. The histone H3 gene on chromosome 3 has multiple potential binding sites for the Sp1 transcription factor. The coding regions show greater than 95% conservation among the four genes. This is due to the strict pattern of codon usage and the presence of two long (greater than 60 base) regions of completely conserved nucleic acid sequence. These conserved regions in the coding sequence may have an important functional role at the mRNA or DNA level.

Amino Acid Sequence↗

Comparative molecular evolution of primary (Buchnera) and secondary symbionts of aphids based on two protein-coding genes.

A+T content, phylogenetic relationships, codon usage, evolutionary rates, and ratio of synonymous versus non-synonymous substitutions have been studied in partial sequences of the atpD and aroQ/pheA genes of primary ( Buchnera) and secondary symbionts of aphids and a set of selected non-symbiotic bacteria, belonging to the five subdivisions of the Proteobacteria. Compared to the homologous genes of the last group, both genes belonging to Buchnera behave in a similar way, showing a higher A+T content, forming a monophyletic group, a loss in codon bias, especially in third base position, an evolutionary acceleration and an increase in the number of non-synonymous substitutions, confirming previous results reported elsewhere for other genes. When available, these properties have been partly observed with the secondary symbionts, but with values that are intermediate between Buchnera and free living Proteobacteria. They show high A+T content, but not as high as Buchnera, a non-solved phylogenetic position between Buchnera, and the other gamma-Proteobacteria, a loss in codon bias, again not as high as in Buchnera and a significant evolutionary acceleration in the case of the three atpD genes, but not when considering aroQ/pheA genes. These results give support to the hypothesis that they are symbionts at different stages of the symbiotic accommodation to the host.

AT Rich Sequence↗

Analysis of sequences from the extremely A + T-rich genome of Plasmodium falciparum.

The genome of the human malaria parasite Plasmodium falciparum has an A + T content of about 82%, higher than any other organism whose DNA has been characterized. Computer analysis of 36 kb of available nucleotide sequences from this species showed that the coding regions, with an A + T content of 69.0%, are flanked by more A + T-rich regions of 86.0% A + T. Within the coding sequences, the A/T ratio was 1.68 in the mRNA sense strand, and overall A + T content in the three codon positions increased in the order 1st-2nd-3rd position. Codons with T or especially A in the third position were strongly preferred. Codon usage among individual parasite genes was very similar compared to genes from other species. Dinucleotide frequencies for the parasite DNA were close to those expected for a random sequence with the known base composition, except that the CpG frequency in the coding sequences was low.

Adenine↗

Heterologous expression of proteins from Plasmodium falciparum: results from 1000 genes.

As part of a structural genomics initiative, 1000 open reading frames from Plasmodium falciparum, the causative agent of the most deadly form of malaria, were tested in an E. coli protein expression system. Three hundred and thirty-seven of these targets were observed to express, although typically the protein was insoluble. Sixty-three of the targets provided soluble protein in yields ranging from 0.9 to 406.6 mg from one liter of rich media. Higher molecular weight, greater protein disorder (segmental analysis, SEG), more basic isoelectric point (pI), and a lack of homology to E. coli proteins were all highly and independently correlated with difficulties in expression. Surprisingly, codon usage and the percentage of adenosines and thymidines (%AT) did not appear to play a significant role. Of those proteins which expressed, high pI and a hypothetical annotation were both strongly and independently correlated with insolubility. The overwhelmingly important role of pI in both expression and solubility appears to be a surprising and fundamental issue in the heterologous expression of P. falciparum proteins in E. coli. Twelve targets which did not express in E. coli from the native gene sequence were codon-optimized through whole gene synthesis, resulting in the (insoluble) expression of three of these proteins. Seventeen targets which were expressed insolubly in E. coli were moved into a baculovirus/Sf-21 system, resulting in the soluble expression of one protein at a high level and six others at a low level. A variety of factors conspire to make the heterologous expression of P. falciparum proteins challenging, and these observations lay the groundwork for a rational approach to prioritizing and, ultimately, eliminating these impediments.

Animals↗

Sequence analysis of an aphid endosymbiont DNA fragment containing rpoB (beta-subunit of RNA polymerase) and portions of rplL and rpoC.

The aphid Schizaphis graminum is dependent on an association with a prokaryotic endosymbiont (Buchnera aphidicola). The nucleotide (nt) sequence of a 5040 base pair (bp) DNA fragment of B. aphidicola, homologous to the rplL-rpoB-rpoC portion of the Escherichia coli beta operon, was determined. The DNA coded for the terminal 35 amino acids of RplL (large ribosomal subunit protein L7/L12), the complete RpoB (beta-subunit of RNA polymerase), and the first 209 amino acids of RpoC (beta'-subunit of RNA polymerase). The deduced sequences of B. aphidicola RplL, RpoB, and RpoC were 71, 84, and 91% identical, respectively, to the homologous proteins of E. coli. The sequences of two portions of the intergenic region between rplL and rpoB were nearly identical in both B. aphidicola and E. coli. One sequence constituted an inverted repeat that could be an RNase III-messenger RNA processing site; the other sequence preceded RpoB. A compilation of the codon usage for RpoB, RpoC, and other B. aphidicola proteins indicated a major preference for A or T in the first and third positions, a result consistent with the low guanine plus cytosine (G + C) content of the DNA of this organism.

Amino Acid Sequence↗

Actin in the oomycetous fungus Phytophthora infestans is the product of several genes.

Actin (ACT) in Phytophthora infestans is encoded by at least two genes, in contrast to unicellular and other filamentous fungi where there is a single gene. These genes (designated actA and actB) have been isolated from a genomic library of P. infestans. The complete nucleotide sequence of both genes has been determined. Unlike the actin-encoding genes (act) of other filamentous fungi, no introns are obvious in the coding region, a feature shared with the act genes of certain protists. Northern blotting and primer extension studies of the mRNA show that actA and actB are actively transcribed in mycelium, sporangia and germinating cysts but only at a low level in the case of actB. Both genes display bias in their codon usage. This is more extreme in actA. The deduced ACTB protein is strikingly similar to that of the Phytophthora megasperma actin and is more diverged from other actins than ACTA.

Actins↗

Genes and regulatory sites of the "host-takeover module" in the terminal redundancy of Bacillus subtilis bacteriophage SPO1.

Early in infection of Bacillus subtilis by bacteriophage SPO1, the synthesis of most host-specific macromolecules is replaced by the corresponding phage-specific biosyntheses. It is believed that this subversion of the host biosynthetic machinery is accomplished primarily by a cluster of early genes in the SPO1 terminal redundancy. Here we analyze the nucleotide sequence of this 11.5-kb "host-takeover module," which appears to be designed for particularly efficient expression. Promoters, ribosome-binding sites, and codon usage statistics all show characteristics known to be associated with efficient function in B. subtilis. The promoters and ribosome-binding sites have additional conserved features which are not characteristic of their host counterparts and which may be important for competition with host genes for the cellular biosynthetic machinery. The module includes 24 genes, tightly packed into 12 operons driven by the previously identified early promoters PE1 to PE12. The genes are smaller than average, with half of them having fewer than 100 codons. Most of their inferred products show little similarity to known proteins, although zinc finger, trans-membrane, and RNA polymerase-binding domains were identified. Transcription-termination and RNase III cleavage sites were found at appropriate locations.

Amino Acid Sequence↗

Extreme differences in charge changes during protein evolution.

The maintenance of a proper distribution of charged amino acid residues might be expected to be an important factor in protein evolution. We therefore compared the inferred changes in charge during the evolution of 43 protein families with the changes expected on the basis of random base substitutions. It was found that certain proteins, like the eye lens crystallins and most histones, display an extreme avoidance of changes in charge. Other proteins, like phospholipase A2 and ferredoxin, apparently have sustained more charged replacements than expected, suggesting a positive selection for changes in charge. Depending on function and structure of a protein, charged residues apparently can be important targets for selective forces in protein evolution. It appears that actual biased codon usage tends to decrease the proportion of charged amino acid replacements. The influence of nonrandomness of mutations is more equivocal. Genes that use the mitochondrial instead of the universal code lower the probability that charge changes will occur in the encoded proteins.

Biological Evolution↗

Cloning and sequence analysis of the phenylalanyl-tRNA synthetase genes (pheST) from Thermus thermophilus.

While crystals suitable for X-ray diffraction analyses are available of phenylalanyl-tRNA synthetase (PheRS) from the thermophilic bacterium Thermus thermophilus, neither the primary structure of its constituent alpha and beta subunits nor the nucleotide sequence of the corresponding pheS and pheT genes were known. Using specific oligonucleotides of conserved pheS regions that were adapted to the T. thermophilus codon usage, we identified, cloned and subsequently sequenced the pheST genes of this bacterium. The sequences reported here will greatly aid in the three-dimensional structure determination of T. thermophilus PheRS, a heterotetrameric (alpha 2 beta 2), class II aminoacyl-tRNA synthetase.

Amino Acid Sequence↗

The genomic nucleotide sequences of two differentially expressed actin-coding genes from the sea star Pisaster ochraceus.

The genomic sequences of two differentially expressed actin genes from the sea star Pisaster ochraceus are reported. The cytoplasmic actin gene (Cy) is expressed in eggs and early development. The muscle actin gene (M) is expressed in tube feet and testes. Both genes contain an 1125-nucleotide coding region interrupted by three introns at codons 41, 121 and 204. Gene M contains two additional introns at codons 150 and 267. The intron position at codon 150, although present in higher vertebrate actins, has not been reported in actin genes from invertebrates. The M gene coding region has 89.5% nucleotide homology to the Cy gene, and differs from the Cy actin gene in 13 of 375 amino acids (aa), 11 of which are found in the C-terminal half of the gene. The C-terminal half of the M gene contains a significant number of muscle isotype codons. Even though there is only 1 aa change in the first 150 codons, there have been limited substitutions at many four-fold degenerate sites which may indicate selection pressure upon the secondary structure of the mRNA and/or a biased codon usage. Variant CCAAT, TATA, and poly(A)-addition signals have been identified in the 5' and 3' flanking regions. The presence of 5' and 3' splice junction sequences in the 5' flanking region of the Cy gene suggests the potential for an intron there.

Actins↗

An analysis of determinants of amino acids substitution rates in bacterial proteins.

The variation of amino acid substitution rates in proteins depends on several variables. Among these, the protein's expression level, functional category, essentiality, or metabolic costs of its amino acid residues may play an important role. However, the relative importance of each variable has not yet been evaluated in comparative analyses. To this aim, we made regression analyses combining data available on these variables and on evolutionary rates, in two well-documented model bacteria, Escherichia coli and Bacillus subtilis. In both bacteria, the level of expression of the protein in the cell was by far the most important driving force constraining the amino acids substitution rate. Subsequent inclusion in the analysis of the other variables added little further information. Furthermore, when the rates of synonymous substitutions were included in the analysis of the E. coli data, only the variable expression levels remained statistically significant. The rate of nonsynonymous substitution was shown to correlate with expression levels independently of the rate of synonymous substitution. These results suggest an important direct influence of expression levels, or at least codon usage bias for translation optimization, on the rates of nonsynonymous substitutions in bacteria. They also indicate that when a control for this variable is included, essentiality plays no significant role in the rate of protein evolution in bacteria, as is the case in eukaryotes.

Amino Acids↗