Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Complete nucleotide sequence of the Campylobacter jejuni 72Dz asd gene.

Campylobacter jejuni asd gene was sequenced. The GC content of the gene coding region is 32.7%. The codon usage is typical for a gene in a genome with low GC content. The structure of the gene regulatory sequences resembles that one used for Escherichia coli gene transcription and translation. The amino acid sequence of the Asd protein exhibits significant homology to asd gene products from other microorganisms.

Amino Acid Sequence↗

Unraveling selection in the mitochondrial genome of Drosophila.

We examine mitochondrial DNA variation at the cytochrome b locus within and between three species of Drosophila to determine whether patterns of variation conform to the predictions of neutral molecular evolution. The entire 1137-bp cytochrome b locus was sequenced in 16 lines of Drosophila melanogaster, 18 lines of Drosophila simulans and 13 lines of Drosophila yakuba. Patterns of variation depart from neutrality by several test criteria. Analysis of the evolutionary clock hypothesis shows unequal rates of change along D. simulans lineages. A comparison within and between species of the ratio of amino acid replacement change to synonymous change reveals a relative excess of amino acid replacement polymorphism compared to the neutral prediction, suggestive of slightly deleterious or diversifying selection. There is evidence for excess homozygosity in our world wide sample of D. melanogaster and D. simulans alleles, as well as a reduction in the number of segregating sites in D. simulans, indicative of selective sweeps. Furthermore, a test of neutrality for codon usage shows the direction of mutations at third positions differs among different topological regions of the gene tree. The analyses indicate that molecular variation and evolution of mtDNA are governed by many of the same selective forces that have been shown to govern nuclear genome evolution and suggest caution be taken in the use of mtDNA as a "neutral" molecular marker.

Animals↗

Forbidden synonymous substitutions in coding regions.

In the evolution of highly conserved genes, a few "synonymous" substitutions at third bases that would not alter the protein sequence are forbidden or very rare, presumably as a result of functional requirements of the gene or the messenger RNA. Another 10% or 20% of codons are significantly less variable by synonymous substitution than are the majority of codons. The changes that occur at the majority of third bases are subject to codon usage restrictions. These usage restrictions control sequence similarities between very distant genes. For example, 70% of third bases are identical in calmodulin genes of man and trypanosome. Third-base similarities of distant genes for conserved proteins are mathematically predicted, on the basis of the G+C composition of third bases. These observations indicate the need for reexamination of methods used to calculate synonymous substitutions.

Actins↗

Recombination and base composition: the case of the highly self-fertilizing plant Arabidopsis thaliana.

BACKGROUND: Rates of recombination can vary among genomic regions in eukaryotes, and this is believed to have major effects on their genome organization in terms of base composition, DNA repeat density, intron size, evolutionary rates and gene order. In highly self-fertilizing species such as Arabidopsis thaliana, however, heterozygosity is expected to be strongly reduced and recombination will be much less effective, so that its influence on genome organization should be greatly reduced. RESULTS: Here we investigated theoretically the joint effects of recombination and self-fertilization on base composition, and tested the predictions with genomic data from the complete A. thaliana genome. We show that, in this species, both codon-usage bias and GC content do not correlate with the local rates of crossing over, in agreement with our theoretical results. CONCLUSIONS: We conclude that levels of inbreeding modulate the effect of recombination on base composition, and possibly other genomic features (for example, transposable element dynamics). We argue that inbreeding should be considered when interpreting patterns of molecular evolution.

Arabidopsis↗

Equal G and C contents in histone genes indicate selection pressures on mRNA secondary structure.

Protein-specific versus taxon-specific patterns of nucleotide frequencies were studied in histone genes. The third positions of codons have a (well-known) taxon-specific G+C level and a histone type-specific G/C ratio. This ratio counterbalances the G/C ratio in the first and second positions so that the overall G and C levels in the coding region become approximately equal. The compensation of the G/C ratio indicates a selection pressure at the mRNA level rather than a selection pressure or mutation bias at the DNA level or a selection pressure on codon usage. The structure of histone mRNAs is compatible with the hypothesis that the G/C compensation is due to selection pressures on mRNA secondary structure. Nevertheless, no specific motifs seem to have been selected, and the free energy of the secondary structures is only slightly lower than that expected on the basis of nucleotide frequencies.

Animals↗

Cloning and cDNA sequence of the rat X-chromosome linked phosphoglycerate kinase.

This paper reports the isolation and the sequence determination of rat phosphoglycerate kinase (PGK) cDNA clones. This cDNA, derived from an X-linked PGK gene transcript, contains a reading frame of 1254 nt and 5' and 3' non coding regions of 40 and 380 nt respectively. Analysis of the nucleotide sequence at the three codon position shows a biased codon usage with a prevalence of the triplet G non G N. Comparison of the inferred rat amino acid sequence with that of other organisms makes possible the calculation of the unit evolutionary period (UEP) for this enzyme, placing it at around 40 million years (My). Thus PGK is one of the oldest housekeeping enzymes.

Amino Acid Sequence↗

Statistical analysis and prediction of the exonic structure of human genes.

Nonhomologous fully sequenced human protein-coding genes were studied. Three sets of exon-exon junctions were formed defined by the intron (shadow) position relative to the reading frame. For the analysis of intron shadow signals in exons, information content and discrimination energy approaches were used with the correction allowing one to ignore the influence of a protein-coding message. The corrected formulas allow one to define the consensuses for the three types of intron shadow signals as a AG/guwn, cAG/GUnn, and cAG/gunU, and provide better recognition than the original formulas. The analysis of the codon usage in the signal positions leads to the conclusion that the prevalence of some amino acids in corresponding protein sites is caused by the signal requirements and not vice versa. The distribution of potential intron shadow signals in exons contradicts the hypothesis of intron insertion into suitable preexisting sites. There exists a correlation between the intron types and/or the exon length modulo 3.

Amino Acid Sequence↗

New molecular markers for phlebotomine sand flies.

Using degenerate-primers PCR we isolated and sequenced fragments from the sand fly Lutzomyia longipalpis homologous to two behavioural genes in Drosophila, cacophony and period. In addition we identified a number of other gene fragments that show homology to genes previously cloned in Drosophila. A codon usage table for L. longipalpis based on these and other genes was calculated. These new molecular markers will be useful in population genetics and evolutionary studies in phlebotomine sand flies and in establishing a preliminary genetic map in these important leishmaniasis vectors.

Amino Acid Sequence↗

Nucleotide sequence of the gene encoding the nitrogenase iron protein (nifH) of Azospirillum brasilense and identification of a region controlling nifH transcription.

The DNA sequence was determined for the Azospirillum brasilense nifH gene and part of the nifD gene. The nifH gene is 885 bp long and encodes 293 amino acid residues. The region upstream of the nifH open reading frame contains a putative promoter whose sequence shows perfect homology with promoters of other diazotrophic bacteria and two putative upstream activator sequences. Experiments with the promoter-probe vector pAF300 showed that this region promotes transcription in response to the nitrogen and oxygen availability of the cell. The amino acid sequence was deduced from the DNA nucleotide sequence of nifH; the polypeptide contains the four cysteine residues highly conserved among other nifH products and an arginine residue at position 101 which could be the site of the modification occurring during the "switch-off" of nitrogenase. The codon usage appears to be very biased reflecting the high G + C content of the Azospirillum nifH gene. In a comparison of the amino acid sequence with the other 18 known nifH gene products, the A. brasilense nifH product showed the highest level of homology with fast-growing Rhizobia suggesting interesting evolutionary implications.

Amino Acid Sequence↗

Distribution of potential type II restriction sites (palindromes) in prokaryotes.

Restriction-modification systems are used as a defensive mechanism against inappropriate invasion of foreign DNA. The recognition sequences for the common type II restriction enzymes and their corresponding methylases are usually palindromes. In this study, we identified the most over- and underrepresented words in DNA of four bacteria: Escherichia coli, Bacillus subtilis, Clostridium perfringens, and Pseudomonas aeruginosa. Using maximum order Markov chain analysis, we found that palindromic words were most often more underrepresented than their non-palindromic counterparts. No strict rule for the intragenic palindrome content could be derived, but for three of the bacteria there was a weak correlation between codon usage bias and palindrome content. A clear drop in palindrome counts was observed in the Shine-Dalgarno region for B. subtilis and C. perfringens, but not in E. coli or P. aeruginosa. It was also shown that palindromes in eubacteria and archaebacteria seem to occur slightly more infrequently than expected on the basis of the genomic GC-content, but some exceptions to this principle exist.

Bacteria↗

The nucleotide sequence coding for major outer membrane protein OmpA of Shigella dysenteriae.

The nucleotide sequence of the ompA gene from Shigella dysenteriae has been determined and the amino acid sequence of the pro-OmpA protein predicted. Sequence comparison between the ompA genes of S.dysenteriae and Escherichia coli showed that features such as mRNA secondary structure and codon usage, as well as polypeptide function, have been conserved during evolution. The pro-OmpA protein of S.dysenteriae consists of 351 residues, as opposed to the 346 of the E.coli protein and also shows several amino acid changes. These changes have been used to interpret differences in the biological activity of the two proteins.

Amino Acid Sequence↗

Homologous nucleotide sequences at the 5' termini of messenger RNAs synthesized from the yeast enolase and glyceraldehyde-3-phosphate dehydrogenase gene families. The primary structure of a third yeast glyceraldehyde-3-phosphate dehydrogenase gene.

Genomic DNA containing a third yeast glyceraldehyde-3-phosphate dehydrogenase structural gene has been isolated on a bacterial plasmid designated pgap11. The complete nucleotide sequence of this structural gene was determined. The gene contains no intervening sequences, codon usage is highly biased, and the nucleotide sequence of the coding portion of this gene is 90% homologous to the other two glyceraldehyde-3-phosphate dehydrogenase genes (Holland, J. P., and Holland, M. J. (1980) J. Biol. Chem. 255, 2596-2605). Based on the extent of nucleotide sequence divergence among the three glyceraldehyde-3-phosphate dehydrogenase genes, it is likely that they arose as a consequence of two duplication events and the gene contained on the hybrid plasmid designated pgap11 is a product of the first duplication event. All three structural genes share extensive nucleotide sequence homology in the 5'-noncoding regions adjacent to the three respective translational initiation codons. The gene contained on pgap11 is not homologous to the others downstream from the respective translational termination codon, however. The 5' termini of messenger RNAs synthesized from the three glyceraldehyde-3-phosphate dehydrogenase and two yeast enolase genes have been mapped to sites ranging from 36 to 82 nucleotides upstream from the respective translational initiation codons. In each case the 5' terminus of the mRNA maps to a region of strong nucleotide sequence homology which is shared by all five structural genes. These latter data confirm that all five structural genes are expressed during vegetative cell growth and further support the hypothesis that a portion of the 5'-noncoding flanking region of the yeast glyceraldehyde-3-phosphate dehydrogenase and enolase genes evolved from a common precursor sequence.

Base Sequence↗

The complete mitochondrial DNA sequence of the Atlantic salmon, Salmo salar.

The complete sequence of the Atlantic salmon (Salmo salar) mitochondrial genome has been determined. The entire sequence is 16665 base pairs (bp) in length, with a gene content (13 protein-coding, two ribosomal RNA [rRNA] and 22 transfer RNA [tRNA] genes) and order conforming to that observed in most other vertebrates. Base composition and codon usage have been detailed. Nucleotide and derived amino acid sequences of the 13 protein-coding genes from Atlantic salmon have been compared with their counterparts in rainbow trout. A putative structure for the origin of L-strand replication (O(L)) is proposed, and sequence features of the control region (D-loop) are described.

Animals↗

Target selection for structural genomics: a single genome approach.

We describe our strategy for selecting targets for protein structure determination in context of structural genomics of a single genome. In the course of target selection, we have studied two of the smallest microbial genomes, Mycoplasma genitalium and Mycoplasma pneumoniae. To our surprise, we found that only 71 Mycoplasma genes or their orthologues can be considered as easy targets for high-throughput structural studies--far fewer than expected. We discuss the methods and criteria used for target selection and the reasons explaining rarity of easy targets. First, despite the common opinion that protein folds can be predicted for only 30-50% of genes, the number of "truly unknown" structures is less than one-third. Second, due to the different codon usage, two thirds of Mycoplasma proteins cannot be directly expressed in E. coli in high-throughput manner and require substitution by their homologues from other organisms. Third, membrane or large multi-domain proteins are difficult targets because of solubility and size issues and often require identification and structure determination of protein domains. Finally, we propose different approaches to address the difficult targets.

Cell Membrane↗

Expression of tetanus toxin fragment C in E. coli: high level expression by removing rare codons.

Tetanus toxin fragment C had been previously expressed in Escherichia coli at 3-4% cell protein. The codon bias for tetanus toxin in Clostridium tetani is very different from that of highly expressed homologous genes in E. coli, resulting in the presence of many rare E. coli codons in the sequence encoding fragment C. We have replaced the coding sequence by sequence optimized for codon usage in E. coli, and show that the expression of fragment C is increased. Although the level of mRNA also increased this appeared to be a secondary consequence of more efficient translation. Complete sequence replacement increased expression to approximately 11-14% cell protein but only after the promoter strength had been improved.

Amino Acid Sequence↗

The mosaic genome of warm-blooded vertebrates.

Most of the nuclear genome of warm-blooded vertebrates is a mosaic of very long (much greater than 200 kilobases) DNA segments, the isochores; these isochores are fairly homogeneous in base composition and belong to a small number of major classes distinguished by differences in guanine-cytosine (GC) content. The families of DNA molecules derived from such classes can be separated and used to study the genome distribution of any sequence which can be probed. This approach has revealed (i) that the distribution of genes, integrated viral sequences, and interspersed repeats is highly nonuniform in the genome, and (ii) that the base composition and ratio of CpG to GpC in both coding and noncoding sequences, as well as codon usage, mainly depend on the GC content of the isochores harboring the sequences. The compositional compartmentalization of the genome of warm-blooded vertebrates is discussed with respect to its evolutionary origin, its causes, and its effects on chromosome structure and function.

Animals↗

Protein secondary structural types are differentially coded on messenger RNA.

Tricodon regions on messenger RNAs corresponding to a set of proteins from Escherichia coli were scrutinized for their translation speed. The fractional frequency values of the individual codons as they occur in mRNAs of highly expressed genes from Escherichia coli were taken as an indicative measure of the translation speed. The tricodons were classified by the sum of the frequency values of the constituent codons. Examination of the conformation of the encoded amino acid residues in the corresponding protein tertiary structures revealed a correlation between codon usage in mRNA and topological features of the encoded proteins. Alpha helices on proteins tend to be preferentially coded by translationally fast mRNA regions while the slow segments often code for beta strands and coil regions. Fast regions correspondingly avoid coding for beta strands and coil regions while the slow regions similarly move away from encoding alpha helices. Structural and mechanistic aspects of the ribosome peptide channel support the relevance of sequence fragment translation and subsequent conformation. A discussion is presented relating the observation to the reported kinetic data on the formation and stabilization of protein secondary structural types during protein folding. The observed absence of such strong positive selection for codons in non-highly expressed genes is compatible with existing theories that mutation pressure may well dominate codon selection in non-highly expressed genes.

Bacterial Proteins↗

Adhesive protein cDNA sequence of the mussel Mytilus coruscus and its evolutionary implications.

cDNA encoding the adhesive protein of the mussel Mytilus coruscus (Mcfp1) was isolated. The coding region encoded 848 amino acids (a.a.) comprising the 20-a.a. signal peptide, the 21-a.a. nonrepetitive linker, and the 805-a.a. repetitive domain. Although the first 204 nucleotides and the 3'-untranslated region of Mcfp1 cDNA were homologous to corresponding parts of M. galloprovincialis adhesive protein (Mgfp1) cDNA, the other parts diverged. The representative repeat motif of the repetitive domain, YKPK(I/P)(S/T)YPP(T/S), was similar but slightly different from the repeat motif of Mgfp1. The codon usage patterns for the same amino acids were different in different positions of the decapeptide motif. Almost identical nucleotide sequences encoding the two to 13 repeats appeared several times in the repetitive region, which suggests that the adhesive protein genes of mussels have evolved through the duplication of these repeat units.

Amino Acid Sequence↗