Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Comparison of codon usage measures and their applicability in prediction of microbial gene expressivity.

BACKGROUND: There are a number of methods (also called: measures) currently in use that quantify codon usage in genes. These measures are often influenced by other sequence properties, such as length. This can introduce strong methodological bias into measurements; therefore we attempted to develop a method free from such dependencies. One of the common applications of codon usage analyses is to quantitatively predict gene expressivity. RESULTS: We compared the performance of several commonly used measures and a novel method we introduce in this paper--Measure Independent of Length and Composition (MILC). Large, randomly generated sequence sets were used to test for dependence on (i) sequence length, (ii) overall amount of codon bias and (iii) codon bias discrepancy in the sequences. A derivative of the method, named MELP (MILC-based Expression Level Predictor) can be used to quantitatively predict gene expression levels from genomic data. It was compared to other similar predictors by examining their correlation with actual, experimentally obtained mRNA or protein abundances. CONCLUSION: We have established that MILC is a generally applicable measure, being resistant to changes in gene length and overall nucleotide composition, and introducing little noise into measurements. Other methods, however, may also be appropriate in certain applications. Our efforts to quantitatively predict gene expression levels in several prokaryotes and unicellular eukaryotes met with varying levels of success, depending on the experimental dataset and predictor used. Out of all methods, MELP and Rainer Merkl's GCB method had the most consistent behaviour. A 'reference set' containing known ribosomal protein genes appears to be a valid starting point for a codon usage-based expressivity prediction.

Chi-Square Distribution↗

The partial mitochondrial genome of the Cephalothrix rufifrons (Nemertea, Palaeonemertea): characterization and implications for the phylogenetic position of Nemertea.

A continuous 10.1kb fragment of the Cephalothrix rufifrons (Nemertea, Palaeonemertea) mitochondrial genome was sequenced and characterized to further assess organization of protostome mitochondrial genomes and evaluate the phylogenetic potential of gene arrangement and amino acid characters. The genome is A-T rich (72%), and this biased base composition is partly reflected in codon usage. Inferred tRNA secondary structures are typical of those reported for other metazoan mitochondrial DNAs. The arrangement of the 26 genes contained in the fragment exhibits marked similarity to those of many protostome taxa, most notably molluscs with highly conserved arrangements and a phoronid. Separate and simultaneous phylogenetic analyses of inferred amino acid sequences and gene adjacencies place the nemertean within the protostomes among coelomate lophotrochozoan taxa, but do not find a well-supported sister taxon link.

Animals↗

The mitochondrial genome of Strongyloides stercoralis (Nematoda) - idiosyncratic gene order and evolutionary implications.

The complete mitochondrial genome sequence of the parasitic nematode Strongyloides stercoralis was determined, and its organisation and structure compared with other nematodes for which complete mitochondrial sequence data were available. The mitochondrial genome of S. stercoralis is 13,758 bp in size and contains 36 genes (all transcribed in the clockwise direction) but lacks the atp8 gene. This genome has a high T content (55.9%) and a low C content (8.3%). Corresponding to this T content, there are 16 (poly-T) tracts of >/=12 Ts distributed across the genome. In protein-coding genes, the T bias is greatest (76.4%) at the third codon position compared with the first and second codon positions. Also, the C content is higher at the first (9.3%) and second (13.4%) codon positions than at the third (2%) position. These nucleotide biases have a significant effect on predicted codon usage patterns and, hence, on amino acid compositions of the mitochondrial proteins. Interestingly, six of the 12 protein-coding genes are predicted to employ a unique initiation codon (TTT), which has not yet been reported for any other animal mitochondrial genome. The secondary structures predicted for the 22 transfer RNA (trn) genes and the two ribosomal RNA (rrn) genes are similar to those of other nematodes. In contrast, the gene arrangement in the mitochondrial genome of S. stercoralis is different from all other nematodes studied to date, revealing only a limited number of shared gene boundaries (atp6-nad2 and cox2-rrnL). Evolutionary analyses of mitochondrial nucleotide and amino acid sequence data sets for S. stercoralis and seven other nematodes demonstrate that the mitochondrial genome provides a rich source of phylogenetically informative characters. In conclusion, the S. stercoralis mitochondrial genome, with its unique gene order and characteristics, should provide a resource for comparative mitochondrial genomics and systematics studies of parasitic nematodes.

Amino Acid Sequence↗

Structure and organization of the mitochondrial genome of the canine heartworm, Dirofilaria immitis.

This study determined the complete mitochondrial (mt) genome sequence of the canine heartworm, Dirofilaria immitis, and compared its structure, organization and other characteristics with Onchocerca volvulus and other secernentean nematodes. The D. immitis mt genome is 13814 bp in size and contains 36 of the 37 genes typical of metazoan organisms, and lacks the ATP synthetase subunit 8 gene. All of the genes are transcribed in the same direction. For the entire genome, the nucleotide contents are approximately 55% (T), approximately 19% (each for A and G) and approximately 7% (C), which is very similar to those of the protein-coding genes. In the latter genes, most (approximately 69%) third codon positions have a T, but rarely (approximately 1-9%) have an A or a C. The C content (8-12%) is higher at the first and second codon positions compared with the third position (approximately 1%). These nucleotide biases have a significant effect on the codon usage patterns and, thus, on the amino acid composition of the proteins. The mt genome organization of D. immitis is essentially the same as that of O. volvulus, but is distinctly different from other secernentean nematodes sequenced thus far. Irrespective of transpositions of transfer RNA (trn) genes and the non-coding, AT-rich region, there are 4 gene- or gene block-translocations between the mt genome of D. immitis and those of Caenorhabditis elegans, Ascaris suum and the 2 human hookworms, Ancylostoma duodenale and Necator americanus. For D. immitis, the 22 trn genes have secondary structures typical of other secernentean nematodes, and possess a TV-replacement loop instead of a TpsiC arm and loop. Like O. volvulus, the mt trnK and trnP of D. immitis use the anticodons CUU and AGG, whereas in other nematodes, UUU and UGG are employed, respectively. Also, the secondary structures of the 2 ribosomal RNA (rrn) genes are similar to the models for other nematodes. Overall, the availability of the complete D. immitis mt genome sequence provides a resource for future studies of the comparative mt genomics and of the population genetics and/or phylogeny of parasitic nematodes.

Amino Acid Sequence↗

Analysis of major ampullate silk cDNAs from two non-orb-weaving spiders.

Compared to other arthropods, spiders are unique in their use of silk throughout their life span and the extraordinary mechanical properties of the silk threads they produce. Studies on orb-weaving spider silk proteins have shown that silk proteins are composed of highly repetitive regions, characterized by alanine and glycine-rich units. We have isolated and sequenced four partial cDNA clones representing major ampullate spider silk gene transcripts from two non-orb weavers: three for Kukulcania hibernalis and one for Agelenopsis aperta. These cDNA sequences were compared to each other, as well as to the previously published orb-weaver silk gene sequences. The results indicate that the repeats encoding conserved amino acid motifs such as polyA and polyGA that are characteristic of some orb-weaving spider silks are also found in some of the cDNAs reported in this study. However, we also found other motifs such as polyGS and polyGV in the cDNA sequences from the two non-orb-weaving spiders. The amino acid composition of the silk gland extracts shows that alanine and glycine are the major components of the silk of these two non-orb weavers as is the case in orb-weaver silks. Sequence alignment shows that A. aperta's cDNA displays a C-terminal encoding region that is about 44% similar to the one present in N. clavipes's MaSp1 cDNA. In addition, as previously observed for spider silk sequences, the analysis of the codon usage for these four cDNAs demonstrates a bias for A or T in the wobble base position.

Amino Acid Sequence↗

The complete mitochondrial DNA sequence of the horseshoe crab Limulus polyphemus.

We determined the complete 14,985-nt sequence of the mitochondrial DNA of the horseshoe crab Limulus polyphemus (Arthropoda: Xiphosura). This mtDNA encodes the 13 protein, 2 rRNA, and 22 tRNA genes typical for metazoans. The arrangement of these genes and about half of the sequence was reported previously; however, the sequence contained a large number of errors, which are corrected here. The two strands of Limulus mtDNA have significantly different nucleotide compositions. The strand encoding most mitochondrial proteins has 1. 25 times as many A's as T's and 2.33 times as many C's as G's. This nucleotide bias correlates with the biases in amino acid content and synonymous codon usage in proteins encoded by different strands and with the number of non-Watson-Crick base pairs in the stem regions of encoded tRNAs. The sizes of most mitochondrial protein genes in Limulus are either identical to or slightly smaller than those of their Drosophila counterparts. The usage of the initiation and termination codons in these genes seems to follow patterns that are conserved among most arthropod and some other metazoan mitochondrial genomes. The noncoding region of Limulus mtDNA contains a potential stem-loop structure, and we found a similar structure in the noncoding region of the published mtDNA of the prostriate tick Ixodes hexagonus. A simulation study was designed to evaluate the significance of these secondary structures; it revealed that they are statistically significant. No significant, comparable structure can be identified for the metastriate ticks Rhipicephalus sanguineus and Boophilus microplus. The latter two animals also share a mitochondrial gene rearrangement and an unusual structure of mt-tRNA(C) that is exactly the same association of changes as previously reported for a group of lizards. This suggests that the changes observed are not independent and that the stem-loop structure found in the noncoding regions of Limulus and Ixodes mtDNA may play the same role as that between trnN and trnC in vertebrates, i.e., the role of lagging strand origin of replication.

Animals↗

Identification, genetic analysis and DNA sequence of a 7.8-kb virulence region of the Salmonella typhimurium virulence plasmid.

The 90-kilobase (kb) virulence plasmid of Salmonella typhimurium is responsible for invasion from the intestines to mesenteric lymph nodes and spleens of orally inoculated mice. We used Tn5 and aminoglycoside phosphotransferase (aph) gene insertion mutagenesis and deletion mutagenesis of a previously identified 14-kb virulence region to reduce this virulence region to 7.8kb. The 7.8-kb virulence region subcloned into a low copy-number vector conferred a wild-type level of splenic infection to virulence plasmid-cured S. typhimurium and conferred essentially a wild-type oral LD50. Insertion mutagenesis identified five loci essential for virulence, and DNA sequence analysis of the virulence region identified six open reading frames. Expected protein products were identified from four of the six genes, with three of the proteins identified as doublet bands in Escherichia coli minicells. Three of the five mutated genes were able to be complemented by clones containing only the corresponding wild-type gene. Only one of the five deduced amino acid sequences, that of the positive regulatory element, SpvR, possessed significant homology to other proteins. The codon usage for the virulence genes showed no codon bias, which is consistent with the low levels of expression observed for the corresponding proteins. Consensus promoters for several different sigma factors were identified upstream of several of the genes, whereas only consensus Rho-dependent termination sequences were observed between certain of the genes. The operon structure of this virulence region therefore appears to be complex. The construction of the cloned 7.8-kb virulence region and the determination of the DNA sequence will aid in the further genetic analysis of the five plasmid-encoded virulence genes of S. typhimurium.

Amino Acid Sequence↗

[Characterization and phylogenetic analysis of the cytochrome oxidase subunit I gene of mitochondrial genome from Takifugu fasciatus].

The Takifugu fasciatus mitochondrial cytochrome oxidase I gene (COI) and its associated tRNA genes were sequenced by PCR. The open reading frame of COI gene contains 1,546 bp nucleotides, encoding a putative protein of 515 amino acid residues. The pattern of codon usage of the COI gene is less biased toward A+T. The COI gene of T. fasciatus shows a high degree of homology with that from the other 14 fish species recorded in the GenBank and has 97.6% homology with Takifugu rubripes, 76.5% with Masturus lanceolatus and 75.4% with Mola mola. The phylogenetic trees show that the relationships based on the homology is consistent with the morphological and taxonomic results. Predicted secondary structures of the tRNA genes suggest that they have the classical cloverleaf structures.

Amino Acid Sequence↗

Cloning, nucleotide sequence, and expression of the DNA ligase-encoding gene from Thermus filiformis.

The gene encoding Thermus filiformis (Tfi) DNA ligase was cloned and its nucleotide sequence was determined by the chain-termination method. The primary structure of Tfi DNA ligase was deduced from its nucleotide sequence. The Tfi DNA ligase comprises of 667 amino acid residues and its molecular mass was determined to be 75,936 Da. The deduced amino acid sequence of Tfi DNA ligase showed a 86.5% homology to Tth DNA ligase and 43.5% to E. coli DNA ligase. The Lys-116 of Lys-Val-Asp-Gly motif was proposed to be the active residue of Tfi DNA ligase. In comparison with the amino acid composition of DNA ligase, Tfi DNA ligase showed a significant increase in the proportion of charged residues, Arg and Glu, compared to E. coli DNA ligase. The G + C content in the first, second, and third positions of the codons used were 70.3%, 40.3%, and 90.3%, respectively. Codon usage in Tfi DNA ligase was heavily biased towards the use of G + C in the third position. Under tac promoter control, Tfi DNA ligase was overproduced to greater than 9% of E. coli BL26Blue cellular proteins.

Amino Acid Sequence↗

The evolution of codon preferences in Drosophila: a maximum-likelihood approach to parameter estimation and hypothesis testing.

Synonymous codon usage in related species may differ as a result of variation in mutation biases, differences in the overall strength and efficiency of selection, and shifts in codon preference-the selective hierarchy of codons within and between amino acids. We have developed a maximum-likelihood method to employ explicit population genetic models to analyze the evolution of parameters determining codon usage. The method is applied to twofold degenerate amino acids in 50 orthologous genes from D. melanogaster and D. virilis. We find that D. virilis has significantly reduced selection on codon usage for all amino acids, but the data are incompatible with a simple model in which there is a single difference in the long-term Ne, or overall strength of selection, between the two species, indicating shifts in codon preference. The strength of selection acting on codon usage in D. melanogaster is estimated to be |Nes| approximately 0.4 for most CT-ending twofold degenerate amino acids, but 1.7 times greater for cysteine and 1.4 times greater for AG-ending codons. In D. virilis, the strength of selection acting on codon usage for most amino acids is only half that acting in D. melanogaster but is considerably greater than half for cysteine, perhaps indicating the dual selection pressures of translational efficiency and accuracy. Selection coefficients in orthologues are highly correlated (rho = 0.46), but a number of genes deviate significantly from this relationship.

Amino Acids↗

Cloning and nucleotide sequence analysis of human embryonic zeta-globin cDNA.

Clones of human embryonic alpha-like zeta-globin cDNA were isolated, by detection using cross-hybridization to human alpha-globin cDNA probes, from a cDNA library derived from the mRNA of the human erythroleukemia cell line K562. Nucleotide sequence analysis of these cDNA clones revealed a coding sequence that corresponds perfectly to the independently derived amino acid sequence of the human zeta-globin chain. Comparison of the nucleotide sequence of human zeta-globin cDNA with that of human alpha-globin cDNA confirmed previous estimates of very distant evolutionary divergence between the human zeta- and alpha-globin genes. Nevertheless, the human zeta-globin cDNA sequence shares a remarkable similarity to that of the alpha-globin gene in its codon usage, high G + C base composition, and lack of bias against usage of CG dinucleotides.

Base Composition↗

Comparative analyses of codon and amino acid usage in symbiotic island and core genome in nitrogen-fixing symbiotic bacterium Bradyrhizobium japonicum.

Genes involved in the symbiotic interactions between the nitrogen-fixing endosymbiont Bradyrhizobium japonicum, and its leguminous host are mostly clustered in a symbiotic island (SI), acquired by the bacterium through a process of horizontal transfer. A comparative analysis of the codon and amino acid usage in core and SI genes/proteins of B. japonicum has been carried out in the present study. The mutational bias, translational selection, and gene length are found to be the major sources of variation in synonymous codon usage in the core genome as well as in SI, the strength of translational selection being higher in core genes than in SI. In core proteins, hydrophobicity is the main source of variation in amino acid usage, expressivity and aromaticity being the second and third important sources. But in SI proteins, aromaticity is the chief source of variation, followed by expressivity and hydrophobicity. In SI proteins, both the mean molecular weight and mean aromaticity of individual proteins exhibit significant positive correlation with gene expressivity, which violate the cost-minimization hypothesis. Investigation of nucleotide substitution patterns in B. japonicum and Mesorhizobium loti orthologous genes reveals that both synonymous and non-synonymous sites of highly expressed genes are more conserved than their lowly expressed counterparts and this conservation is more pronounced in the genes present in core genome than in SI.

Amino Acids↗

Evolution of codon usage and base contents in kinetoplastid protozoans.

In this study we analyze and compare the trends in codon usage in five representative species of kinetoplastid protozoans (Crithidia fasciculata, Leishmania donovani, L. major, Trypanosoma cruzi and T. brucei), with the purpose of investigating the processes underlying these trends. A principal component analysis shows that the G+C content at the third codon position represents the main source of codon-usage variation, both within species (among genes) and among species. The non-Trypanosoma species exhibit narrow distributions in codon usage, while both Trypanosoma species present large within-species heterogeneity. The three non-Trypanosoma species have very similar codon-usage preferences. These codon preferences are also shared by the highly expressed genes of T. cruzi and to a lesser degree by those of T. brucei. This leads to the conclusion that the codon preferences shared by these species are the ancestral ones in the kinetoplastids. On the other hand, the study of noncoding sequences shows that Trypanosoma species exhibit mutational biases toward A + T richness, while the non-Trypanosoma species present mutational pressure in the opposite direction. These data taken together allow us to infer the origin of the different codon-usage distributions observed in the five species studied. In C. fasciculata and Leishmania, both mutational biases and (translational) selection pull toward G + C richness, resulting in a narrow distribution. In Trypanosoma species the mutational pressure toward A + T richness produced a shift in their genomes that differentially affected coding and noncoding sequences. The effect of these pressures on the third codon position of genes seems to have been inversely proportional to the level of gene expression.

Animals↗

Genomic choice of codons in 16 microbial species.

We study the codon usage over whole set of ORFs of 16 unicellular microbial species: eight archaebacteria, seven eubacteria, and one eukarya. We first try to define, for each species, the neutral expected codon usage to better approach subsequently the influence of selection. Overlapping triplets counted from the complete DNA genomic sequence and mean amino acid composition of ORFs allow us to build satisfying expected codon usage for each species. Within species deviation from this neutral model is then studied through Correspondence Analysis and characterization with bias index, N(C)' (effective number of codons reported to neutral model). Our results are compared to previously published ones for three species and let appear good agreement in spite of very different methods. We thus propose set of codons probably preferred by selection for nine other species. In the four last species, no clear preference can be evidenced. Finally, we characterize variation of codon usage over functional categories. We propose that the high degree of bias of proteins involved in translation, ribosomal structure and biogenesis has a positive influence on overexpression of the corresponding genes under optimum growth conditions and is a negative regulator of the same genes when amino acids become limited resources.

Base Composition↗

Molecular structure of the Frankia spp. nifD-K intergenic spacer and design of Frankia genus compatible primer.

The nifD-K intergenic spacer (IGS) of ArI3 and ACoN24d were found to have a length 265 and 199 nucleotides, respectively. They are markedly less conserved than the two neighbouring genes and have, in some instances, a repeated structure reminiscent of an insertion event. The repeated sequence and the IGSs have no detectable homology with sequences in DNA databanks. The IGS has a stem-loop structure with a low folding energy, lower than that between nifH and nifD. No convincing alignment of IGS sequences could be obtained among Frankia strains. Only between ACoN24d and ArI3, which belong to the same genomic species, was the alignment good enough to permit detection of a doubly repeated structure. No promoter could be detected in the IGSs. The putative nifK open reading frame (ORF) in Frankia strain ArI3 has a length of 1587 nucleotides, starting with a GTG codon, preceded by a ribosome binding site of a structure similar to that of nifH (GGAGGN7). The codon usage was similar to that of previously sequenced Frankia genes with a strong bias toward G- and C-ending codons except in the case of glycine where GGT is frequent. Alignment of the three Frankia nifK sequences (EUN1f; ArI3 and ACoN24d) with those of other nitrogen-fixing bacteria permitted detection of a sequence conserved among the three Frankia strains but absent in the other sequences. A primer targeted to that region in combination with FGPD807-85 amplified the nifD-KIGS sequences of all Frankia strains (except the non-nitrogen-fixing Frankia strains CN3 and AgB1-9) and yet failed to amplify DNA of all other nitrogen-fixing bacteria.(ABSTRACT TRUNCATED AT 250 WORDS)

Actinomycetales↗

Cloning, sequence and expression of a beta-tubulin-encoding gene in the homobasidiomycete Schizophyllum commune.

The beta-tubulin (beta Tub)-encoding gene (tub-2) of Schizophyllum commune is the first tubulin gene isolated, cloned and sequenced from higher filamentous fungi (homobasidiomycetes). The S. commune tub-2 gene is organized into nine exons and eight introns. The introns vary from 48 to 107 nt in length, and are distributed throughout the gene. The tub-2 exons code for a protein of 445 amino acids (aa), which shows great homology with beta Tubs of filamentous ascomycetes, plants, and animals, but less homology with yeasts. The codon usage of tub-2 from S. commune is biased, as it is in most beta Tub-encoding genes of filamentous fungi. The S. commune beta Tub shows a conserved aa sequence in the C-terminal domain, which is suggested to interact with microtubule-associated proteins in animals. In contrast, the S. commune beta Tub deviates from most known beta Tubs by having a Cys165 residue, which might be significant for the insensitivity of S. commune haploid strains to the antimicrotubule drug, benomyl. In tub-2 of different haploid strains, sequence polymorphisms occur in the 5' and 3' flanking regions. The expression of tub-2 is high in young mycelium, which has a high number of extending apical cells, but decreases with the aging of the mycelium. No significant difference in the hybridization signal intensity for the tub-2 transcripts was recorded either during intercellular nuclear migration at early mating, or in mycelia with a mutation in the B mating-type gene.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Evidence for translational selection in codon usage in Echinococcus spp.

We analysed the intragenomic variation in codon usage in Echinococcus spp. by correspondence analysis. This approach detected a trend among genes which was correlated with expression levels. Among the (presumed) highly expressed sequences we found an increased usage of a subset of codons, almost all of them G- or C- ending. Since an increase in these bases at the synonymous sites is against the mutational bias (these genomes are slightly A+-T- rich), we conclude that codon usage in Echinococcus is the result of an equilibrium between compositional pressure and selection, the latter acting at the level of translation, mainly on highly expressed genes. This is the first report where translational selection for codon usage is detected among Platyhelminthes.

Animals↗

Nonneutral GC3 and retroelement codon mimicry in Phytophthora.

Phytophthora is a genus entirely comprised of destructive plant pathogens. It belongs to the Stramenopila, a unique branch of eukaryotes, phylogenetically distinct from plants, animals, or fungi. Phytophthora genes show a strong preference for usage of codons ending with G or C (high GC3). The presence of high GC3 in genes can be utilized to differentiate coding regions from noncoding regions in the genome. We found that both selective pressure and mutation bias drive codon bias in Phytophthora. Indicative for selection pressure is the higher GC3 value of highly expressed genes in different Phytophthora species. Lineage specific GC increase of noncoding regions is reminiscent of whole-genome mutation bias, whereas the elevated Phytophthora GC3 is primarily a result of translation efficiency-driven selection. Heterogeneous retrotransposons exist in Phytophthora genomes and many of them vary in their GC content. Interestingly, the most widespread groups of retroelements in Phytophthora show high GC3 and a codon bias that is similar to host genes. Apparently, selection pressure has been exerted on the retroelement's codon usage, and such mimicry of host codon bias might be beneficial for the propagation of retrotransposons.

Base Composition↗