Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Extreme differences in charge changes during protein evolution.

The maintenance of a proper distribution of charged amino acid residues might be expected to be an important factor in protein evolution. We therefore compared the inferred changes in charge during the evolution of 43 protein families with the changes expected on the basis of random base substitutions. It was found that certain proteins, like the eye lens crystallins and most histones, display an extreme avoidance of changes in charge. Other proteins, like phospholipase A2 and ferredoxin, apparently have sustained more charged replacements than expected, suggesting a positive selection for changes in charge. Depending on function and structure of a protein, charged residues apparently can be important targets for selective forces in protein evolution. It appears that actual biased codon usage tends to decrease the proportion of charged amino acid replacements. The influence of nonrandomness of mutations is more equivocal. Genes that use the mitochondrial instead of the universal code lower the probability that charge changes will occur in the encoded proteins.

Biological Evolution↗

Development of Polymorphic EST Markers Suitable for Genetic Linkage Mapping of Catfish.

: Expressed sequence tag (EST) markers are important for gene mapping and for marker-assisted selection (MAS). To develop EST markers for use in catfish gene mapping, 100 randomly picked complementary DNAs from the channel catfish (Ictalurus punctatus) pituitary library were sequenced. The EST sequences were used to design primers to amplify channel catfish and blue catfish (I. furcatus) genomic DNAs. Polymerase chain reaction products of the ESTs were analyzed to determine length polymorphism between the channel catfish and blue catfish. Eleven polymorphic EST markers were identified. Five of the 11 EST markers were from known genes and the other six were from unidentified ESTs. Seven ESTs were found to be associated with microsatellite sequences. Analysis of channel catfish gene sequences indicated highly biased codon usage, with 16 codons being preferably used. These codons were more preferably used in highly expressed ribosomal protein genes and in highly expressed pituitary hormone genes. G/C-rich codons are less used in channel catfish than those in other vertebrates suggesting AT-richness of the channel catfish genome.

Journal Article↗

Correlations between mRNA expression levels and GC contents of coding and untranslated regions of genes in rodents.

Gene expression is regulated by a highly coordinated network of events whose efficiency may constrain the level of expression. Among other factors, natural selection for increased translational efficiency and/or fidelity may shape nucleotide composition and, hence, codon usage during evolution. Previous studies have shown that highly expressed genes in Saccharomyces cerevisiae, Caenorhabditis elegans, and Drosophila melanogaster have relatively higher codon usage biases. However, in the case of mammals, results have been equivocal. In this study, we assessed the correlation between nucleotide composition and mRNA expression levels of rodent genes measured by cDNA microarray and serial analysis of gene expression (SAGE) techniques. We found that mRNA expression levels were correlated with the third nucleotide position GC (GC3) content for both Rattus norvegicus (r = 0.246, p = 0.01; N = 110) and Mus musculus (r = 0.21, p = 0.0026; N = 203) genes. However, no significant correlation was evident between mRNA expression level and GC contents of 5'- and 3'-untranslated regions (UTRs) for either species. This suggests that, in rodents, nucleotide composition of coding sequences and UTRs might evolve differentially when considered along an expression gradient. Accordingly, it is possible that higher GC levels may present the rodent genes with a selective advantage for translational efficiency. However, the increase in GC3 content seems to level off above an expressional threshold (e.g., >or=threefold the median expression for R. norvegicus), suggesting that conflicting demands posed by different aspects of transcriptional and translational machineries (e.g., efficiency versus fidelity) may set an upper limit for GC3.

Animals↗

Phylogeny and the evolution of the Amylase multigenes in the Drosophila montium species subgroup.

To investigate the phylogenetic relationships and molecular evolution of alpha-amylase (Amy) genes in the Drosophila montium species subgroup, we constructed the phylogenetic tree of the Amy genes from 40 species from the montium subgroup. On our tree the sequences of the auraria, kikkawai, and jambulina complexes formed distinct tight clusters. However, there were a few inconsistencies between the clustering pattern of the sequences and taxonomic classification in the kikkawai and jambulina complexes. Sequences of species from other complexes (bocqueti, bakoue, nikananu, and serrata) often did not cluster with their respective taxonomic groups. This suggests that relationships among the Amy genes may be different from those among species due to their particular evolution. Alternatively, the current taxonomy of the investigated species is unreliable. Two types of divergent paralogous Amy genes, the so-called Amy1- and Amy3-type genes, previously identified in the D. kikkawai complex, were common in the montium subgroup, suggesting that the duplication event from which these genes originate is as ancient as the subgroup or it could even predate its differentiation. Thc Amy1-type genes were closer to the Amy genes of D. melanogaster and D. pseudoobscura than to the Amy3-type genes. In the Amy1-type genes, the loss of the ancestral intron occurred independently in the auraria complex and in several Afrotropical species. The GC content at synonymous third codon positions (GC3s) of the Amy1-type genes was higher than that of the Amy3-type genes. Furthermore, the Amy1-type genes had more biased codon usage than the Amy3-type genes. The correlations between GC3s and GC content in the introns (GCi) differed between these two Amy-type genes. These findings suggest that the evolutionary forces that have affected silent sites of the two Amy-type genes in the montium species subgroup may differ.

Amylases↗

Transposable element orientation bias in the Drosophila melanogaster genome.

Nonrandom distributions of transposable elements can be generated by a variety of genomic features. Using the full D. melanogaster genome as a model, we characterize the orientations of different classes of transposable elements in relation to the directionality of genes. DNA-mediated transposable elements are more likely to be in the same orientation as neighboring genes when they occur in the nontranscribed region's that flank genes. However, RNA-mediated transposable elements located in an intron are more often oriented in the direction opposite to that of the host gene. These orientation biases are strongest for genes with highly biased codon usage, probably reflecting the ability of such loci to respond to weak positive or negative selection. The leading hypothesis for selection against transposable elements in the coding orientation proposes that transcription termination poly(A) signal motifs within retroelements interfere with normal gene transcription. However, after accounting for differences in base composition between the strands, we find no evidence for global selection against spurious transcription termination signals in introns. We therefore conclude that premature termination of host gene transcription due to the presence of poly(A) signal motifs in retroelements might only partially explain strand-specific detrimental effects in the D. melanogaster genome.

Animals↗

Mitogenomic and phylogenomic analyses identify a cohesive Western Atlantic lineage within the Narcine complex (Torpediniformes: Narcinidae).

BACKGROUND: Accurate species delimitation within electric rays of the genus Narcine has been hindered by overlapping morphological characters and limited molecular resolution in previous single-locus studies. This study aims to evaluate phylogenetic relationships and species boundaries within the Narcine species complex across the Western Atlantic using complete mitochondrial genomes. METHODS AND RESULTS: Seven complete mitogenomes were newly assembled from individuals representing distinct morphotypes sampled across geographically widespread Western Atlantic localities and analyzed together with publicly available reference sequences. Mitochondrial protein-coding genes (PCGs) were examined using concatenated nucleotide and amino acid datasets under partitioned maximum-likelihood frameworks. Both approaches recovered highly congruent topologies, consistently supporting a single, well-defined western Atlantic mitochondrial lineage with low internal divergence (0.04-2.13%). Species delimitation analyses based on multiple methods yielded partially congruent results but consistently identified a dominant lineage encompassing all Atlantic samples. In contrast, two Colombian reference mitogenomes formed a separate and highly divergent lineage relative to the Atlantic group, despite showing moderate divergence between them. Comparative mitogenomic analyses revealed conserved genome organization, nucleotide composition bias, codon usage, and transfer RNA (tRNA) structures. All PCGs evolved under strong purifying selection, with Ka/Ks ratios well below unity. CONCLUSIONS: These results support mitochondrial genetic continuity across the Western Atlantic Narcine populations and do not provide mitochondrial evidence for multiple evolutionary lineages within the Western Atlantic. The marked mitochondrial divergence of Colombian reference mitogenomes highlights potential issues in sequence attribution and underscores the importance of data curation. Overall, complete mitochondrial genomes provide a robust framework for species delimitation and future integrative taxonomic assessments within Narcine.

Animals↗

Cloning and cDNA sequence of the rat X-chromosome linked phosphoglycerate kinase.

This paper reports the isolation and the sequence determination of rat phosphoglycerate kinase (PGK) cDNA clones. This cDNA, derived from an X-linked PGK gene transcript, contains a reading frame of 1254 nt and 5' and 3' non coding regions of 40 and 380 nt respectively. Analysis of the nucleotide sequence at the three codon position shows a biased codon usage with a prevalence of the triplet G non G N. Comparison of the inferred rat amino acid sequence with that of other organisms makes possible the calculation of the unit evolutionary period (UEP) for this enzyme, placing it at around 40 million years (My). Thus PGK is one of the oldest housekeeping enzymes.

Amino Acid Sequence↗

Sequence, evolution and differential expression of the two genes encoding variant small subunits of ribulose bisphosphate carboxylase/oxygenase in Chlamydomonas reinhardtii.

We have sequenced the two genes for the small subunit of ribulose bisphosphate carboxylase/oxygenase (Rubisco) in Chlamydomonas reinhardtii and analyzed their expression. The two genes encode variant small subunits that differ by four amino acid residues. Both genes are expressed and each is transcribed into an RNA of distinct size. The accumulation of the two RNAs changes depending on the growth conditions, so the small subunit composition of Rubisco may be expected to differ in response to the environment. The C. reinhardtii small subunit sequence is homologous to those of vascular plants or cyanobacteria, but is longer at the amino terminus and in internal positions. The number and location of the intervening sequences in the genes from C. reinhardtii and from other plants differ. In several cases, internal length differences in the polypeptide coincide with the positions of introns in the coding sequence. Thus, changes in the exon structure of the genes during evolution may have been accompanied by substantial changes in the encoded protein. The translation and splicing signals in C. reinhardtii are similar to those of other eukaryotes, but the transcription signals are less conserved and the highly biased codon usage is very unusual.

Amino Acid Sequence↗

DNA sequence and comparative analyses of the equine herpesvirus type 1 immediate early gene.

The immediate early (IE) proteins of herpesviruses are important regulatory factors which control the expression of genes at the transcriptional level. We report the DNA sequence of the immediate early gene of the alphaherpesvirus equine herpesvirus type 1 (EHV-1). This sequence is shown to be extremely rich in guanine and cytosine, resulting in a highly biased codon usage. The IE gene region possesses 38 open reading frames (ORFs) greater than 300 bp in length, 11 of which have coding regions of at least 100 amino acids (aa) following potential translation initiator codons. The largest ORF consists of 1487 codons (4461 bp) starting with the first ATG and would encode a protein of MW 155,000. TATA and CCAAT sequences as well as several potential cis-acting elements lie upstream to the major ORF. The deduced amino acid sequence for the 155,000 protein has a high degree of homology to the herpes simplex virus type 1 (HSV-1) ICP4 protein and its varicella-zoster virus (VZV) homolog. The regions of the EHV-1 IE protein that are homologous with these proteins correspond to the previously determined pattern of homology between the HSV and VZV IE polypeptides. However, there are are a number of differences within these broadly defined regions. It is therefore expected that this comparative study will facilitate the identification of functionally important residues within the amino acid sequence of IE proteins.

Amino Acid Sequence↗

Effect of a rare leucine codon, TTA, on expression of a foreign gene in Streptomyces lividans.

Streptomyces are bacteria with a very high chromosomal G+C composition (> 70 mol%) and extremely biased codon usage. In order to investigate the relationship between codon usage and gene expression in Streptomyces, we used ssi (Streptomyces subtilisin inhibitor) as a reporter gene and monitored its secretory expression in S. lividans. In consequence of alteration of the native codons of Leu, Lys and Ser of ssi to minor ones by site-directed mutagenesis, i.e., Leu79-Leu80: CTG-CTC to TTA-TTA, Lys89: AAG to AAA, Ser108-Ser109: TCG-AGC to TCT-TCT, respectively, the production of SSI was reduced remarkably in the case of TTA codons, while it was slightly increased in the case of AAA and almost the same in TCT codons. This conspicuous decrease found for Leu codon replacement was probably due to the low availability of intracellular tRNA(Leu) (UUA), a product of bldA which has been reported to be expressed only during the late stage of growth.

Amino Acid Sequence↗

Baculovirus expression of human basic fibroblast growth factor from a synthetic gene: role of the Kozak consensus and comparison with bacterial expression.

Synthetic genes encoding the 146 and 155 amino acid forms of human basic fibroblast growth factor (bFGF) were constructed with codon usage biased towards the polyhedrin-encoding gene of Autographa californica nuclear polyhedrosis virus (AcNPV). Expression of both bFGF genes in Spodoptera frugiperda (SF-21) suspension cell culture using a recombinant baculovirus yielded approximately 2.5 mg of mitogenically fully active protein per 10(9) cells following heparin-affinity chromatography. To improve translational efficiency, the Kozak consensus sequence was introduced and it was found that neither the replacement of a pyrimidine by a purine at position -3, nor the nature of the base at position +4 had any noticeable effect on the final levels of bFGF expression in SF-21 cells. The bases at these critical points in the consensus do not therefore play a major role in expression levels of the bFGF synthetic genes. The two synthetic genes were also expressed in Escherichia coli as native proteins using the T7 expression system. 5 mg of mitogenically fully active bFGF were obtained from 1 l of bacterial culture. Both insect cell- and E. coli-derived bFGF were equally mitogenic for Swiss 3T3 fibroblasts.

3T3 Cells↗

Post-transcriptional control in Escherichia coli: translation and degradation of the atp operon mRNA.

An attractive subject for investigations of post-transcriptional control is the atp operon, whose nine genes are differentially expressed. The primary mode of control of atp gene expression is exercised at the translational level. It has been clearly demonstrated for almost all of the atp genes that the primary and secondary structures of their respective translational initiation regions direct translational initiation rates that correspond well to the requirements for these subunits in the cell. The relationship between the structure of the translational initiation region, including bases upstream from the Shine-Dalgarno region and downstream from the start codon, and the rates of initiation that it determines, has been investigated in more detail using various polycistronic and monocistronic systems. No evidence could be found for a role of codon usage bias in controlling overall translation rates. The functional half-lives of atpE and of the other six cistrons downstream from it are similar. The chemical stabilities of the first two cistrons of the polycistronic atp mRNA may, however, be lower, and we are investigating the possibility that there may also be control of atp gene expression exercised at the level of mRNA stability. The effects of manipulations of the intercistronic regions of at least the plasmid borne atp operon are consistent with a model of mRNA decay in which rate control is associated with endonucleolytic cleavages within individual cistrons. The experimental data are discussed in relation to the possible ways in which primary and secondary structures of the mRNA might control translational efficiency and stability.

Blotting, Northern↗

The genomic nucleotide sequences of two differentially expressed actin-coding genes from the sea star Pisaster ochraceus.

The genomic sequences of two differentially expressed actin genes from the sea star Pisaster ochraceus are reported. The cytoplasmic actin gene (Cy) is expressed in eggs and early development. The muscle actin gene (M) is expressed in tube feet and testes. Both genes contain an 1125-nucleotide coding region interrupted by three introns at codons 41, 121 and 204. Gene M contains two additional introns at codons 150 and 267. The intron position at codon 150, although present in higher vertebrate actins, has not been reported in actin genes from invertebrates. The M gene coding region has 89.5% nucleotide homology to the Cy gene, and differs from the Cy actin gene in 13 of 375 amino acids (aa), 11 of which are found in the C-terminal half of the gene. The C-terminal half of the M gene contains a significant number of muscle isotype codons. Even though there is only 1 aa change in the first 150 codons, there have been limited substitutions at many four-fold degenerate sites which may indicate selection pressure upon the secondary structure of the mRNA and/or a biased codon usage. Variant CCAAT, TATA, and poly(A)-addition signals have been identified in the 5' and 3' flanking regions. The presence of 5' and 3' splice junction sequences in the 5' flanking region of the Cy gene suggests the potential for an intron there.

Actins↗

The phosphofructokinase genes of yeast evolved from two duplication events.

Yeast phosphofructokinase (PFK) is an octameric enzyme composed of four alpha-subunits and four beta-subunits, encoded by the genes PFK1 and PFK2, respectively. PFK1 was mapped 23 cM distal to ADE3 on chromosome VII, and PFK2 30 cM proximal to RNA1 on chromosome XIII. The entire nucleotide sequences for the two genes were obtained by sequencing both DNA strands. Only one major open reading frame was found for each gene. They encode 987 aa for PFK1 (Mr 107,984) and 959 aa for PFK2 (Mr 104,589). Both genes show a biased codon usage. The deduced amino acid sequences showed: (i) 20% homology between the N- and the C-terminal halves of each subunit, (ii) 55% homology between the two subunits, and (iii) significant homologies to the PFK sequences from human and rabbit muscle (42%), Escherichia coli (34%), and Bacillus (36%). These data support the view that two gene duplication events occurred in the evolution of the yeast PFK genes. The first duplication event took place soon after the separation of prokaryotic and eukaryotic lineage and the second in Saccharomyces later in the phylogeny. Functional domains in the yeast subunits were deduced by comparison to the rabbit muscle enzyme.

Amino Acid Sequence↗

Chloramphenicol resistance in Campylobacter coli: nucleotide sequence, expression, and cloning vector construction.

A chloramphenicol-resistance determinant (CmR), originally cloned from Campylobacter coli plasmid pNR9589 in Japan, was isolated and the nucleotide sequence determined, which contained an open reading frame of 621 bp. The gene product was identified as Cm acetyltransferase (CAT), which had a putative amino acid sequence that showed 43% to 57% identity with other CAT proteins of both Gram+ and Gram- origin. Although expression of the cat gene was constitutive in both C. coli and Escherichia coli, results of primer extension experiments indicated that transcription was initiated at different sites in these two species. A kanamycin-resistance determinant, identified as the aphA-3 gene, was located downstream from the cat gene. The codon usage of the cat gene is very different from that used in E. coli, however, the CAT polypeptide was synthesized in large amounts in E. coli maxicells. Therefore, the codon usage bias is not one of the obstacles which affects Campylobacter spp. gene expression in E. coli. New Campylobacter cloning vectors were constructed in this study.

Amino Acid Sequence↗

Expression of a foreign gene in Chlamydomonas reinhardtii.

Genomic transformation of Chlamydomonas reinhardtii exposed to glass-bead abrasion was accomplished with a chimeric neomycin phosphotransferaseII (NPTII)-encoding gene (nos::npt) flanked by the nopaline synthase promoter and polyadenylation sequences obtained from the Ti plasmid of Agrobacterium tumefaciens. These sequences were in a plasmid (pGA482) which also contained gene nit1 encoding nitrate reductase of C. reinhardtii. Transformants were selected by their ability to grow on medium containing nitrate, and 52% of these was also resistant to kanamycin. Evidence for nos::npt expression includes: (1) hybridization with probes specific for npt, (2) demonstration of NPTII activity after electrophoresis of extracts, and (3) chromatographic identification of the reaction product of NPTII, kanamycin phosphate. The highly biased codon usage in Chlamydomonas does not preclude expression.

Agrobacterium tumefaciens↗

Cloning, sequencing and analysis of the ggh-A gene encoding a 1,4-beta-D-glucan glucohydrolase from Microbispora bispora.

The ggh-A gene, encoding a 1,4-beta-D-glucan glucohydrolase/beta-glucosidase, of Microbispora bispora (Mb) was subcloned and expressed from a 4.0-kb XhoI DNA fragment. The nucleotide sequence of this fragment was determined. Analysis of the sequence revealed one open reading frame (ORF) which encodes a 986-amino-acid (aa) protein with a calculated molecular weight of 107,510. The ggh-A ORF has features typical of an actinomycete gene including high GC content (70.5%) and corresponding biased codon usage. Comparison of the aa sequence of the Mb 1,4-beta-D-glucan glucohydrolase (Mbggh-A) with other glycosidases reveals high overall homology to several beta-glucosidases and a 1,4-beta-D-glucan glucohydrolase belonging to the glycosyl hydrolase family 3. The aa sequence alignments of Mbggh-A and beta-glucosidases show that the active site region potentially involves two Asp residues. The aa sequence homology studies revealed a potential two-domain structure for Mbggh-A and other beta-glucosidases. Furthermore, Mbggh-A has localized homology to a cellulose-binding domain present in some xylanases. This report is significant, as, to date, 1,4-beta-D-glucan glucohydrolases have rarely been reported, though they are assumed to have a critical role in cellulolysis.

Actinomycetales↗

Nucleotide sequence of the invasion plasmid antigen B and C genes (ipaB and ipaC) of Shigella flexneri.

The nucleotide sequence of a 4.8 kilobase (kb) HindIII fragment from pWR100, the virulence plasmid of Shigella flexneri 5, was determined and analysed. This fragment encodes polypeptides b (62 kilodalton, kD) and c (43 kD) which have already been described as two of the four immunogenic polypeptides of Shigellae. The nucleotide sequence revealed that in addition to the ipaB and ipaC genes encoding polypeptides b and c, a third complete open reading frame was found within the fragment. The gene, named ippI, encoded a 17 kD polypeptide. The deduced amino acids sequence of polypeptides b and c showed no signal peptide but presence of highly hydrophobic domains compatible with a transmembraneous location. The surprising A and T richness of the three genes as compared with the Escherichia coli and Shigella genomes, resulted in a biased codon usage, and raises the question of the origin of the sequences.

Amino Acid Sequence↗