Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Comparative genomics of three strains of Ehrlichia ruminantium: a review.

The tick-borne Rickettsiale Ehrlichia ruminantium (E. ruminantium) is the causative agent of heartwater in Africa and the Caribbean. Heartwater, responsible for major losses on livestock in Africa represents also a threat for the American mainland. Three complete genomes corresponding to two different groups of differing phenotypes, Gardel and Welgevonden, have been recently described. One genome (Erga) represents the Gardel group from Guadeloupe Island and two genomes (Erwo and Erwe) belong to the Welgevonden group. Erwo, isolated in South Africa, is the parental strain of Erwe, which was maintained for 18 years in Guadeloupe under different culture conditions than Erwo. The three strains display genomes of differing sizes with 1,499,920 bp, 1,512,977 bp, and 1,516,355 bp for Erga, Erwe, and Erwo, respectively. Gene sequences and order are highly conserved between the three strains, although several gene truncations could be pinpointed, most of them occurring within three regions of accumulated differences (RAD). E. ruminantium displays a strong leading/lagging compositional bias inducing a strand-specific codon usage. Finally, a striking feature of E. ruminantium is the presence of long intergenic regions containing tandem repeats. These repeats are at the origin of an active process, specific to E. ruminantium, of genome expansion/contraction based on the addition or removal of tandem units.

Animals↗

Molecular analysis and characterization of the Cochliobolus heterostrophus beta-tubulin gene and its possible role in conferring resistance to benomyl.

Cochliobolus heterostrophus Tub1 described here is the first beta-tubulin gene characterized from a naturally occurring benomyl-resistant ascomycete plant pathogen. The gene encodes a protein of 447 amino acids. The coding region of Tub1 is interrupted by three introns, of 116, 55, and 56 nt, situated after codons 4, 12, and 53, respectively. As a result of the preference for pyrimidines in the third position of the codons when a choice exists between purines and pyrimidines, codon usage in the Tub1 gene is biased. Tub1 shows high homology with beta-tubulin genes of other ascomycete species. However, Tub1 is exceptional in having Tyr(167), compared with Phe(167), possessed by beta-tubulin genes of other ascomycetes sequenced thus far. The Tyr(167) residue has been associated with benomyl resistance in other organisms. In contrast, all other benomyl-implicated residues of Tub1 correspond to sensitivity. Based on these results, we suggest that benomyl resistance in the fungus probably is attributed to Tyr(167).

Journal Article↗

Coding sequence divergence between two closely related plant species: Arabidopsis thaliana and Brassica rapa ssp. pekinensis.

To characterize the coding-sequence divergence of closely related genomes, we compared DNA sequence divergence between sequences from a Brassica rapa ssp. pekinensis EST library isolated from flower buds and genomic sequences from Arabidopsis thaliana. The specific objectives were (i) to determine the distribution of and relationship between K(a) and K(s), (ii) to identify genes with the lowest and highest K(a): K(s) values, and (iii) to evaluate how codon usage has diverged between two closely related species. We found that the distribution of K(a): K(s) was unimodal, and that substitution rates were more variable at nonsynonymous than synonymous sites, and detected no evidence that K(a) and K(s) were positively correlated. Several genes had K(a): K(s) values equal to or near zero, as expected for genes that have evolved under strong selective constraint. In contrast, there were no genes with K(a): K(s) >1 and thus we found no strong evidence that any of the 218 sequences we analyzed have evolved in response to positive selection. We detected a stronger codon bias but a lower frequency of GC at synonymous sites in A. thaliana than B. rapa. Moreover, there has been a shift in the profile of most commonly used synonymous codons since these two species diverged from one another. This shift in codon usage may have been caused by stronger selection acting on codon usage or by a shift in the direction of mutational bias in the B. rapa phylogenetic lineage.

Arabidopsis↗

Codon preference in corynebacteria.

The codon usage (CU) of 34 genes from the closely related species, Brevibacterium lactofermentum and Corynebacterium glutamicum (BLCG), was analysed and compared with that of 23 genes from other Brevibacterium and Corynebacterium species. The G+C content of the BLCG genes ranged from 50 to 62%. A wider range was found in other corynebacterial genes (25-71%). The G+C contents of non-coding regions in glutamic acid bacteria are lower than those of the coding regions and both values are lower than the G+C content of ribosomal RNA (rRNA) sequences, suggesting an unusual biased mutation pressure. The CU and synonymous codon usage (SCU) analysis showed several common characteristics among the sequenced corynebacterial genes, consistent with the close relatedness of B. lactofermentum and C. glutamicum. A subset of 25 preferred codons were deduced from the presumably highly expressed genes and they encode most of the amino acid (aa) residues of the BLCG group. An analysis of the effective number of codons (Nc) was carried out in order to check the GC3s (G+C content at the silent third position of sense codons) dependence of the CU in corynebacteria. Nc values showed differences between the BLCG group and other corynebacterial sequences. A comparison of the most used codons for each aa showed a stronger similarity to Streptomyces than to Escherichia coli. The CU/SCU tables of corynebacteria are useful for identification of protein-coding regions, including start codons when they are uncertain, and for designing oligodeoxyribonucleotide probes from an aa sequence.

Base Sequence↗

Expression of human lymphotoxin alpha in Aspergillus niger.

A gene-fusion expression strategy was applied for heterologous expression of human lymphotoxin alpha (LTalpha) in the Aspergillus niger AB1.13 protease-deficient strain. The LTalpha gene was fused with the A. niger glucoamylase GII-form as a carrier-gene, behind its transcription control and secretion signals. Special attention was paid to the influence of different codon usage on secretion of protein. In the case of human tumor necrosis factor alpha (TNFalpha) a dramatic change of secretion has been observed when human cDNA sequence was used instead of synthetic E. coli biased codons. In the case of LTalpha such a change of codon usage brought improvement at the RNA level, however, no increase in the quantity of secreted protein was observed, due to the proteolitic activity of the host organism. The estimated yield of secretion of LTalpha from A. niger into the soya medium was 50 pg l(-1) of culture.

Artificial Gene Fusion↗

Nucleotide sequence of the tcml gene (ribosomal protein L3) of Saccharomyces cerevisiae.

The yeast tcml gene, which codes for ribosomal protein L3, has been isolated by using recombinant DNA and genetic complementation. The DNA fragment carrying this gene has been subcloned and we have determined its DNA sequence. The 20 amino acid residues at the amino terminus as inferred from the nucleotide sequence agreed exactly with the amino acid sequence data. The amino acid composition of the encoded protein agreed with that determined for purified ribosomal protein L3. Codon usage in the tcml gene was strongly biased in the direction found for several other abundant Saccharomyces cerevisiae proteins. The tcml gene has no introns, which appears to be atypical of ribosomal protein structural genes.

Amino Acid Sequence↗

Phylogenetic relationships of the liverworts (Hepaticae), a basal embryophyte lineage, inferred from nucleotide sequence data of the chloroplast gene rbcL.

Sequence data from the chloroplast-encoded gene rbcL were obtained for 24 liverworts, a basal group of embryophytes. Maximum likelihood and parsimony analyses of these data, along with data from other major green plant lineages, confirm hypotheses based on morphological data, such as the paraphyly of bryophytes, and the basal position of liverworts. Molecular data corroborate the deep separation between the complex thalloid and leafy/simple thalloid liverworts implied by morphological data, but the monophyly of liverworts could not be rejected. The effects of accounting for site-to-site rate heterogeneity in these data were examined using maximum likelihood methods. Comparison of trees obtained with and without rate heterogeneity showed that simply allowing for heterogeneity had a greater improvement on likelihood score than optimization of transition/transversion bias. Incorporation of site-to-site rate heterogeneity in the larger analysis, however, did not necessarily change which topology was favored. Properties of rbcL sequences from the two liverwort groups were compared. Significantly different substitution rates were found between leafy/simple thalloid and complex thalloid liverwort taxa, with rates of rbcL sequence evolution in leafy/simple thalloid taxa being higher and more indicative of those of vascular plants, and with those of complex thalloid taxa (such as Marchantia) being slower. Codon usage in rbcL in complex thalloid liverworts was biased toward NNU and NNA, compared to the leafy/simple thalloid liverworts. Although base composition and relative substitution rates differed between the two groups, no significant differences were detected within each of the two groups of liverworts. The signal present in first and second codon sites versus third codon sites was compared. While the third codon positions in rbcL across this taxon sampling are highly variable (with only 15 constant sites of 439), the trees obtained were in general agreement with trees from the entire data set and with trees obtained from independent sources of data. The presence of signal in third codon positions across greater than 400 MY of plant evolution means that definitions of saturation based on pair-wise comparisons of sequences inadequately assess phylogenetic signal.

Chloroplasts↗

Protein encoding genes in an ancient plant: analysis of codon usage, retained genes and splice sites in a moss, Physcomitrella patens.

BACKGROUND: The moss Physcomitrella patens is an emerging plant model system due to its high rate of homologous recombination, haploidy, simple body plan, physiological properties as well as phylogenetic position. Available EST data was clustered and assembled, and provided the basis for a genome-wide analysis of protein encoding genes. RESULTS: We have clustered and assembled Physcomitrella patens EST and CDS data in order to represent the transcriptome of this non-seed plant. Clustering of the publicly available data and subsequent prediction resulted in a total of 19,081 non-redundant ORF. Of these putative transcripts, approximately 30% have a homolog in both rice and Arabidopsis transcriptome. More than 130 transcripts are not present in seed plants but can be found in other kingdoms. These potential "retained genes" might have been lost during seed plant evolution. Functional annotation of these genes reveals unequal distribution among taxonomic groups and intriguing putative functions such as cytotoxicity and nucleic acid repair. Whereas introns in the moss are larger on average than in the seed plant Arabidopsis thaliana, position and amount of introns are approximately the same. Contrary to Arabidopsis, where CDS contain on average 44% G/C, in Physcomitrella the average G/C content is 50%. Interestingly, moss orthologs of Arabidopsis genes show a significant drift of codon fraction usage, towards the seed plant. While averaged codon bias is the same in Physcomitrella and Arabidopsis, the distribution pattern is different, with 15% of moss genes being unbiased. Species-specific, sensitive and selective splice site prediction for Physcomitrella has been developed using a dataset of 368 donor and acceptor sites, utilizing a support vector machine. The prediction accuracy is better than those achieved with tools trained on Arabidopsis data. CONCLUSION: Analysis of the moss transcriptome displays differences in gene structure, codon and splice site usage in comparison with the seed plant Arabidopsis. Putative retained genes exhibit possible functions that might explain the peculiar physiological properties of mosses. Both the transcriptome representation (including a BLAST and retrieval service) and splice site prediction have been made available on http://www.cosmoss.org, setting the basis for assembly and annotation of the Physcomitrella genome, of which draft shotgun sequences will become available in 2005.

Alternative Splicing↗

Ribosome traffic in E. coli and regulation of gene expression.

The ribosome traffic during translation of E. coli coding sequences was simulated, assuming that the rate of translation of individual codons is limited by the cognate tRNA availability. Actual translation rates were taken from Solomovici et al. (J. theor. Biol. 185, 511-521, 1997). The mean translation rates of the 4271 sequences cover a broad, two-fold range, whereas the local rate of translation along messengers varies three-fold on average. The simulation allows one to sketch the ribosome traffic on the polysome, in particular by providing the extent of mRNA sequences uncovered between consecutive ribosomes and the time during which these sequences are exposed. These parameters may participate in the control of mRNA stability and transcriptional polarity. By averaging the translation rates in a 17-codon window, assumed to be the sequence covered by a translating ribosome, and sliding this window along a given coding sequence, the addresses KMAX and KMIN, and the times TMAX and TMIN of respectively the slowest and the fastest translated window were determined. It is shown that under the assumptions made, TMAX sets the number of proteins translated from a given mRNA molecule per unit time, in case the delay between consecutive translation starts is below TMAX. Both windows display two strong biases, one as expected on the usage of codon frequencies, and the other surprisingly on the occurrence of amino acids.

Amino Acids↗

Isolation of cDNAs encoding 6-phosphogluconate dehydrogenase and glucose-6-phosphate dehydrogenase from the mediterranean fruit fly Ceratitis capitata: correlating genetic and physical maps of chromosome 5.

We have isolated and determined the nucleotide sequences for cDNA clones encoding glucose-6-phosphate dehydrogenase (G6PD) and 6-phosphogluconate dehydrogenase (6PGD) from the medfly Ceratitis capitata. The derived amino acid sequences for G6PD and 6PGD are presented and compared with G6PDs and 6PGDs from other species. The codon usage of the cDNA clones has little bias with the notable exceptions of arginine, glycine and leucine. The chromosomal location of the genes for 6PGD and G6PD were determined by in situ hybridization to salivary gland polytene chromosomes. This localization orients a genetic map of enzymatic loci and illustrates a remarkable similarity in the intra chromosomal order of homologous genes between Drosophila melanogaster and medfly.

Amino Acid Sequence↗

Codon and amino acid usage in retroviral genomes is consistent with virus-specific nucleotide pressure.

Retroviral RNA genomes are known to have a biased nucleotide composition. For instance, the plus-strand RNA of human immunodeficiency virus (HIV) is A-rich, and the genome of human T cell leukemia virus (HTLV) is C-rich, and other retroviruses have a U-rich or G-rich genome. The biased composition of these genomes is most likely caused by directional mutational pressure of the respective reverse transcriptase enzymes. Using a set of retroviral genomes with a distinct nucleotide composition, we performed skew analyses of the nucleotide bias along the complete viral genome. Distinct nucleotide signatures were apparent, and these typical patterns were generally conserved across the viral genome. Furthermore, it is demonstrated that this typical nucleotide bias, combined with a profound discrimination against the CpG dinucleotide sequence, strongly influences the codon usage of the retroviruses in a direct manner, and their amino acid usage in an indirect manner. The fact that both codon usage and amino acid usage are so closely entwined with the genome composition has important practical implications. For instance, the typical trends in nucleotide usage could influence the molecular phylogenetic reconstruction of the family Retroviridae.

Amino Acids↗

The strength of translational selection for codon usage varies in the three replicons of Sinorhizobium meliloti.

The genome of the nitrogen-fixing bacterium Sinorhizobium meliloti is composed of three replicons of 3.65 (chromosome), 1.35 (pSymA) and 1.68 Mb (pSymB), respectively. While the chromosome encodes for most of the housekeeping functions, the three elements may contribute to symbiosis, though pSymA is absolutely necessary for nodulation and nitrogen fixation, since it harbours all the characterized nodulation and symbiotic fixation genes. On the other hand, the majority of the sequences located in this megaplasmid are probably not expressed during the free-living stage of the organism. Since most of the sequences located in pSymA are transcribed only at the stage of bacteroids when most probably the fate of the bacterium is to die, the mutations occurring at this stage will not be fixed in the population. Therefore, if natural selection contributes to the codon usage pattern in this species, its effect will be much weaker for the genes placed in pSymA. A codon usage analysis of the genes comprising the three replicons is consistent with the conclusion that selection for translational speed shapes the codon usage of the two replicons which are important for competitive cell growth while the codon usage of the third replicon reflects primarily the mutational bias.

Base Sequence↗

Nucleic acid composition, codon usage, and the rate of synonymous substitution in protein-coding genes.

Based on the rates of synonymous substitution in 42 protein-coding gene pairs from rat and human, a correlation is shown to exist between the frequency of the nucleotides in all positions of the codon and the synonymous substitution rate. The correlation coefficients were positive for A and T and negative for C and G. This means that AT-rich genes accumulate more synonymous substitutions than GC-rich genes. Biased patterns of mutation could not account for this phenomenon. Thus, the variation in synonymous substitution rates and the resulting unequal codon usage must be the consequence of selection against A and T in synonymous positions. Most of the variation in rates of synonymous substitution can be explained by the nucleotide composition in synonymous positions. Codon-anticodon interactions, dinucleotide frequencies, and contextual factors influence neither the rates of synonymous substitution nor codon usage. Interestingly, the nucleotide in the second position of codons (always a nonsynonymous position) was found to affect the rate of synonymous substitution. This finding links the rate of nonsynonymous substitution with the synonymous rate. Consequently, highly conservative proteins are expected to be encoded by genes that evolve slowly in terms of synonymous substitutions, and are consequently highly biased in their codon usage.

Animals↗

The effect of queuosine on tRNA structure and function.

Computational modeling was performed to determine the potential function of the queuosine modification of tRNA found in wobble position 34 of tRNAasp, tRNAasn, tRNAhis, and tRNAtyr. Using the crystal structure of tRNAasp and a tRNA-tRNA-mRNA complex model, we show that the queuosine modification serves as a structurally restrictive base for tRNA anticodon loop flexibility. An extended intraresidue and intramolecular hydrogen bonding network is established by queuosine. The quaternary amine of the 7-aminomethyl side chain hydrogen bonds with the base's carbonyl oxygen. This positions the dihydroxycyclopentenediol ring of queuosine in proper orientation for hydrogen bonding with the backbone of the neighboring uridine 33 residue. The interresidue association stabilizes the formation of a cross-loop hydrogen bond between the uridine 33 base and the phosphoribosyl backbone of the cytosine at position 36. Additional interactions between RNAs in the translation complex were studied with regard to potential codon context and codon bias effects. Neither steric nor electrostatic interaction occurs between aminoacyl- and peptidyl-site tRNA anticodon loops that are modified with queuosine. However, there is a difference in the strength of anticodon/codon associations (codon bias) based on the presence or lack of queuosine in the wobble position of the tRNA. Unmodified (guanosine-containing) tRNAasp forms a very stable association with cytosine (GAC), but is much less stable in complex with a uridine-containing codon (GAU). Queuosine-modified tRNAasp exhibits no bias for either of cognate codons GAC or GAU and demonstrates a lower binding energy similar to the wobble pairing of guanosine-containing tRNA with a GAU codon. This is proposed to be due to the inflexibility of the queuosine-modified anticodon loop to accommodate proper positioning for optimal Watson-Crick type associations. A preliminary survey of codon usage patterns in oncodevelopmental versus housekeeping gene transcripts suggests a significant difference in bias for the queuosine-associated codons. Therefore, the queuosine modification may have the potential to influence cellular growth and differentiation by codon bias-based regulation of protein synthesis for discrete mRNA transcripts.

Anticodon↗

Cloning and characterization of the gene for beta-tubulin from a benomyl-resistant mutant of Neurospora crassa and its use as a dominant selectable marker.

We cloned the beta-tubulin gene of Neurospora crassa from a benomyl-resistant strain and determined its nucleotide sequence. The gene encodes a 447-residue protein which shows strong homology to other beta-tubulins. The coding region is interrupted by six introns, five of which are within the region coding for the first 54 amino acids of the protein. Intron position comparisons between the N. crassa gene and other fungal beta-tubulin genes reveal considerable positional conservation. The mutation responsible for benomyl resistance was determined; it caused a phenylalanine-to-tyrosine change at position 167. Codon usage in the beta-tubulin gene is biased, as has been observed for other abundantly expressed N. crassa genes such as am and the H3 and H4 histone genes. This bias results in pyrimidines in the third positions of 96% of the codons in codon families in which there is a choice between purines and pyrimidines in this position. Bias is also evident by the absence of 19 of the 61 sense codons. We demonstrated that benomyl resistance is due to the cloned beta-tubulin gene of strain Bml511(r)a and that this gene can be used as a dominant selectable marker in N. crassa transformation.

Amino Acid Sequence↗

Divergent evolutionary constraints on mitochondrial and nuclear genomes of malaria parasites.

Genetic variation among malaria parasites has important consequences with regard to drug resistance, pathogenicity, immunity, transmission, and speciation. In this regard, malaria parasites have been shown to display a high degree of inter- and intra-species genetic divergence. The nuclear genomes of Plasmodium falciparum, Plasmodium yoelii, and Plasmodium gallinaceum are vastly divergent yet share a similar codon usage and total A/T content of approximately 82%. This is in contrast to other primate-specific species including P. vivax which have an A/T content of approximately 67%. To assess the effects of this evolutionary divergence on the conservation of gene content, organization, and codon usage in the mitochondrial DNA (mtDNA) of malaria parasites, we have cloned and sequenced the mitochondrial genome of Plasmodium vivax, and compared it with the mtDNAs of P. falciparum, P. yoelii, and P. gallinaceum. The P. vivax mitochondrial genome was found to be 5990 base pairs in length, and displayed a gene organization identical to that of P. falciparum, P. yoelii, and P. gallinaceum. Furthermore, there was a remarkable 90% conservation of sequence identity between the mitochondrial genomes of all four species. As an example of intra-species conservation, comparison of mtDNAs from two independently cloned P. falciparum isolates, Malay Camp and C10, revealed only a single nucleotide substitution. A/T content of the P. vivax mitochondrial genome was found to be identical to other species of Plasmodium, hence, we have postulated that the mitochondrial genomes of malaria parasites were refractory to the evolutionary shifts in nucleotide content seen among the nuclear genomes of malaria parasites. Among different Plasmodium species, the second position of mitochondrial codons were found to be the least prone to substitutions and displayed a significant bias in pyrimidines. These aspects of mitochondrial codon usage were distinct from the nuclear genome and may reflect functional aspects of decoding by the mitochondrial translational system.

Amino Acid Sequence↗

Codon usage in Cryptosporidium parvum differs from that in other Eimeriorina.

Codon usage of Crytosporidium parvum was compared with those of other Eimeriorina Toxoplasma gondii and Eimeria tenella and revealed a biased use of synonymous codons with a preference for NNU (40.0%) and NNA (33.4%). There was no close resemblance of the codon usage of C. parvum to T. gondii (correlation coefficient, r = 0.14) or E. tenella (r = 0.14) but it was similar to Entamoeba histolytica (r = 0.75) and Plasmodium falciparum (r = 0.5). Analysis of the codon usage in homologous gene sequences (actin, beta-tubulin) also failed to reveal a close relationship between C. parvum and T. gondii or E. tenella. The low usage codons in C. parvum were most frequently used codons in T. gondii and E. tenella. These observations are consistent with 18S rRNA sequence analysis which shows no close relationship of Cryptosporidium with other Eimeriorina (Sarcocystis, Toxoplasma and Eimeria) and questions the validity of the current classification of C. parvum.

Actins↗

The biosynthesis of the ubiquinol-cytochrome c reductase complex in yeast. DNA sequence analysis of the nuclear gene coding for the 14-kDa subunit.

The nuclear gene coding for the imported 14-kDa subunit of the ubiquinol-cytochrome c reductase of yeast mitochondria has been sequenced in an attempt to define regulatory and protein topogenic elements. The gene has a length of 381 base pairs and is potentially capable of encoding a polypeptide of 14561 Da. It is transcribed into a single low-abundance RNA of 680 nucleotides whose 5' and 3' termini map, respectively, 30-35 nucleotides upstream and 180-190 nucleotides downstream of the initiator and termination codons. Consistent with the estimated low level of the mRNA, codon usage in the gene is not strongly biased and other features, characteristic of highly expressed genes in yeast, are absent. The 14-kDa protein is predicted to be a predominantly hydrophilic protein, with only a single, short hydrophobic stretch located between positions 19-38. Comparison with other imported mitochondrial proteins so far sequenced has failed to reveal unifying features that might serve as targeting elements. Steady-state levels of the 14-kDa and 11-kDa subunits are reduced in mit- mutants which synthesize truncated forms of apocytochrome b and in these, newly synthesized subunits exhibit a specifically increased turnover rate. We suggest that association of these two subunits with the complex may be mediated or enhanced by interaction with other subunits, in particular cytochrome b.

Amino Acid Sequence↗