Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Amino acids runs and genomic compositional biases in vertebrates.

A compositional analysis of a sample of 50 zebrafish proteins containing at least one alanine run and of their open reading frames (ORFs) has been performed. The sample of poly(Ala) proteins showed a tendency to have runs of other amino acids (His/H, Gln/Q, Ser/S, Pro/P). Their ORFs and the first and second codon positions had higher GC contents than a reference gene set. The "universal" correlation between the GC content of the first+second and third codon positions (GC1+2 vs GC3) does not hold, but I provide an explanation in terms of genomic heterogeneity. Significant correlation between AHQS content and GC3 was obtained, reflecting codon bias favoring G/C at the third codon position of these amino acids. A correspondence analysis (COA) of relative synonymous codon usage showed that the poly(Ala) proteins have a biased distribution according to the second axis of the COA, which correlates with gene expression in zebrafish. A comparison with human is undertaken.

Amino Acids↗

Codon usage in the G+C-rich Streptomyces genome.

The codon usage (CU) patterns of 64 genes from the Gram+ prokaryotic genus Streptomyces were analysed. Despite the extremely high overall G+C content of the Streptomyces genome (estimated at 0.74), individual genes varied in G+C content from 0.610 to 0.797, and had third codon position G+C contents (GC3s) that varied from 0.764 to 0.983. The variation in GC3s explains a significant proportion of the variation in CU patterns. This is consistent with an evolutionary model of the Streptomyces genome where biased mutation pressure has led to a high average G+C content with random variation about the mean, although the variation observed is greater than that expected from a simple binomial model. The only gene in the sample that can be confidently predicted to be highly expressed, EF-Tu of Streptomyces coelicolor A3(2) (GC3s = 0.927), shows a preference for a third position C in several of the four codon families, and for CGY and GGY for Arg and Gly codons, respectively (Y = pyrimidine); similar CU patterns are found in highly expressed genes of the G+C-rich Micrococcus luteus genome. It thus appears that codon usage in Streptomyces is determined predominantly by mutation bias, with weak translational selection operating only in highly expressed genes. We discuss the possible consequences of the extreme codon bias of Streptomyces and consider how it may have evolved. A set of CU tables is provided for use with computer programs that locate protein-coding regions.

Base Composition↗

The Chlorella H+/hexose cotransporter gene.

The complete genomic sequence of the inducible Chlorella kessleri H+/hexose cotransporter (HUP1) has been obtained from two overlapping clones isolated from a lambda gt10 library. The HUP1 gene is interrupted by 14 introns with the first intron being located in the 5'-untranslated part of the gene. The average intron length is 220 bp, yielding a very regular intron/exon pattern in the gene. The codon usage in this gene is strongly biased with a clear preference for C and a strong suppression of A. A consensus sequence for a putative algal polyadenylation sequence is shown and compared with other algal cDNA sequences.

Base Sequence↗

Gene expression and protein length influence codon usage and rates of sequence evolution in Populus tremula.

Codon bias is generally thought to be determined by a balance between mutation, genetic drift, and natural selection on translational efficiency. However, natural selection on codon usage is considered to be a weak evolutionary force and selection on codon usage is expected to be strongest in species with large effective population sizes. In this paper, I study associations between codon usage, gene expression, and molecular evolution at synonymous and nonsynonymous sites in the long-lived, woody perennial plant Populus tremula (Salicaceae). Using expression data for 558 genes derived from expressed sequence tags (EST) libraries from 19 different tissues and developmental stages, I study how gene expression levels within single tissues as well as across tissues affect codon usage and rates sequence evolution at synonymous and nonsynonymous sites. I show that gene expression have direct effects on both codon usage and the level of selective constraint of proteins in P. tremula, although in different ways. Codon usage genes is primarily determined by how highly expressed a genes is, whereas rates of sequence evolution are primarily determined by how widely expressed genes are. In addition to the effects of gene expression, protein length appear to be an important factor influencing virtually all aspects of molecular evolution in P. tremula.

Codon↗

Codon bias evolution in Drosophila. Population genetics of mutation-selection drift.

Although non-random patterns of synonymous codon usage are a prominent feature in the genomes of many organisms, the relatives roles of mutational biases and natural selection in maintaining codon bias remain a contentious issue. In some species, patterns of codon bias and empirical findings on the biology of translation suggest 'major codon preference', a balance among mutation pressure, genetic drift, and weak selection in favor of translationally superior codons. Population genetics theory makes testable predictions to distinguish such a model from a strictly mutational model of codon bias. Major codon preference predicts two fitness classes of synonymous DNA changes: 'preferred' mutations from non-major to major codons and 'unpreferred' changes in the opposite direction. An extension of current statistical methods is employed to reveal differences in the within and between species dynamics of preferred and unpreferred silent mutations in Drosophila simulans. In this lineage, codon bias appears to be maintained under roughly equal magnitudes of natural selection and genetic drift. In the sibling species, D. melanogaster, however, a reduction in N(e)s, the product of effective population size and selection coefficient, appears to have allowed a genome-wide reduction in codon bias.

Animals↗

Evidence for genetic drift in endosymbionts (Buchnera): analyses of protein-coding genes.

Buchnera, the bacterial endosymbionts of aphids, undergo severe population bottlenecks during maternal transmission through their hosts. Previous studies suggest an increased effect of drift within these strictly asexual, small populations, resulting in an increased fixation of slightly deleterious mutations. This study further explores sequence evolution in Buchnera using three approaches. First, patterns of codon usage were compared across several homologous Escherichia coli and Buchnera loci, in order to test the prediction that selection for the use of optimal codons is less effective in small populations. A chi 2-based measure of codon bias was developed to adjust for the overall A + T richness of silent positions in the endosymbionts. In contrast to E. coli homologues, adaptive codon bias across Buchnera loci is markedly low, and patterns of codon usage lack a strong relationship with gene expression level. These data suggest that codon usage in Buchnera has been shaped largely by mutational pressure and drift rather than by selection for translational efficiency. One exception to the overall lack of bias is groEL, which is known to be constitutively overexpressed in Buchnera and other endosymbionts. Second, relative-rate tests show elevated rates of sequence evolution of numerous protein-coding loci across Buchnera, compared to E. coli. Finally, consistently higher ratios of nonsynonymous to synonymous substitutions in Buchnera loci relative to the enteric bacteria strongly suggest the accumulation of nonsynonymous substitutions in endosymbiont lineages. Combined, these results suggest a decreased effectiveness of purifying selection in purging endosymbiont populations of slightly deleterious mutations, particularly those affecting codon usage and amino acid identity.

Animals↗

The two beta-tubulin genes of Chlamydomonas reinhardtii code for identical proteins.

The two beta-tubulin genes of the unicellular green alga Chlamydomonas reinhardtii are expressed coordinately after deflagellation and produce two transcripts of 2.1 and 2.0 kilobases. Full-length cDNA clones corresponding to the transcript of each gene were isolated. DNA sequences were obtained from the cDNA clones and from cloned tubulin gene fragments. Both genes contained 1,332 base pairs of coding sequence, with only 19 nucleotide differences between the genes. Because all the differences occurred at the third base position of a codon and did not change the predicted amino acid sequence, we concluded that both beta-tubulin genes code for the same protein of 443 amino acids. The predicted amino acid sequence is 89 and 72% homologous with beta-tubulins from chicken and yeast cells, respectively. Each gene had three intervening sequences, which occurred at identical positions. Although the first two intervening sequences were not conserved between the two genes, the nucleotide sequence of the third intervening sequence was 89% conserved between the genes. The codon usage in the tubulin genes of C. reinhardtii was very biased: only 37 different codons were used. Striking differences occurred between the codons used in these nuclear genes and C. reinhardtii chloroplast genes.

Amino Acid Sequence↗

The large mitochondrial genome of Syndiclis anlungensis (Lauraceae): Genome structure, comparative analysis, and phylogenetic relationships among Syndiclis species.

The complete mitochondrial genome (mitogenome) of Syndiclis anlungensis, a critically endangered tropical tree, was determined in this study. The mitogenome spans 2,368,454&#xa0;bp across four contigs and harbors 41 protein-coding genes, 22 tRNA genes, and three rRNA genes. Potential mutation regions, including 1317 repeat sequences and 698 simple sequence repeats (SSRs), were accurately located in the S. anlungensis mitogenome. Sixty-five transferred fragments of the repeats were found between its mitochondrial and chloroplast genomes. When compared to three other Laurales mitogenomes, extensive gene order shuffling is evident, leaving only five conserved gene clusters intact. Codon usage analysis reveals a pronounced A/T bias in both mitochondrial and chloroplast genes, and three mitochondrial genes (atp9, rps19, and sdh3) stand out for their high divergence across eleven Syndiclis taxa. Selection analyses indicate strong purifying pressure on rpl2, rpl16, and sdh3 (Ka/Ks&#xa0;<&#xa0;1), with no positive selection detected. Using 41 mitochondrial protein-coding gene sequences from sixteen and three individuals of Syndiclis and Beilschmiedia species, respectively, our phylogenetic tree recovers Syndiclis as monophyletic, with two well-supported clades: one includes S. anlungensis, S. chinensis, S. lotungensis, S. marlipoensis, and a putative new Syndiclis species from Yunnan; the other contains S. furfuracea, S. hongkongensis, S. kwangsiensis, and three putative new Syndiclis species from Guangdong and Vietnam.

Genome, Mitochondrial↗

Duplication-induced mutation of a new Neurospora gene required for acetate utilization: properties of the mutant and predicted amino acid sequence of the protein product.

A cloned Neurospora crassa genomic sequence, selected as preferentially transcribed when acetate was the sole carbon source, was introduced in extra copies at ectopic loci by transformation. Sexual crossing of transformants yielded acetate nonutilizing mutants with methylation and restriction site changes within both the ectopic DNA and the normally located gene. Such changes are typical of the duplication-induced premeiotic disruption (the RIP effect) first described by Selker et al. (E. U. Selker, E. B. Cambareri, B. C. Jensen, and K. R. Haack, Cell 51:741-752, 1987). The mutants had the unusual phenotype of growth on ethanol but not on acetate as the carbon source. In a cross to the wild type of a mutant strain in which the original ectopic gene sequence had been removed by segregation, the acetate nonutilizing phenotype invariably segregated together with a RIP-induced EcoRI site at the normal locus. This mutant was transformed to the ability to use acetate by the cloned sequence. The locus of the mutation, designated acu-8, was mapped between trp-3 and un-15 on linkage group 2. The transcribed portion of the clone, identified by probing with cDNA, was sequenced, and a putative 525-codon open reading frame with two introns was identified. The codon usage was found to be strongly biased in a way typical of most Neurospora genes sequenced so far. The predicted amino acid sequence shows no significant resemblance to anything previously recorded. These results provide a first example of the use of the RIP effect to obtain a mutant phenotype for a gene previously known only as a transcribed wild-type DNA sequence.

Acetates↗

Genomic heterogeneity of background substitutional patterns in Drosophila melanogaster.

Mutation is the underlying force that provides the variation upon which evolutionary forces can act. It is important to understand how mutation rates vary within genomes and how the probabilities of fixation of new mutations vary as well. If substitutional processes across the genome are heterogeneous, then examining patterns of coding sequence evolution without taking these underlying variations into account may be misleading. Here we present the first rigorous test of substitution rate heterogeneity in the Drosophila melanogaster genome using almost 1500 nonfunctional fragments of the transposable element DNAREP1_DM. Not only do our analyses suggest that substitutional patterns in heterochromatic and euchromatic sequences are different, but also they provide support in favor of a recombination-associated substitutional bias toward G and C in this species. The magnitude of this bias is entirely sufficient to explain recombination-associated patterns of codon usage on the autosomes of the D. melanogaster genome. We also document a bias toward lower GC content in the pattern of small insertions and deletions (indels). In addition, the GC content of noncoding DNA in Drosophila is higher than would be predicted on the basis of the pattern of nucleotide substitutions and small indels. However, we argue that the fast turnover of noncoding sequences in Drosophila makes it difficult to assess the importance of the GC biases in nucleotide substitutions and small indels in shaping the base composition of noncoding sequences.

Animals↗

Selection at the amino acid level can influence synonymous codon usage: implications for the study of codon adaptation in plastid genes.

A previously employed method that uses the composition of noncoding DNA as the basis of a test for selection between synonymous codons in plastid genes is reevaluated. The test requires the assumption that in the absence of selective differences between synonymous codons the composition of silent sites in coding sequences will match the composition of noncoding sites. It is demonstrated here that this assumption is not necessarily true and, more generally, that using compositional properties to draw inferences about selection on silent changes in coding sequences is much more problematic than commonly assumed. This is so because selection on nonsynonymous changes can influence the composition of synonymous sites (i.e., codon usage) in a complex manner, meaning that the composition biases of different silent sites, including neutral noncoding DNA, are not comparable. These findings also draw into question the commonly utilized method of investigating how selection to increase translation accuracy influences codon usage. The work then focuses on implications for studies that assess codon adaptation, which is selection on codon usage to enhance translation rate, in plastid genes. A new test that does not require the use of noncoding DNA is proposed and applied. The results of this test suggest that far fewer plastid genes display codon adaptation than previously thought.

Algorithms↗

Evolutionary change of codon usage for the histone gene family in Drosophila melanogaster and Drosophila hydei.

The nucleotide divergence in the protein-coding region for replication-dependent and replication-independent histone 3 and 4 genes of Drosophila melanogaster and Drosophila hydei occurred mostly at the synonymous site. Therefore, the pattern of codon usage was analyzed in the two species, considering the genomic codon bias, which is proposed for estimating the genomic composition pressure in the protein-coding regions. The results indicated that the codon usage in the histone gene family could be explained mostly by the genomic codon bias. However, biases for Ala and Arg were commonly observed for the histone 3 and histone 4 gene families, and biases for Ser, Leu, and Glu were observed in a gene-specific manner. This suggests that both genomic codon bias and gene- or codon-specific bias are responsible for the nucleotide differentiation in the protein-coding region of the histone genes.

Animals↗

Identification of non-catalytic conserved regions in xylanases encoded by the xynB and xynD genes of the cellulolytic rumen anaerobe Ruminococcus flavefaciens.

xynB is one of at least four genes from the cellulolytic rumen anaerobe Ruminococcus flavefaciens 17 that encode xylanase activity. The xynB gene is predicted to encode a 781-amino acid product starting with a signal peptide, followed by an amino-terminal xylanase domain which is identical at 89% and 78% of residues, respectively, to the amino-terminal xylanase domains of the bifunctional XynD and XynA enzymes from the same organism. Two separate regions within the carboxy-terminal 537 amino acids of XynB also show close similarities with domain B of XynD. These regions show no significant homology with cellulose- or xylan-binding domains from other species, or with any other sequences, and their functions are unknown. In addition a 30 to 32-residue threonine-rich region is present in both XynD and XynB. Codon usage shows a consistent pattern of bias in the three xylanase genes from R. flavefaciens that have been sequenced.

Amino Acid Sequence↗

Unusually high-level expression of a foreign gene (hepatitis B virus core antigen) in Saccharomyces cerevisiae.

As a model system for the study of factors affecting gene expression, hepatitis B virus core antigen (HBcAg) has been expressed in the yeast Saccharomyces cerevisiae. The singularly high levels of expression achieved are approx. 40% of the soluble yeast protein. The HBcAg polypeptides are present as 28-nm particles which are morphologically indistinguishable from HBcAg particles in human plasma and are highly immunogenic in mice. The plasmid construction employed to achieve these very high levels of expression utilizes the constitutively active yeast promoter from the GAP491 gene which is fused in a way that all non-translated sequences flanking the HBcAg coding region are yeast-derived. Hybrid constructions containing 3'-nontranslated viral DNA (yeast 5') or 5'-nontranslated viral DNA (yeast 3') as well as a construction with both 5'- and 3'-nontranslated viral DNA also have been made. A comparison of these constructions for levels of HBcAg expression indicates that the strongest contributor to the high levels of protein is the presence of 5'-flanking sequences which are yeast-derived; secondarily, a significant improvement can be achieved if the 3'-flanking sequences also are yeast-derived. The high abundance of HBcAg in the highest producer is explicable in part on the basis of the very high stability in yeast cells of HBcAg polypeptides. Analysis of the HBcAg coding sequence reveals a very low index of codon bias for S. cerevisiae, largely discounting codon usage as a contributor to the high level of protein obtained.

Genes↗

Specific features of immunoglobulin VH genes of the Antarctic teleost Trematomus bernacchii.

The somatic recombination of different germline-encoded gene segments constitutes a principal source of antibody diversity. In order to investigate the diversity in recombined gene segments encoding the immunoglobulin heavy chain of the Antarctic teleost Trematomus bernacchii, a VH library was constructed by 5'-RACE (rapid amplification of cDNA ends) using RNA isolated from the spleen of an individual specimen. Analysis of cDNA sequences of 45 rearranged VH/D/JH segments revealed specific features, such as: high number of repeats, up to 8 bp long, and palindromic sequences, especially in CDRs (complementary determining regions); occurrence of the RGYW consensus, known as mutational hot spot, higher than in other species. Sixty-four percent of single base substitutions was found within this motif. In addition, the usage of serine codons showed a clear bias for AGY in CDRs, particularly in CDR2, and for TCN in FRs (framework regions). In CDRs, the frequency of non-synonymous changes was higher than that of synonymous changes. Diversity generated by insertions/deletions occurred more often than in other species; inserted bases were often repeats of adjacent bases. In particular the CDR2 showed the highest length variability as compared to other species. Alignment of VH sequences indicated that also the gene conversion mechanism may contribute to generating diversity. These data indicate a CDR mutability higher than in other species and provide some insights into the hypermutational events that may also occur in teleosts.

Animals↗

The nucleotide sequence of myosin light chain (L-2A) mRNA from embryonic chicken cardiac muscle tissue.

The nucleotide sequence of a cDNA clone (pML10) for chicken cardiac myosin light chain is described. The cDNA insert contains 613 nucleotides representing the entire coding sequence, with the exception of nine NH2-terminal amino acids, and the full 3'-non-coding region of 146 nucleotides. The missing 5' terminus of the mRNA, not represented in the clone pML10, was obtained by extension of the cDNA using a 43 nucleotide long internal EcoR1 fragment as a primer. The non-coding region contains several direct and inverted repeated sequences and the polyadenylation signal sequence AATAAA. The coding portion exhibits non-random usage of synonymous codons with a strong bias for codons ending in G and C.

Amino Acid Sequence↗

Structure of a Ruminococcus albus endo-1,4-beta-glucanase gene.

A chromosomal DNA fragment encoding an endo-1,4-beta-glucanase I (Eg I) gene from Ruminococcus albus cloned and expressed in Escherichia coli with pUC18 was fully sequenced by the dideoxy-chain termination method. The sequence contained a consensus promoter sequence and a structural amino acid sequence. The initial 43 amino acids of the protein were deduced to be a signal sequence, since they are missing in the mature protein (Eg I). High homology was found when the amino acid sequence of the Eg I was compared with that of endoglucanase E from Clostridium thermocellum. Codon usage of the gene was not biased. These results suggested that the properties of the Eg I gene from R. albus was specified from the known beta-glucanase genes of the other organisms.

Amino Acid Sequence↗

Comparative genomics of three strains of Ehrlichia ruminantium: a review.

The tick-borne Rickettsiale Ehrlichia ruminantium (E. ruminantium) is the causative agent of heartwater in Africa and the Caribbean. Heartwater, responsible for major losses on livestock in Africa represents also a threat for the American mainland. Three complete genomes corresponding to two different groups of differing phenotypes, Gardel and Welgevonden, have been recently described. One genome (Erga) represents the Gardel group from Guadeloupe Island and two genomes (Erwo and Erwe) belong to the Welgevonden group. Erwo, isolated in South Africa, is the parental strain of Erwe, which was maintained for 18 years in Guadeloupe under different culture conditions than Erwo. The three strains display genomes of differing sizes with 1,499,920 bp, 1,512,977 bp, and 1,516,355 bp for Erga, Erwe, and Erwo, respectively. Gene sequences and order are highly conserved between the three strains, although several gene truncations could be pinpointed, most of them occurring within three regions of accumulated differences (RAD). E. ruminantium displays a strong leading/lagging compositional bias inducing a strand-specific codon usage. Finally, a striking feature of E. ruminantium is the presence of long intergenic regions containing tandem repeats. These repeats are at the origin of an active process, specific to E. ruminantium, of genome expansion/contraction based on the addition or removal of tandem units.

Animals↗