Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Molecular population genetics and evolution of a prion-like protein in Saccharomyces cerevisiae.

The prion-like behavior of Sup35p, the eRF3 homolog in the yeast Saccharomyces cerevisiae, mediates the activity of the cytoplasmic nonsense suppressor known as [PSI(+)]. Sup35p is divided into three regions of distinct function. The N-terminal and middle (M) regions are required for the induction and propagation of [PSI(+)] but are not necessary for translation termination or cell viability. The C-terminal region encompasses the termination function. The existence of the N-terminal region in SUP35 homologs of other fungi has led some to suggest that this region has an adaptive function separate from translation termination. To examine this hypothesis, we sequenced portions of SUP35 in 21 strains of S. cerevisiae, including 13 clinical isolates. We analyzed nucleotide polymorphism within this species and compared it to sequence divergence from a sister species, S. paradoxus. The N domain of Sup35p is highly conserved in amino acid sequence and is highly biased in codon usage toward preferred codons. Amino acid changes are under weak purifying selection based on a quantitative analysis of polymorphism and divergence. We also conclude that the clinical strains of S. cerevisiae are not recently derived and that outcrossing between strains in S. cerevisiae may be relatively rare in nature.

Amino Acid Sequence↗

A unique pattern of intrastrand anomalies in base composition of the DNA in hypotrichs.

The 50 non-coding bases immediately internal to the telomeric repeats in the two 5' ends of macronuclear DNA molecules of a group of hypotrichous ciliates are anomalous in composition, consisting of 61% purines and 39% pyrimidines, A>T (ratio of 44:32), and G>C (ratio of 17:7). These ratio imbalances violate parity rule 2, according to which A should equal T and G should equal C within a DNA strand and therefore pyrimidines should equal purines. The purine-rich and base ratio imbalances are in marked contrast to the rest of the non-coding parts of the molecules, which have the theoretically expected purine content of 50%, with A = T and G = C. The ORFs contain an average of 52% purines as a result of bias in codon usage. The 50 bases that flank the 5' ends of macronuclear sequences in micronuclear DNA (12 cases) consist of approximately 50% purines. Thus, the 50 bases in the 5' ends of macronuclear sequences in micronuclear DNA are islands of purine richness in which A>T and G>C. These islands may serve as signals for the excision of macronuclear molecules during macronuclear development. We have found no published reports of coding or non-coding native DNA with such anomalous base composition.

Animals↗

Isolation and characterization of the genomic region from Drosophila kuntzei containing the Adh and Adhr genes.

The nucleotide sequences of the Adh and Adhr genes of Drosophila kuntzei were derived from combined overlapping sequences of clones isolated from a genomic library and from cloned PCR and inverse-PCR fragments. Only a proximal promoter was detected upstream of the Adh gene, indicating that D. kuntzei Adh is regulated by a one-promoter system. Further upstream of the Adh structural gene, an adult enhancer region (AAE) was found that contains most of the regulatory sequences described for AAEs of other Drosophila species. Analysis of the ADH protein showed an amino acid change from valine to threonine in the active site at position 189 which is also found in D. funebris but is otherwise unique among Drosophila. This difference alone may be responsible for the very low ADH activity found in this species and may cause a difference in substrate usage pattern. Codon bias in Adh and Adhr was comparable and found to be very low compared with other species. Phylogenetic analysis showed that D. kuntzei is closest related to D. funebris and D. immigrans. The time of divergence between D. kuntzei and D. funebris was estimated to be 14.2-20.2 Myr and that between D. kuntzei-D. funebris and D. immigrans to be 30.8-44.0 Myr. An analysis of the genetic variation in the Adh gene and upstream sequences of four European strains showed that this gene was highly variable. Overall nucleotide diversity (pi) was 0.0139, which is two times higher than that in D. melanogaster.

Alcohol Dehydrogenase↗

Lateral gene transfer and ancient paralogy of operons containing redundant copies of tryptophan-pathway genes in Xylella species and in heterocystous cyanobacteria.

BACKGROUND: Tryptophan-pathway genes that exist within an apparent operon-like organization were evaluated as examples of multi-genic genomic regions that contain phylogenetically incongruous genes and coexist with genes outside the operon that are congruous. A seven-gene cluster in Xylella fastidiosa includes genes encoding the two subunits of anthranilate synthase, an aryl-CoA synthetase, and trpR. A second gene block, present in the Anabaena/Nostoc lineage, but not in other cyanobacteria, contains a near-complete tryptophan operon nested within an apparent supraoperon containing other aromatic-pathway genes. RESULTS: The gene block in X. fastidiosa exhibits a sharply delineated low-GC content. This, as well as bias of codon usage and 3:1 dinucleotide analysis, strongly implicates lateral gene transfer (LGT). In contrast, parametric studies and protein tree phylogenies did not support the origination of the Anabaena/Nostoc gene block by LGT. CONCLUSIONS: Judging from the apparent minimal amelioration, the low-GC gene block in X. fastidiosa probably originated by LGT at a relatively recent time. The surprising inability to pinpoint a donor lineage still leaves room for alternative, albeit less likely, explanations other than LGT. On the other hand, the large Anabaena/Nostoc gene block does not seem to have arisen by LGT. We suggest that the contemporary Anabaena/Nostoc array of divergent paralogs represents an ancient ancestral state of paralog divergence, with extensive streamlining by gene loss occurring in the lineage of descent representing other (unicellular) cyanobacteria.

Amino Acid Sequence↗

The role of selection in the evolution of human mitochondrial genomes.

High mutation rate in mammalian mitochondrial DNA generates a highly divergent pool of alleles even within species that have dispersed and expanded in size recently. Phylogenetic analysis of 277 human mitochondrial genomes revealed a significant (P < 0.01) excess of rRNA and nonsynonymous base substitutions among hotspots of recurrent mutation. Most hotspots involved transitions from guanine to adenine that, with thymine-to-cytosine transitions, illustrate the asymmetric bias in codon usage at synonymous sites on the heavy-strand DNA. The mitochondrion-encoded tRNAThr varied significantly more than any other tRNA gene. Threonine and valine codons were involved in 259 of the 414 amino acid replacements observed. The ratio of nonsynonymous changes from and to threonine and valine differed significantly (P = 0.003) between populations with neutral (22/58) and populations with significantly negative Tajima's D values (70/76), independent of their geographic location. In contrast to a recent suggestion that the excess of nonsilent mutations is characteristic of Arctic populations, implying their role in cold adaptation, we demonstrate that the surplus of nonsynonymous mutations is a general feature of the young branches of the phylogenetic tree, affecting also those that are found only in Africa. We introduce a new calibration method of the mutation rate of synonymous transitions to estimate the coalescent times of mtDNA haplogroups.

Amino Acid Substitution↗

Nucleotide sequence of the coding portion of human alpha globin messenger RNA.

The nucleotide sequence of the coding portion of human alpha globin mRNA has been determined by sequence analysis using human alpha globin cDNA cloned in bacterial plasmids. The sequence was obtained by a combination of direct sequence analysis of the cloned cDNA and analysis of cDNA obtained by primer extension, using short restriction endonuclease fragments of cloned alpha cDNA that were hybridized to human globin mRNA and elongated on the mRNA template by viral reverse transcriptase. The human alpha globin mRNA has an unexpectedly high G + C base composition (64.7%), similar to that observed for rabbit globin alpha mRNA, and displays a striking bias in the use of synonym codons for various amino acids. The bias in codon usage of human alpha globin mRNA is similar, with some exceptions, to that previously observed for rabbit alpha globin mRNA as well as for human and rabbit beta globin mRNAs. A detailed restriction endonuclease map of the human alpha globin cDNA is presented.

Amino Acid Sequence↗

The atypical codon usage of the plant psbA gene may be the remnant of an ancestral bias.

The psbA gene of the chloroplast genome has a codon usage that is unusual for plant chloroplast genes. In the present study the evolutionary status of this codon usage is tested by reconstructing putative ancestral psbA sequences to determine the pattern of change in codon bias during angiosperm divergence. It is shown that the codon biases of the ancestral genes are much stronger than all extant flowering plant psbA genes. This is related to previous work that demonstrated a significant increase in synonymous substitution in psbA relative to other chloroplast genes. It is suggested, based on the two lines of evidence, that the codon bias of this gene currently is not being maintained by selection. Rather, the atypical codon bias simply may be a remnant of an ancestral codon bias that now is being degraded by the mutation bias of the chloroplast genome, in other words, that the psbA gene is not at equilibrium. A model for the evolution of selective pressure on the codon usage of plant chloroplast genes is discussed.

Base Sequence↗

Codon usage divergence of homologous vertebrate genes and codon usage clock.

This paper is concerned with the divergence of synonymous codon usage and its bias in three homologous genes within vertebrate species. Genetic distances among species are described in terms of synonymous codon usage divergence and the correlation is found between the genetic distances and taxonomic distances among species under study. A codon usage clock is reported in alpha-globin and beta-globin. A method is developed to define the synonymous codon preference bias and it is observed that the bias changes considerably among species.

Animals↗

The close proximity of Escherichia coli genes: consequences for stop codon and synonymous codon use.

It is shown that synonymous codon usage is less biased in favor of those codons preferred by highly expressed genes at the end of Escherichia coli genes than in the middle. This appears to be due to the close proximity of many E. coli genes. It is shown that a substantial number of genes overlap either the Shine-Dalgarno sequence or the coding sequence of the next gene on the chromosome and that the codons that overlap have lower synonymous codon bias than those which do not. It is also shown that there is an increase in the frequency of A-ending codons, and a decrease in the frequency of G-ending codons at the end of E. coli genes that lie close to another gene. It is suggested that these trends in composition could be associated with selection against the formation of mRNA secondary structure near the start of the next gene on the chromosome. Stop codon use is also affected by the close proximity of genes; many genes are forced to use TGA and TAG stop codons because they terminate either within the Shine-Dalgarno or coding sequence of the next gene on the chromosome. The implications these results have for the evolution of synonymous codon use are discussed.

Base Sequence↗

The genome of Campylobacter jejuni: codon and amino acid usage.

The genes from the genome of the AT-rich bacterium Campylobacter jejuni were analysed and characterised with respect to usage and amino acid usage. Codon usage is generally biased for all amino acids having synonymous codons, so that AT-rich synonyms are most frequently used. Markov chain analysis showed that codon bias and over- or underrepresentation of the corresponding tri-letter words are not related. Predicted secondary structure, lipophilicity, codon position within the gene, strand, and position on the (+)-strand were all shown to be determinants of codon usage, and these effects were in part directly explained by compositional phenomena. Codon context and the GC-content at the wobble position of the fourfold degenerate sites exert indirect effects on codon usage. The factors that affect codon usage seem to affect all amino acids, rather than selected amino acids. The usage of amino acids correlates well with the GC-content of genes, i.e. usage of amino acids encoded by GC-rich codons increases with GC-content and vice versa.

Amino Acids↗

First complete mitochondrial genome of Uzelothrips scabrosus (Thysanoptera: Uzelothripidae) provides insights into gene rearrangements and phylogenetic position within Terebrantia.

The family Uzelothripidae is represented by a single genus Uzelothrips and can be distinguished from others by the presence of whip-like antennae, a circular ventral sensorium on antennal segment III, a well-developed tentorium, and a membranous ovipositor. Here, we generated the first complete mitochondrial genome of Uzelothrips scabrosus (15,674&#xa0;bp) using next-generation sequencing to explore the gene rearrangements and phylogenetic relationships. It consists of 13 protein-coding genes, 22 transfer RNAs, two ribosomal RNAs, and two putative control regions. The genome exhibits strong AT bias (71.35%) with negative AT and GC skew. Codon usage analyses indicate a strong bias towards A/U-ending codons and influenced by both natural selection and mutation pressure. All PCGs were under purifying selection, with cox1 being the most conserved and nad4L the most variable. The gene order of the family Uzelothripidae is highly rearranged compared to the ancestral insect gene order. Comparative analysis revealed that gene block B was the most widely conserved, whereas the remaining gene blocks exhibited family or lineage-specific conservation patterns, reflecting extensive mitochondrial gene rearrangements during the evolution of the Thysanoptera. Moreover, 228 synapomorphic and 68 autapomorphic gene boundaries were identified across thysanopteran mitogenomes. Phylogenies indicated that the family Uzelothripidae is in a sister relationship with Stenurothripidae, and the Uzelothripidae&#xa0;+&#xa0;Stenurothripidae clade is sister to Thripidae. This study provides the first mitogenomic insights into Uzelothripidae and highlights the need for broader taxon sampling and nuclear genomic data to resolve deep evolutionary relationships within Thysanoptera.

Comparative analysis↗

Synonymous codon choices in the extremely GC-poor genome of Plasmodium falciparum: compositional constraints and translational selection.

We have analyzed the patterns of synonymous codon preferences of the nuclear genes of Plasmodium falciparum, a unicellular parasite characterized by an extremely GC-poor genome. When all genes are considered, codon usage is strongly biased toward A and T in third codon positions, as expected, but multivariate statistical analysis detects a major trend among genes. At one end genes display codon choices determined mainly by the extreme genome composition of this parasite, and very probably their expression level is low. At the other end a few genes exhibit an increased relative usage of a particular subset of codons, many of which are C-ending. Since the majority of these few genes is putatively highly expressed, we postulate that the increased C-ending codons are translationally optimal. In conclusion, while codon usage of the majority of P. falciparum genes is determined mainly by compositional constraints, a small number of genes exhibit translational selection.

Animals↗

Compositional properties of nuclear genes from Plasmodium falciparum.

We have analyzed the compositional distributions of coding sequences and their different codon positions, as well as the codon usage of the nuclear genes of Plasmodium falciparum, a parasite characterized by an extremely GC-poor genome. As expected, coding sequences are AT-rich, codon usage is strongly biased towards A or T in third codon positions, and some particular amino acids (aa) are especially abundant in the encoded proteins. Remarkably, however, no difference was detected between housekeeping (HK) and antigen (Ag) genes, in spite of differences in expression level and evolutionary constraints. Moreover, all the features found in P. falciparum are very similar to those found in a bacterium characterized by a very GC-poor genome, Staphylococcus aureus. These findings stress the importance of compositional constraints in determining codon usage and aa utilisation.

Amino Acids↗

Compositional pressure and translational selection determine codon usage in the extremely GC-poor unicellular eukaryote Entamoeba histolytica.

It is widely accepted that the compositional pressure is the only factor shaping codon usage in unicellular species displaying extremely biased genomic compositions. This seems to be the case in the prokaryotes Mycoplasma capricolum, Rickettsia prowasekii and Borrelia burgdorferi (GC-poor), and in Micrococcus luteus (GC-rich). However, in the GC-poor unicellular eukaryotes Dictyostelium discoideum and Plasmodium falciparum, there is evidence that selection, acting at the level of translation, influences codon choices. This is a twofold intriguing finding, since (1) the genomic GC levels of the above mentioned eukaryotes are lower than the GC% of any studied bacteria, and (2) bacteria usually have larger effective population sizes than eukaryotes, and hence natural selection is expected to overcome more efficiently the randomizing effects of genetic drift among prokaryotes than among eukaryotes. In order to gain a new insight about this problem, we analysed the patterns of codon preferences of the nuclear genes of Entamoeba histolytica, a unicellular eukaryote characterised by an extremely AT-rich genome (GC = 25%). The overall codon usage is strongly biased towards A and T in the third codon positions, and among the presumed highly expressed sequences, there is an increased relative usage of a subset of codons, many of which are C-ending. Since an increase in C in third codon positions is 'against' the compositional bias, we conclude that codon usage in E. histolytica, as happens in D. discoideum and P. falciparum, is the result of an equilibrium between compositional pressure and selection. These findings raise the question of why strongly compositionally biased eukaryotic cells may be more sensitive to the (presumed) slight differences among synonymous codons than compositionally biased bacteria.

Animals↗

Synonymous codon usage in bacteria.

In most bacteria, synonymous codons are not used with equal frequencies. Different factors have been proposed to contribute to codon usage preference, including translational selection, GC composition, strand-specific mutational bias, amino acid conservation, protein hydropathy, transcriptional selection and even RNA stability. The review discusses these factors and their contribution to bias in synonymous codon usage in bacterial genomes.

Bacteria↗

Coevolution of codon usage and transfer RNA abundance.

The use of synonymous codons is strongly biased in the bacterium Escherichia coli and yeast, comprising both bias between codons recognized by the same transfer RNA and bias between groups of codons recognized by different synonymous tRNAs. A major determinant of the second sort of bias is tRNA content, codons recognized by abundant tRNAs being used more often than those recognised by rare tRNAs, particularly in highly expressed genes, probably owing to selection at the level of translation against codons recognized by rare tRNAs. Conversely, codon usage is likely to exert selection pressure on tRNA abundance. Here I develop a model for the coevolution of codon usage and tRNA abundance which explains why there are unequal abundances of synonymous tRNAs leading to biased usage between groups of codons recognized by them in unicellular organisms.

Biological Evolution↗

A review of protein structure and gene organisation for proteins associated with mineralised tissue and calcium phosphate stabilisation encoded on human chromosome 4.

Several proteins associated with mineralised tissue (teeth and bone) or involved in calcium phosphate stabilisation in the body fluids, milk and saliva have been mapped to the q arm of human chromosome 4. These include the dentine/bone proteins dentine sialophosphoprotein (DSPP), dentine matrix protein 1 (DMP1), bone sialoprotein (BSP), matrix extracellular phosphoglycoprotein, osteopontin (OPN), enamelin, ameloblastin, milk caseins, salivary statherin, and proline-rich proteins. The proposed function of those that are multiphosphorylated is: (i) the stabilisation of calcium phosphate in solution (e.g. casein, statherin) preventing spontaneous precipitation and seeded-crystal growth or (ii) promoting biomineralisation (e.g. the phosphophoryn domain of DSPP), where the protein described as a template macromolecule, is proposed to act as a nucleator/promoter of crystal growth. The genes of these proteins have been subjected to conserved chromosomal synteny during mammalian evolution. The multiphosphorylated proteins statherin, caseins, phosphophoryn, BSP and OPN have been characterised as intrinsically disordered. The codon usage patterns for the amino acid serine reveal a bias for AGC and AGT codons within the human genes dspp, dmp1 and bsp, mouse dspp and dmp1 but not significantly for statherin or caseins. This pattern was also observed in the gene encoding hen phosvitin that also contains stretches of multiphosphorylated serines and in the dmp1 gene sequences of mammalian, reptilian and avian classes. In conclusion, these intrinsically disordered multiphosphorylated proteins are the translation products of genes displaying examples of codon usage bias, internal repeats and conserved chromosomal synteny within the mammalian class.

Animals↗

Sequence diversity and molecular evolution of the merozoite surface antigen 2 of Plasmodium falciparum.

Eleven new alleles of the Plasmodium falciparum merozoite surface antigen 2 (MSA2) from Papua New Guinea were analyzed by direct sequencing of polymerase chain reaction (PCR) products. We have used the sequence information to trace the molecular evolution of MSA2. The repeats of ten alleles belonging to the 3D7 allelic family differed considerably in size, nucleotide sequence, and repeat copy number. In the repeat region of these new alleles, codon usage was extremely biased with an exclusive use of NNT codons. Another new allele sequenced belonged to the FC27 family and confirmed the family-specific conserved structure of 96 and 36 bp repeats. In order to assess sequence microheterogeneity within samples defined as the same genotype by restriction fragment length polymorphism (RFLP), we have analyzed single-strand conformation polymorphism (SSCP) of different samples of the most frequent allele (D10 of the FC27 family) in the study population. No sequence heterogeneity could be detected within the repeat region. Based on analysis of the repeat regions in both allelic families, we discuss the hypothesis of a different evolutionary strategy being represented by each of the allelic families. Kew words: Merozoite surface antigen 2 - Nucleotide sequence comparisons - Molecular evolution

Amino Acid Sequence↗