Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Modeling sequencing errors by combining Hidden Markov models.

Among the largest resources for biological sequence data is the large amount of expressed sequence tags (ESTs) available in public and proprietary databases. ESTs provide information on transcripts but for technical reasons they often contain sequencing errors. Therefore, when analyzing EST sequences computationally, such errors must be taken into account. Earlier attempts to model error prone coding regions have shown good performance in detecting and predicting these while correcting sequencing errors using codon usage frequencies. In the research presented here, we improve the detection of translation start and stop sites by integrating a more complex mRNA model with codon usage bias based error correction into one hidden Markov model (HMM), thus generalizing this error correction approach to more complex HMMs. We show that our method maintains the performance in detecting coding sequences.

Algorithms↗

Expressed sequence tags from immature female sexual organ of a liverwort, Marchantia polymorpha.

A total of 970 expressed sequence tag (EST) clones were generated from immature female sexual organ of a liverwort, Marchantia polymorpha. The 376 ESTs resulted in 123 redundant groups, thus the total number of unique sequences in the EST set was 717. Database search by BLAST algorithm showed that 302 of the unique sequences shared significant similarities to known nucleotide or amino acid sequences. Six unique sequences showed significant similarities to genes that are involved in flower development and sexual reproduction, such as cynarase, fimbriata-associated protein and S-receptor kinase genes. The remaining unique 415 sequences have no significant similarity with any database-registered genes or proteins. The redundant 123 ESTs implied the presence of gene families and abundant transcripts of unknown identity. Analyses of the coding sequences of 61 unique sequences, which contained no ambiguous bases in the predicted coding regions, highly homologous to known sequences at the amino acid level with a similarity score greater than 400, and with stop codons at similar positions as their possible orthologues, indicated the presence of biased codon usage and higher GC content within the coding sequences (50.4%) than that within 3' flanking sequences (41.9%).

Amino Acid Sequence↗

A eubacterial gene conferring spectinomycin resistance on Chlamydomonas reinhardtii: integration into the nuclear genome and gene expression.

We have constructed a dominant selectable marker for nuclear transformation of C. reinhardtii, composed of the coding sequence of the eubacterial aadA gene (conferring spectinomycin resistance) fused to the 5' and 3' untranslated regions of the endogenous RbcS2 gene. Spectinomycin-resistant transformants isolated by direct selection (1) contain the chimeric gene(s) stably integrated into the nuclear genome, (2) show cosegregation of the resistance phenotype with the introduced DNA, and (3) synthesize the expected mRNA and protein. Small linearized plasmids appeared to be inserted into the nuclear genome preferentially through their ends, with relatively few large deletions and/or rearrangements. Multiple copy transformants often integrated concatemers of transforming DNA. Our detailed analysis of the complex integration patterns of plasmid DNA in C. reinhardtii nuclear transformants should be useful for improving the technique of insertional mutagenesis. We also found that the spectinomycin-resistance phenotype was unstable in about half of the transformants. When maintained under nonselective conditions, neither the aadA mRNA nor the AadA protein were detected in these subclones. Moreover, since the integrated transforming DNA was not altered or lost expression of the RbcS2::aadA::RbcS2 gene(s) appears to be repressed. Measurements of transcriptional activity, mRNA accumulation, and mRNA stability suggest that expression of this chimeric gene(s) may also be affected by rapid RNA degradation, presumably due to defects in mRNA processing and, or nuclear export. Thus, both gene silencing and transcript instability, rather than biased codon usage, may explain the difficulties encountered in the expression of foreign genes in the nuclear genome of Chlamydomonas.

Animals↗

Compositional biases and polyalanine runs in humans.

Human proteins containing polyalanine tracts tend to have runs of other amino acids and their open reading frames (ORFs) display a biased codon usage. Their alanine, glycine, proline, and histidine content strongly correlates with the GC content of the third codon base, suggesting that the compositional specificity of these proteins is dictated to a great extent by the evolution of their ORFs.

Codon↗

Nonneutral evolution of the transcribed pseudogene Makorin1-p1 in mice.

Pseudogenes are nonfunctional relics of formerly functional genes and are thought to evolve neutrally. In some pseudogenes, however, the molecular evolutionary patterns are atypical of neutrally evolving sequences, exhibiting sequence conservation, codon-usage bias, and other features associated with functional genes. Makorin1-p1 is a transcribed pseudogene first identified in the mouse Mus musculus. The transcript of Makorin1-p1 can regulate the stability of the transcript of its paralogous functional gene Makorin1. Specifically, the half-life of Makorin1 mRNA increases significantly in the presence of Makorin1-p1 transcript, and targeted deletion of Makorin1-p1 is lethal in mice. Here, we show that Makorin1-p1 originated after the separation of Mus and Rattus but before the divergence of M. musculus and M. pahari. The transcribed region of Makorin1-p1 exhibits rates of point and indel substitutions that are two to four times lower than those in the untranscribed region, suggesting that the transcribed region is under functional constraint and is not evolving neutrally. Although the transcript of Makorin1-p1 likely functions by its sequence similarity to Makorin1, we find no evidence of gene conversion between them, indicating that functional conservation alone is sufficient to maintain their coordinated evolution. A duplication-degeneration model is proposed to explain how Makorin1-p1 was co-opted into the regulatory system of Makorin1. There are over 10,000 pseudogenes in a typical mammalian genome, and it is plausible that many functional but untranslatable pseudogenes exist. Our results illustrate the potential of using evolutionary analysis to identify such pseudogenes from genome sequences.

Animals↗

Evolution of the AID/APOBEC family of polynucleotide (deoxy)cytidine deaminases.

The AID/APOBEC family (comprising AID, APOBEC1, APOBEC2, and APOBEC3 subgroups) contains members that can deaminate cytidine in RNA and/or DNA and exhibit diverse physiological functions (AID and APOBEC3 deaminating DNA to trigger pathways in adaptive and innate immunity; APOBEC1 mediating apolipoprotein B RNA editing). The founder member APOBEC1, which has been used as a paradigm, is an RNA-editing enzyme with proposed antecedents in yeast. Here, we have undertaken phylogenetic analysis to glean insight into the primary physiological function of the AID/APOBEC family. We find that although the family forms part of a larger superfamily of deaminases distributed throughout the biological world, the AID/APOBEC family itself is restricted to vertebrates with homologs of AID (a DNA deaminase that triggers antibody gene diversification) and of APOBEC2 (unknown function) identifiable in sequence databases from bony fish, birds, amphibians, and mammals. The cloning of an AID homolog from dogfish reveals that AID extends at least as far back as cartilaginous fish. Like mammalian AID, the pufferfish AID homolog can trigger deoxycytidine deamination in DNA but, consistent with its cold-blooded origin, is thermolabile. The fine specificity of its mutator activity and the biased codon usage in pufferfish IgV genes appear broadly similar to that of their mammalian counterparts, consistent with a coevolution of the antibody mutator and its substrate for the optimal targeting of somatic mutation during antibody maturation. By contrast, APOBEC1 and APOBEC3 are later evolutionary arrivals with orthologs not found in pufferfish (although synteny with mammals is maintained in respect of the flanking loci). We conclude that AID and APOBEC2 are likely to be the ancestral members of the AID/APOBEC family (going back to the beginning of vertebrate speciation) with both APOBEC1 and APOBEC3 being mammal-specific derivatives of AID and a complex set of domain shuffling underpinning the expansion and evolution of the primate APOBEC3s.

APOBEC-1 Deaminase↗

Alternatively and constitutively spliced exons are subject to different evolutionary forces.

There has been a controversy on whether alternatively spliced exons (ASEs) evolve faster than constitutively spliced exons (CSEs). Although it has been noted that ASEs are subject to weaker selective constraints than CSEs, so they evolve faster, there have also been studies that indicated slower evolution in ASEs than in CSEs. In this study, we retrieve more than 5,000 human-mouse orthologous exons and calculate the synonymous (KS) and nonsynonymous (KA) substitution rates in these exons. Our results show that ASEs have higher KA values and higher KA/KS ratios than CSEs, indicating faster amino acid-level evolution in ASEs. The faster evolution may be in part due to weaker selective constraints. It is also possible that the faster rate is in part due to faster functional evolution in ASEs. On the other hand, the majority of ASEs have lower KS values than CSEs. With reference to the substitution rate in introns, we show that the KS values in ASEs are close to the neutral substitution rate, whereas the synonymous substitution rate in CSEs has likely been accelerated. The elevated synonymous rate in CSEs is not related to CpG dinucleotides or low-complexity regions of protein but may be weakly related to codon usage bias. The overall trends of higher KA and lower KS in ASEs than in CSEs are also observed in human-rat and mouse-rat comparisons. Therefore, our observations hold for mammals of different molecular clocks.

Animals↗

The rate of adaptive evolution in enteric bacteria.

Here we estimate the rate of adaptive substitution in a set of 410 genes that are present in 6 Escherichia coli and 6 Salmonella enterica genomes. We estimate that more than 50% of amino acid substitutions in this set of genes have been fixed by positive selection between the E. coli and S. enterica lineages. We also show that the proportion of adaptive substitutions is uncorrelated with the rate of amino acid substitution or gene function but that it may be correlated with levels of synonymous codon usage bias.

Adaptation, Biological↗

Structural comparison of yeast ribosomal protein genes.

The primary structure of the genes encoding the yeast ribosomal proteins L17a and L25 was determined, as well as the positions of the 5'- and 3'-termini of the corresponding mRNAs. Comparison of the gene sequences to those obtained for various other yeast ribosomal protein genes revealed several similarities. In all split genes the intron is located near the 5'-side of the amino acid coding region. Among the introns a clear pattern of sequence conservation can be observed. In particular the intron-exon boundaries and a region close to the 3'-splice site show sequence homology. Conserved sequences were also found in the leader and trailer regions of the ribosomal protein mRNAs. The 5'-flanking regions of the yeast ribosomal protein genes appeared to contain sequence elements that many but not all ribosomal protein genes have in common, and therefore may be implicated in the coordinate expression of these genes. The amino acid coding sequences of the ribosomal protein genes show a biased codon usage. Like most yeast ribosomal protein molecules, L17a and L25 are particularly basic at their N-terminus.

Amino Acid Sequence↗

Nucleotide sequence of the yeast cell division cycle start genes CDC28, CDC36, CDC37, and CDC39, and a structural analysis of the predicted products.

The nucleotide sequences of the yeast cell division cycle start genes CDC36, CDC37, and CDC39 are presented. An open reading frame corresponding in size and mapped position to the mRNA for each gene was revealed. These sequences, as well as that of the CDC28 gene, were analyzed for the presence of consensus sequences postulated to be transcriptional or translational signals, or to be involved in mRNA processing. In addition, the predicted protein products of the four genes were subjected to a number of structural and statistical analyses including codon usage bias analysis, secondary structure analysis and hydropathicity analysis.

Amino Acid Sequence↗

Sequence and expression of NUC1, the gene encoding the mitochondrial nuclease in Saccharomyces cerevisiae.

The DNA sequence and studies on the expression of the NUC1 gene from Saccharomyces cerevisiae are presented. The NUC1 locus is located in the distal portion of the left arm of Chromosome X and encodes the major nuclease found in mitochondria. The inferred amino acid sequence of NUC1 predicts that the nuclease is basic, rich in prolines, of average hydrophobicity, and has a molecular weight for the primary translation product of 37,209 daltons. NUC1 is very poorly expressed, consistent with the codon usage bias determined from the DNA sequence and our previous determination of the number of enzyme molecules per cell. Mapping of the 5' terminus of the NUC1 mRNA reveals that the mRNA has a long 400 base untranslated leader in which are found three open reading frames, each initiated by an AUG. The possibility that these upstream open reading frames contribute to the poor expression of the NUC1 gene is discussed.

Amino Acid Sequence↗

Behavior of restriction-modification systems as selfish mobile elements and their impact on genome evolution.

Restriction-modification (RM) systems are composed of genes that encode a restriction enzyme and a modification methylase. RM systems sometimes behave as discrete units of life, like viruses and transposons. RM complexes attack invading DNA that has not been properly modified and thus may serve as a tool of defense for bacterial cells. However, any threat to their maintenance, such as a challenge by a competing genetic element (an incompatible plasmid or an allelic homologous stretch of DNA, for example) can lead to cell death through restriction breakage in the genome. This post-segregational or post-disturbance cell killing may provide the RM complexes (and any DNA linked with them) with a competitive advantage. There is evidence that they have undergone extensive horizontal transfer between genomes, as inferred from their sequence homology, codon usage bias and GC content difference. They are often linked with mobile genetic elements such as plasmids, viruses, transposons and integrons. The comparison of closely related bacterial genomes also suggests that, at times, RM genes themselves behave as mobile elements and cause genome rearrangements. Indeed some bacterial genomes that survived post-disturbance attack by an RM gene complex in the laboratory have experienced genome rearrangements. The avoidance of some restriction sites by bacterial genomes may result from selection by past restriction attacks. Both bacteriophages and bacteria also appear to use homologous recombination to cope with the selfish behavior of RM systems. RM systems compete with each other in several ways. One is competition for recognition sequences in post-segregational killing. Another is super-infection exclusion, that is, the killing of the cell carrying an RM system when it is infected with another RM system of the same regulatory specificity but of a different sequence specificity. The capacity of RM systems to act as selfish, mobile genetic elements may underlie the structure and function of RM enzymes.

Base Sequence↗

Extent of gene duplication in the genomes of Drosophila, nematode, and yeast.

We conducted a detailed analysis of duplicate genes in three complete genomes: yeast, Drosophila, and Caenorhabditis elegans. For two proteins belonging to the same family we used the criteria: (1) their similarity is > or =I (I = 30% if L > or = 150 a.a. and I = 0.01n + 4.8L(-0.32(1 + exp(-L/1000))) if L < 150 a.a., where n = 6 and L is the length of the alignable region), and (2) the length of the alignable region between the two sequences is > or = 80% of the longer protein. We found it very important to delete isoforms (caused by alternative splicing), same genes with different names, and proteins derived from repetitive elements. We estimated that there were 530, 674, and 1,219 protein families in yeast, Drosophila, and C. elegans, respectively, so, as expected, yeast has the smallest number of duplicate genes. However, for the duplicate pairs with the number of substitutions per synonymous site (K(S)) < 0.01, Drosophila has only seven pairs, whereas yeast has 58 pairs and nematode has 153 pairs. After considering the possible effects of codon usage bias and gene conversion, these numbers became 6, 55, and 147, respectively. Thus, Drosophila appears to have much fewer young duplicate genes than do yeast and nematode. The larger numbers of duplicate pairs with K(S) < 0.01 in yeast and C. elegans were probably largely caused by block duplications. At any rate, it is clear that the genome of Drosophila melanogaster has undergone few gene duplications in the recent past and has much fewer gene families than C. elegans.

Animals↗

Intraspecific nuclear DNA variation in Drosophila.

We have summarized and analyzed all available nuclear DNA sequence polymorphism studies for three species of Drosophila, D. melanogaster (24 loci), D. simulans (12 loci), and D. pseudoobscura (5 loci). Our major findings are: (1) The average nucleotide heterozygosity ranges from about 0.4% to 2% depending upon species and function of the region, i.e., coding or noncoding. (2) Compared to D. simulans and D. pseudoobscura (which are about equally variable), D. melanogaster displays a low degree of DNA polymorphism. (3) Noncoding introns and 3' and 5' flanking DNA shows less polymorphism than silent sites within coding DNA. (4) X-linked genes are less variable than autosomal genes. (5) Transition (Ts) and transversion (Tv) polymorphisms are about equally frequent in non-coding DNA and at fourfold degenerate sites in coding DNA while Ts polymorphisms outnumber Tv polymorphisms by about 2:1 in total coding DNA. The increased Ts polymorphism in coding regions is likely due to the structure of the genetic code: silent changes are more often Ts's than are replacement substitutions. (6) The proportion of replacement polymorphisms is significantly higher in D. melanogaster than in D. simulans. (7) The level of variation in coding DNA and the adjacent noncoding DNA is significantly correlated indicating regional effects, most notably recombination. (8) Surprisingly, the level of polymorphism at silent coding sites in D. melanogaster is positively correlated with degree of codon usage bias. (9) Three proposed tests of the neutral theory of DNA polymorphisms have been performed on the data: Tajima's test, the HKA test, and the McDonald-Kreitman test. About half of the loci fail to conform to the expectations of neutral theory by one of the tests. We conclude that many variables are affecting levels of DNA polymorphism in Drosophila, from properties of nucleotides to population history and, perhaps, mating structure. No simple, all encompassing explanation satisfactorily accounts for the data.

Animals↗

Preponderance of slightly deleterious polymorphism in mitochondrial DNA: nonsynonymous/synonymous rate ratio is much higher within species than between species.

We estimated synonymous (dN) and nonsynonymous (dS) substitution rates for protein-coding genes of the mitochondrial genome from two individuals each of the species human, chimpanzee, and gorilla. The genes were analyzed both separately and in a combined data set. Pairwise sequence comparisons suggest that the dN/dS rate ratios are about 5-10 times higher in within-species comparisons than in between-species comparisons. This result is confirmed by a more rigorous likelihood ratio test, which rejected the null hypothesis that the dN/dS rate ratios are identical within and between species. The likelihood models account for the genetic code structure, transition/transversion rate ratio, and codon usage bias and are expected to produce more reliable results than the commonly used contingency test. Separate analyses of different genes show that the dN/dS rate ratios are higher within species than between species for all 13 mitochondrial genes, with the difference being statistically significant for all except three small or slowly evolving genes. Furthermore, in conserved genes, nonsynonymous rates within species tend to be higher than the between-species rates by a greater proportion than in fast-changing genes. Our findings confirm and extend earlier results obtained from smaller data sets and suggest the operation of slightly deleterious mutations throughout the mitochondrial genome in the hominoids. Implications of the results for evolutionary studies and, in particular, for studies of the origin of modern humans, are discussed.

Adenosine Triphosphatases↗

Evolution of nucleotide substitutions and gene regulation in the amylase multigenes in Drosophila kikkawai and its sibling species.

In order to determine evolutionary changes in gene regulation and the nucleotide substitution pattern in a multigene family, the amylase multigenes were characterized in Drosophila kikkawai and its sibling species. The nucleotide substitution pattern was investigated. Drosophila kikkawai has four amylase genes. The Amy1 and Amy2 genes are a head-to-head duplication in the middle of the B arm of the second chromosome, while the Amy3 and Amy4 genes are a tail-to-tail duplication near the centromere of the same chromosome. In the sibling species of D. kikkawai (Drosophila bocki, Drosophila leontia, and Drosophila lini), sequencing of the Amy1, Amy2, Amy3, and Amy4 genes revealed that the Amy1 and Amy2 gene group diverged from Amy3 and Amy4 after duplication. In the Amy1 and Amy2 genes, the divergent evolution occurred in the flanking regions; in contrast, the coding regions have evolved in concerted fashion. The electrophoretic pattern of AMY isozymes was also examined. In D. kikkawai and its siblings, two or three electrophoretically different isozymes are encoded by the Amy1 and Amy2 genes (S isozyme) and by the Amy3 and Amy4 genes (F (M) isozymes). The S and F (M) isozymes show different patterns of band intensity when larvae and flies were fed in different media. Amy1 and Amy2, which encode the S isozyme, are more strikingly regulated than Amy3 and Amy4, which encode the F (M) isozyme. The GC content and codon usage bias were higher for the Amy1 and Amy2 genes than for the Amy3 and Amy4 genes. Although the ratio of synonymous and replacement substitutions within the Amy1 and Amy2 gene group was not significantly different from that within the Amy3 and Amy4 gene group, the synonymous substitution rate in the lineage of Amy1 and Amy2 was lower than that of Amy3 and Amy4. In conclusion, after the first duplication but before speciation of four species, the synonymous substitution rate between the two lineages and the electrophoretic pattern of the isozymes encoded by them changed, although we do not know whether there was any evolutionary relationship between the two.

Amino Acid Substitution↗

Positive selection is a general phenomenon in the evolution of abalone sperm lysin.

Lysin is a 16kDa acrosomal protein used by abalone sperm to create a hole in the egg vitelline envelope (VE). The interaction of lysin with the VE is species-selective and is one step in the multistep fertilization process that restricts heterospecific (cross-species) fertilization. For this reason, the evolution of lysin could play a role in establishing prezygotic reproductive isolation between species. Previously, we sequenced sperm lysin cDNAs from seven California abalone species and showed that positive Darwinian selection promotes their divergence. In this paper an additional 13 lysin sequences are presented representing species from Japan, Taiwan, Australia, New Zealand, South Africa, and Europe. The total of 20 sequences represents the most extensive analysis of a fertilization protein to date. The phylogenetic analysis divides the sequences into two major clades, one composed of species from the northern Pacific (California and Japan) and the other composed of species from other parts of the world. Analysis of nucleotide substitution demonstrates that positive selection is a general process in the evolution of this fertilization protein. Analysis of nucleotide and codon usage bias shows that neither parameter can account for the robust data supporting positive selection. The selection pressure responsible for the positive selection on lysin remains unknown.

Amino Acid Sequence↗

Molecular population genetics of Escherichia coli: DNA sequence diversity at the celC, crr, and gutB loci of natural isolates.

The DNA sequences of three genes--celC, crr, and gutB--have been determined for each of 11 or 12 natural isolates of Escherichia coli from the ECOR collection. These genes encode the phosphoenolpyruvate-dependent phosphotransferase-system enzyme III proteins specific for beta-glucoside sugars (celC), glucose (crr), and glucitol (gutB), respectively. There is little evidence of recombination at or among these loci; among these strains, relationships inferred from each gene are largely consistent with each other and with the relationship inferred from multilocus enzyme electrophoresis. DNA sequence diversity is similar for all three genes, particularly when silent (synonymous) sites only are considered. This is surprising because there is much stronger codon usage bias at crr than at celC or gutB. The extent of divergence in the protein sequences encoded by these three genes varies considerably. The constitutively expressed glucose-specific enzyme is completely conserved. It is surprising that the inducible glucitol-specific enzyme, which is functional, is more variable than the cellobiose-specific enzyme, which is cryptic; the latter might be expected to be under less (if any) purifying selection.

Amino Acid Sequence↗