Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,783 records · Page 99Linked to original sources

Thermoadaptation trait revealed by the genome sequence of thermophilic Geobacillus kaustophilus.

We present herein the first complete genome sequence of a thermophilic Bacillus-related species, Geobacillus kaustophilus HTA426, which is composed of a 3.54 Mb chromosome and a 47.9 kb plasmid, along with a comparative analysis with five other mesophilic bacillar genomes. Upon orthologous grouping of the six bacillar sequenced genomes, it was found that 1257 common orthologous groups composed of 1308 genes (37%) are shared by all the bacilli, whereas 839 genes (24%) in the G.kaustophilus genome were found to be unique to that species. We were able to find the first prokaryotic sperm protamine P1 homolog, polyamine synthase, polyamine ABC transporter and RNA methylase in the 839 unique genes; these may contribute to thermophily by stabilizing the nucleic acids. Contrasting results were obtained from the principal component analysis (PCA) of the amino acid composition and synonymous codon usage for highlighting the thermophilic signature of the G.kaustophilus genome. Only in the PCA of the amino acid composition were the Bacillus-related species located near, but were distinguishable from, the borderline distinguishing thermophiles from mesophiles on the second principal axis. Further analysis revealed some asymmetric amino acid substitutions between the thermophiles and the mesophiles, which are possibly associated with the thermoadaptation of the organism.

Adaptation, Physiological↗

The Diatom EST Database.

The Diatom EST database provides integrated access to expressed sequence tag (EST) data from two eukaryotic microalgae of the class Bacillariophyceae, Phaeodactylum tricornutum and Thalassiosira pseudonana. The database currently contains sequences of close to 30,000 ESTs organized into PtDB, the P.tricornutum EST database, and TpDB, the T.pseudonana EST database. The EST sequences were clustered and assembled into a non-redundant set for each organism, and these non-redundant sequences were then subjected to automated annotation using similarity searches against protein and domain databases. EST sequences, clusters of contiguous sequences, their annotation and analysis with reference to the publicly available databases, and a codon usage table derived from a subset of sequences from PtDB and TpDB can all be accessed in the Diatom EST Database. The underlying RDBMS enables queries over the raw and annotated EST data and retrieval of information through a user-friendly web interface, with options to perform keyword and BLAST searches. The EST data can also be retrieved based on Pfam domains, Cluster of Orthologous Groups (COG) and Gene Ontologies (GO) assigned to them by similarity searches. The Database is available at http://avesthagen.sznbowler.com.

DNA, Algal↗

ChloroplastDB: the Chloroplast Genome Database.

The Chloroplast Genome Database (ChloroplastDB) is an interactive, web-based database for fully sequenced plastid genomes, containing genomic, protein, DNA and RNA sequences, gene locations, RNA-editing sites, putative protein families and alignments (http://chloroplast.cbio.psu.edu/). With recent technical advances, the rate of generating new organelle genomes has increased dramatically. However, the established ontology for chloroplast genes and gene features has not been uniformly applied to all chloroplast genomes available in the sequence databases. For example, annotations for some published genome sequences have not evolved with gene naming conventions. ChloroplastDB provides unified annotations, gene name search, BLAST and download functions for chloroplast encoded genes and genomic sequences. A user can retrieve all orthologous sequences with one search regardless of gene names in GenBank. This feature alone greatly facilitates comparative research on sequence evolution including changes in gene content, codon usage, gene structure and post-transcriptional modifications such as RNA editing. Orthologous protein sets are classified by TribeMCL and each set is assigned a standard gene name. Over the next few years, as the number of sequenced chloroplast genomes increases rapidly, the tools available in ChloroplastDB will allow researchers to easily identify and compile target data for comparative analysis of chloroplast genes and genomes.

Chloroplasts↗

The primary structure of phosphoenolpyruvate carboxylase of Escherichia coli. Nucleotide sequence of the ppc gene and deduced amino acid sequence.

The nucleotide sequence of the ppc gene, the structural gene for phosphoenolpyruvate carboxylase [EC 4.1.1.31], of Escherichia coli K-12 was determined. The gene codes for a polypeptide comprising 883 amino acid residues with a calculated molecular weight of 99,061. The amino acid sequence deduced from the nucleotide sequence was entirely consistent with the protein chemical data obtained with the purified enzyme, including the NH2- and COOH-terminal sequences and amino acid composition. The coding region is preceded by two putative ribosome binding sites, and is followed closely by a good representative of rho-independent terminator. The codon usage in the ppc gene suggests a moderate expression of the gene. The secondary structure of the enzyme was predicted from the deduced amino acid sequence.

Amino Acid Sequence↗

A novel mitochondrial gene order in the crinoid echinoderm Florometra serratissima.

The complete nucleotide sequence of the mitochondrial genome of the crinoid Florometra serratissima has been determined. It is a circular DNA molecule, 16,005 bp in length, containing the genes for 13 proteins, small and large ribosomal RNAs, and 22 transfer RNAs (tRNAs). Three regions of unassigned sequence (UAS) greater than 73 bp have been located. The largest, UAS I, is 432 bp long and exhibits sequence similarity to the putative mitochondrial control regions seen in other animals. UAS II (77 bp) and UAS III (73 bp) are located between the 5' ends of coding sequences and may play roles as bidirectional promoters. Analyses of nucleotide composition revealed that the major peptide-encoding strand is high in T and low in C. This bias is reflected in a specific pattern of codon usage. Molecular phylogenetic analyses based on cytochrome c oxidase (COI, COII, and COIII) amino acid and nucleotide sequences did not resolve all the relationships between echinoderm classes. The overall animal mitochondrial gene content has been maintained in the crinoid, but there is extensive rearrangement with respect to both the echinoid and the asteroid mtDNA gene maps. Florometra serratissima has a novel genome organization in a segment containing most of the tRNA genes, large and small rRNA genes, and the NADH dehydrogenase subunit 1 and 2 genes. Potential pathways and mechanisms for gene rearrangements between mitochondrial gene maps of echinoderm classes and vertebrates are discussed as indicators of early deuterostome phylogeny.

Animals↗

Extent of gene duplication in the genomes of Drosophila, nematode, and yeast.

We conducted a detailed analysis of duplicate genes in three complete genomes: yeast, Drosophila, and Caenorhabditis elegans. For two proteins belonging to the same family we used the criteria: (1) their similarity is > or =I (I = 30% if L > or = 150 a.a. and I = 0.01n + 4.8L(-0.32(1 + exp(-L/1000))) if L < 150 a.a., where n = 6 and L is the length of the alignable region), and (2) the length of the alignable region between the two sequences is > or = 80% of the longer protein. We found it very important to delete isoforms (caused by alternative splicing), same genes with different names, and proteins derived from repetitive elements. We estimated that there were 530, 674, and 1,219 protein families in yeast, Drosophila, and C. elegans, respectively, so, as expected, yeast has the smallest number of duplicate genes. However, for the duplicate pairs with the number of substitutions per synonymous site (K(S)) < 0.01, Drosophila has only seven pairs, whereas yeast has 58 pairs and nematode has 153 pairs. After considering the possible effects of codon usage bias and gene conversion, these numbers became 6, 55, and 147, respectively. Thus, Drosophila appears to have much fewer young duplicate genes than do yeast and nematode. The larger numbers of duplicate pairs with K(S) < 0.01 in yeast and C. elegans were probably largely caused by block duplications. At any rate, it is clear that the genome of Drosophila melanogaster has undergone few gene duplications in the recent past and has much fewer gene families than C. elegans.

Animals↗

Patterns of nucleotide substitution among simultaneously duplicated gene pairs in Arabidopsis thaliana.

We characterized rates and patterns of synonymous and nonsynonymous substitution in 242 duplicated gene pairs on chromosomes 2 and 4 of Arabidopsis thaliana. Based on their collinear order along the two chromosomes, the gene pairs were likely duplicated contemporaneously, and therefore comparison of genetic distances among gene pairs provides insights into the distribution of nucleotide substitution rates among plant nuclear genes. Rates of synonymous substitution varied 13.8-fold among the duplicated gene pairs, but 90% of gene pairs differed by less than 2.6-fold. Average nonsynonymous rates were approximately fivefold lower than average synonymous rates; this rate difference is lower than that of previously studied nonplant lineages. The coefficient of variation of rates among genes was 0.65 for nonsynonymous rates and 0.44 for synonymous rates, indicating that synonymous and nonsynonymous rates vary among genes to roughly the same extent. The causes underlying rate variation were explored. Our analyses tentatively suggest an effect of physical location on synonymous substitution rates but no similar effect on nonsynonymous rates. Nonsynonymous substitution rates were negatively correlated with GC content at synonymous third codon positions, and synonymous substitution rates were negatively correlated with codon bias, as observed in other systems. Finally, the 242 gene pairs permitted investigation of the processes underlying divergence between paralogs. We found no evidence of positive selection, little evidence that paralogs evolve at different rates, and no evidence of differential codon usage or third position GC content.

Arabidopsis↗

Intraspecific nuclear DNA variation in Drosophila.

We have summarized and analyzed all available nuclear DNA sequence polymorphism studies for three species of Drosophila, D. melanogaster (24 loci), D. simulans (12 loci), and D. pseudoobscura (5 loci). Our major findings are: (1) The average nucleotide heterozygosity ranges from about 0.4% to 2% depending upon species and function of the region, i.e., coding or noncoding. (2) Compared to D. simulans and D. pseudoobscura (which are about equally variable), D. melanogaster displays a low degree of DNA polymorphism. (3) Noncoding introns and 3' and 5' flanking DNA shows less polymorphism than silent sites within coding DNA. (4) X-linked genes are less variable than autosomal genes. (5) Transition (Ts) and transversion (Tv) polymorphisms are about equally frequent in non-coding DNA and at fourfold degenerate sites in coding DNA while Ts polymorphisms outnumber Tv polymorphisms by about 2:1 in total coding DNA. The increased Ts polymorphism in coding regions is likely due to the structure of the genetic code: silent changes are more often Ts's than are replacement substitutions. (6) The proportion of replacement polymorphisms is significantly higher in D. melanogaster than in D. simulans. (7) The level of variation in coding DNA and the adjacent noncoding DNA is significantly correlated indicating regional effects, most notably recombination. (8) Surprisingly, the level of polymorphism at silent coding sites in D. melanogaster is positively correlated with degree of codon usage bias. (9) Three proposed tests of the neutral theory of DNA polymorphisms have been performed on the data: Tajima's test, the HKA test, and the McDonald-Kreitman test. About half of the loci fail to conform to the expectations of neutral theory by one of the tests. We conclude that many variables are affecting levels of DNA polymorphism in Drosophila, from properties of nucleotides to population history and, perhaps, mating structure. No simple, all encompassing explanation satisfactorily accounts for the data.

Animals↗

Asymmetric substitution patterns in the two DNA strands of bacteria.

Analyses of the genomes of three prokaryotes, Escherichia coli, Bacillus subtilis, and Haemophilus influenzae, revealed a new type of genomic compartmentalization of base frequencies. There was a departure from intrastrand equifrequency between A and T or between C and G, showing that the substitution patterns of the two strands of DNA were asymmetric. The positions of the boundaries between these compartments were found to coincide with the origin and terminus of chromosome replication, and there were more A-T and C-G deviations in intergenic regions and third codon positions, suggesting that a mutational bias was responsible for this asymmetry. The strand asymmetry was found to be due to a difference in base compositions of transcripts in the leading and lagging strands. This difference is sufficient to affect codon usage, but it is small compared to the effects of gene expressivity and amino-acid composition.

Bacillus subtilis↗

Preponderance of slightly deleterious polymorphism in mitochondrial DNA: nonsynonymous/synonymous rate ratio is much higher within species than between species.

We estimated synonymous (dN) and nonsynonymous (dS) substitution rates for protein-coding genes of the mitochondrial genome from two individuals each of the species human, chimpanzee, and gorilla. The genes were analyzed both separately and in a combined data set. Pairwise sequence comparisons suggest that the dN/dS rate ratios are about 5-10 times higher in within-species comparisons than in between-species comparisons. This result is confirmed by a more rigorous likelihood ratio test, which rejected the null hypothesis that the dN/dS rate ratios are identical within and between species. The likelihood models account for the genetic code structure, transition/transversion rate ratio, and codon usage bias and are expected to produce more reliable results than the commonly used contingency test. Separate analyses of different genes show that the dN/dS rate ratios are higher within species than between species for all 13 mitochondrial genes, with the difference being statistically significant for all except three small or slowly evolving genes. Furthermore, in conserved genes, nonsynonymous rates within species tend to be higher than the between-species rates by a greater proportion than in fast-changing genes. Our findings confirm and extend earlier results obtained from smaller data sets and suggest the operation of slightly deleterious mutations throughout the mitochondrial genome in the hominoids. Implications of the results for evolutionary studies and, in particular, for studies of the origin of modern humans, are discussed.

Adenosine Triphosphatases↗

Determinants of substitution rates in mammalian genes: expression pattern affects selection intensity but not mutation rate.

To determine whether gene expression patterns affect mutation rates and/or selection intensity in mammalian genes, we studied the relationships between substitution rates and tissue distribution of gene expression. For this purpose, we analyzed 2,400 human/rodent and 834 mouse/rat orthologous genes, and we measured (using expressed sequence tag data) their expression patterns in 19 tissues from three development states. We show that substitution rates at nonsynonymous sites are strongly negatively correlated with tissue distribution breadth: almost threefold lower in ubiquitous than in tissue-specific genes. Nonsynonymous substitution rates also vary considerably according to the tissues: the average rate is twofold lower in brain-, muscle-, retina- and neuron-specific genes than in lymphocyte-, lung-, and liver-specific genes. Interestingly, 5' and 3' untranslated regions (UTRs) show exactly the same trend. These results demonstrate that the expression pattern is an essential factor in determining the selective pressure on functional sites in both coding and noncoding regions. Conversely, silent substitution rates do not vary with expression pattern, even in ubiquitously expressed genes. This latter result thus suggests that synonymous codon usage is not constrained by selection in mammals. Furthermore, this result also indicates that there is no reduction of mutation rates in genes expressed in the germ line, contrary to what had been hypothesized based on the fact that transcribed DNA is more efficiently repaired than nontranscribed DNA.

Animals↗

Horizontal gene transfer of glycosyl hydrolases of the rumen fungi.

By combining analyses of G + C content and patterns of codon usage and constructing phylogenetic trees, we describe the gene transfer of an endoglucanase (celA) from the rumen bacteria Fibrobacter succinogenes to the rumen fungi Orpinomyces joyonii. The strong similarity between different glycosyl hydrolases of rumen fungi and bacteria suggests that most, if not all, of the glycosyl hydrolases of rumen fungi that play an important role in the degradation of cellulose and other plant polysaccharides were acquired by horizontal gene transfer events. This acquisition allows fungi to establish a habitat within a new environmental niche: the rumen of the herbivorous mammals for which cellulose and plant hemicellulose constitute the main raw nutritive substrate.

Animals↗

Evolution of nucleotide substitutions and gene regulation in the amylase multigenes in Drosophila kikkawai and its sibling species.

In order to determine evolutionary changes in gene regulation and the nucleotide substitution pattern in a multigene family, the amylase multigenes were characterized in Drosophila kikkawai and its sibling species. The nucleotide substitution pattern was investigated. Drosophila kikkawai has four amylase genes. The Amy1 and Amy2 genes are a head-to-head duplication in the middle of the B arm of the second chromosome, while the Amy3 and Amy4 genes are a tail-to-tail duplication near the centromere of the same chromosome. In the sibling species of D. kikkawai (Drosophila bocki, Drosophila leontia, and Drosophila lini), sequencing of the Amy1, Amy2, Amy3, and Amy4 genes revealed that the Amy1 and Amy2 gene group diverged from Amy3 and Amy4 after duplication. In the Amy1 and Amy2 genes, the divergent evolution occurred in the flanking regions; in contrast, the coding regions have evolved in concerted fashion. The electrophoretic pattern of AMY isozymes was also examined. In D. kikkawai and its siblings, two or three electrophoretically different isozymes are encoded by the Amy1 and Amy2 genes (S isozyme) and by the Amy3 and Amy4 genes (F (M) isozymes). The S and F (M) isozymes show different patterns of band intensity when larvae and flies were fed in different media. Amy1 and Amy2, which encode the S isozyme, are more strikingly regulated than Amy3 and Amy4, which encode the F (M) isozyme. The GC content and codon usage bias were higher for the Amy1 and Amy2 genes than for the Amy3 and Amy4 genes. Although the ratio of synonymous and replacement substitutions within the Amy1 and Amy2 gene group was not significantly different from that within the Amy3 and Amy4 gene group, the synonymous substitution rate in the lineage of Amy1 and Amy2 was lower than that of Amy3 and Amy4. In conclusion, after the first duplication but before speciation of four species, the synonymous substitution rate between the two lineages and the electrophoretic pattern of the isozymes encoded by them changed, although we do not know whether there was any evolutionary relationship between the two.

Amino Acid Substitution↗

Dramatic mitochondrial gene rearrangements in the hermit crab Pagurus longicarpus (Crustacea, anomura).

The entire mitochondrial gene order of the crustacean Pagurus longicarpus was determined by sequencing all but approximately 300 bp of the mitochondrial genome. We report the first major gene rearrangements found in the clade including Crustacea and Insecta. At least eight mitochondrial gene rearrangements have dramatically altered the gene order of the hermit crab P. longicarpus relative to the putatively ancestral crustacean gene order. These include two rearrangements of protein-coding genes, the first reported for any nonchelicerate arthropod. Codon usage and amino acid sequences do not deviate substantially from those reported for other crustaceans. Investigating the phylogenetic distribution of these eight rearrangements will add additional characters to help resolve decapod phylogeny.

Animals↗

Positive selection is a general phenomenon in the evolution of abalone sperm lysin.

Lysin is a 16kDa acrosomal protein used by abalone sperm to create a hole in the egg vitelline envelope (VE). The interaction of lysin with the VE is species-selective and is one step in the multistep fertilization process that restricts heterospecific (cross-species) fertilization. For this reason, the evolution of lysin could play a role in establishing prezygotic reproductive isolation between species. Previously, we sequenced sperm lysin cDNAs from seven California abalone species and showed that positive Darwinian selection promotes their divergence. In this paper an additional 13 lysin sequences are presented representing species from Japan, Taiwan, Australia, New Zealand, South Africa, and Europe. The total of 20 sequences represents the most extensive analysis of a fertilization protein to date. The phylogenetic analysis divides the sequences into two major clades, one composed of species from the northern Pacific (California and Japan) and the other composed of species from other parts of the world. Analysis of nucleotide substitution demonstrates that positive selection is a general process in the evolution of this fertilization protein. Analysis of nucleotide and codon usage bias shows that neither parameter can account for the robust data supporting positive selection. The selection pressure responsible for the positive selection on lysin remains unknown.

Amino Acid Sequence↗

Molecular population genetics of Escherichia coli: DNA sequence diversity at the celC, crr, and gutB loci of natural isolates.

The DNA sequences of three genes--celC, crr, and gutB--have been determined for each of 11 or 12 natural isolates of Escherichia coli from the ECOR collection. These genes encode the phosphoenolpyruvate-dependent phosphotransferase-system enzyme III proteins specific for beta-glucoside sugars (celC), glucose (crr), and glucitol (gutB), respectively. There is little evidence of recombination at or among these loci; among these strains, relationships inferred from each gene are largely consistent with each other and with the relationship inferred from multilocus enzyme electrophoresis. DNA sequence diversity is similar for all three genes, particularly when silent (synonymous) sites only are considered. This is surprising because there is much stronger codon usage bias at crr than at celC or gutB. The extent of divergence in the protein sequences encoded by these three genes varies considerably. The constitutively expressed glucose-specific enzyme is completely conserved. It is surprising that the inducible glucitol-specific enzyme, which is functional, is more variable than the cellobiose-specific enzyme, which is cryptic; the latter might be expected to be under less (if any) purifying selection.

Amino Acid Sequence↗

PCR-based gene synthesis as an efficient approach for expression of the A+T-rich malaria genome.

The A+T-rich genome of the human malaria parasite Plasmodium falciparum encodes genes of biological importance that cannot be expressed efficiently in heterologous eukaryotic systems, owing to an extremely biased codon usage and the presence of numerous cryptic polyadenylation sites. In this work we have optimized an assembly polymerase chain reaction (PCR) method for the fast and extremely accurate synthesis of a 2.1 kb Plasmodium falciparum gene (pfsub-1) encoding a subtilisin-like protease. A total of 104 oligonucleotides, designed with the aid of dedicated computer software, were assembled in a single-step PCR. The assembly was then further amplified by PCR to produce a synthetic gene which has been cloned and successfully expressed in both Pichia pastoris and recombinant baculovirus-infected High Five(TM) cells. We believe this strategy to be of special interest as it is simple, accessible and has no limitation with respect to the size of the gene to be synthesized. Used as a systematic approach for the malarial genome or any other A + T-rich organism, the method allows the rapid synthesis of a nucleotide sequence optimized for expression in the system of choice and production of sufficiently large amounts of biological material for complete molecular and structural characterization.

Amino Acid Sequence↗

Synthesis, cloning and expression of the single-chain Fv gene of the HPr-specific monoclonal antibody, Jel42. Determination of binding constants with wild-type and mutant HPrs.

The monoclonal antibody Jel42 is specific for the Escherichia coli histidine-containing protein, HPr, which is an 85 amino acid phosphocarrier protein of the phosphoenolpyruvate:sugar phosphotransferase system. The binding domain (Fv) has been produced as a single chain Fv (scFv). The scFv gene was synthesized in vitro and coded for pelB leader peptide-heavy chain-linker-light chain-(His)(5) tail. The linker is three repeats from the C-terminal repetitive sequence of eukaryotic RNA polymerase II. This linker acts as a tag; it is the antigen for the monoclonal antibody Jel352. The codon usage was maximized for E.coli expression, and many unique restriction endonuclease sites were incorporated. The scFv gene incorporated into pT7-7 was highly expressed, yielding 10-30% of the cell protein as the scFv, which was found in inclusion bodies with the leader peptide cleaved. Jel42 scFv was purified by denaturation/renaturation yielding preparations with K(d) values from 20 to 175 nM. However, based upon an assessment of the amount of active refolded scFv, the binding dissociation constant was estimated to be 2.7 +/- 2.0 nM compared with 2.8 +/- 1.6 and 3.7 +/- 0.3 nM previously determined for the Jel42 antibody and Fab fragment respectively. The effect of mutation of the antigen HPr on the binding constant of the scFv was very similar to the properties determined for the antibody and the Fab fragment. It was concluded that the small percentage ( approximately 6%) of refolded scFv is a true mimic of the Jel42 binding domain and that the incorrectly folded scFv cannot be detected in the binding assay.

Amino Acid Sequence↗