Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,621 records · Page 90Linked to original sources

The evolution of genomic anatomy.

Just as Darwin applied his theory of natural selection to understand the details of natural history, so others have applied the idea to almost every aspect of biology from morphology to medicine. Can we similarly comprehend the rapidly accumulating details of the natural history of genomes or is selection not that strong a force? Recent case histories indicate that selection can affect everything from minuscule details, such as codon usage, to broader scale patterns, such as the linkage arrangement of genes, their chromosomal position and copy number. Although we should not assume that the structure of genomes is exclusively the result of history and chance, few generalities are presently possible because evidence is largely restricted to case-by-case analyses.

Journal Article↗

Glyceraldehyde-3-phosphate dehydrogenase from Tetrahymena pyriformis: enzyme purification and characterization of a gapC gene with primitive eukaryotic features.

Glyceraldehyde-3-phosphate dehydrogenase (GAPDH, EC.1.2.1.12) was purified to electrophoretic homogeneity from an amicronucleated strain of the ciliate Tetrahymena pyriformis using a three-step procedure. The native enzyme is an homotetramer of 145 kDa exhibiting absolute specificity for NAD. In its catalytic properties it is similar to other glycolytic GAPDHs. Chromatofocusing analysis showed the presence of only one basic GAPDH isoform with an isoelectric point of 8.8. Western blots using a monospecific polyclonal antibody raised against the T. pyriformis GAPDH showed a single 36-kDa band corresponding to the enzyme subunit in the cytosolic protein fraction of this strain and the closely related species, both from the class Oligohymenophorea, Paramecium tetraurelia. No bands were immunodetected in the ciliate Colpoda inflata (class Colpodea) and in the diverse eukaryotes and eubacteria tested. A 0.5-kb DNA fragment which corresponds to an internal region of a gapC gene was generated by polymerase chain reaction using cDNA of T. pyriformis as template. This gene codes for a basic GAPDH protein with eukaryotic-diplomonad signatures and exhibits a codon usage biased in the manner typical for T. pyriformis genes. Southern blots performed both under homologous and heterologous conditions using this amplified cDNA fragment as a probe, indicated that it should be the only gapC gene present in the macronuclear genome of this ciliate, its expression being confirmed by Northern blot analysis. These results are discussed in connection with the peculiar genomic organization of ciliates and in the context of protist evolution.

Amino Acid Sequence↗

Detecting pathogenicity islands and anomalous gene clusters by iterative discriminant analysis.

We present a simple method to detect pathogenicity islands and anomalous gene clusters in bacterial genomes. The method uses iterative discriminant analysis to define genomic regions that deviate most from the rest of the genome in three compositional criteria: G+C content, dinucleotide frequency and codon usage. Using this method, we identify many virulence-related gene islands, e.g. encoding protein secretion systems, adhesins, toxins, and other anomalous gene clusters, such as prophages. The program and the whole dataset, including the catalogs of genes in the detected anomalous segments, are publicly available at http://compbio.sibsnet.org/projects/pai-ida/. This program can be used in searching for virulence-related factors in newly sequenced bacterial genomes.

Algorithms↗

msDNA-St85, a multicopy single-stranded DNA isolated from Salmonella enterica serovar Typhimurium LT2 with the genomic analysis of its retron.

Bacterial reverse transcriptase is responsible for the production of a small satellite DNA-RNA complex called multicopy single-stranded DNA (msDNA) that has been found in a wide variety of Gram-negative bacteria. Here we describe the isolation and characterization of a novel msDNA, msDNA-St85, from Salmonella enterica serovar Typhimurium LT2. We determined the nucleotide sequence of msDNA-St85 and the location of retron-St85 on the chromosome that is responsible for msDNA-St85 production by analyzing the complete genomic sequence of S. typhimurium LT2. It was found that the G+C content and the codon usage of retron-St85 were significantly different from those of the S. typhimurium genome, indicating that retron-St85 was probably acquired recently in this bacterium. This is the first report for identification of an msDNA in the genus Salmonella with the complete description and analysis of its retron.

Amino Acid Sequence↗

The gene encoding pyolysin, the pore-forming toxin of Arcanobacterium pyogenes, resides within a genomic islet flanked by essential genes.

The plo gene, encoding the Arcanobacterium pyogenes cholesterol-dependent cytolysin, pyolysin (PLO), was localized to a 2.7-kb genomic islet of reduced %G+C content and alternate codon usage frequency. This islet, conserved among isolates from diverse hosts and geographical locations, separated the housekeeping genes smc and ftsY, which are found adjacent in many prokaryotes. The ftsY and ffh genes, located downstream of the plo islet, encode components of the signal recognition particle. Mutational analysis suggested that these genes were essential for viability in A. pyogenes. The A. pyogenes ffh gene was unable to complement a conditional ffh mutant of Escherichia coli and its overexpression was toxic in E. coli. Mutagenesis of the islet-encoded orf121 did not affect plo expression, indicating that it may not be involved directly in the regulation of plo expression. Regardless, the presence of the plo gene as part of a genomic islet inserted between genes essential for normal growth may provide selective pressure for the retention of this important virulence factor.

Bacterial Proteins↗

Identification and recombinant expression of glyceraldehyde-3-phosphate dehydrogenase of Plasmodium falciparum.

The gene coding for the cytosolic glyceraldehyde-3-phosphate dehydrogenase (GAPDH; EC 1.2.1.12) was isolated from Plasmodium falciparum. The gene contains 1 intron and the A+T content is characteristic for the codon usage of P. falciparum. The predicted open reading frame codes for 337 amino acids (36651Da) and is 63.5% identical to the human erythrocytic GAPDH. GAPDH sequences from several field isolates of P. falciparum displayed 100% conservation. Phylogenetic analysis supports the hypothesis that dinoflagellates and Plasmodium are closely related. The protein encoded by the pfGAPDH was expressed recombinantly in Escherichia coli and exhibited enzymatic activity with NAD(+) but not with NADP(+) as cofactor. Antiserum raised against the recombinantly expressed enzyme detected specifically all developmental stages of cultured P. falciparum blood-stage parasites.

Amino Acid Sequence↗

Sequence and expression of the SerJ immobilization antigen gene of Tetrahymena thermophila regulated by dominant epistasis.

In ciliates, variable surface protein genes encoding the immobilization antigen (-ag) are expressed under different environmental conditions, including temperature and salt stress. These i-ags are GPI-linked and coat the entire external surface of the cell, including the cilia. In Tetrahymena thermophila-ag in natural isolates is the result of dominant epistasis masking the expression of the H i-ag ordinarily expressed at 20-36 degrees C. This report describes the expression and sequence of the Ser-ag. J is present on the cell surface up to 38 degrees C; above 38 degrees C SerSeranked by an A-rich 5' UTR and a 3' UTR containing putative mRNA destabilization motifs. The encoded J polypeptide consists of 438 amino acids and is rich in alanine, cysteine, serine and threonine. The N- and resemble signal peptide and GPI-anchor addition sites, respectively. The majority of the molecule consists of four imperfect repeats with 10 periodic cysteines per repeat in the pattern CX(6)CX(2)CX(21)CX(4)CX(13-15)CX(2)CX(18)CX(3)CX(11)CX(9-10). Although H i-ags encoded by paralogous SerH genes have 3.5 imperfect repeats with eight periodic cysteines per repeat, J nevertheless resembles H with respect to amino acid composition, codon usage, N- and C-termini, the arrangement of the cysteine periods, and regulation by mRNA stability. However, despite these similarities and epistasis, the evolutionary relationship between SerH and SerJ is unclear.

Amino Acid Motifs↗

Covariation of GC content and the silent site substitution rate in rodents: implications for methodology and for the evolution of isochores.

Many attempts to test selectionist and neutralist models employ estimates of synonymous (Ks) and non-synonymous (Ka) substitution rates of orthologous genes. For example, a stronger Ka-Ks correlation than expected under neutrality has been argued to indicate a role for selection and the absence of a Ks-GC4 correlation has been argued to be inconsistent with neutral models for isochore evolution. However, both of these results, we have shown previously, are sensitive to the method by which Ka and Ks are estimated. Using a maximum likelihood (ML) estimator (GY94) we found a positive correlation between Ks and GC4 and only a weak correlation between Ka and Ks, lower than expected under neutral expectations. This ML method is computationally slow. Recently, a new ad hoc approximation of this ML method has been provided (YN00). This is effectively an extension of Li's protocol but that also allows for codon usage bias. This method is computationally near-instantaneous and therefore potentially of great utility for analysis of large datasets. Here we ask whether this method might have such applicability. To this end we ask whether it too recovers the two unusual results. We report that when the ML and earlier ad hoc methods disagree, YN00 recovers the results described by the ML methods, i.e. a positive correlation between GC4 and Ks and only a weak correlation between Ks and Ka. If the ML method can be trusted, then YN00 can also be considered an adequately reliable method for analysis of large datasets. Assuming this to be so we also analyze further the patterns. We show, for example, that the positive correlation between GC4 and Ks is probably in part a mutational bias, there being more methyl induced CpG-->TpG mutations in GC rich regions. As regards the evolution of isochores, it seems inappropriate to use the claimed lack of a correlation between GC and Ks as definitive evidence either against or for any model. If the positive correlation is real then, we argue, this is hard to reconcile with the biased gene conversion model for isochore formation as this predicts a negative correlation.

Animals↗

Functional screening of a retroviral peptide library for MHC class I presentation.

We have used retroviral vector technology to develop a method for functional screening of combinatorial peptide libraries expressed inside mammalian cells with the ultimate goal of identifying new drug targets. The method was validated in a library screening experiment based on antigen presentation of small peptides. A library encoding SIXNXEKX-peptides, where X designates randomised positions corresponding to major histocompatibility (MHC) class I anchor residues, was generated in a retroviral vector. The library was transduced into a population of antigen presenting cells (APCs) known to mediate MHC class I restricted presentation of the SIINFEKL peptide. The cellular library was screened by using an antigen presentation assay in which a T cell hybridoma recognising the MHC class I/SIINFEKL peptide complex was employed. Using this experimental model, we identified two positive cellular clones both encoding SIINFEKL peptides with identical codon usage. This number corresponded well to the expected frequency of SIINFEKL in the library. The lack of identification of other peptides capable of activating the T-hybridoma supports previous findings of a high degree of specificity at the level of peptide-loading of MHC-molecules. The result further demonstrates the potential of using combinatorial libraries for functional screening and selection of effector peptides stably expressed in mammalian cells.

Amino Acid Sequence↗

Mitochondrial DNA in metazoa: degree of freedom in a frozen event.

The mitochondrial genome (mtDNA), due to its peculiar features such as exclusive presence of orthologous genes, uniparental inheritance, lack of recombination, small size and constant gene content, certainly represents a major model system in studies on evolutionary genomics in metazoan. In 800 million years of evolution the gene content of metazoan mitochondrial genomes has remained practically frozen but several evolutionary processes have taken place. These processes, reviewed here, include rearrangements of gene order, changes in base composition and arising of compositional asymmetry between the two strands, variations in the genetic code and evolution of codon usage, lineage-specific nucleotide substitution rates and evolutionary patterns of mtDNA control regions.

Animals↗

The use of synthetic genes for the expression of ciliate proteins in heterologous systems.

The common fish parasite, Ichthyophthirius multifiliis, expresses abundant glycosylated phosphatidylinositol (GPI)-anchored membrane proteins known as immobilization antigens, or i-antigens. These proteins are targets of the host immune response, and have been identified as potential candidates for recombinant subunit vaccine development. Nevertheless, because Ichthyophthirius utilizes a non-standard genetic code, expression of the corresponding gene products, either as subunit antigens in conventional protein expression systems, or as vector-encoded antigens in the case of DNA vaccines, is far from straightforward. To overcome this problem, we utilized 'assembly polymerase chain reaction' to manufacture synthetic versions of two genes (designated IAG52A[G5/CC] and IAG52B[G5/CC]) encoding approximately 52/55 kDa i-antigens from parasite strain G5. This approach made it possible to eliminate unwanted stop codons and substitute the preferred codon usage of channel catfish for the native sequences of the genes. To determine whether the synthetic alleles could be expressed in cells that use the standard genetic code, we introduced IAG52A[G5/CC] into a variety of heterologous cell types and tested for expression either by immunofluorescence light microscopy or Western blotting. When cloned downstream of appropriate promoters, IAG52A[G5/CC] was expressed in Escherichia coli, mammalian COS-7 cells, and channel catfish where it elicited antigen-specific immune responses. Interestingly, the localization pattern of the corresponding gene product in COS-7 cells indicated that while the protein was correctly folded, it was not present on the cell membrane, suggesting that the signal peptides required for GPI-anchor addition differ in ciliate and mammalian systems. Construction of synthetic alleles should have practical utility in the development of vaccines against Ichthyophthirius, and at the same time, provide a general method for the expression of ciliate genes in heterologous systems.

Animals↗

An algorithm for detecting directional and non-directional positive selection, neutrality and negative selection in protein coding DNA sequences.

Positive selection or adaptive evolution is thought to be responsible, at least some of the time, for the rapid accumulation of advantageous changes in protein-coding genes. The origin of new enzymatic functions, erection of barriers to heterospecific fertilization, and evasion of host response by pathogens, among other things, are thought to be instances of adaptive evolution. Detecting positive selection in protein-coding genes is fraught with difficulties. Saturation for sequence change, codon usage bias, ephemeral selection events and differential selective pressures on amino acids all contribute to the problem. A number of solutions have been proposed with varying degrees of success, however they suffer from limitations of not being accurate enough or being prohibitively computationally intensive. We have developed a character-based method of identifying lineages that undergo positive selection. In our method we assess the possibility that for each internal branch of a phylogenetic tree an event occurred that subsequently gave rise to a greater number of replacement substitutions than might be expected. We classify these replacement substitutions into two categories - whether they subsequently became invariable or changed again in at least one descendent lineage. The former situation indicates that the new character state is under strong selection to preserve its new identity (directional selection), while the latter situation indicates that there is a persistent pressure to change identity (non-directional selection). The method is fast and accurate, easy to implement, sensitive to short-lived selection events and robust with respect to sampling density and proportion of sites under the influence of positive selection.

Algorithms↗

Genetic characterization of the styrene lower catabolic pathway of Pseudomonas sp. strain Y2.

Pseudomonas sp. strain Y2 is a styrene-degrading bacterium, which initiates the catabolism of this compound via its transformation into phenylacetate by the sequential oxidation of the vinyl side chain. The styrene upper catabolic gene cluster (sty genes) had been localized in a 9.2-kb chromosomal region. This report describes the isolation, sequencing and analysis of an adjacent 20.5-kb chromosomal region that contains the genes of the styrene lower degradative pathway (paa genes), which are involved in the transformation of phenylacetate into aliphatic compounds that can enter the Krebs cycle. Hence, Pseudomonas sp. strain Y2 becomes the first microorganism whose entire styrene catabolic cluster has been completely characterized. Analysis of the paa gene cluster has revealed the presence of 17 open reading frames as well as gene duplications and gene reorganizations that are absent in other phenylacetate catabolic clusters described so far. The functionality of these genes has been proved by means of both complementation experiments on Pseudomonas putida mutants and in vitro enzymatic assays. Moreover, a DNA cassette encoding the whole styrene lower pathway has been constructed and has been used to expand the ability of Pseudomonas strains to degrade phenylacetic acid. For the first time, two functional phenylacetate-CoA ligases have been identified in an aerobic phenylacetic acid degradation pathway. Although the upper and lower styrene catabolic clusters are adjacent in the Pseudomonas sp. strain Y2 chromosome, their particular base composition and codon usage suggest a distinct evolutionary history.

Base Sequence↗

Bacteriophage B103: complete DNA sequence of its genome and relationship to other Bacillus phages.

The genome of Bacillus subtilis bacteriophage B103 consists of double-stranded linear DNA 18,630 bp long. The DNA was sequenced, and the sequence was compared with DNA sequences of closely related phages, namely the members of the phage phi29 family. Among them, phage Nf was shown to be the most closely related to B103. Comparisons of several open reading frames (ORFs) among the family members helped to identify genes 1 and 5. A cluster of ORFs between genes 16 and 17 contains two ORFs with partial homology with two phi29 ORFs located in the same region. There are three more ORFs in this region of B103 with good ribosome binding sites (RBS) and optimal codon usage that are not homologous to any of the phi29 ORFs. The function of these five ORFs remains unexplained. It was shown that major promoters characterized in phi29 are retained in B103. Where many substitutions occur in the vicinity of a promoter, at least the -10 and -35 boxes are conserved.

Amino Acid Sequence↗

Cloning and heterologous expression of Entamoeba histolytica adenylate kinase and uridylate/cytidylate kinase.

We have isolated two cDNA clones encoding Entamoeba histolytica nucleotide kinases, EhAK and EhUK, expressed them in E. coli and performed functional studies of the recombinant enzymes. Nucleotide sequence analysis showed that EhAK and EhUK genes exhibited the features characteristic of E. histolytica genes, such as transcripts with relatively short 5' and 3' untranslated flanking regions containing the conserved E. histolytica transcription promoter elements located 5' to the initiation codon and a polyadenylation signal in the 3' UTR, a distinctive codon usage bias for A or T in the third position and an AT bias greater than 75% in the flanking regions of the transcripts. At the protein level, both enzymes belong to the short variant nucleoside monophosphate (NMP) kinases, which lack a 29amino acid LID region present in the long variant isoenzymes. EhAK was 30-38% identical to the members of the adenylate kinase (AK) family while EhUK was more similar (48-49% identity) to UMP/CMP kinases. Both enzymes used ATP as preferred phosphate-group donor but each one exhibited strict specificity for the acceptor NMP, EhAK for AMP and EhUK for the pyrimidine nucleoside monophosphates UMP and CMP. Biochemical characterization of the enzymes and phylogenetic reconstruction showed that EhUK is an authentic and well conserved member of the UMP/CMP kinase group while EhAK is the most divergent member known of the AK1 isoenzymes.

Adenylate Kinase↗

Cloning and expression of the gene coding for FtsH protease from Mycobacterium tuberculosis H37Rv.

This study was aimed at the molecular cloning and expression of the gene coding for FtsH protease of Mycobacterium tuberculosis H37Rv (virulent). PCR on the genomic DNA of M. tuberculosis H37Ra (non-virulent) using the oligodeoxynucleotide primers, which were designed based on the codon usage pattern of M. tuberculosis and against the nucleotide (nt) sequence corresponding to two conserved domains of the FtsH protein of Escherichia coli, yielded a 363-bp product. The amino-acid sequence, deduced from the nt sequence of the PCR product, revealed the presence of two ATP-binding motifs and the AAA Signature motif (Second Region of Homology) that are characteristic features found conserved in the FtsH molecules from eubacteria, archaebacteria, and eukaryotes. Southern hybridisation of the NheI digest of the cosmid SCY6F7 containing part of the genomic DNA of M. tuberculosis H37Rv using the PCR fragment as the probe identified the full-length ftsH gene in the 7.2-kb fragment. The gene was subcloned into pBS (SK+) vector, and the FtsH product that was expressed in E. coli transformed with the vector was identified as an 85-kDa protein localised in the membrane.

ATP-Dependent Proteases↗

A compact gene cluster in Drosophila: the unrelated Cs gene is compressed between duplicated amd and Ddc.

Cs, a gene with unknown function, and amd and Ddc, which encode decarboxylases, are among the most closely spaced genes in D. melanogaster. Untranslated 3' ends of the convergently transcribed genes Cs and Ddc are known to overlap by 88bp. A number of questions arise about the organization of this tightly-packed gene region and about the evolution and function of the Cs gene. We have now investigated this three-gene cluster in Scaptodrosophila lebanonensis (which diverged from D. melanogaster 60-65 MYA), as well as in D. melanogaster and D. simulans. Gene order and direction of transcription is the same in all three species. The Cs gene codes, in Scaptodrosophila, for a polypeptide of 544 amino acids; in D. melanogaster, it consists of 504 amino acids, which is twice as long as previously suggested, which makes the gene density even more spectacular. The Cs sequences exhibit higher number of non-synonymous substitutions between species, higher ratios of non-synonymous to synonymous substitutions, and lower codon usage bias than other genes, suggesting that Cs is less functionally constrained than the other genes. This is consistent with the failure of inducing phenotypic mutations in D. melanogaster. The function of Cs remains to be identified, but a high degree of similarity indicates that it is homologous to genes coding for a corticosteroid-binding protein in yeast and a polyamine oxidase in maize.

Amino Acid Sequence↗

Identification of the gene immediately downstream of the murine INK4a/ARF locus.

The tumor suppressor gene ARF is formed by three exons, namely exons 1 beta, 2 and 3. Here, we show that embryo fibroblasts from mice genetically deficient in exons 2 and 3 (Delta 2,3) express a transcript formed by exon 1 beta followed by the 3'-terminal exon of the gene immediately downstream of the INK4a/ARF locus, which we have called NTp16 (Next-To-p16). The chimeric ARF-NTp16 transcript is not detectable in wild-type fibroblasts but its expression level in Delta 2,3 fibroblasts is 30% compared to the level of the normal ARF transcript in wild-type cells. Expression of the ARF-NTp16 transcript in Delta 2,3 cells is subject to normal regulatory features, such as upregulation by the accumulation of cell doublings, and by the presence of oncogenic Ras or E1a. The chimeric ARF-NTp16 transcript has the potential to encode a 17kDa peptide; however, this peptide is not accumulated in cells at detectable levels, probably reflecting poor codon usage or protein instability. We conclude that Delta 2,3 cells do not retain ARF functionality, at least to a significant extent. Interestingly, the expression pattern of the full-length NTp16 gene is altered in several tissues by the presence of the Delta 2,3 mutation. Finally, these data identify the gene immediately downstream of the INK4a/ARF locus, a region that has been previously proposed to contain another tumor suppressor different from the INK4a/ARF genes.

Amino Acid Sequence↗