Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Analysis of a 14-kb fragment containing a putative cell wall gene and a candidate for the ARA1, arabinose kinase, gene from chromosome IV of Arabidopsis thaliana.

An Arabidopsis thaliana genomic DNA fragment of 14kb has been characterized in the framework of the E.S.S.A. programme. Computational and molecular approaches identified three novel gene sequences coding, respectively, for a protein of unknown function, a putative membrane-anchored cell wall protein and an arabinose kinase gene corresponding to the locus ARA1. The latter two genes named AtSEB1 and AtISA1 have been characterized in detail. They are very different in their organization, codon usage and level of expression. Homologues of AtSEB1 and AtISA1 have been identified. Sequence comparisons showed that the former genes contained a long 5' extension coding for an N-terminal domain probably specifying subcellular localization. Cloning and sequencing of the cognate cDNA for the AtISA1 homologue in A. thaliana, named GAL1, indicate that it encodes for a galactokinase-like protein. Our results highlight the integrative outcome of a systematic sequencing project in which links between biochemically and genetically characterized mutants, ESTs and genomic sequence data are generated.

Amino Acid Sequence↗

Ratios of radical to conservative amino acid replacement are affected by mutational and compositional factors and may not be indicative of positive Darwinian selection.

The ratio of radical to conservative amino acid replacements is frequently used to infer positive Darwinian selection. This method is based on the assumption that radical replacements are more likely than conservative replacements to improve the function of a protein. Therefore, if positive selection plays a major role in the evolution of a protein, one would expect the radical-conservative ratio to exceed the expectation under neutrality. Here, we investigate the possibility that factors unrelated to selection, i.e., transition-transversion ratio, codon usage, genetic code, and amino acid composition, influence the radical-conservative replacement ratio. All factors that have been studied were found to affect the radical-conservative replacement ratio. In particular, amino acid composition and transition-transversion ratio are shown to have the most profound effects. Because none of the studied factors had anything to do with selection (positive or otherwise) and also because all of them (singly or in combination) affected a measure that was supposed to be indicative of positive selection, we conclude that selectional inferences based on radical-conservative replacement ratios should be treated with suspicion.

Amino Acid Substitution↗

The molecular phylogeny of five eastern north Pacific octopus species.

The DNA sequence of a 612-nucleotide fragment of the mitochondrial cytochrome oxidase subunit III gene from five Octopus species has been determined. The COIII gene in these species shows an extreme bias against G in the DNA sense strand with a moderately low C composition. The bias against G and C severely restricts the codon usage in Octopus COIII genes. The aligned DNA sequences were subjected to distance, maximum-likelihood, and parsimony analyses to ascertain the phylogenetic relationship of the species. The results of all of the analyses were concordant. The analyses indicate that O. bimaculoides and O. bimaculatus are the least diverged of the species and fall into a separate clade from O. dolfleini and O. californicus, which are also closely related. O. rubescens is about equally removed from the other species, but parsimony, distance, maximum-likelihood, and logdet analyses suggest that it is more closely aligned with the O. bimaculoides/O. bimaculatus lineage.

Amino Acid Sequence↗

Optimisation of secretion of recombinant HBsAg virus-like particles: Impact on the development of HIV-1/HBV bivalent vaccines.

The hepatitis B surface antigen (HBsAg) assembles into virus-like particles (VLPs) that can be used as carrier of immunogenic peptides for the development of bivalent vaccine candidates. It is shown here that by respecting certain qualitative features of mammalian preS1 and preS2 protein domains upstream of HBsAg, foreign sequences can be inserted in their place while maintaining efficient secretion of VLPs. A polyepitope bearing HIV-1 epitopes restricted to the HLA-A*0201 class I allele was optimised for secretion as an HBsAg fusion protein by counterbalancing the generally hydrophobic class I epitopes with hydrophilic spacers, eliminating epitopes bearing cysteine residues, limiting the number of internal methionine residues to a minimum and adopting Homo sapiens codon usage. The optimised HIV-1 polyepitope-HBsAg recombinant protein with up to 138 residues assembled into efficiently secreted recombinant VLPs. DNA immunisation in HLA-A*0201 and HLA-A*0201/HLA-DR1 transgenic mice resulted in the recovery of humoral response against the carrier and enhanced levels of HIV-1 specific CD8(+) T lymphocyte activation. Efficient self-assembly of recombinant HBsAg VLPs opens up the possibility of making efficient bivalent HBV/HIV vaccine candidates, which is particularly apposite given that the two viruses are frequently associated.

AIDS Vaccines↗

Cloning and sequencing of the homologues of both the bacterial and eukaryotic initiation factor genes (hIF-2 and heIF-2 gamma) from archaeal Halobacterium halobium.

The cloning and sequencing of the genes encoding the translational initiation factors (hIF-2 and heIF-2 gamma) was performed by screening the halophilic archaeon Halobacterium halobium genomic library with a probe constructed from the peptide IGHVDHGK that is conserved in archaeal GTP-binding elongation factors. The codon usage by the hIF-2 and heIF-2 gamma genes showed a preference for triplets ending in G or C. This characteristic is almost identical to that of other H. halobium genes. The translated protein of hIF-2 and heIF-2 gamma genes is made of 414 and 583 amino acid residues, respectively, and contains the sequence motif for the binding of GTP. The sequence of hIF-2 shows a strong similarity to the initiation factor IF-2 from Bacteria whereas heIF-2 gamma shows a strong similarity to the initiation factor eIF-2 gamma from Eucarya.

Amino Acid Sequence↗

OGRe: a relational database for comparative analysis of mitochondrial genomes.

Organellar Genome Retrieval (OGRe) is a relational database of complete mitochondrial genome sequences for over 250 Metazoan species. OGRe provides a resource for the comparative analysis of mitochondrial genomes at several levels. At the sequence level, OGRe allows the retrieval of any selected set of mitochondrial genes from any selected set of species. Species are classified using a taxonomic system that allows easy selection of related groups of species. Sequence alignments are also available for some species. At the level of individual nucleotides, the system contains information on base frequencies and codon usage frequencies that can be compared between organisms. At the level of whole genomes, OGRe provides several ways of visualizing information on gene order. Diagrams illustrating the genome arrangement can be generated for any selected set of species automatically from the information in the database. Searches can be done based on gene arrangement to find sets of species that have the same order as one another. Diagrams for pairwise comparison of species can be produced that show the positions of break-points in the gene order and use colour to highlight the sections of the genome that have moved. OGRe is available from http://www.bioinf.man.ac.uk/ogre.

Animals↗

Sequence and phylogenetic analysis of the complete mitochondrial genome of the flour beetle Tribolium castanaeum.

We describe the first complete mitochondrial genome sequence from a representative of the insect order Coleoptera, the flour beetle Tribolium castaneum. The 15,881 bp long Tribolium mitochondrial genome encodes 13 putative proteins, two ribosomal RNAs and 22 tRNAs canonical for animal mitochondrial genomes. Their arrangement is identical to that in Drosophila melanogaster, which is considered ancestral for insects and crustaceans (Boore et al., 1998; Hwang, et al., 2001a). Nucleotide composition, amino acid composition, and codon usage fall within the range of values observed in other insect mitochondrial genomes. Most notable features are the use of TCT as tRNA(Ser(AGN)) anticodon instead of GCT, which is used in most other arthropod species, and the relative scarcity of special sequence motifs in the 1431 bp long control region. Phylogenetic analysis confirmed resolving power in the conserved regions of the mitochondrial proteome regarding diversification events, which predate the emergence of pterygote insects, while little resolution was obtained at the level of basal perygote diversification. The partition of faster evolving amino acid sites harbored strong support for joining Lepidoptera with Diptera, which is consistent with a monophyletic Mecopterida.

Animals↗

A microarray-based comparative analysis of gene expression profiles during grain development in transgenic and wild type wheat.

Global, comparative gene expression analysis is potentially a very powerful tool in the safety assessment of transgenic plants since it allows for the detection of differences in gene expression patterns between a transgenic line and the mother variety. In the present study, we compared the gene expression profile in developing seeds of wild type wheat and wheat transformed for endosperm-specific expression of an Aspergillus fumigatus phytase. High-level expression of the phytase gene was ensured by codon modification towards the prevalent codon usage of wheat genes and by using the wheat 1DX5HMW glutenin promoter for driving transgene expression. A 9K wheat unigene cDNA microarray was produced from cDNA libraries prepared mainly from developing wheat seed. The arrays were hybridised to flourescently labelled cDNA prepared from developing seeds of the transgenic wheat line and the mother variety, Bobwhite, at three developmental stages. Comparisons and statistical analyses of the gene expression profiles of the transgenic line vs. that of the mother line revealed only slight differences at the three developmental stages. In the few cases where differential expression was indicated by the statistical analysis it was primarily genes that were strongly expressed over a shorter interval of seed development such as genes encoding storage proteins. Accordingly, we interpret these differences in gene expression levels to result from minor asynchrony in seed development between the transgenic line and the mother line. In support of this, real time PCR validation of results from selected genes at the late developmental stage could not confirm differential expression of these genes. We conclude that the expression of the codon-modified A. fumigatus phytase gene in the wheat seed had no significant effects on the overall gene expression patterns in the developing seed.

6-Phytase↗

Cloning and sequencing of the gene encoding thermostable elongation factor 2 in Sulfolobus solfataricus.

The gene (aEF-2) coding for the translation elongation factor 2 (aEF-2) in the thermoacidophilic archaebacterium, Sulfolobus solfataricus, has been cloned and sequenced. The deduced primary structure of aEF-2 is composed of 735 amino acids (aa), excluding the Met start residue. There are no Cys residues and the calculated M(r) is 81,699. In the coding region of aEF-2, the high A + T content greatly influences the codon usage. From the alignment of the primary structure of aEF-2 with that of the analogous factors from the three kingdoms, aa identities were derived. The greatest identity (82%) was found with EF-2 from Sulfolobus acidocaldarius; lower values were observed with other archaebacterial EF-2 (45-47%), eukaryotic EF-2 (38-40%) and with the functional eubacterial analogue EF-G (28-31%). aEF-2 possesses the consensus sequences required for a GTP-binding protein and the four regions which are supposed to be involved in the functional regulation of EF-2/EF-G. These data should have phylogenetic implications.

Amino Acid Sequence↗

Primary structure of heat-labile enterotoxin produced by Escherichia coli pathogenic for humans.

Heat-labile enterotoxin of Escherichia coli pathogenic for humans (LTh) or for piglets (LTp) and Vibrio cholerae enterotoxin (CT) are structurally and functionally similar toxins. We have determined the complete nucleotide sequence of the toxA gene which encodes the subunit A of LTh (LTh A). The deduced amino acid sequence consists of 258 residues including a signal peptide of 18 residues. According to the previously completed LTh B sequence (103 residues), the predicted holotoxin (1A5B) of LTh comprises 755 residues and has Mr = 87,866. With respect to LTh A and LTh B, secondary structures, local hydrophilicity, and sites for antigenic determinants were predicted. Both codon usage and G + C content of the toxA gene and the LTh B gene (toxB) were markedly different from those observed with several E. coli chromosomal genes. Its relatively low G + C content was rather close to that of the V. cholerae chromosome. Although the toxA gene shares a common ancestor with the LTp A gene (eltA), the two genes are apparently distinguishable from each other in their sequences. Like LTh B reported previously, the predicted sequence of the catalytic fragment LTh A1 also showed more homology to that of CT A1 than did that of LTp A1. In contrast, unique sequences were found in LTh A2.

Amino Acid Sequence↗

The mitochondrial plasmid from Neurospora intermedia strain Labelle-1b contains a long open reading frame with blocks of amino acids characteristic of reverse transcriptases and related proteins.

We have determined the DNA sequence of the 4070 base pair mitochondrial plasmid from the Labelle-1b strain of Neurospora intermedia. Analysis of the sequence revealed that the plasmid contains a long open reading frame (ORF) that could encode a protein of up to 1151 amino acids. Codon usage in the long ORF shows no clear relationship to Neurospora mitochondrial genes, nuclear genes, nor to the ORF of a different Neurospora mitochondrial plasmid. The long ORF contains regions of similarity to yeast mitochondrial RNA polymerase as well as blocks of amino acids that are characteristic of reverse transcriptases and the ORFs of certain group II mtDNA introns (Michel and Lang, (1985) Nature 316,641). The plasmid gives rise to specific transcripts, some of which may be unit length, and which carry the information for expression of the long ORF. The genetic organization and content of the plasmid suggest that it is related to mobile genetic elements.

Amino Acid Sequence↗

The complete sequence of the mitochondrial genome of Daphnia pulex (Cladocera: Crustacea).

The sequence of the mitochondrial DNA (mtDNA) of the branchiopod crustacean Daphnia pulex has been completed. It is 15333bp with an A+T content of 62.3%, and contains the typical complement of 13 protein-coding, 22 transfer RNA (tRNA) and two ribosomal RNA (rRNA) genes. Comparison of this sequence with the sequences of the other eight completely sequenced arthropod mtDNAs showed that gene order and orientation are identical to that of Drosophila but different from Artemia due to the rearrangement of two tRNA genes. Nucleotide composition, codon usage, and amino acid composition are very similar in the crustaceans, but divergent from insects and chelicerates which show a much higher bias towards A+T. However, with few exceptions, the mitochondrial proteins of Daphnia are more similar to those of the dipteran insects (Drosophila and Anopheles) than to those of Artemia, at both the nucleotide and amino acid levels, suggesting that Artemia mtDNA is evolving at an accelerated rate. These results also show that sequence evolution and the evolution of nucleotide composition can be decoupled. Analysis of nucleotide substitution patterns in COII showed that there has been an unbiased acceleration of the overall substitution rate in Artemia. In contrast, the accelerated substitution rate in Apis is due partly to extreme A+T mutation pressure. Secondary structures are proposed for the Daphnia tRNAs and rRNAs. The tRNAs are similar to those of other arthropods but tend to have TPsiC arms that are only 4bp long. The rRNA secondary structures are similar to those proposed for insects except for the absence of a small number of helices in Daphnia. Phylogenetic analysis of second codon positions grouped Daphnia with Artemia, as expected, despite the latter's accelerated divergence rate. In contrast, the unusual pattern of mtDNA divergence in Apis led to a topology in which the holometabolous insects (Anopheles, Drosophila, Apis) appeared to be paraphyletic with respect to the hemimetabolous insect, Locusta, due to the early branching of Apis.

Animals↗

Transposable element orientation bias in the Drosophila melanogaster genome.

Nonrandom distributions of transposable elements can be generated by a variety of genomic features. Using the full D. melanogaster genome as a model, we characterize the orientations of different classes of transposable elements in relation to the directionality of genes. DNA-mediated transposable elements are more likely to be in the same orientation as neighboring genes when they occur in the nontranscribed region's that flank genes. However, RNA-mediated transposable elements located in an intron are more often oriented in the direction opposite to that of the host gene. These orientation biases are strongest for genes with highly biased codon usage, probably reflecting the ability of such loci to respond to weak positive or negative selection. The leading hypothesis for selection against transposable elements in the coding orientation proposes that transcription termination poly(A) signal motifs within retroelements interfere with normal gene transcription. However, after accounting for differences in base composition between the strands, we find no evidence for global selection against spurious transcription termination signals in introns. We therefore conclude that premature termination of host gene transcription due to the presence of poly(A) signal motifs in retroelements might only partially explain strand-specific detrimental effects in the D. melanogaster genome.

Animals↗

The complete mitochondrial genome of the articulate brachiopod Terebratalia transversa.

We sequenced the complete mitochondrial DNA (mtDNA) of the articulate brachiopod Terebratalia transversa. The circular genome is 14,291 bp in size, relatively small compared with other published metazoan mtDNAs. The 37 genes commonly found in animal mtDNA are present; the size decrease is due to the truncation of several tRNA, rRNA, and protein genes, to some nucleotide overlaps, and to a paucity of noncoding nucleotides. Although the gene arrangement differs radically from those reported for other metazoans, some gene junctions are shared with two other articulate brachiopods, Laqueus rubellus and Terebratulina retusa. All genes in the T. transversa mtDNA, unlike those in most metazoan mtDNAs reported, are encoded by the same strand. The A+T content (59.1%) is low for a metazoan mtDNA, and there is a high propensity for homopolymer runs and a strong base-compositional strand bias. The coding strand is quite G+T-rich, a skew that is shared by the confamilial (laqueid) species L. rubellus but is the opposite of that found in T. retusa, a cancellothyridid. These compositional skews are strongly reflected in the codon usage patterns and the amino acid compositions of the mitochondrial proteins, with markedly different usages being observed between T. retusa and the two laqueids. This observation, plus the similarity of the laqueid noncoding regions to the reverse complement of the noncoding region of the cancellothyridid, suggests that an inversion that resulted in a reversal in the direction of first-strand replication has occurred in one of the two lineages. In addition to the presence of one noncoding region in T. transversa that is comparable with those in the other brachiopod mtDNAs, there are two others with the potential to form secondary structures; one or both of these may be involved in the process of transcript cleavage.

Amino Acid Sequence↗

ANGLE: a sequencing errors resistant program for predicting protein coding regions in unfinished cDNA.

In the process of making full-length cDNA, predicting protein coding regions helps both in the preliminary analysis of genes and in any succeeding process. However, unfinished cDNA contains artifacts including many sequencing errors, which hinder the correct evaluation of coding sequences. Especially, predictions of short sequences are difficult because they provide little information for evaluating coding potential. In this paper, we describe ANGLE, a new program for predicting coding sequences in low quality cDNA. To achieve error-tolerant prediction, ANGLE uses a machine-learning approach, which makes better expression of coding sequence maximizing the use of limited information from input sequences. Our method utilizes not only codon usage, but also protein structure information which is difficult to be used for stochastic model-based algorithms, and optimizes limited information from a short segment when deciding coding potential, with the result that predictive accuracy does not depend on the length of an input sequence. The performance of ANGLE is compared with ESTSCAN on four dataset each of them having a different error rate (one frame-shift error or one substitution error per 200-500 nucleotides) and on one dataset which has no error. ANGLE outperforms ESTSCAN by 9.26% in average Matthews's correlation coefficient on short sequence dataset (< 1000 bases). On long sequence dataset, ANGLE achieves comparable performance.

Algorithms↗

Comparative analysis of genes encoding methyl coenzyme M reductase in methanogenic bacteria.

The sequence of the gene cluster encoding the methyl coenzyme M reductase (MCR) in Methanococcus voltae was determined. It contains five open reading frames (ORF), three of which encode the known enzyme subunits. Putative ribosome binding sites were found in front of all ORFs. They differ in their degrees of complementarity to the 3' end of the 16 S rRNA, which is discussed in terms of different translation efficiencies of the respective genes. The codon usage bias is different in the subunit encoding genes compared with the two other ORFs in the cluster and two other known genes of Mc. voltae. This is interpreted in terms of increased translational accuracy of the highly expressed MCR subunit genes. The derived polypeptide sequences encoded by the five ORFs of the MCR cluster were compared to those of the respective genes in Methanobacterium thermoautotrophicum Marburg and Methanosarcina barkeri. Conserved regions were detected in the enzyme subunits, which are candidates for factor binding domains. Conserved hydrophobic sequences found in the alpha and beta subunits are discussed with respect to the membrane association of the enzyme.

Amino Acid Sequence↗

Removal of a cryptic intron and subcellular localization of green fluorescent protein are required to mark transgenic Arabidopsis plants brightly.

The green fluorescent protein (GFP) from the jellyfish Aequorea victoria is finding wide use as a genetic marker that can be directly visualized in the living cells of many heterologous organisms. We have sought to express GFP in the model plant Arabidopsis thaliana, but have found that proper expression of GFP is curtailed due to aberrant mRNA processing. An 84-nt cryptic intron is efficiently recognized and excised from transcripts of the GFP coding sequence. The cryptic intron contains sequences similar to those required for recognition of normal plant introns. We have modified the codon usage of the gfp gene to mutate the intron and to restore proper expression in Arabidopsis. GFP is mainly localized within the nucleoplasm and cytoplasm of transformed Arabidopsis cells and can give rise to high levels of fluorescence, but it proved difficult to efficiently regenerate transgenic plants from such highly fluorescent cells. However, when GFP is targeted to the endoplasmic reticulum, transformed cells regenerate routinely to give highly fluorescent plants. These modified forms of the gfp gene are useful for directly monitoring gene expression and protein localization and dynamics at high resolution, and as a simply scored genetic marker in living plants.

Agrobacterium tumefaciens↗

Comparative complete genome sequence analysis of the amino acid replacements responsible for the thermostability of Corynebacterium efficiens.

Corynebacterium efficiens is the closest relative of Corynebacterium glutamicum, a species widely used for the industrial production of amino acids. C. efficiens but not C. glutamicum can grow above 40 degrees C. We sequenced the complete C. efficiens genome to investigate the basis of its thermostability by comparing its genome with that of C. glutamicum. The difference in GC content between the species was reflected in codon usage and nucleotide substitutions. Our comparative genomic study clearly showed that there was tremendous bias in amino acid substitutions in all orthologous ORFs. Analysis of the direction of the amino acid substitutions suggested that three substitutions are important for the stability of the C. efficiens proteins: from lysine to arginine, serine to alanine, and serine to threonine. Our results strongly suggest that the accumulation of these three types of amino acid substitutions correlates with the acquisition of thermostability and is responsible for the greater GC content of C. efficiens.

Amino Acid Sequence↗