Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Translational control of oskar generates short OSK, the isoform that induces pole plasma assembly.

At the posterior pole of the Drosophila oocyte, oskar induces a tightly localized assembly of pole plasm. This spatial restriction of oskar activity has been thought to be achieved by the localization of oskar mRNA, since mislocalization of the RNA to the anterior induces anterior pole plasm. However, ectopic pole plasm does not form in mutant ovaries where oskar mRNA is not localized, suggesting that the unlocalized mRNA is inactive. As a first step towards understanding how oskar activity is restricted to the posterior pole, we analyzed oskar translation in wild type and mutants. We show that the targeting of oskar activity to the posterior pole involves two steps of spatial restriction, cytoskeleton-dependent localization of the mRNA and localization-dependent translation. Furthermore, our experiments demonstrate that two isoforms of Oskar protein are produced by alternative start codon usage. The short isoform, which is translated from the second in-frame AUG of the mRNA, has full oskar activity. Finally, we show that when oskar RNA is localized, accumulation of Oskar protein requires the functions of vasa and tudor, as well as oskar itself, suggesting a positive feedback mechanism in the induction of pole plasm by oskar.

Animals↗

Nucleotide sequence of the coding portion of human alpha globin messenger RNA.

The nucleotide sequence of the coding portion of human alpha globin mRNA has been determined by sequence analysis using human alpha globin cDNA cloned in bacterial plasmids. The sequence was obtained by a combination of direct sequence analysis of the cloned cDNA and analysis of cDNA obtained by primer extension, using short restriction endonuclease fragments of cloned alpha cDNA that were hybridized to human globin mRNA and elongated on the mRNA template by viral reverse transcriptase. The human alpha globin mRNA has an unexpectedly high G + C base composition (64.7%), similar to that observed for rabbit globin alpha mRNA, and displays a striking bias in the use of synonym codons for various amino acids. The bias in codon usage of human alpha globin mRNA is similar, with some exceptions, to that previously observed for rabbit alpha globin mRNA as well as for human and rabbit beta globin mRNAs. A detailed restriction endonuclease map of the human alpha globin cDNA is presented.

Amino Acid Sequence↗

Nucleotide sequence of the genomic region encompassing Adh and Adh-dup genes of D. lebanonensis (Scaptodrosophila): gene expression and evolutionary relationships.

The region of the genome of D. lebanonensis that contains the Adh gene and the downstream Adh-dup gene was sequenced. The structure of the two genes is the same as has been described for D. melanogaster. Adh has two promoters and Adh-dup has only one putative promoter. The levels of expression of the two genes in this species are dramatically different. Hybridizing the same Northern blots with a specific probe for Adh-dup, we did not find transcripts for this gene in D. lebanonensis. The level of Adh distal transcript in adults of D. lebanonensis is five times greater than that of D. melanogaster adults. The maximum levels of proximal transcript are attained at different larval stages in the two species, being three times higher in D. melanogaster late-second-instar larvae than in D. lebanonensis first-instar larvae. The level of Adh transcripts allowed us to determine distal and proximal initiation transcription sites, the position of the first intron, the use of two polyadenylation signals, and the heterogeneity of polyadenylation sites. Temporal and spatial expression profiles of the Adh gene of D. lebanonensis show qualitative differences compared with D. melanogaster. Adh and Adh-dup evolve differently as shown by the synonymous and nonsynonymous substitution rates for the coding region of both genes when compared across two species of the melanogaster group, two of the obscura group of the subgenus Sophophora and D. lebanonensis of the victoria group of the subgenus Scaptodrsophila. Synonymous rates for Adh are approximately half those for Adh-dup, while nonsynonymous rates for Adh are generally higher than those for Adh-dup. Adh shows 76.8% identities at the protein level and 70.2% identities at the nucleotide level while Adh-dup shows 83.7% identities at the protein level and 67.5% identities at the nucleotide level. Codon usage for Adh-dup is shown to be less biased than for Adh, which could explain the higher synonymous rates and the generally lower nonsynonymous substitution rates in Adh-dup compared with Adh. Phylogenetic trees reconstructed by distance matrix and parsimony methods show that Sophophora and Scaptodrosophila subgenera diverged shortly after the separation from the Drosophila subgenus.

Alcohol Dehydrogenase↗

Endosymbiotic origin and codon bias of the nuclear gene for chloroplast glyceraldehyde-3-phosphate dehydrogenase from maize.

The nuclei of plant cells harbor genes for two types of glyceraldehyde-3-phosphate dehydrogenases (GAPDH) displaying a sequence divergence corresponding to the prokaryote/eukaryote separation. This strongly supports the endosymbiotic theory of chloroplast evolution and in particular the gene transfer hypothesis suggesting that the gene for the chloroplast enzyme, initially located in the genome of the endosymbiotic chloroplast progenitor, was transferred during the course of evolution into the nuclear genome of the endosymbiotic host. Codon usage in the gene for chloroplast GAPDH of maize is radically different from that employed by present-day chloroplasts and from that of the cytosolic (glycolytic) enzyme from the same cell. This reveals the presence of subcellular selective pressures which appear to be involved in the optimization of gene expression in the economically important graminaceous monocots.

Amino Acid Sequence↗

Nucleotide sequence and analysis of the lethal factor gene (lef) from Bacillus anthracis.

The nucleotide sequence of the Bacillus anthracis lethal factor (LF) gene (lef) has been determined. LF is part of the tripartite protein exotoxin of B. anthracis along with protective antigen (PA) and edema factor (EF). The apparent ATG start codon, which is located immediately upstream from codons which specify the first 16 amino acids (aa) of the mature secreted LF, is preceded by an AAAGGAG sequence, which is its probable ribosome-binding site. This ATG codon begins a continuous 2427-bp open reading frame which encodes the 809-aa LF-precursor protein with an Mr of 93,798. The mature secreted protein (776 aa; Mr 90,237) was preceded by a 33-aa signal peptide which has characteristics in common with leader peptides for other secreted proteins of the Bacillus species. The codon usage of the LF gene reflects its high (70%) A + T content. The N-terminus of LF (first 300 aa) shared extensive homology with the N-terminus of the anthrax EF protein. Since LF and EF each bind PA at the same site, these homologous regions probably represent their common PA-binding domains.

Amino Acid Sequence↗

crp genes of Shigella flexneri, Salmonella typhimurium, and Escherichia coli.

The complete nucleotide sequences of the Salmonella typhimurium LT2 and Shigella flexneri 2B crp genes were determined and compared with those of the Escherichia coli K-12 crp gene. The Shigella flexneri gene was almost like the E. coli crp gene, with only four silent base pair changes. The S. typhimurium and E. coli crp genes presented a higher degree of divergence in their nucleotide sequence with 77 changes, but the corresponding amino acid sequences presented only one amino acid difference. The nucleotide sequences of the crp genes diverged to the same extent as in the other genes, trp, ompA, metJ, and araC, which are structural or regulatory genes. An analysis of the amino acid divergence, however, revealed that the catabolite gene activator protein, the crp gene product, is the most conserved protein observed so far. Comparison of codon usage in S. typhimurium and E. coli for all genes sequenced in both organisms showed that their patterns were similar. Comparison of the regulatory regions of the S. typhimurium and E. coli crp genes showed that the most conserved sequences were those known to be essential for the expression of E. coli crp.

Amino Acid Sequence↗

Polymorphism of the IGHA gene in sheep.

Genetic variation in immunoglobulin A, the most abundant immunoglobulin in mammalian cells, has not been reported in ruminants. In this study, variation in the immunoglobulin heavy alpha chain constant gene (IGHA) of sheep was investigated by amplification of a fragment that included the hinge coding sequence, followed by single-strand conformational polymorphism (SSCP) analysis and DNA sequencing. Three novel sequences, each characterized by unique SSCP banding patterns, were identified. One or two sequences were detected in individual sheep and all the sequences identified shared high homology to the published ovine and bovine IGHA sequences, suggesting that these sequences represent allelic variants of the IGHA gene in sheep. Sequence alignment showed that these sequences differed mainly in the 3' end of exon 1 and in the coding sequence of the hinge region. There was either a deletion or an insertion of two codons in the hinge coding region in these allelic variants. Codon usage in the hinge coding region was quite different from that in the non-hinge coding regions of the gene, suggesting different evolution of the IGHA hinge sequence. Three novel amino acid sequences of ovine IGHA were also predicted, and variation in these sequences might not only affect antigen recognition but also susceptibility to cleavage by bacterial or parasitic proteases.

Alleles↗

Peptidase D gene (pepD) of Escherichia coli K-12: nucleotide sequence, transcript mapping, and comparison with other peptidase genes.

The nucleotide sequence of a 2.3-kilobase-pair DNA fragment of Escherichia coli that contains the transcription signals and the coding region of the pepD gene specifying aminopeptidase D was determined. The location and extent of the open reading frame were verified by partial amino acid sequencing of the purified pepD product. By use of a promoter-screening vector, initiation signals for pepD transcription were located in the 5'-flanking region of the open reading frame. Analysis of pepD transcripts by S1 mapping, primer extension, and Northern (RNA) hybridization revealed two species of monocistronic mRNA with different 5' ends and a common 3' end. Calculation of the degree of codon usage bias in the coding region suggested that the efficiency of pepD translation is relatively low. As deduced from the predicted amino acid sequence, peptidase D is a slightly hydrophilic protein of 485 amino acid residues that contains no extended domains of marked hydrophobicity. Structural and functional features of the pepD gene are discussed and compared with other already sequenced peptidase genes of E. coli.

Amino Acid Sequence↗

The amino acid sequence of a crystal protein from Bacillus thuringiensis deduced from the DNA base sequence.

We have determined the nucleotide sequence of a 4222-base segment of DNA which contains the promoter, the coding region, and the terminator of a crystal protein gene cloned from a Bacillus thuringiensis plasmid. A sequence of 1176 amino acids encoding a Mr 133,500 peptide was deduced from the single open reading frame. This protein-coding region was analyzed for codon usage, predicted hydropathy, and predicted secondary structure. Examination of the base sequence revealed the presence of several inverted and direct repeats located in both the coding and noncoding regions. S1 nuclease mapping was used to locate the transcription termination point at a site following a potentially very stable stem-and-loop structure.

Amino Acid Sequence↗

Analysis of five presumptive protein-coding sequences clustered between the primosome genes, 41 and 61, of bacteriophages T4, T2, and T6.

In bacteriophage T4, there is a strong tendency for genes that encode interacting proteins to be clustered on the chromosome. There is 1.6 kb of DNA between the DNA helicase (gene 41) and the DNA primase (gene 61) genes of this virus. The DNA sequence of this region suggests that it contains five genes, designated as open reading frames (ORFs) 61.1 to 61.5, predicted to encode proteins ranging in size from 5.94 to 22.88 kDa. Are these ORFs actually genes? As one test, we compared the DNA sequence of this region in bacteriophages T2, T4, and T6 and found that ORFs 61.1, 61.3, 61.4, and 61.5 are highly conserved among the three closely related viruses. In contrast, ORF 61.2 is conserved between phages T4 and T6 yet is absent from phage T2, where it is replaced by another ORF, T2 ORF 61.2, which is not found in the T4 and T6 genomes. As a second, independent test for coding sequences, we calculated the codon base position preferences for all ORFs in this region that could encode proteins that contain at least 30 amino acids. Both the T4/T6 and T2 versions of ORF 61.2, as well as the other ORFs, have codon base position preferences that are indistinguishable from those of known T4 genes (coefficients of 0.81 to 0.94); the six other possible ORFs of at least 90 bp in this region are ruled out as genes by this test (coefficients less than zero). Thus, both evolutionary conservation and codon usage patterns lead us to conclude that ORFs 61.1 to 61.5 represent important protein-coding sequences for this family of bacteriophages. Because they are located between the genes that encode the two interacting proteins of the T4 primosome (DNA helicase plus DNA primase), one or more may function in DNA replication by modulating primosome function.

Amino Acid Sequence↗

Gene prediction using the Self-Organizing Map: automatic generation of multiple gene models.

BACKGROUND: Many current gene prediction methods use only one model to represent protein-coding regions in a genome, and so are less likely to predict the location of genes that have an atypical sequence composition. It is likely that future improvements in gene finding will involve the development of methods that can adequately deal with intra-genomic compositional variation. RESULTS: This work explores a new approach to gene-prediction, based on the Self-Organizing Map, which has the ability to automatically identify multiple gene models within a genome. The current implementation, named RescueNet, uses relative synonymous codon usage as the indicator of protein-coding potential. CONCLUSIONS: While its raw accuracy rate can be less than other methods, RescueNet consistently identifies some genes that other methods do not, and should therefore be of interest to gene-prediction software developers and genome annotation teams alike. RescueNet is recommended for use in conjunction with, or as a complement to, other gene prediction methods.

Chromosome Mapping↗

Isolation and characterization of a beta-tubulin gene from Candida albicans.

We report the isolation and nucleotide sequence determination of a beta-tubulin gene (TUB2) from the pathogenic dimorphic fungus Candida albicans. Nucleotide sequence analysis revealed that TUB2 encodes a protein of 449 amino acids (aa) with considerable sequence homology to beta-tubulins isolated from other fungal species. The nucleotide sequence of the C. albicans gene is 70% homologous to that of the Saccharomyces cerevisiae gene. The coding region for the C. albicans beta-tubulin gene is interrupted by two introns. The first intron occurs after the 4th aa and the second intron occurs after the 13th aa. A comparison with other fungal beta-tubulin genes indicates that the intron locations are highly conserved. Codon usage in the C. albicans TUB2 gene is nonrandom, as has been observed for other fungal beta-tubulin genes. The C. albicans TUB2 gene is transcribed to yield a 1.8-kb mRNA species. On the basis of genomic Southern-blot analysis, we conclude that C. albicans most likely possesses a single beta-tubulin gene.

Amino Acid Sequence↗

PCR-RFLP and sequence analysis of a non-ribosomal fragment for genetic characterization of European stone fruit yellows phytoplasmas infecting various Prunus species.

A 927 bp non-ribosomal fragment was used to assess the genetic variability of the European stone fruit yellows (ESFY) phytoplasma infecting 14 different Prunus species. For this, 175 isolates originating from four different Mediterranean countries were tested by PCR-RFLP analysis with seven restriction enzymes. No polymorphism among the ESFY phytoplasma could be observed but 12 out of 18 restriction sites differed between the homologous fragments of ESFY and apple proliferation (AP) phytoplasmas. An 846 bp fragment of a French ESFY isolate was sequenced, it included the 3'-end of a putative nitroreductase gene, an intergenic region and a truncated open reading frame. This ESFY phytoplasma sequence showed 89.7% identity with the equivalent AP phytoplasma nucleotide sequence (83. 9% identity at the amino acid level). The G+C content of the entire sequence was extremely low (15.4%) and A+T-rich codons were highly preferred in codon usage. In this paper, we report the presence of the ESFY phytoplasma for the first time in Turkey and in five Prunus hosts never reported previously. Our results also indicate that the ESFY phytoplasma isolates affecting various Prunus species are genetically homogenous but can be distinguished from the AP phytoplasma. Therefore, they are likely to represent different taxons.

Amino Acid Sequence↗

Rates of mitochondrial DNA evolution in sharks are slow compared with mammals.

The rate of mitochondrial DNA (mtDNA) evolution has been carefully calibrated only in primates. Similarity between the primate calibration and rates estimated for other vertebrates has led to widespread assumption of a constant molecular clock in vertebrates even though this has never been rigorously tested. We report here the examination of mtDNA sequence variation for 13 species of sharks from two orders that are well represented in the fossil record to test the constancy hypothesis. Nucleotide substitution rates in the cytochrome b and cytochrome oxidase I genes in sharks are seven- to eightfold slower than in primates or ungulates. This difference in substitution rate cannot be explained by nucleotide composition bias, codon-usage bias, selection, or choice of genes sequenced, and was confirmed by comparing species recently separated by the rise of the Isthmus of Panama. Such differences in mtDNA substitution rates among taxa indicate that it is inappropriate to use a calibration for one group to estimate divergence times or demographic parameters for another group. High-resolution studies of molecular evolutionary rates require taxon-specific calibrations.

Animals↗

Molecular features of mollicutes.

It is now firmly established that the mollicutes are true eubacteria. They have evolved regressively (i.e., by genome reduction) from gram-positive bacterial ancestors with a low content of guanine plus cytosine in DNA--more specifically, from certain clostridia. Many of their properties, such as small genome size, small number of rRNA operons and tRNA genes, lack of a cell wall, fastidious growth, and limited metabolic activities, are seen as the result of this evolution. Other properties, such as the anaerobiosis of their earliest evolving members (anaeroplasmas and asteroleplasmas), the high adenine-plus-thymine content of their DNA, their lack of sensitivity to rifampin, and the regulatory signals for the transcription of their DNA, have been inherited from their eubacterial ancestors. However, the mollicutes are not simply wall-less gram-positive bacteria. They have properties of their own. High adenine-thymine pressure has resulted in a particular codon usage, where, for instance, UGA is read as tryptophan and not as stop. These organisms occupy unique ecological niches and have developed peculiar systems for pathogenicity, cell adhesion, antigenic variation, and (in the case of the spiroplasmas) helical morphology and motility. The putative role of certain mollicutes as cofactors in the development of AIDS may involve their mitogenicity, their superantigenicity, and their ability to induce cytokines.

Bacterial Outer Membrane Proteins↗

NRSub: a non-redundant data base for the Bacillus subtilis genome.

We have organized the DNA sequences of Bacillus subtillis from the EMBL collection to build the NRSub data base. This data base is free from duplications and all detected overlapping sequences are merged into contigs. Data on gene mapping and codon usage are also included. NRSub is publically available through anonymous FTP in flat file format or structured on the form of an ACNUC data base. Under this format, it is possible to use NRSub with the retrieval program Query--win. This program integrates a graphical interface and may be installed on any kind of UNX computer under X Window and on which the Vibrant and Motif libraries are available.

Bacillus subtilis↗

Duplication-induced mutation of a new Neurospora gene required for acetate utilization: properties of the mutant and predicted amino acid sequence of the protein product.

A cloned Neurospora crassa genomic sequence, selected as preferentially transcribed when acetate was the sole carbon source, was introduced in extra copies at ectopic loci by transformation. Sexual crossing of transformants yielded acetate nonutilizing mutants with methylation and restriction site changes within both the ectopic DNA and the normally located gene. Such changes are typical of the duplication-induced premeiotic disruption (the RIP effect) first described by Selker et al. (E. U. Selker, E. B. Cambareri, B. C. Jensen, and K. R. Haack, Cell 51:741-752, 1987). The mutants had the unusual phenotype of growth on ethanol but not on acetate as the carbon source. In a cross to the wild type of a mutant strain in which the original ectopic gene sequence had been removed by segregation, the acetate nonutilizing phenotype invariably segregated together with a RIP-induced EcoRI site at the normal locus. This mutant was transformed to the ability to use acetate by the cloned sequence. The locus of the mutation, designated acu-8, was mapped between trp-3 and un-15 on linkage group 2. The transcribed portion of the clone, identified by probing with cDNA, was sequenced, and a putative 525-codon open reading frame with two introns was identified. The codon usage was found to be strongly biased in a way typical of most Neurospora genes sequenced so far. The predicted amino acid sequence shows no significant resemblance to anything previously recorded. These results provide a first example of the use of the RIP effect to obtain a mutant phenotype for a gene previously known only as a transcribed wild-type DNA sequence.

Acetates↗

Concerted evolution at a multicopy locus in the protozoan parasite Theileria parva: extreme divergence of potential protein-coding sequences.

Concerted evolution of multicopy gene families in vertebrates is recognized as an important force in the generation of biological novelty but has not been documented for the multicopy genes of protozoa. A multicopy locus, Tpr, which consists of tandemly arrayed open reading frames (ORFs) containing several repeated elements has been described for Theileria parva. Herein we show that probes derived from the 5'/N-terminal ends of ORFs in the genomic DNAs of T. parva Uganda (1,108 codons) and Boleni (699 codons) hybridized with multicopy sequences in homologous DNA but did not detect similar sequences in the DNA of 14 heterologous T. parva stocks and clones. The probe sequences were, however, protein coding according to predictive algorithms and codon usage. The 3'/C-terminal ends of the Uganda and Boleni ORFs exhibited 75% similarity and identity, respectively, to the previously identified Tpr1 and Tpr2 repetitive elements of T. parva Muguga. Tpr1-homologous sequences were detected in two additional species of Theileria. Eight different Tpr1-homologous transcripts were present in piroplasm mRNA from a single T. parva Muguga-infected animal. The Tpr1 and Tpr2 amino acid sequences contained six predicted membrane-associated segments. The ratio of synonymous to nonsynonymous substitutions indicates that Tpr1 evolves like protein-encoding DNA. The previously determined nucleotide sequence of the gene encoding the p67 antigen is completely identical in T. parva Muguga, Boleni, and Uganda, including the third base in codons. The data suggest that concerted evolution can lead to the radical divergence of coding sequences and that this can be a mechanism for the generation of novel genes.

Amino Acid Sequence↗