Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,765 records · Page 98Linked to original sources

A group II intron in the Neurospora mitochondrial coI gene: nucleotide sequence and implications for splicing and molecular evolution.

The temperature-sensitive Neurospora nuclear mutant cyt18-1 is deficient in splicing many Group I mitochondrial introns when grown at its non-permissive temperature; however, splicing of intron 1 in the coI gene of the Adiopodoume (formerly called North Africa) strain is unaffected (R.A. Collins and A.M. Lambowitz, J. Mol. Biol. 184: 413-428, 1985). Here we show that coI intron 1 is a typical Group II intron, the only one identified to date in Neurospora. The differential effect of the cyt18-1 mutation suggests that splicing of certain introns could be regulated independently of others by nuclear-encoded proteins. The intron contains a long open reading frame (ORF) resembling that of the Neurospora Mauriceville mitochondrial plasmid. The intron and plasmid ORFs share unusual features of codon usage that suggest both evolved outside of the Neurospora mitochondrial genetic system.

Amino Acid Sequence↗

A comparison of snRNP-associated Sm-autoantigens: human N, rat N and human B/B'.

N is a tissue-specific, Sm-epitope bearing, snRNP-associated protein found predominantly in brain. The cDNA sequence encoding human N is compared to those for rat N and human B/B'. The amino acid sequences of human and rat N are 100% conserved. Although the amino acid sequences of N and B/B' are very similar to each other, B/B' contains 50 amino acids which are not present in N. On Northern blots the cDNAs encoding N and B/B' recognize two different RNA species. A comparison of the codon usage, as specified by the open reading frames of N and B/B' as well as results from Southern blots, show that N and B/B' are derived from different genes.

Amino Acid Sequence↗

Oligonucleotide correlations between infector and host genomes hint at evolutionary relationships.

The frequencies of oligonucleotides of length 3-6 were studied in 211 sequences of human DNA (659 kilobases), 22 sequences of DNA of human viruses (120 kbs), in 181 sequences of E. coli (442 kbs), and in 42 sequences of phages of E. coli (137 kbs). The sequences were obtained from Genbank(R) 48. The observed frequencies (O) were compared to the expected frequencies (E) obtained in two ways: 1) according to nucleotide composition for each series, and 2) according to first order Markow chains for triplets, second order for quadruplets, and third order for quintuplets and sextuplets. The ratio O/E was obtained for each oligonucleotide. Then, the correlation between the ratio O/E in a pair of series was calculated. Strong correlations were observed for sequences of man and human viruses, and for E. coli and its phages. Other correlations were small. For higher order Markov chains, there is indication of some correlation also between viruses and phages. It was concluded that through analysis of parallel oligonucleotide series it may be possible to infer some of the complex evolutionary relationships existing between cells and their infectors beyond the level of codon usage.

Base Composition↗

Homology of lysS and lysU, the two Escherichia coli genes encoding distinct lysyl-tRNA synthetase species.

In Escherichia coli, two distinct lysyl-tRNA synthetase species are encoded by two genes: the constitutive lysS gene and the thermoinducible lysU gene. These two genes have been isolated and sequenced. Their nucleotide and deduced amino acid sequences show 79% and 88% identity, respectively. Codon usage analysis indicates the lysS product being more efficiently translated than the lysU one. In addition, the lysS sequence exactly coincides with the sequence of herC, a gene which is part of the prfB-herC operon. In contrast to the recent proposal of Gampel and Tzagoloff (1989, Proc. Natl. Acad. Sci. USA 86, 6023-6027), the lysU sequence is distinct from the open reading frame located adjacent to frdA, although large homologies are shared by these two genes.

Amino Acid Sequence↗

Gene distribution and isochore organization in the nuclear genome of plants.

The genomic distribution of 23 nuclear genes from three dicotyledons (pea, sunflower, tobacco) and five monocotyledons of the Gramineae family (barley, maize, rice, oat, wheat) was studied by localizing these genes in DNA fractions obtained by preparative centrifugation in Cs2SO4/BAMD density gradients. Each one of these genes (and of many other related genes and pseudogenes) was found to be located in DNA fragments (50-100 Kb in size) that were less than 1-2% GC apart from each other. This definitively demonstrates the existence of isochores in plant genomes, namely of compositionally homogeneous DNA regions at least 100-200 Kb in size. Moreover, the GC levels of the 23 coding sequences studied, of their first, second and third codon positions, and of the corresponding introns were found to be linearly correlated with the GC levels of the isochores harboring those genes. Compositional correlations displayed increasing slopes when going from second to first to third codon position with obvious effects on codon usage. Coding sequences for seed storage proteins and phytochrome of Gramineae deviate from the compositional correlations just described. Finally, CpG doublets of coding sequences were characterized by a shortage that decreased and vanished with increasing GC levels of the sequences. A number of these findings bear a striking similarity with results previously obtained for vertebrate genes.

Animals↗

Retroviral-type zinc fingers and glycine-rich repeats in a protein encoded by cnjB, a Tetrahymena gene active during meiosis.

We have determined the nucleotide sequence of the cnjB gene from the ciliate Tetrahymena thermophila. This gene is transcriptionally active only during early conjugation, peaking in meiotic prophase. It contains 13 introns, four transcription start points and codes for a putative polypeptide (CnjB) of 1748 amino acids with a calculated molecular weight of 200 kilodaltons and a pl of 7.9. The coding region of cnjB has a low GC content (32% GC) and unusual codon usage. The C-terminal one-third of CnjB consists of three repetitive domains. Introns were absent in this region of cnjB. One of the repetitive domains consists of seven CCHC or retroviral-type zinc fingers, a motif found in one or two copies in retroviral nucleocapsid proteins. This motif has also been found recently in seven copies in the human nucleic-acid binding protein CNBP, in an apparent CNBP homologue in Schizosaccharomyces pombe and in one copy in a Xenopus gene active in early embryos. The other two domains are on either side of the zinc finger domain and contain a repeated glycine-rich motif seen in the heterogeneous nuclear ribonuclear proteins A1 and A2/B1 as well as other proteins. Both CCHC zinc fingers and glycine-rich repeats have been found in proteins with single-stranded nucleic acid-binding activity as well as strand-annealing activity. CnjB is, to our knowledge, the first protein found to contain both types of motifs.

Amino Acid Sequence↗

Identification of coding regions in genomic DNA sequences: an application of dynamic programming and neural networks.

Dynamic programming (DP) is applied to the problem of precisely identifying internal exons and introns in genomic DNA sequences. The program GeneParser first scores the sequence of interest for splice sites and for these intron- and exon-specific content measures: codon usage, local compositional complexity, 6-tuple frequency, length distribution and periodic asymmetry. This information is then organized for interpretation by DP. GeneParser employs the DP algorithm to enforce the constraints that introns and exons must be adjacent and non-overlapping and finds the highest scoring combination of introns and exons subject to these constraints. Weights for the various classification procedures are determined by training a simple feed-forward neural network to maximize the number of correct predictions. In a pilot study, the system has been trained on a set of 56 human gene fragments containing 150 internal exons in a total of 158,691 bps of genomic sequence. When tested against the training data, GeneParser precisely identifies 75% of the exons and correctly predicts 86% of coding nucleotides as coding while only 13% of non-exon bps were predicted to be coding. This corresponds to a correlation coefficient for exon prediction of 0.85. Because of the simplicity of the network weighting scheme, generalization performance is nearly as good as with the training set.

Algorithms↗

Singular over-representation of an octameric palindrome, HIP1, in DNA from many cyanobacteria.

An octameric palindrome (5'-GCGATCGC-3') is abundant in cyanobacterial sequences within databases (GenBank/EMBL) and was designated HIP1 (highly iterated palindrome). The frequency of occurrence of all 256 octameric palindromes has now been determined in sub-databases revealing large and unique over-representation of HIP1 in cyanobacterial entries. DNA sequences from other bacteria were searched for any over-represented octameric palindromes analogous to HIP1. Only two sequences were identified, in the genomes of a thermophile and halophilic archaebacteria, although these were less abundant than HIP1 in cyanobacteria and relate to codon usage. To test the proposed widespread distribution of HIP1 in DNA from the cyanobacterium Synechococcus PCC 6301, randomly selected genomic clones were partly sequenced. HIP1 constituted 2.5% of the novel sequences, equivalent to a site on average once every 320 nucleotides. An oligonucleotide including HIP1 was also tested in PCR. Multiple products were obtained using template DNA from cyanobacterial strains in which HIP1 is abundant in known sequences, and some strains generated characteristic HIP-PCR banding patterns. However, analysis of DNA from one strain (not previously represented in databases) by random sequencing, HIP-PCR and Pvul digestion, confirms that not all cyanobacterial genomes are rich in HIP1.

Base Sequence↗

NRSub: a non-redundant database for Bacillus subtilis.

In the context of the international project aimed at sequencing the whole genome of Bacillus subtilis we have developed a non-redundant, fully annotated database of sequences from this organism. Starting from the B.subtilis sequences available in the EMBL, GenBank and DDBJ collections we have removed all encountered duplications and then added extra annotations to the sequences (e.g. accession numbers for the genes, locations on the genetic map, codon usage, etc.) We have also added cross-references to the EMBL, MEDLINE, SWISS-PROT and ENZYME data banks. The present system results from merging of the NRSub and SubtiList databases and the sequence contigs used in the two systems are identical. NRSub is distributed as a flatfile in EMBL format (which is supported by most sequence analysis software packages) and as an ACNUC database, while SubtiList is distributed as a relational database under 4th Dimension. It is possible to access the data through two dedicated World Wide Web servers located in France and Japan.

Bacillus subtilis↗

The translational signal database, TransTerm: more organisms, complete genomes.

TransTerm is a database of initiation and termination sequence contexts from more than 250 organisms listed in GenBank, including the four complete genomes:Haemophilus influenzae, Methanococcus jannaschii, Mycoplasma genitalium,and Saccharomyces cerevisiae. For the current release, more than 60 000 coding sequences were analysed. The tabulated data include initiation and termination contexts organised by species along with quantitative parameters about individual coding sequences (length, %GC, GC3, Nc and CAI). There are also tables of initiation- and termination-region nucleotide-frequencies, codon usage tables and summaries of stop signal usage. TransTerm is available on the World Wide Web at: http://biochem.otago.ac.nz:800/Transterm/homepage.h tml

Base Sequence↗

The NRSub database: update 1997.

In the context of the international project aiming at sequencing the whole genome of Bacillus subtilis we have developed NRSub, a non-redundant database of sequences from this organism. Starting from the B.subtilis sequences available in the repository collections we have removed all encountered duplications, then we have added extra annotations to the sequences (e.g. accession numbers for the genes, locations on the genetic map, codon usage index). We have also added cross-references with EMBL/GenBank/DDBJ, MEDLINE, SWISS-PROT and ENZYME databases. NRSub is distributed through anonymous FTP as a text file in EMBL format and as an ACNUC database. It is also possible to access the database through two dedicated World Wide Web servers located in France (http://acnuc.univ-lyon1.fr/nrsub/nrsub.++ +html ) and in Japan (http://ddbjs4h.genes.nig.ac.jp/ ).

Academies and Institutes↗

An Integrated Sequence-Structure Database incorporating matching mRNA sequence, amino acid sequence and protein three-dimensional structure data.

We have constructed a non-homologous database, termed the Integrated Sequence-Structure Database (ISSD) which comprises the coding sequences of genes, amino acid sequences of the corresponding proteins, their secondary structure and straight phi,psi angles assignments, and polypeptide backbone coordinates. Each protein entry in the database holds the alignment of nucleotide sequence, amino acid sequence and the PDB three-dimensional structure data. The nucleotide and amino acid sequences for each entry are selected on the basis of exact matches of the source organism and cell environment. The current version 1.0 of ISSD is available on the WWW at http://www.protein.bio.msu.su/issd/ and includes 107 non-homologous mammalian proteins, of which 80 are human proteins. The database has been used by us for the analysis of synonymous codon usage patterns in mRNA sequences showing their correlation with the three-dimensional structure features in the encoded proteins. Possible ISSD applications include optimisation of protein expression, improvement of the protein structure prediction accuracy, and analysis of evolutionary aspects of the nucleotide sequence-protein structure relationship.

Algorithms↗

EMGLib: the enhanced microbial genomes library (update 2000).

As the number of complete microbial genomes publicly available is still growing, the problem of annotation quality in these very large sequences remains unsolved. Indeed, the number of annotations associated with complete genomes is usually lower than those of the shorter entries encountered in the repository collections. Moreover, classical sequence database management systems have difficulties in handling entries of such size. In this context, the Enhanced Microbial Genomes Library (EMGLib) was developed to try to alleviate these problems. This library contains all the complete genomes from prokaryotes (bacteria and archaea) already sequenced and the yeast genome in GenBank format. The annotations are improved by the introduction of data on codon usage, gene orientation on the chromosome and gene families. It is possible to access EMGLib through two database systems set up on WWW servers: the PBIL server at http://pbil.univ-lyon1.fr/emglib.html and the MICADO server at http://locus.jouy.inra.fr/micado

Base Sequence↗

A unique pattern of intrastrand anomalies in base composition of the DNA in hypotrichs.

The 50 non-coding bases immediately internal to the telomeric repeats in the two 5' ends of macronuclear DNA molecules of a group of hypotrichous ciliates are anomalous in composition, consisting of 61% purines and 39% pyrimidines, A>T (ratio of 44:32), and G>C (ratio of 17:7). These ratio imbalances violate parity rule 2, according to which A should equal T and G should equal C within a DNA strand and therefore pyrimidines should equal purines. The purine-rich and base ratio imbalances are in marked contrast to the rest of the non-coding parts of the molecules, which have the theoretically expected purine content of 50%, with A = T and G = C. The ORFs contain an average of 52% purines as a result of bias in codon usage. The 50 bases that flank the 5' ends of macronuclear sequences in micronuclear DNA (12 cases) consist of approximately 50% purines. Thus, the 50 bases in the 5' ends of macronuclear sequences in micronuclear DNA are islands of purine richness in which A>T and G>C. These islands may serve as signals for the excision of macronuclear molecules during macronuclear development. We have found no published reports of coding or non-coding native DNA with such anomalous base composition.

Animals↗

Behavior of restriction-modification systems as selfish mobile elements and their impact on genome evolution.

Restriction-modification (RM) systems are composed of genes that encode a restriction enzyme and a modification methylase. RM systems sometimes behave as discrete units of life, like viruses and transposons. RM complexes attack invading DNA that has not been properly modified and thus may serve as a tool of defense for bacterial cells. However, any threat to their maintenance, such as a challenge by a competing genetic element (an incompatible plasmid or an allelic homologous stretch of DNA, for example) can lead to cell death through restriction breakage in the genome. This post-segregational or post-disturbance cell killing may provide the RM complexes (and any DNA linked with them) with a competitive advantage. There is evidence that they have undergone extensive horizontal transfer between genomes, as inferred from their sequence homology, codon usage bias and GC content difference. They are often linked with mobile genetic elements such as plasmids, viruses, transposons and integrons. The comparison of closely related bacterial genomes also suggests that, at times, RM genes themselves behave as mobile elements and cause genome rearrangements. Indeed some bacterial genomes that survived post-disturbance attack by an RM gene complex in the laboratory have experienced genome rearrangements. The avoidance of some restriction sites by bacterial genomes may result from selection by past restriction attacks. Both bacteriophages and bacteria also appear to use homologous recombination to cope with the selfish behavior of RM systems. RM systems compete with each other in several ways. One is competition for recognition sequences in post-segregational killing. Another is super-infection exclusion, that is, the killing of the cell carrying an RM system when it is infected with another RM system of the same regulatory specificity but of a different sequence specificity. The capacity of RM systems to act as selfish, mobile genetic elements may underlie the structure and function of RM enzymes.

Base Sequence↗

Mining Bacillus subtilis chromosome heterogeneities using hidden Markov models.

We present here the use of a new statistical segmentation method on the Bacillus subtilis chromosome sequence. Maximum likelihood parameter estimation of a hidden Markov model, based on the expectation-maximization algorithm, enables one to segment the DNA sequence according to its local composition. This approach is not based on sliding windows; it enables different compositional classes to be separated without prior knowledge of their content, size and localization. We compared these compositional classes, obtained from the sequence, with the annotated DNA physical map, sequence homologies and repeat regions. The first heterogeneity revealed discriminates between the two coding strands and the non-coding regions. Other main heterogeneities arise; some are related to horizontal gene transfer, some to t-enriched composition of hydrophobic protein coding strands, and others to the codon usage fitness of highly expressed genes. Concerning potential and established gene transfers, we found 9 of the 10 known prophages, plus 14 new regions of atypical composition. Some of them are surrounded by repeats, most of their genes have unknown function or possess homology to genes involved in secondary catabolism, metal and antibiotic resistance. Surprisingly, we notice that all of these detected regions are a + t-richer than the host genome, raising the question of their remote sources.

Bacillus subtilis↗

Sequence analysis of cloned cDNA encoding part of an immunoglobulin heavy chain.

The recombinant plasmid pH21-1 consists of mouse-derived complementary DNA (cDNA) in the E. coli plasmid pMB9. The mouse insertion has been completely sequenced, and encodes the CH3 domain and half the CH2 domain of the immunoglobulin gamma1 heavy chain. The predicted amino acid sequence differs at several positions from that previously published for this protein. The pattern of codon usage resembles that in some other eukaryotic messenger RNAs. A computer program has been used to predict the optimum secondary structure for the mRNA encoding the CH3 domain and the inter-domain junction.

Animals↗

Nucleotide sequence of the R.meliloti nitrogenase reductase (nifH) gene.

The nucleotide sequence of the structural gene (nifH) of nitrogenase reductase (Fe protein) from R.meliloti 41 with its flanking ends is reported. The amino acid sequence of nitrogenase reductase was deduced from the DNA sequence. The predicted R.meliloti nitrogenase reductase protein consists of 297 amino acid residues, has a molecular weight of 32,740 daltons and contains 5 cysteine residues. The codon usage in the nifH gene is presented. In the 5' flanking region, sequences resembling to consensus sequences of bacterial control regions were found. Comparison of the R.meliloti nifH nucleotide and amino acid sequences with those from different nitrogen-fixing organisms showed that the amino acid sequences are more conserved than the nucleotide sequences. This structural conservation of nitrogenase reductase may be related to its function and may explain the conservation of the nifH gene during evolution.

Amino Acid Sequence↗