Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

Mammalian gene evolution: nucleotide sequence divergence between mouse and rat.

As a paradigm of mammalian gene evolution, the nature and extent of DNA sequence divergence between homologous protein-coding genes from mouse and rat have been investigated. The data set examined includes 363 genes totalling 411 kilobases, making this by far the largest comparison conducted between a single pair of species. Mouse and rat genes are on average 93.4% identical in nucleotide sequence and 93.9% identical in amino acid sequence. Individual genes vary substantially in the extent of nonsynonymous nucleotide substitution, as expected from protein evolution studies; here the variation is characterized. The extent of synonymous (or silent) substitution also varies considerably among genes, though the coefficient of variation is about four times smaller than for nonsynonymous substitutions. A small number of genes mapped to the X-chromosome have a slower rate of molecular evolution than average, as predicted if molecular evolution is "male-driven." Base composition at silent sites varies from 33% to 95% G+C in different genes; mouse and rat homologues differ on average by only 1.7% in silent-site G+C, but it is shown that this is not necessarily due to any selective constraint on their base composition. Synonymous substitution rates and silent site base composition appear to be related (genes at intermediate G+C have on average higher rates), but the relationship is not as strong as in our earlier analyses. Rates of synonymous and nonsynonymous substitution are correlated, apparently because of an excess of substitutions involving adjacent pairs of nucleotides. Several factors suggest that synonymous codon usage in rodent genes is not subject to selection.

Amino Acid Sequence↗

Identification of non-catalytic conserved regions in xylanases encoded by the xynB and xynD genes of the cellulolytic rumen anaerobe Ruminococcus flavefaciens.

xynB is one of at least four genes from the cellulolytic rumen anaerobe Ruminococcus flavefaciens 17 that encode xylanase activity. The xynB gene is predicted to encode a 781-amino acid product starting with a signal peptide, followed by an amino-terminal xylanase domain which is identical at 89% and 78% of residues, respectively, to the amino-terminal xylanase domains of the bifunctional XynD and XynA enzymes from the same organism. Two separate regions within the carboxy-terminal 537 amino acids of XynB also show close similarities with domain B of XynD. These regions show no significant homology with cellulose- or xylan-binding domains from other species, or with any other sequences, and their functions are unknown. In addition a 30 to 32-residue threonine-rich region is present in both XynD and XynB. Codon usage shows a consistent pattern of bias in the three xylanase genes from R. flavefaciens that have been sequenced.

Amino Acid Sequence↗

Analysis of putative DNA amplification genes in the element AUD1 of Streptomyces lividans 66.

The amplifiable AUD1 element of Streptomyces lividans 66 consists of two copies of a 4.7 kb sequence flanked by three copies of a 1 kb sequence. The DNA sequences of the three 1 kb repeats were determined. Two copies (left and middle repeats) were identical: (1009 bp in length) and the right repeat was 1012 bp long and differed at 63 positions. The repeats code for open reading frames (ORFs) with typical Streptomyces codon usage, which would encode proteins of about 36 kD molecular weight. The sequences of these ORFs suggest that they specify DNA-binding proteins and potential palindromic binding sites are found adjacent to the genes. The putative amplification protein encoded by the right repeat was expressed in Escherichia coli.

Base Sequence↗

The kalilo linear senescence-inducing plasmid of Neurospora is an invertron and encodes DNA and RNA polymerases.

The nucleotide sequence of kalilo, a linear plasmid that induces senescence in Neurospora by integrating into the mitochondrial chromosome, reveals structural and genetic features germane to the unique properties of this element. Prominent features include: (1) very long perfect terminal inverted repeats of nucleotide sequences which are devoid of obvious genetic functions, but are unusually GC-rich near both ends of the linear DNA; (2) small imperfect palindromes that are situated at the termini of the plasmid and are cognate with the active sites for plasmid integration into mtDNA; (3) two large, non-overlapping open-reading frames, ORF-1 and ORF-2, which are located on opposite strands of the plasmid and potentially encode RNA and DNA polymerases, respectively, and (4) a set of imperfect palindromes that coincide with similar structures that have been detected at more or less identical locations in the nucleotide sequences of other linear mitochondrial plasmids. The nucleotide sequence does not reveal a distinct gene that codes for the protein that is attached to the ends of the plasmid. However, a 335-amino acid, cryptic, N-terminal domain of the putative DNA polymerase might function as the terminal protein. Although the plasmid has been co-purified with nuclei and mitochondria, its nucleotide composition and codon usage indicate that it is a mitochondrial genetic element.

Amino Acid Sequence↗

The nucleotide sequence of the mercuric resistance operons of plasmid R100 and transposon Tn501: further evidence for mer genes which enhance the activity of the mercuric ion detoxification system.

The DNA sequences of the mercuric resistance determinants of plasmid R100 and transposon Tn501 distal to the gene (merA) coding for mercuric reductase have been determined. These 1.4 kilobase (kb) regions show 79% identity in their nucleotide sequence, and in both sequences two common potential coding sequences have been identified. In R100, the end of the homologous sequence is disrupted by an 11.2 kb segment of DNA which encodes the sulfonamide and streptomycin resistance determinants of Tn21. This insert contains terminal inverted repeat sequences and is flanked by a 5 base pair (bp) direct repeat. The first of the common potential coding sequences is likely to be that of the merD gene. Induction experiments and mercury volatilization studies demonstrate an enhancing but non-essential role for these merA-distal coding sequences in mercury resistance and volatilization. The potential coding sequences have predicted codon usages similar to those found in other Tn501 and R100 mer genes.

Amino Acid Sequence↗

The Escherichia coli dam gene is expressed as a distal gene of a new operon.

DNA containing the Escherichia coli dam gene and sequences upstream from this gene were cloned from the Clarke-Carbon plasmids pLC29-47 and pLC13-42. Promoter activity was localized using pKO expression vectors and galactokinase assays to two regions, one 1650-2100 bp and the other beyond 2400 bp upstream of the dam gene. No promoter activity was detected immediately in front of this gene; plasmid pDam118, from which the nucleotide sequence of the dam gene was determined, is shown to contain the pBR322 promoter for the primer RNA from the pBR322 rep region present on a 76 bp Sau3A fragment inserted upstream of the dam gene in the correct orientation for dam expression. The nucleotide sequence upstream of dam has been determined. An open reading frame (ORF) is present between the nearest promoter region and the dam gene. Codon usage and base frequency analysis indicate that this is expressed as a protein of predicted size 46 kDa. A protein of size close to 46 kDa is expressed from this region, detected using minicell analysis. No function has been determined for this protein, and no significant homology exist between it and sequences in the PIR protein or GenBank DNA databases. This unidentified reading frame (URF) is termed urf-74.3, since it is an URF located at 74.3 min on the E. coli chromosome. Sequence comparisons between the regions upstream of urf-74.3 and the aroB gene show that the aroB gene is located immediately upstream of urf-74.3, and that the promoter activity nearest to dam is found within the aroB structural gene. This activity is relatively weak (about 15% of that of the E. coli gal operon promoter). The promoter activity detected beyond 2400 bp upstream of dam is likely to be that of the aroB gene, and is 3 to 4 times stronger than that found within the aroB gene. Three potential DnaA binding sites, each with homology of 8 of 9 bp, are present, two in the aroB promoter region and one just upstream of the dam gene. Expression through the site adjacent to the dam gene is enhanced 2- to 4-fold in dnaA mutants at 38 degrees C. Restriction site comparisons map these regions precisely on the Clarke-Carbon plasmids pLC13-42 and pLC29-47, and show that the E. coli ponA (mrcA) gene resides about 6 kb upstream of aroB.

Alleles↗

Comparison of the haemolysin secretion protein HlyB from Proteus vulgaris and Escherichia coli; site-directed mutagenesis causing impairment of export function.

The hlyB secretion genes of Proteus vulgaris and Escherichia coli showed 81% nucleotide homology and similar E. coli-atypical codon usage. The deduced protein sequences differed in 54 of 707 residues and shared a previously unreported sequence which corresponds to the ATP-binding motif characteristic of protein kinases. The motif was also conserved in the HlyB of Morganella morganii. Of 4 oligonucleotide-directed substitutions introduced into the putative E. coli HlyB motif, 2 non-conservative changes caused radical reductions in the export of active haemolysin protein.

Amino Acid Sequence↗

Identification and nucleotide sequence of the Acinetobacter calcoaceticus encoded trpE gene.

The trpE gene from Acinetobacter calcoaceticus encoding the anthranilate synthase component I was cloned, identified by deletion analysis and sequenced. It encodes a predicted polypeptide of 497 amino acids with a calculated molecular weight of 55,323. Its primary structure shows 49% identical amino acids with the enzyme from Clostridium thermocellum, 45% with that of Thermus thermophilus and only 35% with that of Escherichia coli. The codon usage of the trpE genes encoding the most homologous enzymes differs greatly indicating selection for amino acid maintainance. The homologies are clustered in the C-terminal 200 amino acids of the sequences indicating that this part is important for enzymic activity.

Acinetobacter↗

The subunit I of the respiratory-chain NADH dehydrogenase from Cephalosporium acremonium: the evolution of a mitochondrial gene.

A Cephalosporium acremonium mitochondrial gene equivalent to human URF1 has been identified. The primary structure of the protein is highly homologous to its human (39%) and A. nidulans (66%) counterparts. Hydrophobicity profiles and predicted secondary structures are also very similar suggesting that this gene codes for the subunit I of the respiratory-chain NADH dehydrogenase. The nucleotide sequence of the gene, 70% homologous to the A. nidulans one, presents a high AT content (72%) and this fact is reflected in the codon usage.

Acremonium↗

The maize plastid psbB-psbF-petB-petD gene cluster: spliced and unspliced petB and petD RNAs encode alternative products.

The chloroplast psbB, psbF, petB, and petD genes are cotranscribed and give rise to many overlapping RNAs. The mechanism and significance of this mode of expression are of interest, particularly because the accumulation of the psb and pet gene products respond differently to both light and, in C4 species such as maize, developmental signals. We present an analysis of the maize psbB, psbF, petB, and petD genes and intergenic regions. The genes are organized similarly in maize (a C4 species) and in several C3 species. Functional class II-like introns interrupt the 5' ends of petB and petD. Both spliced and unspliced RNAs accumulate; these encode alternative forms of the petB and petD proteins, differing at their N-termini. Promoter-like elements between psbF and petB, and biased codon usage suggest that the differential regulation of the psb and pet genes might be achieved at both the transcriptional and translational levels.

Amino Acid Sequence↗

The DNA sequence of the gene for the secreted Bacillus subtilis enzyme levansucrase and its genetic control sites.

We present the sequence of a 2 kb fragment of the Bacillus subtilis Marburg genome containing sacB, the structural gene of levansucrase, a secreted enzyme inducible by sucrose. The peptide sequence deduced for the secreted enzyme is very similar to that directly determined by Delfour (1981) for levansucrase of the non-Marburg strain BS5. The peptide sequence is preceded by a 29 amino acid signal peptide. Codon usage in sacB is rather different from that in the sequenced genes of other secreted enzymes in B. subtilis, especially alpha-amylase. Genetic evidence has shown that the sacB promotor is rather far from the beginning of sacB (200 bp or more). The 200 bp region preceding sacB shows some of the features of an attenuator. A preliminary discussion of the putative workings and roles of this attenuator-like structure is proposed. sacRc mutations, which allow constitutive expression of levansucrase, have been located within the 450 bp upstream of sacB. It is shown that sacRc and sacR+ alleles control in cis the expression of the adjacent sacB gene.

Amino Acid Sequence↗

Sequence analysis of the DdPYR5-6 gene coding for UMP synthase in Dictyostelium discoideum and comparison with orotate phosphoribosyl transferases and OMP decarboxylases.

A Dictyostelium discoideum DNA fragment that complements the ura3 and the ura5 mutants of Saccharomyces cerevisiae has been sequenced. It contains an open reading frame of 478 codons capable of encoding a polypeptide of molecular weight 52475. This gene, named DdPYR5-6, encodes a bifunctional protein composed of the orotate phosphoribosyl transferase (OPRTase) and the orotidine-5'-phosphate decarboxylase (OMPdecase) domains described for UMP synthase in mammals. The existence of separate domains for the two activities was suspected because deletion of the N-terminal coding segment of the gene eliminated the ura5 but not the ura3 complementing activity. We have now confirmed that the two parts of the open reading frame share homology with known OPRTase and OMPdecase sequences. Several blocks of sequence are conserved among OPRTase from bacteria, fungi and slime mold and one of them corresponds to the consensus sequence for phosphoribosylbinding sites. The OMPdecase domain shows extensive similarity with the yeast and Neurospora crassa enzymes, suggesting that they have evolved from an ancestral gene which was fused to the OPRTase gene in D. discoideum. It is less related to the bacterial enzyme but all these sequences present conserved blocks of homology which could identify the active site. The codon usage is strongly biased in a manner similar to that found for other D. discoideum genes. The flanking DNA contains homopolymers of A and T and alternating sequences that are characteristic of the gene organization in D. discoideum.

Amino Acid Sequence↗

Accurate mapping of the Escherichia coli pepD gene by sequence analysis of its 5' flanking region.

A cloned DNA fragment, carrying the gene for peptidase D (pepD) of Escherichia coli, was partially sequenced. By purification of peptidase D and sequence determination of an amino-terminal oligopeptide the reading frame of the pepD gene, starting with a GTG initiator codon, was unambiguously identified. An overlap of the established nucleotide sequence with the previously sequenced 5' flanking region of the gpt gene allowed the exact distance between pepD and gpt to be calculated. The two genes are pointing towards each other and are separated by 260 bp. A search for open reading frames (ORFs) and the analysis of possible codon usage in the intercistronic region indicate the absence of an additional gene (lpcA) between pepD and gpt.

Amino Acid Sequence↗

Strategy for identifying the gene encoding the DNA polymerase of molluscum contagiosum virus type 1.

Molluscum contagiosum virus (MCV) is a member of the family Poxviridae and pathogenic to humans. MCV causes benign epidermal tumors mainly in children and young adults and is a common pathogen in immunecompromised individuals. The viral DNA polymerase is the essential enzyme involved in the replication of the genome of DNA viruses. The identification and characterization of the gene encoding the DNA polymerase of molluscum contagiosum virus type 1 (MCV-1) was carried out by PCR technology and nucleotide sequence analysis. Computer-aided analysis of known amino acid sequences of DNA polymerases from two members of the poxvirus family revealed a high amino acid sequence homology of about 49.7% as detected between the DNA polymerases of vaccinia virus (genus Orthopoxvirus) and fowlpoxvirus (genus Avipoxvirus). Specific oligonucleotide primers were designed and synthesized according to the distinct conserved regions of amino acid sequences of the DNA polymerases in which the codon usage of the MCV-1 genome was considered. Using this technology a 228 bp DNA fragment was amplified and used as hybridization probe for identifying the corresponding gene of the MCV-1 genome. It was found that the PCR product was able to hybridize to the BamHI MCV-1 DNA fragment G (9.2 kbp, 0.284 to 0.332 map units). The nucleotide sequence of this particular region of the MCV-1 genome (7267 bp) between map coordinates 0.284 and 0.315 was determined. The analysis of the DNA sequences revealed the presence of 22 open reading frames (ORFs-1 to -22). ORF-13 (3012 bp; nucleotide positions 6624 to 3612) codes for a putative protein of a predicted size of 115 kDa (1004 aa) which shows 40.1% identity and 35% similarity to the amino acid sequences of the DNA polymerases of vaccinia, variola, and fowlpoxvirus. In addition significant homologies (30% to 55%) were found between the amino acid sequences of the ORFs 3, -5, -9, and -14 and the amino acid sequences of the E6R, E8R, E10R, and a 7.3 kDa protein of vaccinia and variola virus, respectively. Comparative analysis of the genomic positions of the loci of the detected viral genes including the DNA polymerases of MCV-1, vaccinia, and variola virus revealed a similar gene organization and arrangement.

Amino Acid Sequence↗

A DNA-polymerase-related reading frame (pol-r) in the mtDNA of Secale cereale.

Mitochondrial (mt)DNA of Secale cereale contains an open reading frame (pol-r), the potential translation product of which shows significant homology to the type-B DNA polymerase encoded by the S1 plasmid of Zea mays; it contains the highly-conserved domains IIa to V of family B polymerases. The pol-r ORF is transcribed, as proven by RT-PCR, but the transcript is not edited. Upstream of the putative start codon a potential promoter motif was detected, fitting well into the postulated consensus sequence of the transcription initiation regions of Z. mays and Triticum aestivum. The pol-r ORF occurs in mtDNA of the fertile rye variety "Halo" and the cytoplasmic male-sterile (CMS) line "Pampa". Both ORFs are almost identical, apart from the 3' terminus; pol-r from Halo can code for 289 amino acids, pol-r from Pampa for 312 amino acids. Based on codon usage and the lack of editing, pol-r is considered to be a "young" gene, probably introduced in the mtDNA of rye by recombination with an mt plasmid.

Amino Acid Sequence↗

On concerted origin of transfer RNAs with complementary anticodons.

Pairs of antiparallely oriented consensus tRNAs with complementary anticodons show surprisingly small numbers of mispairings within the 17-bp- long anticodon stem and loop region. Even smaller such complementary distances are shown by illegitimately complementary anticodons, i.e. those with allowed pairing between G and U bases. Accordingly, we suppose that transfer RNAs have emerged concertedly as complementary strands of primordial double helix-like RNA molecules. Replication of such molecules with illegitimately complementary anticodons might generate new synonymous codons for the same pair of amino acids. Logically, the idea of tRNA concerted origin dictates very ancient establishment of direct links between anticodons and the type of amino acids with which pre-tRNAs were to be charged. More specifically, anticodons (first of all, the 2nd base) could selectively target 'their' amino acids, reaction of acylating itself being performed by another non-specific site of pre-tRNA or even by another ribozyme. In all, the above findings and speculations are consistent to the hypercyclic concept (Eigen and Schuster, 1979), and throw new light on the genetic code origin and associated problems. Also favoring this idea are data on complementary codon usage patterns in different genomes.

Amino Acids↗

Cloning and characterization of the white and topaz eye color genes from the sheep blowfly Lucilia cuprina.

Clones carrying the white and topaz eye color genes have been isolated from genomic DNA libraries of the blowfly Lucilia cuprina using cloned DNA from the homologous white and scarlet genes, respectively, of Drosophila melanogaster as probes. On the basis of hybridization studies using adjacent restriction fragments, homologous fragments were found to be colinear between the genes from the two species. The nucleotide sequence of a short region of the white gene of L. cuprina has been determined, and the homology to the corresponding region of D. melanogaster is 72%; at the derived amino acid level the homology is greater (84%) due to a marked difference in codon usage between the species. A major difference in genome organization between the two species is that whereas the DNA encompassing the D. melanogaster genes is free of repeated sequences, that encompassing their L. cuprina counterparts contains substantial amounts of repeated sequences. This suggests that the genome of L. cuprina is organized on the short period interspersion pattern. Repeated sequence DNA elements, which appear generally to be short (less than 1 kb) and which vary in repetitive frequency in the genome from greater than 10(4) copies to less than 10(2) copies, are found in at least two different locations in the clones carrying these genes. One type of repeat structure, found by sequencing, consists of tandemly repeating short sequences. Restriction site and restriction fragment length polymorphisms involving both the white and topaz gene regions are found within and between populations of L. cuprina.

Amino Acid Sequence↗

Selection against dam methylation sites in the genomes of DNA of enterobacteriophages.

Postreplicative methylation of adenine in Escherichia coli DNA to produce G6m ATC (where 6mA is 6-methyladenine) has been associated with preferential daughter-strand repair and possibly regulation of replication. An analysis was undertaken to determine if these, or other, as yet unknown roles of GATC, have had an effect on the frequency of GATC in E. coli or bacteriophage DNA. It was first ascertained that the most accurate predictions of GATC frequency were based on the observed frequencies of GAT and ATC, which would be expected since these predictors take into account preferences in codon usage. The predicted frequencies were compared with observed GATC frequencies in all available bacterial and phage nucleotide sequences. The frequency of GATC was close to the predicted frequency in most genes of E. coli and its RNA bacteriophages and in the genes of nonenteric bacteria and their bacteriophages. However, for DNA enterobacteriophages the observed frequency of GATC was generally significantly lower than predicted when assessed by the chi square test. No elevation in the rate of mutation of 6mA in GATC relative to other bases was found when pairs of DNA sequences from closely related phages or pairs of homologous genes from enterobacteria were compared, nor was any preferred pathway for mutation of 6mA evident in the E. coli DNA bacteriophages. This situation contrasts with that of 5-methylcytosine, which is hypermutable, with a preferred pathway to thymine. Thus, the low level of GATC in enterobacteriophages is probably due not to 6mA hypermutability, but no selection against GATC in order to bypass a GATC-mediated host function.(ABSTRACT TRUNCATED AT 250 WORDS)

Bacteriophages↗