Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,747 records · Page 97Linked to original sources

H3 and H4 histone cDNA sequences from Xenopus: a sequence comparison of H4 genes.

Ovarian poly (A) + RNA from Xenopus laevis and Xenopus borealis was used to construct two cDNA libraries which were screened for histone sequences. cDNA clones to H4 mRNA were obtained from both species and an H3 cDNA clone from Xenopus laevis. The complete DNA sequences of these clones have been determined and are presented. These new sequences are compared with other H3 and H4 DNA sequences both in the coding and 3' noncoding regions. We find that there is considerable non-random codon usage in ten H4 genes. In addition there are some sequence similarities in the 3' noncoding regions of H3 and H4 genes.

Animals↗

The genes coding for histone H3 and H4 in Neurospora crassa are unique and contain intervening sequences.

Sequences coding for histone H3 and H4 of Neurospora crassa could be identified in genomic digests with the use of the corresponding genes from sea urchin and X. laevis as hybridization probes. A 2.6 kb HindIII-generated N. crassa DNA fragment, showing homology with the heterologous histone H3-gene probes was cloned in a charon 21A vector. Using DNA from this clone as a homologous hybridization probe a 6.9 kb SalI-generated DNA fragment was isolated which in addition to the histone H3-gene also contains the gene coding for histone H4. Several lines of evidence demonstrate the presence of only a single histone H3- as well as a single histone H4-gene in N. crassa. The two genes are physically linked on the genome. DNA sequencing of the N. crassa histone H3- and H4-genes confirmed their identity and, in addition, revealed the presence of one short intron (67 bp) within the coding sequence of the H3-gene and even two introns (68 and 69 bp) within the H4-gene. The amino acid sequences of the N. crassa histones H3 and H4, as deduced from the DNA sequences, and those of the corresponding yeast histones differ only at a few positions. Much larger sequence differences, however, are observed at the DNA level, reflecting a diverging codon usage in the two lower eukaryotes.

Amino Acid Sequence↗

Statistical characterization of nucleic acid sequence functional domains.

It has long been recognized that various genome classes were distinguishable on the basis of base composition and nearest neighbor frequencies. In addition Grantham et al. (8) have recently presented evidence that these distinctions are preserved at the level of codon usage. As discussed in this report it is now clear that these and related statistics can uniquely characterize the various functional domains of the genome. In particular peptide coding, intervening segments, structural RNA coding and mitochondrial domains of the vertebrate genome are uniquely characterizable. The statistical measures not only reflect understood functional differences among these domains but suggest others. The ability of these simple statistics of nucleic acid sequences to reflect so much of the encoded complex pattern information and/or effects of selective constraints is somewhat surprising. Here, we investigated the statistical measures most distinctive of the various domains and then linked them to our current understandings in so far as possible.

Animals↗

Molecular structure and function of the bacteriocin gene and bacteriocin protein of plasmid Clo DF13.

In this paper we present the complete nucleotide sequence of the bacteriocin gene of plasmid Clo DF13. According to the predicted aminoacid sequence the bacteriocin, cloacin DF13, consists of 561 aminoacids and has a molecular weight of 59,293 D. To obtain insight into the structure and function of specific parts of the cloacin molecule, we constructed a hydration profile and we predicted the secondary structure of the protein. According to our predictions, the N-terminus of cloacin DF13 (corresponding to the first 150-180 aminoacids) is relatively hydrophobic and is rich in glycine residues. The data obtained support previous findings that the N-terminal part of cloacin DF13 is involved in translocation of this protein across the cell membrane. The C-terminal part of the cloacin protein is rich in positively charged aminoacids; this might reflect the RNase activity located within this domain. A comparison of the bacteriocin genes and corresponding proteins of Clo DF13 and Col E1 did not reveal any homology at the level of either the nucleotide or the aminoacid sequence. The codon usage of both genes, however, exhibits striking similarities. The sequence data obtained during this study enabled us to present the nucleotide sequence of the entire cloacin operon. The structure of this operon and the regulation of expression of the genes, located within this operon, is discussed.

Amino Acid Sequence↗

Structural comparison of yeast ribosomal protein genes.

The primary structure of the genes encoding the yeast ribosomal proteins L17a and L25 was determined, as well as the positions of the 5'- and 3'-termini of the corresponding mRNAs. Comparison of the gene sequences to those obtained for various other yeast ribosomal protein genes revealed several similarities. In all split genes the intron is located near the 5'-side of the amino acid coding region. Among the introns a clear pattern of sequence conservation can be observed. In particular the intron-exon boundaries and a region close to the 3'-splice site show sequence homology. Conserved sequences were also found in the leader and trailer regions of the ribosomal protein mRNAs. The 5'-flanking regions of the yeast ribosomal protein genes appeared to contain sequence elements that many but not all ribosomal protein genes have in common, and therefore may be implicated in the coordinate expression of these genes. The amino acid coding sequences of the ribosomal protein genes show a biased codon usage. Like most yeast ribosomal protein molecules, L17a and L25 are particularly basic at their N-terminus.

Amino Acid Sequence↗

The nucleotide sequence of the B gene of bacteriophage Mu.

Bacteriophage Mu is a highly efficient transposon which requires the products of the Mu A and B genes in order to transpose at a normal frequency. We have determined the nucleotide sequence of the B gene as well as that of the A-B intergenic region upstream of B. The protein product of the gene contains 312 amino acids and has a predicted molecular weight of 35,061. As expected, there do not appear to be any potential promoter sequences in the intergenic region prior to the gene, but it is preceded by a strong Shine-Dalgarno sequence. The intergenic region does not contain any obvious transcription termination sequences. The frequency of optimal codon usage is similar to that for other transposon and phage genes, and the amino acid composition is comparable to that of an "average" E. coli protein. A region near the amino terminus of the protein resembles the highly conserved bihelical fold which is involved in DNA contact and sequence specific recognition in a number of DNA binding proteins.

Amino Acid Sequence↗

Structural features of the hisT operon of Escherichia coli K-12.

The DNA sequence of a 2,3-kilobase segment of the E. coli hisT operon was determined. Analysis of the sequence indicated that the upstream gene in the operon encodes a 36,364-dalton polypeptide, which runs aberrantly on SDS-polyacrylamide gels. The distal hisT gene encodes the tRNA modification enzyme, pseudouridine synthase I, which was shown to have a polypeptide molecular mass of 30,399 daltons. The DNA sequence was consistent with the phenotypes and hisT expression of mutant operons. Analysis of the sequence and genetic complementation experiments demonstrated that the upstream and hisT genes are evolutionarily, structurally, and functionally unrelated; however, translation signals for the two genes overlap, which is consistent with genetic evidence suggesting translational coupling. Codon usage in the upstream gene is radically different from the hisT gene and may underlie the differential expression observed from the operon. Gene-inactivation experiments and S1-mapping of in vivo transcripts indicated that the operon contains an additional upstream gene. S1-mapping experiments also confirmed the presence of an internal promoter, which might be stringently controlled. Taken together, these results show that the structure of the hisT operon is complex and suggest that the operon might be regulated at several levels.

Amino Acid Sequence↗

A general DNA analysis program for the Hewlett-Packard Model 86/87 microcomputer.

A program is described to perform general DNA sequence analysis on the Hewlett-Packard Model 86/87 microcomputer operating on 128 K of RAM. The following analytical procedures can be performed: 1. display of the sequence, in whole or part, or its complement; 2. search for specified sequences e.g. restriction sites, and in the case of the latter give fragment sizes; 3. perform a comprehensive search for all known restriction enzyme sites; 4. map sites graphically; 5. perform editing functions; 6. base frequency analysis; 7. search for repeated sequences; 8. search for open reading frames or translate into the amino acid sequence and analyse for basic and acidic amino acids, hydrophobicity, and codon usage. Two sequences, or parts thereof, can be merged in various orientations to mimic recombination strategies, or can be compared for homologies. The program is written in HP BASIC and is designed principally as a tool for the laboratory investigator manipulating a defined set of vectors and recombinant DNA constructs.

Amino Acid Sequence↗

Acinetobacter calcoaceticus encoded mutarotase: nucleotide sequence analysis of the gene and characterization of its secretion in Escherichia coli.

The nucleotide sequence of the mutarotase gene from Acinetobacter calcoaceticus has been determined. It reveals an open reading frame of 381 amino acids. The codon usage of A. calcoaceticus for this gene is similar to E. coli except for the amino acids Leu, Ala, Glu, and Arg where major differences exist. This did not interfere drastically with high level expression in E. coli. The regulatory sequences for the initiation of translation are similar to the ones described for E. coli. The N-terminal 20 amino acids, which are not found in the mature enzyme, show homology to signal sequences of exported proteins. In A. calcoaceticus and E. coli mutarotase is specifically secreted into the periplasmic space. Processing of the signal sequence occurs at identical sites in both organisms. The mature mutarotase consists of 361 amino acids and has a calculated molecular weight of 38457 Da. Expression of mutarotase at a high level in a recombinant E. coli destabilizes the outer membrane. This results in coordinated leakage of mutarotase and beta-lactamase into the culture broth.

Acinetobacter↗

The trfB region of broad host range plasmid RK2: the nucleotide sequence reveals incC and key regulatory gene trfB/korA/korD as overlapping genes.

We report the nucleotide sequence of the trfB region of broad host range plasmid RK2. This region encodes the following loci: trfB, identical to korA and korD, which encodes a key transcriptional repressor of certain RK2 operons; incC, which appears to be involved in plasmid maintenance, possibley through post-transcriptional regulation of trfA product levels; the start of korB, which encodes a second transcriptional repressor of operons involved in stable inheritance of RK2. These loci are expressed as part of the trfB operon. In combination with deletion analysis, transcriptional and translation fusions and 'maxicell' analysis of polypeptides, the DNA sequence allows a number of conclusions to be drawn. First, the korB ORF start codon overlaps the incC ORF stop codon, suggesting the possibility of translational coupling between these two genes. Second, the trfB ORF lies entirely within the first third of the incC ORF using a different phase. Third, the incC ORF appears to contain a second transcriptional start whose function appears to be coupled to translation of the trfB ORF. Analysis of codon usage in the region of overlap between incC and trfB suggests that the incC gene may have evolved before the trfB gene. Determination of the DNA sequence of a mutant in which the product of trfB is rendered defective for transcriptional repression reveals an amino acid alteration within a region of this polypeptide which exhibits homology to the alpha helix-turn-alpha helix motif characteristic of many DNA binding proteins, and which is probably responsible for recognition of the trfB operator by this protein.

Amino Acid Sequence↗

Nucleotide sequence of the yeast cell division cycle start genes CDC28, CDC36, CDC37, and CDC39, and a structural analysis of the predicted products.

The nucleotide sequences of the yeast cell division cycle start genes CDC36, CDC37, and CDC39 are presented. An open reading frame corresponding in size and mapped position to the mRNA for each gene was revealed. These sequences, as well as that of the CDC28 gene, were analyzed for the presence of consensus sequences postulated to be transcriptional or translational signals, or to be involved in mRNA processing. In addition, the predicted protein products of the four genes were subjected to a number of structural and statistical analyses including codon usage bias analysis, secondary structure analysis and hydropathicity analysis.

Amino Acid Sequence↗

The ILV5 gene of Saccharomyces cerevisiae is highly expressed.

The nucleotide sequence of the yeast ILV5 gene, which codes for the branched-chain amino acid biosynthesis enzyme acetohydroxyacid reductoisomerase, has been determined. The ILV5 coding region is 1,185 nucleotides, corresponding to a polypeptide with a molecular weight of 44,280. Transcription of the ILV5 mRNA initiates at position -81 upstream from the ATG translation start codon and terminates between 218 and 222 bases downstream from the stop codon. Consensus sequences have been identified for initiation and termination of transcription, and for general control of amino acid biosynthesis, as well as repression by leucine. The ILV5 gene is regulated slightly by general amino acid control. Codon usage of the ILV5 gene has the strong bias observed in yeast genes that are highly expressed. In agreement with this, the reductoisomerase monomer, with an apparent molecular weight of 40,000, has been identified in an SDS polyacrylamide gel pattern of total soluble yeast proteins as a gene dosage dependent band.

2-Acetolactate Mutase↗

Unusual features of transcribed and translated regions of the histone H4 gene family of Tetrahymena thermophila.

The complete DNA sequence is presented of H4-II, the second of the pair of histone H4 genes of the ciliated protozoan, Tetrahymena thermophila. Both H4 genes code for the same protein. Codon usage in these and other Tetrahymena genes is severely restricted and is similar to that in yeast. Flanking regions are AT-rich (greater than or equal to 75%), relative to coding sequences (approximately 45% GC). Except for small, similarly positioned homologies, flanking sequences of the two genes are different. Canonical sequences in higher eukaryotic promoters are not obvious in these genes. Instead, short, localized, base composition eccentricities characterize the 5' flanking sequences of all Tetrahymena genes analyzed. The consensus, P yP u(A)3-4 ATGG initiates translation in these and all other known Tetrahymena genes. Nuclear transcripts and messages of both growing and starved cells begin at multiple sites, mainly at the first or second A residue following a pyrimidine. The palindrome typical of histone message 3' termini in higher organisms is not present. Downstream of both genes are sequences similar to the processing/polyadenylation signal of higher eukaryotes, although the unique 3' ends are not those predicted by the location of the signals.

Amino Acid Sequence↗

Isolation and characterization of a Neurospora crassa ribosomal protein gene homologous to CYH2 of yeast.

We have isolated and characterized a Neurospora crassa gene homologous to the yeast CYH2 gene encoding L29, a cycloheximide sensitivity-conferring protein of the cytoplasmic ribosome. The cloned Neurospora gene was isolated by cross-hybridization to CYH2. It was sequenced from both cDNA and genomic clones. The coding region is interrupted by seven intervening sequences. Its deduced amino acid sequence shows 70% homology to that of yeast ribosomal protein L29 and 60% homology to that of mammalian ribosomal protein L27', suggesting that the protein has an important role in ribosomal function. The pattern of codon usage is highly biased, consistent with high translation efficiency. There is a single copy of this gene in N. crassa, and R. Metzenberg and coworkers have mapped its genetic location to the vicinity of the cyh-2 locus.

Animals↗

Sequence and properties of the message encoding Tetrahymena hv1, a highly evolutionarily conserved histone H2A variant that is associated with active genes.

hv1 is a histone H2A variant found in the transcriptionally active Tetrahymena macronucleus, but not in the transcriptionally inert micronucleus. hv1 also contains antigenic determinants conserved in the histone complements of representatives of all four eukaryotic kingdoms. A cDNA clone encoding hv1 has been isolated and sequenced. Comparison of the derived protein sequence of hv1 with that of the chicken variant H2A.F and the sea urchin variant H2A.F/Z reveals remarkable homology in all but the extreme amino- and carboxy-termini and a small region in the conserved core. Putative regions of conserved antigenicity are discussed. Evidence is presented that suggests that hv1 is a single-copy, intron-containing gene that encodes a polyadenylated message. Unusual features in the 3' flanking sequence and in codon usage are also described. Evidence is also presented showing that hv1 message amounts are ten-fold greater in growing cells than in starved cells.

Amino Acid Sequence↗

Primary structure and functional organization of Drosophila 1731 retrotransposon.

We have determined the nucleotide sequence of the Drosophila retrotransposon 1731. 1731 is 4648 bp long and is flanked by 336 bp terminal repeats (LTRs) previously described as being reminiscent of provirus LTRs. The 1731 genome consists of two long open reading frames (ORFs 1 and 2) which slightly overlap each other. The ORF 1 and 2 present similarities with retroviral gag and pol genes respectively as shown by computer analysis. The pol gene exhibits several enzymatic activities in the following order: protease, endonuclease and reverse transcriptase. It is possible that 1731 also encompasses a ribonuclease H activity located between the endonuclease and reverse transcriptase domains. Moreover, comparison of the 1731 pol gene with the pol region of copia shows similarities extending over the protease, endonuclease and reverse transcriptase domains. We show that codon usage in the two retrotransposons is different. Finally, no ORF able to encode an env gene is detected in 1731.

Animals↗

Sequence and expression of NUC1, the gene encoding the mitochondrial nuclease in Saccharomyces cerevisiae.

The DNA sequence and studies on the expression of the NUC1 gene from Saccharomyces cerevisiae are presented. The NUC1 locus is located in the distal portion of the left arm of Chromosome X and encodes the major nuclease found in mitochondria. The inferred amino acid sequence of NUC1 predicts that the nuclease is basic, rich in prolines, of average hydrophobicity, and has a molecular weight for the primary translation product of 37,209 daltons. NUC1 is very poorly expressed, consistent with the codon usage bias determined from the DNA sequence and our previous determination of the number of enzyme molecules per cell. Mapping of the 5' terminus of the NUC1 mRNA reveals that the mRNA has a long 400 base untranslated leader in which are found three open reading frames, each initiated by an AUG. The possibility that these upstream open reading frames contribute to the poor expression of the NUC1 gene is discussed.

Amino Acid Sequence↗

Isolation and characterization of the gene coding for Escherichia coli arginyl-tRNA synthetase.

The gene coding for Escherichia coli arginyl-tRNA synthetase (argS) was isolated as a fragment of 2.4 kb after analysis and subcloning of recombinant plasmids from the Clarke and Carbon library. The clone bearing the gene overproduces arginyl-tRNA synthetase by a factor 100. This means that the enzyme represents more than 20% of the cellular total protein content. Sequencing revealed that the fragment contains a unique open reading frame of 1734 bp flanked at its 5' and 3' ends respectively by 247 bp and 397 bp. The length of the corresponding protein (577 aa) is well consistent with earlier Mr determination (about 70 kd). Primer extension analysis of the ArgRS mRNA by reverse transcriptase, located its 5' end respectively at 8 and 30 nucleotides downstream of a TATA and a TTGAC like element (CTGAC) and 60 nucleotides upstream of the unusual translation initiation codon GUG; nuclease S1 analysis located the 3'-end at 48 bp downstream of the translation termination codon. argS has a codon usage pattern typical for highly expressed E. coli genes. With the exception of the presence of a HVGH sequence similar to the HIGH consensus element, ArgRS has no relevant sequence homologies with other aminoacyl-tRNA synthetases.

Amino Acid Sequence↗