Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

The use of an improved transposon mutagenesis system for DNA sequencing leads to the characterization of a new insertion sequence of Streptomyces lividans 66.

A DNA sequencing strategy was developed based on the tetracycline resistance transposon Tn1721. A universal M13 primer binding site (UP) for DNA sequencing and restriction sites for mapping were inserted near one end of Tn1721 and the new derivative, Tn5491, introduced onto a conjugative F' plasmid. The target sequence is inserted between two inverted resolution sites (res) of Tn1721 present on the high-copy plasmid pJOE2114. Due to the inviability of long palindromic sequences in Escherichia coli insertions between the inversely orientated res sites of pJOE2114 are positively selected. Transposition of Tn5491 into the target sequence is selected by cointegrate formation of Tn5491 during transposition, mating and transfer of the nonconjugative sequencing vector. After cointegrate resolution, the additional res sites in the vector result in a second site-specific recombination removing most of the transposon (except of 136 bp) and part of the target sequence. The reduced plasmid sizes and the use of the universal primer improved the quality of the sequencing results obtained on an automated fluorescent sequencer. A 3.35-kb EcoRI fragment from the 30-kb terminal inverted repeats (TIR) of the Streptomyces lividans chromosome was sequenced by this method. A 1304-bp sequence was found on this fragment with the features of insertion elements. The element called IS1372 had 27-bp IR and two potential open reading frames. The predicted gene products had similar sizes and high similarity to gene products encoded by insertion sequences of the IS3 family. Furthermore, a potential signal stimulating ribosomal shifts and typical for members of the IS3 family was identified. Five to seven copies of IS1372 were found in different strains of S. lividans but none in other Streptomyces species tested.

Amino Acid Sequence↗

MultiTag: multiple error-tolerant sequence tag search for the sequence-similarity identification of proteins by mass spectrometry.

The characterization of proteomes by mass spectrometry is largely limited to organisms with sequenced genomes. To identify proteins from organisms with unsequenced genomes, database sequences from related species must be employed for sequence-similarity protein identifications. Peptide sequence tags (Mann, 1994) have been used successfully for the identification of proteins in sequence databases using partially interpreted tandem mass spectra of tryptic peptides. We have extended the ability of sequence tag searching to the identification of proteins whose sequences are yet unknown but are homologous to known database entries. The MultiTag method presented here assigns statistical significance to matches of multiple error-tolerant sequence tags to a database entry and ranks alignments by their significance. The MultiTag approach has the distinct advantage over other sequence-similarity approaches of being able to perform sequence-similarity identifications using only very short (2-4) amino acid residue stretches of peptide sequences, rather than complete peptide sequences deduced by de novo interpretation of tandem mass spectra. This feature facilitates the identification of low abundance proteins, since noisy and low-intensity tandem mass spectra can be utilized.

Alcohol Dehydrogenase↗

The genomic sequences bound to special AT-rich sequence-binding protein 1 (SATB1) in vivo in Jurkat T cells are tightly associated with the nuclear matrix at the bases of the chromatin loops.

Special AT-rich sequence-binding protein 1 (SATB1), a DNA-binding protein expressed predominantly in thymocytes, recognizes an ATC sequence context that consists of a cluster of sequence stretches with well-mixed A's, T's, and C's without G's on one strand. Such regions confer a high propensity for stable base unpairing. Using an in vivo cross-linking strategy, specialized genomic sequences (0.1-1. 1 kbp) that bind to SATB1 in human lymphoblastic cell line Jurkat cells were individually isolated and characterized. All in vivo SATB1-binding sequences examined contained typical ATC sequence contexts, with some exhibiting homology to autonomously replicating sequences from the yeast Saccharomyces cerevisiae that function as replication origins in yeast cells. In addition, LINE 1 elements, satellite 2 sequences, and CpG island-containing DNA were identified. To examine the higher-order packaging of these in vivo SATB1-binding sequences, high-resolution in situ fluorescence hybridization was performed with both nuclear "halos" with distended loops and the nuclear matrix after the majority of DNA had been removed by nuclease digestion. In vivo SATB1-binding sequences hybridized to genomic DNA as single spots within the residual nucleus circumscribed by the halo of DNA and remained as single spots in the nuclear matrix, indicating that these sequences are localized at the base of chromatin loops. In human breast cancer SK-BR-3 cells that do not express SATB1, at least one such sequence was found not anchored onto the nuclear matrix. These findings provide the first evidence that a cell type-specific factor such as SATB1 binds to the base of chromatin loops in vivo and suggests that a specific chromatin loop domain structure is involved in T cell-specific gene regulation.

Base Composition↗

Prediction of the coding sequences of unidentified human genes. I. The coding sequences of 40 new genes (KIAA0001-KIAA0040) deduced by analysis of randomly sampled cDNA clones from human immature myeloid cell line KG-1.

We established a protocol for the prediction of the coding sequences of unidentified human genes based on the double selection and sequence analysis of cDNA clones with inserts carrying unreported 5'-terminal sequences and with insert sizes corresponding to nearly full-length transcripts. By applying the protocol, cDNA clones with inserts longer than 2 kb were isolated from a cDNA library of human immature myeloid cell line KG-1, and the coding sequences of 40 new genes were predicted. A computer search of the sequences indicated that 20 genes contained sequences similar to known genes in the GenBank/EMBL databases. The sequences of the remaining 20 genes were entirely new, and characteristic protein motifs or domains were identified in 32 genes. Other sequence features noted were that the coding sequences of 23 genes were followed by relatively long stretches of 3'-untranslated sequences and that 5 genes contained repetitive sequences in their 3'-untranslated regions. The chromosomal location of these genes has been determined. By increasing the scale of the above analysis, the coding sequences of many unidentified genes can be predicted.

Base Sequence↗

Complete nucleotide sequences of bovine alpha S2- and beta-casein cDNAs: comparisons with related sequences in other species.

The nucleotide sequences corresponding to bovine alpha S2- and beta-casein mRNAs have been determined by cDNA analysis. Both sequences appear to be complete at their 5' ends. The nucleotide sequence of alpha S2-casein, when compared with the corresponding cavine A sequence, helps to define the boundaries of a large amino acid repeat (approximately 80 residues) whereas comparisons with the nucleotide sequences of rat gamma- and mouse epsilon-casein mRNAs also reveal extensive sequence similarities. An alignment of these four sequences shows that the divergence of their translated regions has been characterized by the duplication and deletion of discrete segments of sequence that probably correspond to exons. A high degree of nucleotide substitution is also found when the four sequences are compared, except for well-conserved leader-peptide and phosphorylation-site sequences and, to a lesser extent, the 5'-untranslated regions. Similar comparison of the bovine and rat beta-caseins shows that their divergence has involved a high rate of nucleotide substitution but that no major insertions or deletions of sequence have occurred. The several splice sites that have veen defined in the rat beta-casein gene are likely to have been conserved in the bovine. The contrasting evolutionary histories of the alpha- and beta-casein coding sequences correlate with the distinctive functions of these proteins in the casein micelle system in milk.

Animals↗

Sequencing and analysis of a genomic fragment provide an insight into the Dunaliella viridis genomic sequence.

Dunaliella is a genus of wall-less unicellular eukaryotic green alga. Its exceptional resistances to salt and various other stresses have made it an ideal model for stress tolerance study. However, very little is known about its genome and genomic sequences. In this study, we sequenced and analyzed a 29,268 bp genomic fragment from Dunaliella viridis. The fragment showed low sequence homology to the GenBank database. At the nucleotide level, only a segment with significant sequence homology to 18S rRNA was found. The fragment contained six putative genes, but only one gene showed significant homology at the protein level to GenBank database. The average GC content of this sequence was 51.1%, which was much lower than that of close related green algae Chlamydomonas (65.7%). Significant segmental duplications were found within this fragment. The duplicated sequences accounted for about 35.7% of the entire region. Large amounts of simple sequence repeats (microsatellites) were found, with strong bias towards (AC)(n) type (76%). Analysis of other Dunaliella genomic sequences in the GenBank database (total 25,749 bp) was in agreement with these findings. These sequence features made it difficult to sequence Dunaliella genomic sequences. Further investigation should be made to reveal the biological significance of these unique sequence features.

Animals↗

Universal sequence map (USM) of arbitrary discrete sequences.

BACKGROUND: For over a decade the idea of representing biological sequences in a continuous coordinate space has maintained its appeal but not been fully realized. The basic idea is that any sequence of symbols may define trajectories in the continuous space conserving all its statistical properties. Ideally, such a representation would allow scale independent sequence analysis--without the context of fixed memory length. A simple example would consist on being able to infer the homology between two sequences solely by comparing the coordinates of any two homologous units. RESULTS: We have successfully identified such an iterative function for bijective mapping psi of discrete sequences into objects of continuous state space that enable scale-independent sequence analysis. The technique, named Universal Sequence Mapping (USM), is applicable to sequences with an arbitrary length and arbitrary number of unique units and generates a representation where map distance estimates sequence similarity. The novel USM procedure is based on earlier work by these and other authors on the properties of Chaos Game Representation (CGR). The latter enables the representation of 4 unit type sequences (like DNA) as an order free Markov chain transition table. The properties of USM are illustrated with test data and can be verified for other data by using the accompanying web-based tool:http://bioinformatics.musc.edu/~jonas/usm/. CONCLUSIONS: USM is shown to enable a statistical mechanics approach to sequence analysis. The scale independent representation frees sequence analysis from the need to assume a memory length in the investigation of syntactic rules.

Algorithms↗

Novel multigene families encoding highly repetitive peptide sequences. Sequence analyses of rat and mouse proline-rich protein cDNAs.

Multigene families encode the proline-rich proteins that are so prominent in human saliva and are dramatically induced in mouse and rat salivary glands by isoproterenol treatment and by feeding tannins. A cDNA encoding an acidic proline-rich protein of rat has been sequenced (Ziemer, M. A., Swain, W. F., Rutter, W. J., Clements, S., Ann, D. K., and Carlson D. M. (1984) J. Biol. Chem. 259, 10475-10480). This study presents the nucleotide sequences of five additional proline-rich protein cDNAs complementary to both mouse and rat parotid and submandibular gland mRNAs. Amino acid compositions deduced from the nucleotide sequences are typical for proline-rich proteins: 25-45% proline, 18-22% glycine, and 18-22% glutamine and generally an absence of sulfur-containing amino acids except for the initiator methionine. These proline-rich proteins display unusual repeating peptide sequences of 14-19 amino acids. The derived amino acid sequence of the cDNA insert of plasmid pMP1 from mouse has a 19-amino acid sequence which is repeated four times. The inserts of plasmids pUMP40 and pUMP4 also from mouse encode for 12 and 11 repeats of a 14-amino acid peptide, respectively. These repetitive sequences, and others from rat and mouse cDNAs and from human genomic clones, all show very high homologies and likely evolved from duplication of internal portions of an ancestral gene. Gene conversion could account for the high degree of conservation of nucleotide sequences of the repeat regions. Protein derived from the nucleotide sequences are all characterized by four general regions: a putative signal peptide, a transition region, the repetitive region, and a carboxyl-terminal region. The 5'-flanking sequences and sequences encoding the putative signal peptides are highly conserved (greater than 94%) in all six cDNAs. This sequence conservation may be important in the regulation of the biosynthesis of these unusual proteins.

Amino Acid Sequence↗

An alternative method for direct sequencing of PCR products, for epidemiological studies performed by nucleic sequence comparison. Application to rabbit haemorrhagic disease virus.

A sequencing strategy based on the use of commercially available fluorescent-labeled universal primers for directly sequencing polymerase chain reaction (PCR) amplified material has been developed. The PCR reactions were performed with hybrid primers, specific to the viral sequence and possessing the sequence of the standard sequencing primers at their 5' end (M13 and reverse primers). These amplified fragments were sequenced by a classical dye-primer kit on an automated sequencer. As opposed to the use of fluorescent dideoxynucleotides, this sequencing method yielded accurate, high-grade sequences and had several advantages. First, the intensity of the sequencing peaks was much more homogeneous with the dye primer method. In addition, the problem of altered electrophoretic mobility, which may occur during the sequencing of custom-synthesised fluorescent primers, was avoided. This method was successfully reproduced using several different sets of chimeric primers. We believe that it is suitable for epidemiological studies conducted by nucleic sequence comparison, as in the case of rabbit haemorrhagic disease virus, as well as in other systems.

Animals↗

An integrated approach to the analysis and modeling of protein sequences and structures. III. A comparative study of sequence conservation in protein structural families using multiple structural alignments.

The information required to generate a protein structure is contained in its amino acid sequence, but how three-dimensional information is mapped onto a linear sequence is still incompletely understood. Multiple structure alignments of similar protein structures have been used to investigate conserved sequence features but contradictory results have been obtained, due, in large part, to the absence of subjective criteria to be used in the construction of sequence profiles and in the quantitative comparison of alignment results. Here, we report a new procedure for multiple structure alignment and use it to construct structure-based sequence profiles for similar proteins. The definition of "similar" is based on the structural alignment procedure and on the protein structural distance (PSD) described in paper I of this series, which offers an objective measure for protein structure relationships. Our approach is tested in two well-studied groups of proteins; serine proteases and Ig-like proteins. It is demonstrated that the quality of a sequence profile generated by a multiple structure alignment is quite sensitive to the PSD used as a threshold for the inclusion of proteins in the alignment. Specifically, if the proteins included in the aligned set are too distant in structure from one another, there will be a dilution of information and patterns that are relevant to a subset of the proteins are likely to be lost. In order to understand better how the same three-dimensional information can be encoded in seemingly unrelated sequences, structure-based sequence profiles are constructed for subsets of proteins belonging to nine superfolds. We identify patterns of relatively conserved residues in each subset of proteins. It is demonstrated that the most conserved residues are generally located in the regions where tertiary interactions occur and that are relatively conserved in structure. Nevertheless, the conservation patterns are relatively weak in all cases studied, indicating that structure-determining factors that do not require a particular sequential arrangement of amino acids, such as secondary structure propensities and hydrophobic interactions, are important in encoding protein fold information. In general, we find that similar structures can fold without having a set of highly conserved residue clusters or a well-conserved sequence profile; indeed, in some cases there is no apparent conservation pattern common to structures with the same fold. Thus, when a group of proteins exhibits a common and well-defined sequence pattern, it is more likely that these sequences have a close evolutionary relationship rather than the similarities having arisen from the structural requirements of a given fold.

Algorithms↗

Multiple sequence alignments of partially coding nucleic acid sequences.

BACKGROUND: High quality sequence alignments of RNA and DNA sequences are an important prerequisite for the comparative analysis of genomic sequence data. Nucleic acid sequences, however, exhibit a much larger sequence heterogeneity compared to their encoded protein sequences due to the redundancy of the genetic code. It is desirable, therefore, to make use of the amino acid sequence when aligning coding nucleic acid sequences. In many cases, however, only a part of the sequence of interest is translated. On the other hand, overlapping reading frames may encode multiple alternative proteins, possibly with intermittent non-coding parts. Examples are, in particular, RNA virus genomes. RESULTS: The standard scoring scheme for nucleic acid alignments can be extended to incorporate simultaneously information on translation products in one or more reading frames. Here we present a multiple alignment tool, codaln, that implements a combined nucleic acid plus amino acid scoring model for pairwise and progressive multiple alignments that allows arbitrary weighting for almost all scoring parameters. Resource requirements of codaln are comparable with those of standard tools such as ClustalW. CONCLUSION: We demonstrate the applicability of codaln to various biologically relevant types of sequences (bacteriophage Levivirus and Vertebrate Hox clusters) and show that the combination of nucleic acid and amino acid sequence information leads to improved alignments. These, in turn, increase the performance of analysis tools that depend strictly on good input alignments such as methods for detecting conserved RNA secondary structure elements.

Algorithms↗

FBSA: feature-based sequence alignment technique for very large sequences.

The ability to align pairs of very large molecular sequences is essential for a range of comparative genomic studies. However, given the complexity of genomic sequences, it has been difficult to devise a systematic method that can align - even within the same species - pairs of large sequences. Most existing approaches typically attempt to align nucleotide sequences while ignoring valuable features contained within them, eg they filter out low-complexity regions and retroelements before aligning the sequences. However, features are then added post-alignment for visualisation and analysis purposes. We argue that repetitive elements and other features (such as genes, exons and regulatory elements) should be part of the alignment process. A hierarchical approach that aligns the biologically relevant features before aligning the detailed nucleotide sequences has a number of interesting characteristics: (1) features define 'alignment anchor points' that can guide meaningful nucleotide alignment; (2) features can be weighted; (3) a hierarchical approach would identify only meaningful regions to be aligned; (4) nucleotide sequences can be described as sequences of features and non-features, providing a natural mechanism to divide the sequences for processing; and (5) computational speed is significantly faster than other approaches. In this paper, we describe and discuss a feature-based approach to aligning large genome sequences. We refer to this as 'feature-based sequence alignment'.

Algorithms↗

Ordered shotgun sequencing, a strategy for integrated mapping and sequencing of YAC clones.

Ordered shotgun sequencing proposes to organize the mapping and sequencing of YACs with a hierarchical strategy that incorporates a feedback loop. Building on current protocols, a YAC is subcloned into plasmids, plasmid insert ends are sequenced, and the sequences are overlapped to create a partial map. Complete sequencing then starts with plasmids whose end-sequence tracts have overlapped, but to a minimal extent. The next plasmids to be sequenced are again selected for least overlap, striking out progressively to span the YAC with minimal directed gap-filling. Simulations support its feasibility and indicate that during the generation of the complete sequence, the approach facilitates the early choice of regions for selective sequencing, for example, for coding units. The sequencing of plasmids would also require less redundancy, and discriminate repetitive sequences more easily, than random sequencing across larger clones. The overall effort scales with YAC size and can be further reduced by additional mapping information.

Chromosome Mapping↗

Sequence comparisons using multiple sequences detect three times as many remote homologues as pairwise methods.

The sequences of related proteins can diverge beyond the point where their relationship can be recognised by pairwise sequence comparisons. In attempts to overcome this limitation, methods have been developed that use as a query, not a single sequence, but sets of related sequences or a representation of the characteristics shared by related sequences. Here we describe an assessment of three of these methods: the SAM-T98 implementation of a hidden Markov model procedure; PSI-BLAST; and the intermediate sequence search (ISS) procedure. We determined the extent to which these procedures can detect evolutionary relationships between the members of the sequence database PDBD40-J. This database, derived from the structural classification of proteins (SCOP), contains the sequences of proteins of known structure whose sequence identities with each other are 40% or less. The evolutionary relationships that exist between those that have low sequence identities were found by the examination of their structural details and, in many cases, their functional features. For nine false positive predictions out of a possible 432,680, i.e. at a false positive rate of about 1/50,000, SAM-T98 found 35% of the true homologous relationships in PDBD40-J, whilst PSI-BLAST found 30% and ISS found 25%. Overall, this is about twice the number of PDBD40-J relations that can be detected by the pairwise comparison procedures FASTA (17%) and GAP-BLAST (15%). For distantly related sequences in PDBD40-J, those pairs whose sequence identity is less than 30%, SAM-T98 and PSI-BLAST detect three times the number of relationships found by the pairwise methods.

Databases, Factual↗

Correction of the cDNA-derived protein sequence of prostatic spermine binding protein: pivotal role of tandem mass spectrometry in sequence analysis.

Spermine binding protein (SBP) is a rat ventral prostate protein that binds various polyamines, and the level of this protein and its mRNA is regulated by androgens. Previously, the cDNA for SBP was cloned and sequenced and an amino acid sequence deduced from the cDNA. Data from cloned and sequenced and an amino acid sequence deduced from the cDNA. Data from partial amino acid sequencing of the purified protein were consistent with the amino acid sequence deduced from the cDNA. However, the amino terminus of the protein was blocked, and therefore, direct protein sequence information confirming the cDNA reading frame of this region could not be obtained by Edman degradation. We have now employed an integrated approach using fast atom bombardment mass spectrometry, tandem mass spectrometry, and conventional sequencing methodologies to establish the amino-terminal sequence of the protein and to identify an amino acid sequence (35 residues) present in the purified protein but missing from the amino acid sequence deduced from cDNA clones for this protein. The missing piece of cDNA corresponds to an exon found in mouse genomic clones for a protein similar to rat SBP. Therefore, the cDNA clones for rat SBP may represent splicing variants that lack the sequence information of one exon. The blocked amino terminus of the protein was identified as 5-oxopyrrolidine-2-carboxylic acid. Mass spectrometry also provided evidence regarding glycosylation of the protein. The first of two potential glycosylation sites clearly carries carbohydrate; the second site is, at most, only partially glycosylated.

Amino Acid Sequence↗

Measurement of the effectiveness of transitive sequence comparison, through a third 'intermediate' sequence.

MOTIVATION: Transitive sequence matching expands the scope of sequence comparison by re-running the results of a given query against the databank as a new query. This sometimes results in the initial query sequence (Q) being related to a final match (M) indirectly, through a third, 'intermediate' sequence (Q --> I --> M ). This approach has often been suggested as providing greater sensitivity in sequence comparison; however, it has not yet been possible to gauge its improvement precisely. RESULTS: Here, this improvement is comprehensively measured by seeing what fraction of the known structural relationships transitive sequence matching can uncover beyond that found by normal pairwise comparison (i.e. direct linkage). The structural relationships are taken from a well-characterized test set, the scop classification of protein structure. Specifically, 2055 known structural similarities (called 'pairs') between distantly related proteins constitute the basic test set. To make the measurement of transitive matching properly, special data sets, called 'baseline sets', are derived from this. They consist of pairs of sequences that have a clear structural relationship that cannot be found by normal sequence comparison (i.e. they cannot be directly linked). Specifically, using standard sequence comparison protocols (FASTA with an e-value cut-off of 0. 001), it is found that the baseline set consists of 1742 pairs. A third intermediate sequence can link 86 of these indirectly (5%), where this third sequence is drawn from the entire, current universe of protein sequences. The number of false positives is minimal. Furthermore, when one considers only the relationships within the test set that correspond to a close structural alignment, the coverage increases considerably. In particular, 862 of the baseline set pairs fit to better than 2.6 A RMS, and transitive matching can find 62 of these (9%). AVAILABILITY: All the test data, including precise similarity values calculated from structural alignment, are available in tabular format over the Web from http://bioinfo.mbb. yale.edu/align. CONTACT: Mark.Gerstein@yale.edu

Amino Acid Sequence↗

From first base: the sequence of the tip of the X chromosome of Drosophila melanogaster, a comparison of two sequencing strategies.

We present the sequence of a contiguous 2.63 Mb of DNA extending from the tip of the X chromosome of Drosophila melanogaster. Within this sequence, we predict 277 protein coding genes, of which 94 had been sequenced already in the course of studying the biology of their gene products, and examples of 12 different transposable elements. We show that an interval between bands 3A2 and 3C2, believed in the 1970s to show a correlation between the number of bands on the polytene chromosomes and the 20 genes identified by conventional genetics, is predicted to contain 45 genes from its DNA sequence. We have determined the insertion sites of P-elements from 111 mutant lines, about half of which are in a position likely to affect the expression of novel predicted genes, thus representing a resource for subsequent functional genomic analysis. We compare the European Drosophila Genome Project sequence with the corresponding part of the independently assembled and annotated Joint Sequence determined through "shotgun" sequencing. Discounting differences in the distribution of known transposable elements between the strains sequenced in the two projects, we detected three major sequence differences, two of which are probably explained by errors in assembly; the origin of the third major difference is unclear. In addition there are eight sequence gaps within the Joint Sequence. At least six of these eight gaps are likely to be sites of transposable elements; the other two are complex. Of the 275 genes in common to both projects, 60% are identical within 1% of their predicted amino-acid sequence and 31% show minor differences such as in choice of translation initiation or termination codons; the remaining 9% show major differences in interpretation.

Animals↗

Nucleotide sequence and characteristics of the gene for L-lactate dehydrogenase of Thermus caldophilus GK24 and the deduced amino-acid sequence of the enzyme.

The gene for L-lactate dehydrogenase (LDH) (EC 1.1.1.27) of Thermus caldophilus GK24 was cloned in Escherichia coli using synthetic oligonucleotides as hybridization probes. The nucleotide sequence of the cloned DNA was determined. The primary structure of the LDH was deduced from the nucleotide sequence. The deduced amino acid sequence agreed with the NH2-terminal and COOH-terminal sequences previously reported and the determined amino acid sequences of the peptides obtained from trypsin-digested T. caldophilus LDH. The LDH comprised 310 amino acid residues and its molecular mass was determined to be 32,808. On alignment of the whole amino acid sequences, the T. caldophilus LDH showed about 40% identity with the Bacillus stearothermophilus, Lactobacillus casei and dogfish muscle LDHs. The T. caldophilus LDH gene was expressed with the E. coli lac promoter in E. coli, which resulted in the production of the thermophilic LDH. The gene for the T. caldophilus LDH showed more than 40% identity with those for the human and mouse muscle LDHs on alignment of the whole nucleotide sequences. The G + C content of the coding region for the T. caldophilus LDH was 74.1%, which was higher than that of the chromosomal DNA (67.2%). The G + C contents in the first, second and third positions of the codons used were 77.7%, 48.1% and 95.5% respectively. The high G + C content in the third base caused extremely non-random codon usage in the LDH gene. About half (48.7%) the codons in the LDH gene started with G, and hence there were relatively high contents of Val, Ala, Glu and Gly in the LDH. The contents of Pro, Arg, Ala and Gly, which have high G + C contents in their codons, were also high. Rare codons with U or A as the third base were sometimes used to avoid the TCGA sequence, the recognition site for the restriction endonuclease, TaqI. Two TCGA sequences were found only in the sequence of CTCGAG (XhoI site) in the sequenced region of the T. caldophilus DNA. There were three segments with similar sequences in the two 5' non-coding regions, probably the promoter and ribosome-binding regions, of the genes for the T. caldophilus LDH and the Thermus thermophilus 3-isopropylmalate dehydrogenase.

Amino Acid Sequence↗