Using the FASTA program to search protein and DNA sequence databases.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to W R Pearson.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
We describe the identification of the GSTM1 null, GSTM1 A, GSTM1 B and GSTM1 A,B polymorphisms at the glutathione S-transferase GSTM1 locus using a single-step PCR method. Target DNA was amplified using primers to intron 6 and exon 7 with site-directed mutagenesis being used to introduce a restriction site in DNA amplified from GSTM1 *A, thereby allowing differentiation of this allele and GSTM1 *B. The accuracy of this approach in identifying the GSTM1 A, GSTM1 B, GSTM1 A,B and GSTM1 null polymorphisms was confirmed by comparison with, firstly, an established PCR method that distinguishes GSTM1 *0 homozygotes from individuals with the other GSTM1 genotypes and, secondly, GSTM1 phenotypes determined using chromatofocusing.
Explore the source record for details and available documents.
We report the sequences of two coordinately induced murine glutathione transferase genes, mGSTM1 (GT8.7, Yb1) and mGSTM3 (GT9.3). Genomic clones covering the entire mGSTM1 gene were isolated; comparison of the mGSTM1 gene with genomic sequences from rat class-mu glutathione transferase genes suggests that the mGSTM1 gene is orthologous to the rGSTM1 (rat3, Yb1) gene. The start of mGSTM1 mRNA transcription was mapped by primer extension and RNase protection to 37 nucleotides upstream from the initiation codon. The 160 nucleotides 5'-proximal to the start of transcription match exactly the 5'-end of the class-mu glutathione transferase cDNA clone pmGT10. An mRNA transcript was found approximately 2.0 kb upstream from the start of mGSTM1 transcription; its sequence does not show significant similarity to other sequences in the DNA or protein sequence databases. The mGSTM1 gene contains a TATAAA sequence at -31 nucleotides upstream from the start of transcription, but no exact match to the antioxidant response element (RGTGACNNNGC), the xenobiotic response element (TNGCGTG), or the AP-1 consensus (TGASTMA) is found in the 5'-flanking region, although near matches are found in the 5'-flanking region, in intron 1, and in other parts of the gene. A genomic clone containing the first five exons of the mGSTM3 gene was also isolated. The mGSTM3 gene contains several repetitive elements--two upstream from the start of transcription and one within intron 2--that disrupt its similarity with the mGSTM1 gene. The 5'-flanking sequence of the mGSTM3 gene does not contain a TATAAA sequence or any exact matches to ARE or XRE consensus sequences, although an Sp1 binding site is found at -66. mGSTM3 and mGSTM1 diverge substantially outside their exons and share less than 60% sequence identity in the 5'-flanking region. Thus, it is likely that the mGSTM1-mGSTM3 gene duplication predates the rat-mouse divergence. The strongest region of conservation between the mGSTM1 gene and the mGSTM3 gene occurs in exon 3, intron 3, and exon 4; this region also shares strong similarity with the rGSTM1 (rat3, Yb1) and rGSTM2 (rat4, Yb2) genes.
Levels of mRNAs encoding class-alpha glutathione transferases, class-mu glutathione transferases, quinone reductase, and cytochrome P450 1A were measured after xenobiotic induction in murine tissues and in the Hepa1c1c7 murine hepatoma cell line. RNA levels in liver and intestinal mucosa were determined after induction with phenobarbital, butylated hydroxyanisole, beta-naphthoflavone, isosafrole, or combinations of these compounds. The tissue culture cells were presented with combinations of butylated hydroxyanisole, tert-butyl-hydroquinone, and beta-naphthoflavone. In murine liver and intestinal mucosa, the greatest induction (5-15-fold) of glutathione transferases and quinone reductase was seen with butylated hydroxyanisole. Administration of phenobarbital or beta-naphthoflavone has only a modest effect (2-3-fold). In contrast, cytochrome P450 1A mRNA levels increase only slightly after BHA induction but are induced dramatically by beta-naphthoflavone. The pattern of induction is different in Hepa1c1c7 cells; there the greatest induction of all mRNAs occurred with beta-naphthoflavone. Administration of antioxidants with other xenobiotics increases mRNA levels only slightly over the levels obtained with BHA in murine tissues, or with beta-naphthoflavone in Hepa1c1c7 cells. mGSTM1 (GT8.7, Yb1), the most abundant glutathione transferase mRNA in murine liver, is also the most abundant glutathione transferase mRNA in both normal and induced Hepa1c1c7 cells. Our results suggest that BHA induction in murine liver and intestinal mucosa of class-mu and class-alpha glutathione transferases may involve regulatory elements and mediators that function poorly in Hepa1c1c7 cells.
The GSTM1, GSTM2, GSTM3, GSTM4, and GSTM5 glutathione transferase genes have been mapped to human chromosome 1 by using locus-specific PCR primer pairs spanning exon 6, intron 6, and exon 7, as probes on DNA from human/hamster somatic cell hybrids. For GSTM1, the assignment was confirmed by Southern blot hybridization to a pair of 12.5/2.4-kb HindIII fragments. The GSTM1-specific primer pairs can be used to identify individuals carrying non-null GSTM1 alleles. The organization of these five genes was confirmed by the isolation of a yeast artificial chromosome clone (GSTM-YAC2) that contains all five genes. With this clone, the location of the GSTM1-GSTM5 gene cluster on chromosome 1 was confirmed by fluorescence in situ hybridization. Both regional assignment using the fractional length method and examination of probe signal with reference to R-banded chromosomes induced by BrdU places the gene cluster in or near the 1p13.3 region. The close physical proximity of the GSTM1 and GSTM2 loci, which share 99% nucleotide sequence identity over 460 nucleotides of 3'-untranslated mRNA, suggests that the GSTM1-null allele may result from unequal crossing-over.
Explore the source record for details and available documents.
Efficient dynamic programming algorithms are available for a broad class of protein and DNA sequence comparison problems. These algorithms require computer time proportional to the product of the lengths of the two sequences being compared [O(N2)] but require memory space proportional only to the sum of these lengths [O(N)]. Although the requirement for O(N2) time limits use of the algorithms to the largest computers when searching protein and DNA sequence databases, many other applications of these algorithms, such as calculation of distances for evolutionary trees and comparison of a new sequence to a library of sequence profiles, are well within the capabilities of desktop computers. In particular, the results of library searches with rapid searching programs, such as FASTA or BLAST, should be confirmed by performing a rigorous optimal alignment. Whereas rapid methods do not overlook significant sequence similarities, FASTA limits the number of gaps that can be inserted into an alignment, so that a rigorous alignment may extend the alignment substantially in some cases. BLAST does not allow gaps in the local regions that it reports; a calculation that allows gaps is very likely to extend the alignment substantially. Although a Monte Carlo evaluation of the statistical significance of a similarity score with a rigorous algorithm is much slower than the heuristic approach used by the RDF2 program, the dynamic programming approach should take less than 1 hr on a 386-based PC or desktop Unix workstation. For descriptive purposes, we have limited our discussion to methods for calculating similarity scores and distances that use gap penalties of the form g = rk. Nevertheless, programs for the more general case (g = q+rk) are readily available. Versions of these programs that run either on Unix workstations, IBM-PC class computers, or the Macintosh can be obtained from either of the authors.
A platform program that performs biological sequence comparison provides a case study to compare the relative advantages of a machine-independent approach to parallel computation versus a machine-specific approach. The program consists of two routines: (i) PSCANLIB, which compares a single biological sequence against a database of sequences, and (ii) PCOMPLIB, which compares a database of sequences against another database of sequences, or against itself. The program was first parallelized to run on the Intel Hypercube parallel computer using native Hypercube commands to coordinate the parallel computation. The parallelization logic of the program was then translated into a machine-independent parallel programming language, Linda. These two approaches to parallelization are contrasted in terms of: (i) the expressive power of the logic that coordinates the parallel computation, (ii) the portability of the machine-independent version to other parallel machines and (iii) the relative efficiency of the two versions of the program. In the benchmark tests reported, the benefits of the machine-independent approach were achieved with only a modest sacrifice in efficiency.
We describe an algorithm for aligning two sequences within a diagonal band that requires only O(NW) computation time and O(N) space, where N is the length of the shorter of the two sequences and W is the width of the band. The basic algorithm can be used to calculate either local or global alignment scores. Local alignments are produced by finding the beginning and end of a best local alignment in the band, and then applying the global alignment algorithm between those points. This algorithm has been incorporated into the FASTA program package, where it has decreased the amount of memory required to calculate local alignments from O(NW) to O(N) and decreased the time required to calculate optimized scores for every sequence in a protein sequence database by 40%. On computers with limited memory, such as the IBM-PC, this improvement both allows longer sequences to be aligned and allows optimization within wider bands, which can include longer gaps.
Explore the source record for details and available documents.
cDNA encoding the more acidic form, glutathione transferase (GST) psi, of the polymorphic Mu-class GSTs discovered in liver, was mutated in the 5'-end to create an NcoI site, facilitating cloning into the expression plasmid pKK233-2. The protein expressed from this construct has a point mutation Pro-2----Ala-2, but gives a catalytically functional protein. Back-mutation of the codon for amino acid residue 2 gave rise to a plasmid expressing the wild-type enzyme GST psi, or GST Mu1b-1b. A variant cDNA, differing only in specifying lysine rather than asparagine in position 173 of the coding region, was generated by site-directed mutagenesis. The variant sequence corresponds to another cDNA clone isolated from a human liver cDNA library and expresses the near-neutral GST mu, or GST Mu1a-1a. The two recombinant proteins GST Mu1a-1a and GST Mu1b-1b, by physicochemical as well as kinetic criteria, were found to be indistinguishable from GST mu and GST psi respectively, isolated from human liver. It is therefore concluded that the recombinant proteins correspond to the allelic variants observed in the human population. The two forms have different isoelectric points and correspond to the allelic variants observed in the human population. The two forms have different isoelectric points and their protein subunits can be separated by h.p.l.c. on a reverse-phase column. With standard substrates and inhibitors no differences in kinetic parameters between the two variants were detected. The mutated GST Mu1b-1b (Pro-2----Ala) was not significantly different in catalytic properties from the wild-type enzyme, even though Pro-2 is a well conserved amino acid residue in the known Mu-class GSTs.
A class-mu glutathione transferase cDNA clone, GTHMUS, was isolated from human myoblasts and its sequence was determined. The sequence predicts a protein of molecular weight 25,599 whose 24 amino-terminal residues are identical to those of the class-mu isoenzyme expressed from the GST4 locus. The GTHMUS cDNA shares 93.7% nucleotide sequence identity with a human liver cDNA clone, GTH411, that is encoded at the GST1 locus. Comparison of the liver and muscle cDNA sequences shows two regions of remarkable sequence conservation: a 140-nucleotide region in the 5' coding portion of the molecule that has a single silent nucleotide substitution, and a 550-nucleotide region, including the entire 3' noncoding region, that has only three nucleotide substitutions or deletions. This sequence conservation suggests that gene conversion has occurred between the human GST1 and GST4 glutathione transferase gene loci. The human muscle and liver glutathione transferase clones GTHMUS and GTH411 have been expressed in Escherichia coli. The kinetic mechanism of the muscle enzyme was examined in product inhibition studies. The inhibition patterns are best modeled by a steady-state ordered bi-bi reaction mechanism. Glutathione is the first substrate bound and chloride ion is the first product released. Chloride ion inhibits the muscle enzyme.
Three 'alpha 1-adrenoceptors' and three 'alpha 2-adrenoceptors' have now been cloned. How closely do these receptors match the native receptors that have been identified pharmacologically? What are the properties of these receptors, and how do they relate to other members of the cationic amine receptor family? Kevin Lynch and his colleagues discuss these questions in this review.
The sensitivity and selectivity of the FASTA and the Smith-Waterman protein sequence comparison algorithms were evaluated using the superfamily classification provided in the National Biomedical Research Foundation/Protein Identification Resource (PIR) protein sequence database. Sequences from each of the 34 superfamilies in the PIR database with 20 or more members were compared against the protein sequence database. The similarity scores of the related and unrelated sequences were determined using either the FASTA program or the Smith-Waterman local similarity algorithm. These two sets of similarity scores were used to evaluate the ability of the two comparison algorithms to identify distantly related protein sequences. The FASTA program using the ktup = 2 sensitivity setting performed as well as the Smith-Waterman algorithm for 19 of the 34 superfamilies. Increasing the sensitivity by setting ktup = 1 allowed FASTA to perform as well as Smith-Waterman on an additional 7 superfamilies. The rigorous Smith-Waterman method performed better than FASTA with ktup = 1 on 8 superfamilies, including the globins, immunoglobulin variable regions, calmodulins, and plastocyanins. Several strategies for improving the sensitivity of FASTA were examined. The greatest improvement in sensitivity was achieved by optimizing a band around the best initial region found for every library sequence. For every superfamily except the globins and immunoglobulin variable regions, this strategy was as sensitive as a full Smith-Waterman. For some sequences, additional sensitivity was achieved by including conserved but nonidentical residues in the lookup table used to identify the initial region.
We have written two programs for searching biological sequence databases that run on Intel hypercube computers. PSCANLIB compares a single sequence against a sequence library, and PCOMPLIB compares all the entries in one sequence library against a second library. The programs provide a general framework for similarity searching; they include functions for reading in query sequences, search parameters and library entries, and reporting the results of a search. We have isolated the code for the specific function that calculates the similarity score between the query and library sequence; alternative searching algorithms can be implemented by editing two files. We have implemented the rapid FASTA sequence comparison algorithm and the more rigorous Smith-Waterman algorithm within this framework. The PSCANLIB program on a 16 node iPSC/2 80386-based hypercube can compare a 229 amino acid protein sequence with a 3.4 million residue sequence library in approximately 16 s with the FASTA algorithm. Using the Smith-Waterman algorithm, the same search takes 35 min. The PCOMPLIB program can compare a 0.8 million amino acid protein sequence library with itself in 5.3 min with FASTA on a third-generation 32 node Intel iPSC/860 hypercube.
The FASTA program can search the NBRF protein sequence library (2.5 million residues) in less than 20 min on an IBM-PC microcomputer and unambiguously detect proteins that shared a common ancestor billions of years in the past. FASTA is both fast and selective because it initially considers only amino acid identities. Its sensitivity is increased not only by using the PAM250 matrix to score and rescore regions with large numbers of identities but also by joining initial regions. The results of searches with FASTA compare favorably with results using NWS-based programs that are 100 times slower. FASTA is slightly less sensitive but considerably more selective. It is not clear that NWS-based programs would be more successful in finding distantly related members of the G-protein-coupled receptor family. The joining step by FASTA to calculate the initn score is especially useful for sequences that share regions of sequence similarity that are separated by variable-length loops. FASTP and FASTA were designed to identify protein sequences that have descended from a common ancestor, and they have proved very useful for this task. In many cases, a FASTA sequence search will result in a list of high scoring library sequences that are homologous to the query sequence, or the search will result in a list of sequences with similarity scores that cannot be distinguished from the bulk of the library. In either case, the question of whether there are sequences in the library that are clearly related to the query sequence has been answered unambiguously. Unfortunately, the results often will not be so clear-cut, and careful analysis of similarity scores, statistical significance, the actual aligned residues, and the biological context are required. In the course of analyzing the G-protein-coupled receptor family, several proteins were found that, because of a high initn score and a low init1 score that increased almost 2-fold with optimization, appeared to be members of this family which were not previously recognized. RDF2 analysis showed borderline z values, and only a careful examination of the sequence alignments that focused on the conserved residues provided convincing evidence that the high scores were fortuitous. As sequence comparison methods become more powerful by becoming more sensitive, they become more likely to mislead, and even greater care is required.