Search PubMedSearch

Biomedical subjects

T D Wu

Publications and source records attributed to T D Wu.

10 recordsLinked to original sources

Highly specific protein sequence motifs for genome analysis.

We present a method for discovering conserved sequence motifs from families of aligned protein sequences. The method has been implemented as a computer program called EMOTIF (http://motif. stanford.edu/emotif). Given an aligned set of protein sequences, EMOTIF generates a set of motifs with a wide range of specificities and sensitivities. EMOTIF also can generate motifs that describe possible subfamilies of a protein superfamily. A disjunction of such motifs often can represent the entire superfamily with high specificity and sensitivity. We have used EMOTIF to generate sets of motifs from all 7,000 protein alignments in the BLOCKS and PRINTS databases. The resulting database, called IDENTIFY (http://motif. stanford.edu/identify), contains more than 50,000 motifs. For each alignment, the database contains several motifs having a probability of matching a false positive that range from 10(-10) to 10(-5). Highly specific motifs are well suited for searching entire proteomes, while generating very few false predictions. IDENTIFY assigns biological functions to 25-30% of all proteins encoded by the Saccharomyces cerevisiae genome and by several bacterial genomes. In particular, IDENTIFY assigned functions to 172 of proteins of unknown function in the yeast genome.

Amino Acid Sequence

Regression analysis of multiple protein structures.

A general framework is presented for analyzing multiple protein structures using statistical regression methods. The regression approach can superimpose protein structures rigidly or with shear. Also, this approach can superimpose multiple structures explicitly, without resorting to pairwise superpositions. The algorithm alternates between matching corresponding landmarks among the protein structures and superimposing these landmarks. Matching is performed using a robust dynamic programming technique that uses gap penalties that adapt to the given data. Superposition is performed using either orthogonal transformations, which impose the rigid-body assumption, or affine transformations, which allow shear. The resulting regression model of a protein family measures the amount of structural variability at each landmark. A variation of our algorithm permits a separate weight for each landmark, thereby allowing one to emphasize particular segments of a protein structure or to compensate for variances that differ at various positions in a structure. In addition, a method is introduced for finding an initial correspondence, by measuring the discrete curvature along each protein backbone. Discrete curvature also characterizes the secondary structure of a protein backbone, distinguishing among helical, strand, and loop regions. An example is presented involving a set of seven globin structures. Regression analysis, using both affine and orthogonal transformations, reveals that globins are most strongly conserved structurally in helical regions, particularly in the mid-regions of the E, F, and G helices.

Algorithms

Characterization of a desiccation-related protein in lily pollen during development and stress.

This work characterizes a lily (Lilium longiflorum Thunb. cv. Snow Queen) anther (LLA) protein associated with desiccation. Peptide mapping analysis revealed that the abundant LLA-23 doublet contained similar polypeptides, having an isoelectric point of 6.1. Immunoblots of pollen protein from developing anther/pollen confirmed that the LLA-23 protein accumulated only at the later stage of pollen maturation and that the levels remained steady in mature and vital pollen. The accumulation of LLA-23 proteins was correlated with desiccation that naturally occurred in pollen. Subcellular fractionation of pollen proteins revealed that the protein was located in the cytoplasmic fraction. Premature drying of developing pollen confirmed that the concomitant accumulation of LLA-23 was associated with desiccation. Peptide sequence analysis demonstrates similarities between the lily LLA-23 and a family of water-deficit/ripening-induced proteins including LP3 of pine, DS2 of potato, and Asr of tomato and pummelo. In addition, the concomitant accumulation of LLA-23 can be experimentally manipulated by methyl jasmonate (Me-JA) and salicylic acid (SA) as well as by mannitol and methyl viologen. The LLA-23 represents a novel member of the water-deficit/ripening-induced proteins.

Acetates

Clinical course of gamma-hydroxybutyrate overdose.

STUDY OBJECTIVE: To describe the clinical characteristics and course of gamma-hydroxybutyrate (GHB) overdose. METHODS: We assembled a retrospective series of all cases of GHB ingestion see in an urban public-hospital emergency department and entered in a computerized database January 1993 through December 1996. From these cases we extracted demographic information, concurrent drug use, vital signs, Glasgow Coma Scale (GCS) score, laboratory values, and clinical course. RESULTS: Sixty-one (69%) of the 88 patients were male. The mean age was 28 years. Thirty-four cases (39%) involved coingestion of ethanol, and 25 (28%) involved coingestion of another drug, most commonly amphetamines. Twenty-five cases (28%) had a GCS score of 3, and 28 (33%) had scores ranging from 4 through 8. The mean time to regained consciousness from initial presentation among nonintubated patients with an initial GCS of 13 or less was 146 minutes (range, 16-389). Twenty-two patients (31%) had an initial temperature of 35 degrees C or less. Thirty-two (36%) had asymptomatic bradycardia; in 29 of these cases, the initial GCS score was 8 or less. Ten patients (11%) presented with hypotension (systolic blood pressure < or = 90 mm Hg); 6 of these patients also demonstrated concurrent bradycardia. Arterial blood gases were measured in 30 patients; 21 had a PCO2 of 45 or greater, with pH ranging from 7.24 to 7.34, consistent with mild acute respiratory acidosis. Twenty-six patients (30%) had an episode of emesis; in 22 of these cases, the initial GCS was 8 or less. CONCLUSION: In our study population, patients who overdosed on GHB presented with a markedly decreased level of consciousness. Coingestion of ethanol or other drugs is common, as are bradycardia, hypothermia, respiratory acidosis, and emesis. Hypotension occurs occasionally. Patients typically regain consciousness spontaneously within 5 hours of the ingestion.

Adjuvants, Anesthesia

Modeling and superposition of multiple protein structures using affine transformations: analysis of the globins.

A novel approach for analyzing multiple protein structures is presented. A family of related protein structures may be characterized by an affine model, obtained by applying transformation matrices that permit both rotation and shear. The affine model and transformation matrices can be computed efficiently using a single eigen-decomposition. A novel method for finding correspondences is also introduced. This method matches curvatures along the protein backbone. The algorithm is applied to analyze a set of seven globin structures. Our method identifies 100 corresponding landmarks across all seven structures. Results show that most helices in globins can be identified by high curvature, with the exception of the C and D helices. Analysis of the superposition reveals that globins are most strongly conserved structurally in the mid-regions of the E and G helices.

Algorithms

Enumerating and ranking discrete motifs.

Discrete motifs that discriminate functional classes of proteins are useful for classifying new sequences, capturing structural constraints, and identifying protein subclasses. Despite the fact that the space of such motifs can grow exponentially with sequence length and number, we show that in practice it usually does not, and we describe a technique that infers motifs from aligned protein sequences by exhaustively searching this space. Our method generates sequence motifs over a wide range of recall and precision, and chooses a representative motif based on a score that we derive from both statistical and information-theoretic frameworks. Finally, we show that the selected motifs perform well in practice, classifying unseen sequences with extremely high precision, and infer protein subclasses that correspond to known biochemical classes.

Algorithms

A segment-based dynamic programming algorithm for predicting gene structure.

An algorithm called segment-based dynamic programming is described for predicting gene structure from a sequence of genomic DNA. The algorithm explores the space of gene structures that satisfy junctional and frame constraints and finds the gene structure that optimizes the sum of junctional and segmental scoring functions. Junctional constraints specify acceptable sites of initiation, termination, and splicing, whereas frame constraints ensure that the total exon length is a multiple of three and that no in-frame stop codons occur within exons or at exon-exon junctions. By computing over segments, segment-based dynamic programming maintains reading frame and phase information for each segment, it can assemble exons in-frame as well as score them in-frame. The algorithm is used to quantify the computational power of constraints. Experimental results show that frame constraints reduce the size of the search space by several orders of magnitude and that cardinality constraints place an asymptotic limit on the size of the search space. The algorithm is also used to compare the accuracy of different methods for assembly and scoring. A scoring scheme based on fifth-order Markov hexamer frequencies is presented and used in three objective functions, corresponding to in-frame, frame-independent, and frame-maximal scoring strategies. Experimental results show that in-frame assembly improves specificity only slightly over frame-independent assembly, whereas in-frame scoring improves specificity substantially over frame-independent and frame-maximal scoring.

Algorithms

Discovering empirically conserved amino acid substitution groups in databases of protein families.

This paper introduces a method for identifying empirically conserved amino acid substitution groups. In contrast with existing approaches that view amino acid substitution as a pairwise phenomenon, the method presented here identifies conserved groups of amino acids using a data structure called a conditional distribution matrix. The conditional distribution matrix extends the concept of a pairwise substitution matrix by changing the context of substitution from a single amino acid to a group of amino acids. The matrix tabulates information from a database of protein families that contains numerous aligned positions. Each row in the matrix contains the distribution of amino acids in those aligned positions that contain a given conditioning group of amino acids. The method converts a database of protein families into a conditional distribution matrix and then examines each possible substitution group for evidence of conservation. The algorithm is applied to the BLOCKS and HSSP databases. Twenty amino acid substitution groups are found to be conserved empirically in both databases. These groups provide insight into biochemical properties that are conserved in protein evolution.

Algorithms

Identification of protein motifs using conserved amino acid properties and partitioning techniques.

Analyzing a set of protein sequences involves a fundamental relationship between the coherency of the set and the specificity of the motif that describes it. Motifs may be obscured by training sets that contain incoherent sequences, in part due to protein subclasses, contamination, or errors. We develop an algorithm for motif identification that systematically explores possible patterns of coherency within a set of protein sequences. Our algorithm constructs alternative partitions of the training set data, where one subset of each partition is presumed to contain coherent data and is used for forming a motif. The motif is represented by multiple overlapping amino acid groups based on evolutionary, biochemical, or physical properties. We demonstrate our method on a training set of reverse transcriptases that contains subclasses, sequence errors, misalignments, and contaminating sequences. Despite these complications, our program identifies a novel motif for the subclass of retroviral and retrovirus-related reverse transcriptases. This motif has a much higher specificity than previously reported motifs and suggests the importance of conserved hydrophilic and hydrophobic residues in the structure of reverse transcriptases.

Algorithms

A problem decomposition method for efficient diagnosis and interpretation of multiple disorders.

Diagnosis of multiple disorders can be made more efficient by reasoning explicitly about problem decompositions. A diagnostic problem can be decomposed by hypothesizing about common and disjoint cause relationships among the given symptoms. The resulting structure exploits computational principles of causal intersection, subproblem independence, and minimal factorability to increase efficiency. By assigning structure to a problem, the symptom decomposition approach offers a new type of decision-support task called symptom interpretation. Experimental results indicate that symptom decomposition yields substantial increases in performance compared to existing methods for multidisorder diagnosis.

Algorithms