Search PubMed⌕ Search

Biomedical subjects

Michael Gribskov

Publications and source records attributed to Michael Gribskov.

11 recordsLinked to original sources

2HAPI: a microarray data analysis system.

SUMMARY: 2HAPI (version 2 of High density Array Pattern Interpreter) is a web-based, publicly-available analytical tool designed to aid researchers in microarray data analysis. 2HAPI includes tools for searching, manipulating, visualizing, and clustering the large sets of data generated by microarray experiments. Other features include association of genes with NCBI information and linkage to external data resources. Unique to 2HAPI is the ability to retrieve upstream sequences of co-regulated genes for promoter analysis using MEME (Multiple Expectation-maximization for Motif Elicitation) AVAILABILITY: 2HAPI is freely available at http://array.sdsc.edu. Users can try 2HAPI anonymously with pre-loaded data or they can register as a 2HAPI user and upload their data.

Algorithms↗

The PlantsP and PlantsT Functional Genomics Databases.

PlantsP and PlantsT allow users to quickly gain a global understanding of plant phosphoproteins and plant membrane transporters, respectively, from evolutionary relationships to biochemical function as well as a deep understanding of the molecular biology of individual genes and their products. As one database with two functionally different web interfaces, PlantsP and PlantsT are curated plant-specific databases that combine sequence-derived information with experimental functional-genomics data. PlantsP focuses on proteins involved in the phosphorylation process (i.e., kinases and phosphatases), whereas PlantsT focuses on membrane transport proteins. Experimentally, PlantsP provides a resource for information on a collection of T-DNA insertion mutants (knockouts) in each kinase and phosphatase, primarily in Arabidopsis thaliana, and PlantsT uniquely combines experimental data regarding mineral composition (derived from inductively coupled plasma atomic emission spectroscopy) of mutant and wild-type strains. Both databases provide extensive information on motifs and domains, detailed information contributed by individual experts in their respective fields, and descriptive information drawn directly from the literature. The databases incorporate a unique user annotation and review feature aimed at acquiring expert annotation directly from the plant biology community. PlantsP is available at http://plantsp.sdsc.edu and PlantsT is available at http://plantst.sdsc.edu.

Arabidopsis↗

MODULEWRITER: a program for automatic generation of database interfaces.

MODULEWRITER is a PERL object relational mapping (ORM) tool that automatically generates database specific application programming interfaces (APIs) for SQL databases. The APIs consist of a package of modules providing access to each table row and column. Methods for retrieving, updating and saving entries are provided, as well as other generally useful methods (such as retrieval of the highest numbered entry in a table). MODULEWRITER provides for the inclusion of user-written code, which can be preserved across multiple runs of the MODULEWRITER program.

Computational Biology↗

Challenges in data management for functional genomics.

Biological databases face challenges in four main areas: (1). integration, interoperation and federation; (2). ontologies and definitions of semantics; (3). community annotation; and (4). integration of data analysis tools with databases. Each of these areas provides interesting targets for research and development.

Computational Biology↗

The Arabidopsis CDPK-SnRK superfamily of protein kinases.

The CDPK-SnRK superfamily consists of seven types of serine-threonine protein kinases: calcium-dependent protein kinase (CDPKs), CDPK-related kinases (CRKs), phosphoenolpyruvate carboxylase kinases (PPCKs), PEP carboxylase kinase-related kinases (PEPRKs), calmodulin-dependent protein kinases (CaMKs), calcium and calmodulin-dependent protein kinases (CCaMKs), and SnRKs. Within this superfamily, individual isoforms and subfamilies contain distinct regulatory domains, subcellular targeting information, and substrate specificities. Our analysis of the Arabidopsis genome identified 34 CDPKs, eight CRKs, two PPCKs, two PEPRKs, and 38 SnRKs. No definitive examples were found for a CCaMK similar to those previously identified in lily (Lilium longiflorum) and tobacco (Nicotiana tabacum) or for a CaMK similar to those in animals or yeast. CDPKs are present in plants and a specific subgroup of protists, but CRKs, PPCKs, PEPRKs, and two of the SnRK subgroups have been found only in plants. CDPKs and at least one SnRK have been implicated in decoding calcium signals in Arabidopsis. Analysis of intron placements supports the hypothesis that CDPKs, CRKs, PPCKs and PEPRKs have a common evolutionary origin; however there are no conserved intron positions between these kinases and the SnRK subgroup. CDPKs and SnRKs are found on all five Arabidopsis chromosomes. The presence of closely related kinases in regions of the genome known to have arisen by genome duplication indicates that these kinases probably arose by divergence from common ancestors. The PlantsP database provides a resource of continuously updated information on protein kinases from Arabidopsis and other plants.

Arabidopsis↗

Arabidopsis proteins containing similarity to the universal stress protein domain of bacteria.

We have collected a set of 44 Arabidopsis proteins with similarity to the USPA (universal stress protein A of Escherichia coli) domain of bacteria. The USPA domain is found either in small proteins, or it makes up the N-terminal portion of a larger protein, usually a protein kinase. Phylogenetic tree analysis based upon a multiple sequence alignment of the USPA domains shows that these domains of protein kinases 1.3.1 and 1.3.2 form distinct groups, as do the protein kinases 1.4.1. This indicates that their USPA domain structures have diverged appreciably and suggests that they may subserve distinct cellular functions. Two USPA fold classes have been proposed: one based on Methanococcus jannaschii MJ0577 (1MJH) that binds ATP, and the other based on the Haemophilus influenzae universal stress protein (1JMV), highly similar to E. coli UspA, which does not bind ATP. A set of common residues involved in ATP binding in 1MJH and conserved in similar bacterial sequences is also found in a distinct cluster of Arabidopsis sequences. Threading analysis, which examines aspects of secondary and tertiary structure, confirms this Arabidopsis sequence cluster as highly similar to 1MJH. This structural approach can distinguish between the characteristic fold differences of 1MJH-like and 1JMV-like bacterial proteins and was used to assign the complete set of candidate Arabidopsis proteins to one of these fold classes. It is clear that all the plant sequences have arisen from a 1MJH-like ancestor.

Adenosine Triphosphate↗

Systematic trans-genomic comparison of protein kinases between Arabidopsis and Saccharomyces cerevisiae.

The genome of the budding yeast (Saccharomyces cerevisiae) provides an important paradigm for transgenomic comparisons with other eukaryotic species. Here, we report a systematic comparison of the protein kinases of yeast (119 kinases) and a reference plant Arabidopsis (1,019 kinases). Using a whole-protein-based, hierarchical clustering approach, the complete set of protein kinases from both species were clustered. We validated our clustering by three observations: (a) clustering pattern of functional orthologs proven in genetic complementation experiments, (b) consistency with reported classifications of yeast kinases, and (c) consistency with the biochemical properties of those Arabidopsis kinases already experimentally characterized. The clustering pattern identified no overlap between yeast kinases and the receptor-like kinases (RLKs) of Arabidopsis. Ten more kinase families were found to be specific for one of the two species. Among them, the calcium-dependent protein kinase and phosphoenolpyruvate carboxylase kinase families are specific for plants, whereas the Ca(2+)/calmodulin-dependent protein kinase and provirus insertion in mouse-like kinase families were found only in yeast and animals. Three yeast kinase families, nitrogen permease reactivator/halotolerance-5), polyamine transport kinase, and negative regulator of sexual conjugation and meiosis, are absent in both plants and animals. The majority of yeast kinase families (21 of 26) display Arabidopsis counterparts, and all are mapped into Arabidopsis families of intracellular kinases that are not related to RLKs. Representatives from 11 of the common families (54 kinases from Arabidopsis and 17 from yeast) share an extremely high degree of similarity (blast E value < 10(-80)), suggesting the likelihood of orthologous functions. Selective expansion of yeast kinase families was observed in Arabidopsis. This is most evident for yeast genes CBK1, HRR25, and SNF1 and the kinase family S6K. Reduction of kinase families was also observed, as in the case of the NEK-like family. The distinguishing features between the two sets of kinases are the selective expansion of yeast families and the generation of a limited number of new kinase families for new functionality in Arabidopsis, most notably, the Arabidopsis RLKs that constitute important components of plant intercellular communication apparatus.

Arabidopsis↗

Genomic comparison of P-type ATPase ion pumps in Arabidopsis and rice.

Members of the P-type ATPase ion pump superfamily are found in all three branches of life. Forty-six P-type ATPase genes were identified in Arabidopsis, the largest number yet identified in any organism. The recent completion of two draft sequences of the rice (Oryza sativa) genome allows for comparison of the full complement of P-type ATPases in two different plant species. Here, we identify a similar number (43) in rice, despite the rice genome being more than three times the size of Arabidopsis. The similarly large families suggest that both dicots and monocots have evolved with a large preexisting repertoire of P-type ATPases. Both Arabidopsis and rice have representative members in all five major subfamilies of P-type ATPases: heavy-metal ATPases (P1B), Ca2+-ATPases (endoplasmic reticulum-type Ca2+-ATPase and autoinhibited Ca2+-ATPase, P2A and P2B), H+-ATPases (autoinhibited H+-ATPase, P3A), putative aminophospholipid ATPases (ALA, P4), and a branch with unknown specificity (P5). The close pairing of similar isoforms in rice and Arabidopsis suggests potential orthologous relationships for all 43 rice P-type ATPases. A phylogenetic comparison of protein sequences and intron positions indicates that the common angiosperm ancestor had at least 23 P-type ATPases. Although little is known about unique and common features of related pumps, clear differences between some members of the calcium pumps indicate that evolutionarily conserved clusters may distinguish pumps with either different subcellular locations or biochemical functions.

Adenosine Triphosphatases↗

Homophila: human disease gene cognates in Drosophila.

Although many human genes have been associated with genetic diseases, knowing which mutations result in disease phenotypes often does not explain the etiology of a specific disease. Drosophila melanogaster provides a powerful system in which to use genetic and molecular approaches to investigate human genetic diseases. Homophila is an intergenomic resource linking the human and fly genomes in order to stimulate functional genomic investigations in Drosophila that address questions about genetic disease in humans. Homophila provides a comprehensive linkage between the disease genes compiled in Online Mendelian Inheritance in Man (OMIM) and the complete Drosophila genomic sequence. Homophila is a relational database that allows searching based on human disease descriptions, OMIM number, human or fly gene names, and sequence similarity, and can be accessed at http://homophila.sdsc.edu.

Animals↗

Estimating and evaluating the statistics of gapped local-alignment scores.

We present a novel maximum-likelihood-based algorithm for estimating the distribution of alignment scores from the scores of unrelated sequences in a database search. Using a new method for measuring the accuracy of p-values, we show that our maximum-likelihood-based algorithm is more accurate than existing regression-based and lookup table methods. We explore a more sophisticated way of modeling and estimating the score distributions (using a two-component mixture model and expectation maximization), but conclude that this does not improve significantly over simply ignoring scores with small E-values during estimation. Finally, we measure the classification accuracy of p-values estimated in different ways and observe that inaccurate p-values can, somewhat paradoxically, lead to higher classification accuracy. We explain this paradox and argue that statistical accuracy, not classification accuracy, should be the primary criterion in comparisons of similarity search methods that return p-values that adjust for target sequence length.

Algorithms↗

The complement of protein phosphatase catalytic subunits encoded in the genome of Arabidopsis.

Reversible protein phosphorylation is critically important in the modulation of a wide variety of cellular functions. Several families of protein phosphatases remove phosphate groups placed on key cellular proteins by protein kinases. The complete genomic sequence of the model plant Arabidopsis permits a comprehensive survey of the phosphatases encoded by this organism. Several errors in the sequencing project gene models were found via analysis of predicted phosphatase coding sequences. Structural sequence probes from aligned and unaligned sequence models, and all-against-all BLAST searches, were used to identify 112 phosphatase catalytic subunit sequences, distributed among the serine (Ser)/threonine (Thr) phosphatases (STs) of the protein phosphatase P (PPP) family, STs of the protein phosphatase M (PPM) family (protein phosphatases 2C [PP2Cs] subfamily), protein tyrosine (Tyr) phosphatases (PTPs), low-M(r) protein Tyr phosphatases, and dual-specificity (Tyr and Ser/Thr) phosphatases (DSPs). The Arabidopsis genome contains an abundance of PP2Cs (69) and a dearth of PTPs (one). Eight sequences were identified as new protein phosphatase candidates: five dual-specificity phosphatases and three PP2Cs. We used phylogenetic analyses to infer clustering patterns reflecting sequence similarity and evolutionary ancestry. These clusters, particularly for the largely unexplored PP2C set, will be a rich source of material for plant biologists, allowing the systematic sampling of protein function by genetic and biochemical means.

Animals↗