Search PubMed⌕ Search

Biomedical subjects

Mark Craven

Publications and source records attributed to Mark Craven.

6 recordsLinked to original sources

Learning statistical models for annotating proteins with function information using biomedical text.

BACKGROUND: The BioCreative text mining evaluation investigated the application of text mining methods to the task of automatically extracting information from text in biomedical research articles. We participated in Task 2 of the evaluation. For this task, we built a system to automatically annotate a given protein with codes from the Gene Ontology (GO) using the text of an article from the biomedical literature as evidence. METHODS: Our system relies on simple statistical analyses of the full text article provided. We learn n-gram models for each GO code using statistical methods and use these models to hypothesize annotations. We also learn a set of Naïve Bayes models that identify textual clues of possible connections between the given protein and a hypothesized annotation. These models are used to filter and rank the predictions of the n-gram models. RESULTS: We report experiments evaluating the utility of various components of our system on a set of data held out during development, and experiments evaluating the utility of external data sources that we used to learn our models. Finally, we report our evaluation results from the BioCreative organizers. CONCLUSION: We observe that, on the test data, our system performs quite well relative to the other systems submitted to the evaluation. From other experiments on the held-out data, we observe that (i) the Naïve Bayes models were effective in filtering and ranking the initially hypothesized annotations, and (ii) our learned models were significantly more accurate when external data sources were used during learning.

Bayes Theorem↗

EDGE: a centralized resource for the comparison, analysis, and distribution of toxicogenomic information.

Transcriptional profiling via microarrays holds great promise for toxicant classification and hazard prediction. Unfortunately, the use of different microarray platforms, protocols, and informatics often hinders the meaningful comparison of transcriptional profiling data across laboratories. One solution to this problem is to provide a low-cost and centralized resource that enables researchers to share toxicogenomic data that has been generated on a common platform. In an effort to create such a resource, we developed a standardized set of microarray reagents and reproducible protocols to simplify the analysis of liver gene expression in the mouse model. This resource, referred to as EDGE, was then used to generate a training set of 117 publicly accessible transcriptional profiles that can be accessed at http://edge.oncology.wisc.edu/. The Web-accessible database was also linked to an informatics suite that allows on-line clustering and K-means analyses as well as Boolean and sequence-based searches of the data. We propose that EDGE can serve as a prototype resource for the sharing of toxicogenomics information and be used to develop algorithms for efficient chemical classification and hazard prediction.

Animals↗

Interaction networks in yeast define and enumerate the signaling steps of the vertebrate aryl hydrocarbon receptor.

The aryl hydrocarbon receptor (AHR) is a vertebrate protein that mediates the toxic and adaptive responses to dioxins and related environmental pollutants. In an effort to better understand the details of this signal transduction pathway, we employed the yeast S. cerevisiae as a model system. Through the use of arrayed yeast strains harboring ordered deletions of open reading frames, we determined that 54 out of the 4,507 yeast genes examined significantly influence AHR signal transduction. In an effort to describe the relationship between these modifying genes, we constructed a network map based upon their known protein and genetic interactions. Monte Carlo simulations demonstrated that this network represented a description of AHR signaling that was distinct from those generated by random chance. The network map was then explored with a number of computational and experimental annotations. These analyses revealed that the AHR signaling pathway is defined by at least five distinct signaling steps that are regulated by functional modules of interacting modifiers. These modules can be described as mediating receptor folding, nuclear translocation, transcriptional activation, receptor level, and a previously undescribed nuclear step related to the receptor's Per-Arnt-Sim domain.

Active Transport, Cell Nucleus↗

A Bayesian network approach to operon prediction.

MOTIVATION: In order to understand transcription regulation in a given prokaryotic genome, it is critical to identify operons, the fundamental units of transcription, in such species. While there are a growing number of organisms whose sequence and gene coordinates are known, by and large their operons are not known. RESULTS: We present a probabilistic approach to predicting operons using Bayesian networks. Our approach exploits diverse evidence sources such as sequence and expression data. We evaluate our approach on the Escherichia coli K-12 genome where our results indicate we are able to identify over 78% of its operons at a 10% false positive rate. Also, empirical evaluation using a reduced set of data sources suggests that our approach may have significant value for organisms that do not have as rich of evidence sources as E.coli. AVAILABILITY: Our E.coli K-12 operon predictions are available at http://www.biostat.wisc.edu/gene-regulation.

Algorithms↗

Exposure and sensitization to indoor allergens: association with lung function, bronchial reactivity, and exhaled nitric oxide measures in asthma.

BACKGROUND: Exposure to high levels of allergens in sensitized asthmatic patients causes worsening of pulmonary function in experimental studies. Chronic exposure to lower, naturally occurring levels of allergens might increase the severity of asthma. OBJECTIVE: We sought to study the associations between sensitization and exposure to common indoor allergens (dust mite, cat, and dog) in the home on pulmonary function, exhaled nitric oxide (eNO), and airway reactivity in asthmatic patients. METHODS: Dust samples were collected from the living room carpet and mattress of 311 subject's homes, and Der p 1, Fel d 1, and Can f 1 concentrations were measured by using ELISAs. Spirometry, nonspecific bronchial reactivity, and eNO were measured. RESULTS: Subjects both sensitized and exposed to high levels of sensitizing allergen had significantly lower FEV(1) percent predicted values (mean, 83.7% vs 89.3%; mean difference, 5.6%; 95% CI, 0.6%-10.6%; P =.03), higher eNO values (geometric mean [GM], 12.8 vs 8.7 ppb; GM ratio, 0.7; 95% CI, 0.5-0.8; P =.001), and more severe airways reactivity (PD(20) GM, 0.25 vs 0.73 mg; GM ratio, 2.9; 95% CI, 1.6-5.0; P <.001) compared with subjects not sensitized and exposed. No significant effect of the interaction between sensitization and exposure was found for FEV(1) percent predicted and eNO values. However, there was a significant effect of the interaction between sensitization and exposure to any allergen (P =.05) and between sensitization and exposure to cat allergen (P =.04) for nonspecific bronchial reactivity. CONCLUSION: Asthmatic subjects who are exposed in their homes to allergens to which they are sensitized have a more severe form of the disease.

Adolescent↗

Predicting bacterial transcription units using sequence and expression data.

MOTIVATION: A key aspect of elucidating gene regulation in bacterial genomes is identifying the basic units of transcription. We present a method, based on probabilistic language models, that we apply to predict operons, promoters and terminators in the genome of Escherichia coli K-12. Our approach has two key properties: (i) it provides a coherent set of predictions for related regulatory elements of various types and (ii) it takes advantage of both DNA sequence and gene expression data, including expression measurements from inter-genic probes. RESULTS: Our experimental results show that we are able to predict operons and localize promoters and terminators with high accuracy. Moreover, our models that use both sequence and expression data are more accurate than those that use only one of these two data sources.

Algorithms↗