Search PubMed⌕ Search

Biomedical subjects

Michael Krauthammer

Publications and source records attributed to Michael Krauthammer.

9 recordsLinked to original sources

Genomewide linkage scan for nicotine dependence: identification of a chromosome 5 risk locus.

BACKGROUND: Nicotine dependence (ND) is costly to societies worldwide, moderately heritable, and genetically complex. Risk loci can be identified with genetic linkage analysis independent of prior physiological hypotheses. METHODS: We completed a genomewide linkage scan to map loci increasing risk for DSM-IV ND and for a quantitative assessment of ND as measured by the Fagerstrom Test for Nicotine Dependence (FTND) in a set of 634 small nuclear families ascertained on the basis of multiple individuals affected with cocaine or opioid dependence. Of these, 507 had at least two subjects affected with ND. There are two distinct populations within this sample, European-Americans (EAs) and African-Americans (AAs). RESULTS: A region on chromosome 5 was identified as containing a gene that affects risk for ND on the basis of FTND score in the AA part of our sample (logarithm of the odds [lod] score 3.04; empirically determined to be genomewide-significant, p = .0374; point p = .0001). The highest lod score observed in the EA part of the sample was on chromosome 7 (lod score 2.73). Several other "possible" risk loci were identified in either AA or EA subjects, with many of these in proximity to previously suggested risk loci from other clinical samples. Three nominally significant single-nucleotide polymorphism associations were found at the peptidylglycine alpha-amidating monooxygenase (PAM) locus under the chromosome 5 linkage peak, also in the AA part of the sample. CONCLUSIONS: These data add to the growing evidence for locations for ND risk loci, add a novel statistically significant locus important in AAs, and suggest a gene that might be contributing to this linkage signal.

Chromosome Mapping↗

Towards semantic role labeling & IE in the medical literature.

INTRODUCTION: In this work, we introduce the concept of semantic role labeling to the medical domain. We report first results of porting and adapting an existing resource, Propbank, to the medical field. Propbank is an adjunct to Penn Treebank that provides semantic annotation of predicates and the roles played by their arguments. The main aim of this work is the applicability of the Propbank frame files to predicates typically encountered in the medical literature. METHODS: We analyzed a target corpus of 610,100 abstracts, which was selected by searching for publication type "case reports". From this target corpus, we randomly selected 10,000 sample abstracts to estimate the predicate distribution, and matched the predicates from this sample to the predicates in Propbank. RESULTS: Of the 1998 unique verbs in our sample, 76% were represented in Propbank. This included the 40 most frequent verbs, which represented 49% of all predicate instances in our sample and which matched the Propbank usage in a study of representative sentences. We propose extensions to Propbank that handle medical predicates, which are not adequately covered by Propbank. CONCLUSION: We believe that semantic role labeling using Propbank is a valid approach to capture predicate relations in the medical literature.

Abstracting and Indexing↗

Molecular triangulation: bridging linkage and molecular-network information for identifying candidate genes in Alzheimer's disease.

A major challenge in human genetics is identifying the molecular basis of common heritable disorders. In contrast to rare single-gene diseases, multifactorial disorders are thought to arise from the combined effect of multiple gene variants, such that any single variant may have only a modest effect on disease susceptibility. We present a method to identify genes that may harbor a significant proportion of the genetic variation that predisposes individuals to a given multifactorial disorder. First, we perform an automated literature analysis that predicts physical interactions (edges) among candidate disease genes (seed nodes, selected on the basis of prior information) and other molecular entities. We derive models of molecular networks from this analysis and map the seed nodes to them. We then compute the graph-theoretic distance (the minimum number of edges that must be traversed) between the seed nodes and all other nodes in the network. We assume that nodes that are found in close proximity to multiple seed nodes are the best disease-related candidate genes. To evaluate this approach, we selected four seed genes, each with a proven role in Alzheimer's disease (AD). The method performed well in predicting additional network nodes that match AD gene candidates identified manually by an expert. We also show that the method prioritizes among the seed nodes themselves, rejecting false-positive seeds that are derived from (noisy) whole-genome genetic-linkage scans. We propose that this strategy will provide a valuable means to bridge genetic and genomic knowledge in the search for genetic determinants of multifactorial disorders.

Algorithms↗

Probabilistic inference of molecular networks from noisy data sources.

Information on molecular networks, such as networks of interacting proteins, comes from diverse sources that contain remarkable differences in distribution and quantity of errors. Here, we introduce a probabilistic model useful for predicting protein interactions from heterogeneous data sources. The model describes stochastic generation of protein-protein interaction networks with real-world properties, as well as generation of two heterogeneous sources of protein-interaction information: research results automatically extracted from the literature and yeast two-hybrid experiments. Based on the domain composition of proteins, we use the model to predict protein interactions for pairs of proteins for which no experimental data are available. We further explore the prediction limits, given experimental data that cover only part of the underlying protein networks. This approach can be extended naturally to include other types of biological data sources.

Algorithms↗

GeneWays: a system for extracting, analyzing, visualizing, and integrating molecular pathway data.

The immense growth in the volume of research literature and experimental data in the field of molecular biology calls for efficient automatic methods to capture and store information. In recent years, several groups have worked on specific problems in this area, such as automated selection of articles pertinent to molecular biology, or automated extraction of information using natural-language processing, information visualization, and generation of specialized knowledge bases for molecular biology. GeneWays is an integrated system that combines several such subtasks. It analyzes interactions between molecular substances, drawing on multiple sources of information to infer a consensus view of molecular networks. GeneWays is designed as an open platform, allowing researchers to query, review, and critique stored information.

Artificial Intelligence↗

Term identification in the biomedical literature.

Sophisticated information technologies are needed for effective data acquisition and integration from a growing body of the biomedical literature. Successful term identification is key to getting access to the stored literature information, as it is the terms (and their relationships) that convey knowledge across scientific articles. Due to the complexities of a dynamically changing biomedical terminology, term identification has been recognized as the current bottleneck in text mining, and--as a consequence--has become an important research topic both in natural language processing and biomedical communities. This article overviews state-of-the-art approaches in term identification. The process of identifying terms is analysed through three steps: term recognition, term classification, and term mapping. For each step, main approaches and general trends, along with the major problems, are discussed. By assessing previous work in context of the overall term identification process, the review also tries to delineate needs for future work in the field.

Abbreviations as Topic↗

A native XML database design for clinical document research.

Health-care institutions are gaining an increasing interest in exploiting the data that are gathered through electronic medical records. Narrative data, generated by transcription or direct entry, represents a far greater challenge for analytic tasks. Moreover, a small number of institutions are beginning to explore deeper structuring of narrative data using natural language processing (NLP). The data produced by NLP systems has a complex, nested structure. Current electronic medical records do not have the ability to store and retrieve data of this complexity in a suitable way.

Database Management Systems↗

Of truth and pathways: chasing bits of information through myriads of articles.

Knowledge on interactions between molecules in living cells is indispensable for theoretical analysis and practical applications in modern genomics and molecular biology. Building such networks relies on the assumption that the correct molecular interactions are known or can be identified by reading a few research articles. However, this assumption does not necessarily hold, as truth is rather an emerging property based on many potentially conflicting facts. This paper explores the processes of knowledge generation and publishing in the molecular biology literature using modelling and analysis of real molecular interaction data. The data analysed in this article were automatically extracted from 50000 research articles in molecular biology using a computer system called GeneWays containing a natural language processing module. The paper indicates that truthfulness of statements is associated in the minds of scientists with the relative importance (connectedness) of substances under study, revealing a potential selection bias in the reporting of research results. Aiming at understanding the statistical properties of the life cycle of biological facts reported in research articles, we formulate a stochastic model describing generation and propagation of knowledge about molecular interactions through scientific publications. We hope that in the future such a model can be useful for automatically producing consensus views of molecular interaction data.

Algorithms↗

Representing nested semantic information in a linear string of text using XML.

XML has been widely adopted as an important data interchange language. The structure of XML enables sharing of data elements with variable degrees of nesting as long as the elements are grouped in a strict tree-like fashion. This requirement potentially restricts the usefulness of XML for marking up written text, which often includes features that do not properly nest within other features. We encountered this problem while marking up medical text with structured semantic information from a Natural Language Processor. Traditional approaches to this problem separate the structured information from the actual text mark up. This paper introduces an alternative solution, which tightly integrates the semantic structure with the text. The resulting XML markup preserves the linearity of the medical texts and can therefore be easily expanded with additional types of information.

Programming Languages↗