Search PubMed⌕ Search

Biomedical subjects

Olivier Bodenreider

Publications and source records attributed to Olivier Bodenreider.

At least 19 recordsLinked to original sources

Mapping data elements to terminological resources for integrating biomedical data sources.

BACKGROUND: Data integration is a crucial task in the biomedical domain and integrating data sources is one approach to integrating data. Data elements (DEs) in particular play an important role in data integration. We combine schema- and instance-based approaches to mapping DEs to terminological resources in order to facilitate data sources integration. METHODS: We extracted DEs from eleven disparate biomedical sources. We compared these DEs to concepts and/or terms in biomedical controlled vocabularies and to reference DEs. We also exploited DE values to disambiguate underspecified DEs and to identify additional mappings. RESULTS: 82.5% of the 474 DEs studied are mapped to entries of a terminological resource and 74.7% of the whole set can be associated with reference DEs. Only 6.6% of the DEs had values that could be semantically typed. CONCLUSION: Our study suggests that the integration of biomedical sources can be achieved automatically with limited precision and largely facilitated by mapping DEs to terminological resources.

Abstracting and Indexing↗

Global similarity and local divergence in human and mouse gene co-expression networks.

BACKGROUND: A genome-wide comparative analysis of human and mouse gene expression patterns was performed in order to evaluate the evolutionary divergence of mammalian gene expression. Tissue-specific expression profiles were analyzed for 9,105 human-mouse orthologous gene pairs across 28 tissues. Expression profiles were resolved into species-specific coexpression networks, and the topological properties of the networks were compared between species. RESULTS: At the global level, the topological properties of the human and mouse gene coexpression networks are, essentially, identical. For instance, both networks have topologies with small-world and scale-free properties as well as closely similar average node degrees, clustering coefficients, and path lengths. However, the human and mouse coexpression networks are highly divergent at the local level: only a small fraction (<10%) of coexpressed gene pair relationships are conserved between the two species. A series of controls for experimental and biological variance show that most of this divergence does not result from experimental noise. We further show that, while the expression divergence between species is genuinely rapid, expression does not evolve free from selective (functional) constraint. Indeed, the coexpression networks analyzed here are demonstrably functionally coherent as indicated by the functional similarity of coexpressed gene pairs, and this pattern is most pronounced in the conserved human-mouse intersection network. Numerous dense network clusters show evidence of dedicated functions, such as spermatogenesis and immune response, that are clearly consistent with the coherence of the expression patterns of their constituent gene members. CONCLUSION: The dissonance between global versus local network divergence suggests that the interspecies similarity of the global network properties is of limited biological significance, at best, and that the biologically relevant aspects of the architectures of gene coexpression are specific and particular, rather than universal. Nevertheless, there is substantial evolutionary conservation of the local network structure which is compatible with the notion that gene coexpression networks are subject to purifying selection.

Animals↗

Bio-ontologies: current trends and future directions.

In recent years, as a knowledge-based discipline, bioinformatics has been made more computationally amenable. After its beginnings as a technology advocated by computer scientists to overcome problems of heterogeneity, ontology has been taken up by biologists themselves as a means to consistently annotate features from genotype to phenotype. In medical informatics, artifacts called ontologies have been used for a longer period of time to produce controlled lexicons for coding schemes. In this article, we review the current position in ontologies and how they have become institutionalized within biomedicine. As the field has matured, the much older philosophical aspects of ontology have come into play. With this and the institutionalization of ontology has come greater formality. We review this trend and what benefits it might bring to ontologies and their use within biomedicine.

Animals↗

Experience in reasoning with the foundational model of anatomy in OWL DL.

The objective of this study is to compare description logics (DLs) and frames for representing large-scale biomedical ontologies and reasoning with them. The ontology under investigation is the Foundational Model of Anatomy (FMA). We converted it from its frame-based representation in Protégé into OWL DL. The OWL reasoner Racer helped identify unsatisfiable classes in the FMA. Support for consistency checking is clearly an advantage of using DLs rather than frames. The interest of reclassification was limited, due to the difficulty of defining necessary and sufficient conditions for anatomical entities. The sheer size and complexity of the FMA was also an issue.

Computational Biology↗

Law and order: assessing and enforcing compliance with ontological modeling principles in the Foundational Model of Anatomy.

The objective of this study is to provide an operational definition of principles with which well-formed ontologies should comply. We define 15 such principles, related to classification (e.g., no hierarchical cycles are allowed; concepts have a reasonable number of children), incompatible relationships (e.g., two concepts cannot stand both in a taxonomic and partitive relation), dependence among concepts, and the co-dependence of equivalent sets of relations. Implicit relations--embedded in concept names or inferred from a combination of explicit relations--are used in this process in addition to the relations explicitly represented. As a case study, we investigate the degree to which the Foundational Model of Anatomy (FMA)--a large ontology of anatomy--complies with these 15 principles. The FMA succeeds in complying with all the principles: totally with one and mostly with the others. Reasons for non-compliance are analyzed and suggestions are made for implementing effective enforcement mechanisms in ontology development environments. The limitations of this study are also discussed.

Anatomy↗

Non-lexical approaches to identifying associative relations in the gene ontology.

The Gene Ontology (GO) is a controlled vocabulary widely used for the annotation of gene products. GO is organized in three hierarchies for molecular functions, cellular components, and biological processes but no relations are provided among terms across hierarchies. The objective of this study is to investigate three non-lexical approaches to identifying such associative relations in GO and compare them among themselves and to lexical approaches. The three approaches are: computing similarity in a vector space model, statistical analysis of co-occurrence of GO terms in annotation databases, and association rule mining. Five annotation databases (FlyBase, the Human subset of GOA, MGI, SGD, and WormBase) are used in this study. A total of 7,665 associations were identified by at least one of the three non-lexical approaches. Of these, 12% were identified by more than one approach. While there are almost 6,000 lexical relations among GO terms, only 203 associations were identified by both non-lexical and lexical approaches. The associations identified in this study could serve as the starting point for adding associative relations across hierarchies to GO, but would require manual curation. The application to quality assurance of annotation databases is also discussed.

Alzheimer Disease↗

Genestrace: phenomic knowledge discovery via structured terminology.

The era of applied genomic medicine is quickly approaching accompanied by the increasing availability of detailed genetic information. Understanding the genetic etiology behind complex, multi-gene diseases remains an important challenge. In order to uncover the putative genetic etiology of complex diseases, we designed a method that explores the relationships between two major terminological and ontological resources: the Unified Medical Language System (UMLS) and the Gene Ontology (GO). The UMLS has a mainly clinical emphasis; Gene Ontology has become the standard for biological annotations of genes and gene products. Using statistical and semantic relationships within and between the two resources, we are able to infer relationships between disease concepts in the UMLS and gene products annotated using GO and its associated databases. We validated our inferences by comparing them to the known gene-disease relationships, as defined in the Online Mendelian Inheritance in Man's morbidmap (OMIM). The proof-of-concept methods presented here are unique in that they bypass the ambiguity of the direct extraction of gene or disease term from MEDLINE. Additionally, our methods provide direct links to clinically significant diseases through established terminologies or ontologies. The preliminary results presented here indicate the potential utility of exploiting the existing, manually curated relationships in biomedical resources as a tool for the discovery of potentially valuable new gene-disease relationships.

Computational Biology↗

Issues in the classification of disease instances with ontologies.

Ontologies define classes of entities and their interrelations. They are used to organize data according to a theory of the domain. Towards that end, ontologies provide class definitions (i.e., the necessary and sufficient conditions for defining class membership). In medical ontologies, it is often difficult to establish such definitions for diseases. We use three examples (anemia, leukemia and schizophrenia) to illustrate the limitations of ontologies as classification resources. We show that eligibility criteria are often more useful than the Aristotelian definitions traditionally used in ontologies. Examples of eligibility criteria for diseases include complex predicates such as ' x is an instance of the class C when at least n criteria among m are verified' and 'symptoms must last at least one month if not treated, but less than one month, if effectively treated'. References to normality and abnormality are often found in disease definitions, but the operational definition of these references (i.e., the statistical and contextual information necessary to define them) is rarely provided. We conclude that knowledge bases that include probabilistic and statistical knowledge as well as rule-based criteria are more useful than Aristotelian definitions for representing the predicates defined by necessary and sufficient conditions. Rich knowledge bases are needed to clarify the relations between individuals and classes in various studies and applications. However, as ontologies represent relations among classes, they can play a supporting role in disease classification services built primarily on knowledge bases.

Humans↗

Of mice and men: aligning mouse and human anatomies.

This paper reports on the alignment between mouse and human anatomies, a critical resource for comparative science as diseases in mice are used as mod-els of human disease. The two ontologies under investigation are the NCI Thesaurus (human anatomy) and the Adult Mouse Anatomical Dictionary, each comprising about 2500 anatomical concepts. This study compares two approaches to aligning ontologies. One is fully automatic, based on a combination of lexical and structural similarity; the other is manual. The resulting mappings were evaluated by an expert. 715 and 781 mappings were identified by each method respectively, of which 639 are common to both and all valid. The applications of the map-ping are discussed from the perspective of biology and from that of ontology.

Anatomy↗

Classifying diseases with respect to anatomy: a study in SNOMED CT.

Anatomy is a major organizing principle for dis-eases. In the formal definitions provided by SNOMED CT, for example, the role 'finding site' relates disorders to anatomical entities. This study investigates SNOMED CT and compares the anatomy-based classification of diseases supported by the role finding site to the anatomy-based classification of diseases provided by subsumption (is-a) relations between diseases. For each of the 3,540 anatomical entities associated with disorders,, we compared two sets of disorders: first, the set of disorders associated with any descendant of the anatomical entity under investigation (ANAT); second, the set of dis-orders corresponding to the union of the descendants of the disorders associated with the anatomical entity under investigation (TAXO). The ANAT and TAXO sets were different for 1,231 anatomical entities (35%). In 607 cases, the overlap between ANAT and TAXO was less than 50%. When a difference was found, the TAXO set was always a subset of the ANAT set. Among the 1,025,904 subsumption relations among disorders generated by the ANAT approach, 40% were not present in TAXO. This approach helps identify missing classes and taxonomic relations in existing ontologies. It can be generalized to other kinds of partitions of biomedical ontologies.

Anatomy↗

Utilizing the UMLS for semantic mapping between terminologies.

An algorithm was derived to find candidate mappings between any two terminologies inside the UMLS, making use of synonymy, explicit mapping relations and hierarchical relationships among UMLS concepts. Using an existing set of mappings from SNOMED CT to ICD9CM as our gold standard, we managed to find candidate mappings for 86% of SNOMED CT terms, with recall of 42% and precision of 20%. Among the various methods used, mapping by UMLS synonymy was particularly accurate and could potentially be useful as a quality assurance tool in the creation of mapping sets or in the UMLS editing process. Other strengths and weaknesses of the algorithm are discussed.

Algorithms↗

Approaches to eliminating cycles in the UMLS Metathesaurus: naïve vs. formal.

Applications exploiting the hierarchical relations recorded in the Unified Medical Language System (UMLS) Metathesaurus suffer from the presence of inconsistencies in these relations. A formal approach to identifying and eliminating circular hierarchical relations has been proposed in previous work, leading to the creation of a directed acyclic Metathesaurus graph. However, this approach is at best semi-automatic and its implementation is far from trivial. A simpler, alternative approach consists in avoiding loops while traversing the Metathesaurus graph by preventing nodes from being visited twice. Our objective is to evaluate the benefit of the formal approach to eliminating cycles over a naïve approach to avoiding them. To this end, we compared the size and semantic coherence of sets of descendants obtained by both approaches. 12% of the concepts with descendants exhibit some differences. The formal approach significantly reduces the number of descendants in these cases. The benefits in terms of semantic coherence are more subtle.

Semantics↗

Alignment of multiple ontologies of anatomy: deriving indirect mappings from direct mappings to a reference.

OBJECTIVE: To investigate the indirect alignment of two anatomical ontologies through a reference ontology and to compare it to direct alignment between these two ontologies. The ontologies under investigation are the Adult Mouse Anatomical Dictionary (MA) and the NCI Thesaurus (NCI). The Foundational Model of Anatomy serves as reference ontology. METHODS: The direct alignment employs a combination of lexical and structural similarity. The indirect alignment simply derives mappings from direct alignments to the reference ontology. RESULTS: The indirect MA-NCI alignment yielded 703 mappings and the direct alignment 715, 654 of which are common to both. The mappings specific to one approach were analyzed. CONCLUSIONS: When a reference ontology exists, indirect alignment of multiple ontologies through a reference represents a valid, cost-effective alternative to pairwise alignment.

Anatomy↗

The Unified Medical Language System (UMLS): integrating biomedical terminology.

The Unified Medical Language System (http://umlsks.nlm.nih.gov) is a repository of biomedical vocabularies developed by the US National Library of Medicine. The UMLS integrates over 2 million names for some 900,000 concepts from more than 60 families of biomedical vocabularies, as well as 12 million relations among these concepts. Vocabularies integrated in the UMLS Metathesaurus include the NCBI taxonomy, Gene Ontology, the Medical Subject Headings (MeSH), OMIM and the Digital Anatomist Symbolic Knowledge Base. UMLS concepts are not only inter-related, but may also be linked to external resources such as GenBank. In addition to data, the UMLS includes tools for customizing the Metathesaurus (MetamorphoSys), for generating lexical variants of concept names (lvg) and for extracting UMLS concepts from text (MetaMap). The UMLS knowledge sources are updated quarterly. All vocabularies are available at no fee for research purposes within an institution, but UMLS users are required to sign a license agreement. The UMLS knowledge sources are distributed on CD-ROM and by FTP.

Animals↗

Aligning knowledge sources in the UMLS: methods, quantitative results, and applications.

The UMLS Semantic Network and Metathesaurus are two complementary knowledge sources. While many studies compare relationships across the two structures, their alignment has never been attempted. We applied two methods based on lexical and conceptual similarity to aligning the Semantic Network with the UMLS Metathesaurus. Approximately two thirds of the semantic types could be aligned by lexical similarity. Conceptual similarity suggested mappings in all but ten cases. Potential applications enabled by the alignment are discussed, namely auditing the consistency between the Semantic Network and the Metathesaurus and extending the Semantic Network downwards. The relative contribution and limitations of the two methods used for the alignment are also discussed.

Semantics↗

Comparing associative relationships among equivalent concepts across ontologies.

Methods for comparing associative relationships across ontologies often rely solely on lexical similarity between the names of the relationships, which may lead to missed matches and inaccurate matches. In this paper, we propose a novel method based on the analysis of paths between equivalent concepts across ontologies. Patterns of relationships are identified for each associative relationship. The most frequent patterns indicate a correspondence between an associative relationship in one ontology and one relationship (or combination thereof) in the other. We applied this method to two ontologies of anatomy. Our method was able to identify the correspondence between relationships even in the absence of lexical similarity between relationship names. The various types of matches identified are discussed as well as the application of this method to detecting inconsistencies across the ontologies.

Anatomy↗