Search PubMed⌕ Search

Biomedical subjects

Anita Burgun

Publications and source records attributed to Anita Burgun.

At least 19 recordsLinked to original sources

Mapping data elements to terminological resources for integrating biomedical data sources.

BACKGROUND: Data integration is a crucial task in the biomedical domain and integrating data sources is one approach to integrating data. Data elements (DEs) in particular play an important role in data integration. We combine schema- and instance-based approaches to mapping DEs to terminological resources in order to facilitate data sources integration. METHODS: We extracted DEs from eleven disparate biomedical sources. We compared these DEs to concepts and/or terms in biomedical controlled vocabularies and to reference DEs. We also exploited DE values to disambiguate underspecified DEs and to identify additional mappings. RESULTS: 82.5% of the 474 DEs studied are mapped to entries of a terminological resource and 74.7% of the whole set can be associated with reference DEs. Only 6.6% of the DEs had values that could be semantically typed. CONCLUSION: Our study suggests that the integration of biomedical sources can be achieved automatically with limited precision and largely facilitated by mapping DEs to terminological resources.

Abstracting and Indexing↗

Combining evidence, biomedical literature and statistical dependence: new insights for functional annotation of gene sets.

BACKGROUND: Large-scale genomic studies based on transcriptome technologies provide clusters of genes that need to be functionally annotated. The Gene Ontology (GO) implements a controlled vocabulary organised into three hierarchies: cellular components, molecular functions and biological processes. This terminology allows a coherent and consistent description of the knowledge about gene functions. The GO terms related to genes come primarily from semi-automatic annotations made by trained biologists (annotation based on evidence) or text-mining of the published scientific literature (literature profiling). RESULTS: We report an original functional annotation method based on a combination of evidence and literature that overcomes the weaknesses and the limitations of each approach. It relies on the Gene Ontology Annotation database (GOA Human) and the PubGene biomedical literature index. We support these annotations with statistically associated GO terms and retrieve associative relations across the three GO hierarchies to emphasise the major pathways involved by a gene cluster. Both annotation methods and associative relations were quantitatively evaluated with a reference set of 7397 genes and a multi-cluster study of 14 clusters. We also validated the biological appropriateness of our hybrid method with the annotation of a single gene (cdc2) and that of a down-regulated cluster of 37 genes identified by a transcriptome study of an in vitro enterocyte differentiation model (CaCo-2 cells). CONCLUSION: The combination of both approaches is more informative than either separate approach: literature mining can enrich an annotation based only on evidence. Text-mining of the literature can also find valuable associated MEDLINE references that confirm the relevance of the annotation. Eventually, GO terms networks can be built with associative relations in order to highlight cooperative and competitive pathways and their connected molecular functions.

Algorithms↗

Evidence in pharmacovigilance: extracting adverse drug reactions articles from MEDLINE to link them to case databases.

Literature, specifically MEDLINE, is among the main sources of information used to detect whether a drug may be responsible for Adverse Drug Reactions cases. The aim of our work is to automate the search of publications that correspond to a given Adverse Drug Reactions case: (i) by defining a general pattern for the queries used to search MEDLINE and (ii) by determining the threshold number of publications capable to confirm or infirm the Adverse Drug Reaction. We applied our algorithm to a set of 620 cases from a French pharmacovigilance database. We obtained a precision of 93%, recall 70%. We determined a threshold of 3 publications to confirm an Adverse Drug Reaction case.

Adverse Drug Reaction Reporting Systems↗

Aligning biomedical ontologies using lexical methods and the UMLS: the case of disease ontologies.

The process of aligning ontologies comprises two major steps: i) mapping concepts and ii) characterizing the relations between the concepts. In this paper, we present an alignment method based on a hybrid approach that reuses the UMLS knowledge base and aims at identifying patterns to characterize the relations. The proposed method consist in four steps: 1) exact matching, 2) searching for terms from one ontology that are included in terms from the other ontology, 3) identifying direct relations through the UMLS and 4) extracting syntactico-semantic patterns to infer novel alignments. This method has been applied to aligning the Human Disease ontology and the Mouse Pathology ontology resulting in 48 exact matches and 3,697 pairs of concepts for which one term is included in a term from the other ontology. 1,270 alignments are present in the UMLS. Among these, 903 are characterized by a semantic attribute. Based on these alignments, a study of the syntactic patterns has been done. Not surprisingly, the distribution of the different syntactic patterns is not sufficient to discriminate the different types of relationships found in the UMLS alignments. We have used the semantic categorization of the concepts provided by the UMLS to extract syntactico-semantic patterns. 87 novel alignments based on 6 syntactico-semantic patterns associated with isa and has associated morphology have been inferred.

Animals↗

Desiderata for domain reference ontologies in biomedicine.

Domain reference ontologies represent knowledge about a particular part of the world in a way that is independent from specific objectives, through a theory of the domain. An example of reference ontology in biomedical informatics is the Foundational Model of Anatomy (FMA), an ontology of anatomy that covers the entire range of macroscopic, microscopic, and subcellular anatomy. The purpose of this paper is to explore how two domain reference ontologies--the FMA and the Chemical Entities of Biological Interest (ChEBI) ontology, can be used (i) to align existing terminologies, (ii) to infer new knowledge in ontologies of more complex entities, and (iii) to manage and help reasoning about individual data. We analyze those kinds of usages of these two domain reference ontologies and suggest desiderata for reference ontologies in biomedicine. While a number of groups and communities have investigated general requirements for ontology design and desiderata for controlled medical vocabularies, we are focusing on application purposes. We suggest five desirable characteristics for reference ontologies: good lexical coverage, good coverage in terms of relations, compatibility with standards, modularity, and ability to represent variation in reality.

Anatomy↗

Amplification of Terminologia anatomica by French language terms using Latin terms matching algorithm: a prototype for other language.

OBJECTIVE: Terminologia anatomica is the new standard in anatomical terminology. This terminology is available only in Latin and English and its worldwide adoption is subject to the addition of terms from others languages. On the other hand, Nomina anatomica, the previous standard, has been widely translated. Aim of this work was to append foreign terms to Terminologia by using similarity-matching algorithm between its Latin terms and those from Nomina. METHODS: A semi-automatic matching of Latin terms from Terminologia with those of Nomina was performed using a string-to-string distance algorithm and manual assessment. We used a French-Latin version of Nomina together with Terminologia and we suggested French terms for Terminologia. Coverage was evaluated by the number of exact and approximate matches. A target of 78% was set due to the higher number of terms in Terminologia compared to Nomina. Relevance was estimated by manually comparing the meanings of the English and French terms related to the same Latin term. The question was whether they refer to the same anatomical structure. RESULTS: Exact or approximate matches were found for 5982 terms (76.5%) of Terminologia. Our results indicated that more than 75% of the terms from Terminologia came from Nomina, most of them were left unchanged and all were used with the same meaning. CONCLUSION: This method produces relevant results, reaching our 78% target. The method is based only on Latin terms and can be used for other languages. We consider this work as a starting point for adding terms to other knowledge sources, such as the foundational model of anatomy or the Unified Medical Language System (UMLS).

Algorithms↗

Problem-based learning in medical informatics for undergraduate medical students: an experiment in two medical schools.

PURPOSE: The objective of this work was to assess problem-based learning (PBL) as a method for teaching information and communication technology in medical informatics (MI) courses. A study was conducted in the Schools of Medicine of Rennes and Rouen (France) with third-year medical students. METHODS: The "PBL-in-MI" sessions included a first tutorial group meeting, then personal work, followed by a second tutorial group meeting. A problem that simulated practice and was focused on information technology was discussed. In Rouen, the students were familiar with PBL, and they enrolled on a voluntary basis, while in Rennes, the students were first-ever participants in PBL courses, and the program was mandatory. One hundred and seventy-seven students participated in the PBL-in-MI sessions and were given a questionnaire in order to evaluate qualitatively the sessions. RESULTS AND DISCUSSION: The response rate was 92.1%. The overall opinion of the students was good. 69.8% responded positively to the program. In Rouen, where the students participated in PBL-in-MI sessions on a voluntary basis, the students were significantly more enthusiastic about PBL-in-MI. Moreover, attitudes and opinions of students are plausibly related to differences in previous PBL skills. The fact that the naïve group had two tutors, one trained and one naïve as the students, has been investigated. Teacher naivety was an explanatory factor for the differences between Rennes and Rouen.

Consumer Behavior↗

UMLF: a unified medical lexicon for French.

Medical Informatics has a constant need for basic medical language processing tasks, e.g. for coding into controlled vocabularies, free text indexing and information retrieval. Most of these tasks involve term matching and rely on lexical resources: lists of words with attached information, including inflected forms and derived words, etc. Such resources are publicly available for the English language with the UMLS Specialist Lexicon, but not in other languages. For the French language, several teams have worked on the subject and built local lexical resources. The goal of the present work is to pool and unify these resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a unified medical lexicon for French (UMLF). This paper exposes the issues raised by such an objective, describes the methods on which the project relies and illustrates them with experimental results.

Abstracting and Indexing↗

Modelling a decision-support system for oncology using rule-based and case-based reasoning methodologies.

In most hospital medical units, multidisciplinary committees meet weekly to discuss their patients' cases. The medical experts base their decisions on three sources of information. First, they check if their patient complies with existing guidelines. Failing these, the medical experts will base their therapeutic decisions on the cases of similar patients that they have treated in the past. We propose a multi-modal reasoning decision-support system based on both guideline and case series, which will automatically compare the patient's case to the corresponding guideline, then to other cases, and retrieve similar cases. The general structure of the system is presented here, the domain of application being oncology. As the patients' records are not currently stored in a database in a format which is directly accessible, an object-oriented model is proposed, which includes prognosis factors currently tested in clinical trials, well-established ones, and a description of the illness episodes. The system is designed to be a data warehouse. Such a system does not exist in the literature. Future work will be needed to define the similarity measures, and to connect the system to the current database.

Computer Simulation↗

Non-lexical approaches to identifying associative relations in the gene ontology.

The Gene Ontology (GO) is a controlled vocabulary widely used for the annotation of gene products. GO is organized in three hierarchies for molecular functions, cellular components, and biological processes but no relations are provided among terms across hierarchies. The objective of this study is to investigate three non-lexical approaches to identifying such associative relations in GO and compare them among themselves and to lexical approaches. The three approaches are: computing similarity in a vector space model, statistical analysis of co-occurrence of GO terms in annotation databases, and association rule mining. Five annotation databases (FlyBase, the Human subset of GOA, MGI, SGD, and WormBase) are used in this study. A total of 7,665 associations were identified by at least one of the three non-lexical approaches. Of these, 12% were identified by more than one approach. While there are almost 6,000 lexical relations among GO terms, only 203 associations were identified by both non-lexical and lexical approaches. The associations identified in this study could serve as the starting point for adding associative relations across hierarchies to GO, but would require manual curation. The application to quality assurance of annotation databases is also discussed.

Alzheimer Disease↗

Toward a unified representation of findings in clinical radiology.

The representations of findings in clinical radiology are heterogeneous. Motivations for developing a unified representation include the semantic integration of medical reports based on DICOM-SR(Digital Image Communication in Medicine Structured Reporting), bibliographic databases in the context of evidence-based medicine, and teaching resources. In this work, we propose a unified representation integrating the representations of findings in the UMLS, the GAMUTS in Radiology and the DICOM-SR. We analyse the UMLS and the DCMR (DICOM Content Mapping Resource) of DICOM SR to figure out their own representation of findings. Then we set up a syntax between the UMLS concepts using DICOM-SR relations in order to rewrite the GAMUTS sentences. The translation of the whole GAMUTS using the UMLS concepts and the DICOM SR syntax could be a method to create or supplement the DCMR TIDs (Template ID : Identifier of a Template) and CIDs (Context ID : Identifier of a Context Group) in the field of description of findings in medical imaging. This method could also enable to give an ontologic dimension to the DICOM SR representation system of information. The meaning of the CIDs would then be enhanced far beyond the simple use of the SNOMED vocabulary.

Diagnostic Imaging↗

Issues in the classification of disease instances with ontologies.

Ontologies define classes of entities and their interrelations. They are used to organize data according to a theory of the domain. Towards that end, ontologies provide class definitions (i.e., the necessary and sufficient conditions for defining class membership). In medical ontologies, it is often difficult to establish such definitions for diseases. We use three examples (anemia, leukemia and schizophrenia) to illustrate the limitations of ontologies as classification resources. We show that eligibility criteria are often more useful than the Aristotelian definitions traditionally used in ontologies. Examples of eligibility criteria for diseases include complex predicates such as ' x is an instance of the class C when at least n criteria among m are verified' and 'symptoms must last at least one month if not treated, but less than one month, if effectively treated'. References to normality and abnormality are often found in disease definitions, but the operational definition of these references (i.e., the statistical and contextual information necessary to define them) is rarely provided. We conclude that knowledge bases that include probabilistic and statistical knowledge as well as rule-based criteria are more useful than Aristotelian definitions for representing the predicates defined by necessary and sufficient conditions. Rich knowledge bases are needed to clarify the relations between individuals and classes in various studies and applications. However, as ontologies represent relations among classes, they can play a supporting role in disease classification services built primarily on knowledge bases.

Humans↗

Classifying diseases with respect to anatomy: a study in SNOMED CT.

Anatomy is a major organizing principle for dis-eases. In the formal definitions provided by SNOMED CT, for example, the role 'finding site' relates disorders to anatomical entities. This study investigates SNOMED CT and compares the anatomy-based classification of diseases supported by the role finding site to the anatomy-based classification of diseases provided by subsumption (is-a) relations between diseases. For each of the 3,540 anatomical entities associated with disorders,, we compared two sets of disorders: first, the set of disorders associated with any descendant of the anatomical entity under investigation (ANAT); second, the set of dis-orders corresponding to the union of the descendants of the disorders associated with the anatomical entity under investigation (TAXO). The ANAT and TAXO sets were different for 1,231 anatomical entities (35%). In 607 cases, the overlap between ANAT and TAXO was less than 50%. When a difference was found, the TAXO set was always a subset of the ANAT set. Among the 1,025,904 subsumption relations among disorders generated by the ANAT approach, 40% were not present in TAXO. This approach helps identify missing classes and taxonomic relations in existing ontologies. It can be generalized to other kinds of partitions of biomedical ontologies.

Anatomy↗

Aligning knowledge sources in the UMLS: methods, quantitative results, and applications.

The UMLS Semantic Network and Metathesaurus are two complementary knowledge sources. While many studies compare relationships across the two structures, their alignment has never been attempted. We applied two methods based on lexical and conceptual similarity to aligning the Semantic Network with the UMLS Metathesaurus. Approximately two thirds of the semantic types could be aligned by lexical similarity. Conceptual similarity suggested mappings in all but ten cases. Potential applications enabled by the alignment are discussed, namely auditing the consistency between the Semantic Network and the Metathesaurus and extending the Semantic Network downwards. The relative contribution and limitations of the two methods used for the alignment are also discussed.

Semantics↗

Towards the automatic generation of biomedical sources schema.

Biologists and physicians need to access biological and medical data for their experimentations and researches. This information is available on the Internet and is scattered over many heterogeneous data sources. Collecting information is consequently tedious, time consuming and must be improved. To cope with this difficulty, our overall objective is to realize a mediator-based system to integrate heterogeneous biomedical data sources. This requires first an automatic generation of source schema, which is the goal of this work. For that, we describe an algorithm which is based on information extraction. It consists of the extraction of meta-information from each source to infer their schema. Our system enables users to access relevant and specific data, which are up-to-date. To solve the semantic heterogeneity of data sources, we are considering the creation of an ontology. Finally, the management of source evolution is discussed

Algorithms↗

Experiments in cross-language medical information retrieval using a mixing translation module.

Given the ever-increasing scale and diversity of medical literature widely published in English on the Internet, improving the performance of information retrieval by cross-language is an urgent research objective. Cross-language medical information retrieval (CLMIR) consists of providing a query in one language and searching medical document collections in one or more different languages. Our users of CLMIR are users who are able to read biomedical texts in English, but have difficulty formulating English queries. This paper proposes a French/English CLMIR system as a mixing model for supporting the retrieval of English medical documents. Methods fall into the category of query translation approach in which we use a hybrid machine translation that combines a pattern-based module with a rule-based translator and includes three steps from pre- to- post-translation. In parallel to this hybrid machine translation, we use multilingual UMLS Methasaurus as a complementary translator. The results show that using a mixing translation module outperforms machine translation-based method and thesaurus-based method used separately.

Information Storage and Retrieval↗

Evolving from standard vocabularies to formal ontology for an information system dedicated to organ transplantation.

Semantic heterogeneity is a key issue in the deployment of medical applications. In this paper, we examine solutions to address semantic heterogeneity for a complex organ transplantation information system. The information system that has been developed by the French agency for transplantation (EfG) has to gather medical information concerning patients and donors for organ allocation, epidemiological studies, and public health decisions. We analyze in this context the limits of the traditional approach based on a standard vocabulary and a relational database for the transplantation domain. We present its evolution towards a system that combines a terminology server and a data warehouse. The perspective is now to build a formal ontology that may support semantic integration in an advanced transplantation information system. Two open issues in formal ontology design are discussed. First, providing definitions of all medical concepts is not an obvious task if we want those definitions be meaningful, formal, and compatible with other ontologies. Second, propagation of properties along relations raises the needs for extending description logics with rules.

France↗

Collaborative environment for clinical reasoning and distance learning sessions.

BACKGROUND: The medical curriculum has changed with the adoption of the student-centered learning paradigm. Clinical reasoning learning (CRL) is used in order to develop and improve students' clinical reasoning and problem-solving skills. PURPOSE: We have observed that, in complement to traditional CRL sessions, students commonly consult resources available on the internet. Based on this observation, our objective is to create computer tools to coordinate CRL sessions at distance, integrating these electronic resources at every step of the reasoning process. MATERIAL AND METHODS: In order to create the system, we elaborated an object-oriented model of a computer-supported collaborative learning environment. The proposed system includes a local web-server to store electronic resources and a relational database to store their electronic addresses (urls). JAVA was used as the programming language. RESULTS: We developed a set of cooperative platform-independent tools. This environment includes a communication tool. Multimedia data exchange is possible. Information is shared thanks to an electronic notepad and whiteboard tools. PERSPECTIVES: This learning environment will be integrated in the French Virtual Medical University project, and is intended to be used for undergraduate, internships, residency or continuing medical education.

Computers, Handheld↗