Search PubMed⌕ Search

Biomedical subjects

Robert Baud

Publications and source records attributed to Robert Baud.

At least 19 recordsLinked to original sources

Defining and relating biomedical terms: towards a cross-language morphosemantics-based system.

This paper addresses the issue of how semantic information can be automatically assigned to compound terms, i.e. both a definition and a set of semantic relations. This is particularly crucial when elaborating multilingual databases and when developing cross-language information retrieval systems. The paper shows how morphosemantics can contribute in the constitution of multilingual lexical networks in biomedical corpora. It presents a system capable of labelling terms with morphologically related words, i.e. providing them with a definition, and grouping them according to synonymy, hyponymy and proximity relations. The approach requires the interaction of three techniques: (1) a language-specific morphosemantic parser, (2) a multilingual table defining basic relations between word roots and (3) a set of language-independent rules to draw up the list of related terms. This approach has been fully implemented for French, on an about 29,000 terms biomedical lexicon, resulting to more than 3000 lexical families. A validation of the results against a manually annotated file by experts of the domain is presented, followed by a discussion of our method.

France↗

An Ontology driven collaborative development for biomedical terminologies: from the French CCAM to the Australian ICHI coding system.

The CCAM French coding system of clinical procedures was developed between 1994 and 2004 using, in parallel, a traditional domain expert's consensus method on one hand, and advanced methodologies of ontology driven semantic representation and multilingual generation on the other hand. These advanced methodologies were applied under the framework of an European Union collaborative research project named GALEN and produced a new generation of biomedical terminology. Following the interest in several countries and in WHO, the GALEN network has tested the application of the ontology driven tools to the existing reduced Australian ICHI coding system for interventions presently under investigation by WHO to check its ability and appropriateness to become the reference international coding system for procedures. The initial results are presented and discussed in terms of feasibility and quality assurance for sharing and maintaining consistent medical knowledge and allowing diversity in linguistic expressiveness of end users.

Australia↗

Health search engine with e-document analysis for reliable search results.

OBJECTIVE: After a review of the existing practical solution available to the citizen to retrieve eHealth document, the paper describes an original specialized search engine WRAPIN. METHOD: WRAPIN uses advanced cross lingual information retrieval technologies to check information quality by synthesizing medical concepts, conclusions and references contained in the health literature, to identify accurate, relevant sources. Thanks to MeSH terminology [1] (Medical Subject Headings from the U.S. National Library of Medicine) and advanced approaches such as conclusion extraction from structured document, reformulation of the query, WRAPIN offers to the user a privileged access to navigate through multilingual documents without language or medical prerequisites. RESULTS: The results of an evaluation conducted on the WRAPIN prototype show that results of the WRAPIN search engine are perceived as informative 65% (59% for a general-purpose search engine), reliable and trustworthy 72% (41% for the other engine) by users. But it leaves room for improvement such as the increase of database coverage, the explanation of the original functionalities and an audience adaptability. CONCLUSION: Thanks to evaluation outcomes, WRAPIN is now in exploitation on the HON web site (http://www.healthonnet.org), free of charge. Intended to the citizen it is a good alternative to general-purpose search engines when the user looks up trustworthy health and medical information or wants to check automatically a doubtful content of a Web page.

Europe↗

Methodology to ease the construction of a terminology of problems.

INTRODUCTION: Problem lists summarize an aspect of the patient's medical history and provide an important way to implement entry points for clinical pathways and guideline-oriented care. However, in order to automate processes based on problem lists, the use of controlled vocabularies is required. We developed a methodology to extract a collection of standardized problem-related terms from medical documents entered in free text by physicians. METHODS: We extracted a corpus of sentences describing problems from a randomized selection of admission notes collected at the University Hospitals of Geneva. Theses sentences underwent manual and automatic normalization processes, and a statistical clustering, in order to build a set of terms. RESULTS: We obtained 17,805 sentences from 5000 admission notes. We refined them into 1546 terms, 88.6% of which could be related to a relevant problem statement. DISCUSSION: A clinically relevant problems terminology was derived from clinical admission notes in free-text using a few methodical steps with a reasonable investment of human resources. Such an approach will ease the development and the use of problem lists better suited to user needs.

Medical Records, Problem-Oriented↗

Amplification of Terminologia anatomica by French language terms using Latin terms matching algorithm: a prototype for other language.

OBJECTIVE: Terminologia anatomica is the new standard in anatomical terminology. This terminology is available only in Latin and English and its worldwide adoption is subject to the addition of terms from others languages. On the other hand, Nomina anatomica, the previous standard, has been widely translated. Aim of this work was to append foreign terms to Terminologia by using similarity-matching algorithm between its Latin terms and those from Nomina. METHODS: A semi-automatic matching of Latin terms from Terminologia with those of Nomina was performed using a string-to-string distance algorithm and manual assessment. We used a French-Latin version of Nomina together with Terminologia and we suggested French terms for Terminologia. Coverage was evaluated by the number of exact and approximate matches. A target of 78% was set due to the higher number of terms in Terminologia compared to Nomina. Relevance was estimated by manually comparing the meanings of the English and French terms related to the same Latin term. The question was whether they refer to the same anatomical structure. RESULTS: Exact or approximate matches were found for 5982 terms (76.5%) of Terminologia. Our results indicated that more than 75% of the terms from Terminologia came from Nomina, most of them were left unchanged and all were used with the same meaning. CONCLUSION: This method produces relevant results, reaching our 78% target. The method is based only on Latin terms and can be used for other languages. We consider this work as a starting point for adding terms to other knowledge sources, such as the foundational model of anatomy or the Unified Medical Language System (UMLS).

Algorithms↗

Recent advances in natural language processing for biomedical applications.

We survey a set a recent advances in natural language processing applied to biomedical applications, which were presented in Geneva, Switzerland, in 2004 at an international workshop. While text mining applied to molecular biology and biomedical literature can report several interesting achievements, we observe that studies applied to clinical contents are still rare. In general, we argue that clinical corpora, including electronic patient records, must be made available to fill the gap between bioinformatics and medical informatics.

Abstracting and Indexing↗

UMLF: a unified medical lexicon for French.

Medical Informatics has a constant need for basic medical language processing tasks, e.g. for coding into controlled vocabularies, free text indexing and information retrieval. Most of these tasks involve term matching and rely on lexical resources: lists of words with attached information, including inflected forms and derived words, etc. Such resources are publicly available for the English language with the UMLS Specialist Lexicon, but not in other languages. For the French language, several teams have worked on the subject and built local lexical resources. The goal of the present work is to pool and unify these resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a unified medical lexicon for French (UMLF). This paper exposes the issues raised by such an objective, describes the methods on which the project relies and illustrates them with experimental results.

Abstracting and Indexing↗

Towards a multilingual version of terminologia anatomica.

OBJECTIVE: Terminologia Anatomica (TA) is the new standard in anatomical terminology. This terminology is available only in Latin and English and its worldwide adoption is subdued to the addition of terms from others languages. On the other hand Nomina Anatomica (NA), the previous standard, has been widely translated. Aim of this work was to append foreign terms to TA by using similarity matching algorithm between its Latin terms and those from NA. METHODS: A semi-automatic matching of Latin terms from TA with those of NA was performed using a string-to-string distance algorithm and manual assessment. We used a French - Latin version of NA together with TA and we suggested French terms for TA. Coverage was evaluated by the number of exact and approximate matches. A target of 80% was set due to the superior number of terms in TA compared to NA. Relevance was estimated by manually comparing the meanings of the English and French terms related to the same Latin term. The question was whether they refer to the same anatomical structure. RESULTS: Exact or approximate matches were found for 5,982 terms (76.5%) of TA. Our results outlined that more than 75% of the terms from TA came from NA, most of them were left unchanged and all were used with the same meaning. CONCLUSION: This method produces relevant results, reaching our 80% target. The method is based only on Latin terms and can be used for other languages and for others terminologies including Latin terms.

Algorithms↗

Predicting Lexical Relations between Biomedical Terms: towards a Multilingual Morphosemantics-based System.

This paper addresses the issue of how semantic information can be automatically assigned to compound terms, i.e. both a definition and a set of semantic relations. This issue is particularly crucial when elaborating multilingual databases and when developing cross-language information retrieval systems. The paper shows how morpho-semantics can contribute in the constitution of multilingual lexical networks in biomedical corpora. It presents a system capable of labelling terms with morphologically related words, i.e. providing them with a definition, and grouping them according to synonymy, hyponymy and proximity relations. The approach requires the interaction of three techniques: (1) a la morphosemantic parser, (2) a multilingual table defining basic relations between word roots, and (3) a set of language-independant rules to draw up the list of related terms. This approach has been fully implemented for French, on an about 29,000 terms biomedical lexicon, resulting to more than 3,000 lexical families.

Language↗

Extracting key sentences with latent argumentative structuring.

PROBLEM: Key word assignment has been largely used in MEDLINE to provide an indicative "gist" of the content of articles. Abstracts are also used for this purpose. However with usually more than 300 words, abstracts can still be regarded as long documents; therefore we design a system to select a unique key sentence. This key sentence must be indicative of the article's content and we assume that abstract's conclusions are good candidates. We design and assess the performance of an automatic key sentence selector, which classifies sentences into 4 argumentative moves: PURPOSE, METHODS, RESULTS and CONCLUSION. METHODS: We rely on Bayesian classifiers trained on automatically acquired data. Features representation, selection and weighting are reported and classification effectiveness is evaluated on the four classes using confusion matrices. We also explore the use of simple heuristics to take the position of sentences into account. Recall, precision and F-scores are computed for the CONCLUSION class. For the CONCLUSION class, the F-score reaches 84%. Automatic argumentative classification is feasible on MEDLINE abstracts and should help user navigation in such repositories.

Bayes Theorem↗

A natural language based search engine for ICD10 diagnosis encoding.

We have developed a multiple step process for implementing an ICD10 search engine. The complexity of the task has been shown and we recommend collecting adequate expertise before starting any implementation. Underestimation of the expert time and inadequate data resources are probable reasons for failure. We also claim that when all conditions are met in term of resource and availability of the expertise, the benefits of a responsive ICD10 search engine will be present and the investment will be successful.

Databases as Topic↗

XML as standard for communicating in a document-based electronic patient record: a 3 years experiment.

During the past few years, the eXtensible Markup Language (XML) has progressively become a gold standard for accessing, representing and exchanging information, especially in the health care environment. This paper presents an implementation of the use of XML for the electronic patient record (EPR) and discusses more specifically its growing use in two areas of the EPR: first, as a format for the exchange of structured messages, and second, as a comprehensible way of representing patient documents. These statements rely on a 3 years experiment conducted at the Geneva University Hospital as part of its document-centered EPR.

Delivery of Health Care↗

Towards a unified medical lexicon for French.

Medical Informatics has a constant need for basic Medical Language Processing tasks, e.g., for coding into controlled vocabularies, free text indexing and information retrieval. Most of these tasks involve term matching and rely on lexical resources: lists of words with attached information, including inflected forms and derived words, etc. Such resources are publicly available for the English language with the UMLS Specialist Lexicon, but not in other languages. For the French language, several teams have worked on the subject and built local lexical resources. The goal of the present work is to pool and unify these resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a unified medical lexicon for French (UMLF). This paper exposes the issues raised by such an objective, describes the methods on which the project relies and illustrates them with experimental results.

Algorithms↗

A frame-based representation of ICD-10.

UNLABELLED: Physicians are required to code information concerning a patient's stay in order to measure the medical activity in hospitals. They use the International Statistical Classification of Diseases and Related Health Problems, Tenth Revision (ICD-10). Coding is usually performed manually and computerized tools may be useful in speeding up and facilitating the tedious task of coding patient information. The aim of this work is to build a surface semantic model of ICD-10 in order to ameliorate a coding help system. METHODS: This work was focused on chapter XI of the ICD-10, Diseases of the Digestive System. Each term from both analytical and alphabetical indexes about this chapter were submitted to a morphological analysis in order to extract the medical concepts within. After a statistical analysis of these concepts and the way they connect themselves, a semantic model based on a "semantic frame" approach was built. RESULTS: Although this model could represent a reasonable amount of medical knowledge within chapter XI of the ICD-10 in a quite satisfactory way, it shows lack of efficiency for some other chapters. CONCLUSION: Difficulties have to be overcome when modelling a classification meant for manual utilisation, and a lot of work still has to be done to obtain an effective coding help system using the ICD-10.

Forms and Records Control↗

Clinical documents: attribute-values entity representation,context, page layout and communication.

This paper presents how acquisition, storage and communication of clinical documents is implemented at the University Hospitals of Geneva. Careful attention has been given to user-interfaces, in order to support complex layouts, spell checking, and templates management with automatic prefilling. A dual architecture has been developed for storage using an entity-attribute-value unified database and a consolidated, patient-centered, layout-respectful file-based storage, providing both representation power and speed of access. This architecture allows a great flexibility for storing a continuum of data types, ranging from simple typed values to complex clinical reports. Finally, communication is entirely based on HTTP-XML internally, and a HL-7 CDA interface V2 is currently studied for external communication. Some of the problems encountered, mostly related to the typology of documents and the ontology of clinical attributes are evoked.

Computer Systems↗

UMLF: a Unified Medical Lexicon for French.

Lexical resources for medical language, such as lists of words with inflectional and derivational information, are publicly available for the English lantuate with the UMLS Specialist Lexicon. The goal of the UMLF project is to pool and unify existing resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a Unified Medical Lexicon for French. We present here the current status of the project.

France↗

XML as standard for communicating in a document-based electronic patient record: a three years experiment.

During the past few years, the eXtensible Markup Language (XML) has experienced a growing use for accessing, representing and exchanging information, especially in the health care environment. This paper discusses the potentials of the use of XML for the electronic patient record (EPR) in two ways: first, as a format for the exchange of structured messages, and second, as a comprehensible way of representing patient documents. These statements rely on a three years experiment conducted at the Geneva University Hospital as part of its document-centred EPR.

Medical Records Systems, Computerized↗