Search PubMed⌕ Search

Biomedical subjects

Yves A Lussier

Publications and source records attributed to Yves A Lussier.

18 recordsLinked to original sources

Worldwide Innovative Network Consortium: Building a Common Global Cancer Database.

This review shares the ongoing work of the global Worldwide Innovative Network (WIN) Consortium for Precision Medicine to synthesize emerging cancer treatment data and to define the requirements for a common global cancer database that can truly support precision oncology. We performed a narrative review of emerging cancer treatment data, molecular profiling technologies, and existing clinicogenomic databases, focusing on how tumors are characterized, how subgroups are defined, and how demographic, lifestyle, and environmental factors are captured. The growth in molecular profiling technologies and the development of new targeted therapies are transforming cancer care. Tumors, regardless of tissue origin, are increasingly defined as composites of multiple, often rare, subgroups, each with distinct biology and likely response to specific therapies, based on multidimensional profiling of the tumor and its microenvironment. The solution lies in building vast databases that capture racial and ethnic diversity, reflected in genomic data, as well as diet and lifestyle factors that may have epigenetic impact on gene expression and post-translational modifications. A truly inclusive and informative data set must reflect global diversity, and there are multiple examples of demography-dependent differences in genomic signals. With members caring for and studying patients with cancer across five continents, WIN is actively exploring pathways to create a global cancer database, rich in clinical and molecular detail, granular enough for precise analysis, and large enough to power artificial intelligence-driven insights, provided appropriate data quality, validation, and governance frameworks are in place. This review surveys the current landscape and outlines practical paths forward to achieve this goal.

Humans↗

Bio-Ontology and text: bridging the modeling gap.

MOTIVATION: Natural language processing (NLP) techniques are increasingly being used in biology to automate the capture of new biological discoveries in text, which are being reported at a rapid rate. Yet, information represented in NLP data structures is classically very different from information organized with ontologies as found in model organisms or genetic databases. To facilitate the computational reuse and integration of information buried in unstructured text with that of genetic databases, we propose and evaluate a translational schema that represents a comprehensive set of phenotypic and genetic entities, as well as their closely related biomedical entities and relations as expressed in natural language. In addition, the schema connects different scales of biological information, and provides mappings from the textual information to existing ontologies, which are essential in biology for integration, organization, dissemination and knowledge management of heterogeneous phenotypic information. A common comprehensive representation for otherwise heterogeneous phenotypic and genetic datasets, such as the one proposed, is critical for advancing systems biology because it enables acquisition and reuse of unprecedented volumes of diverse types of knowledge and information from text. RESULTS: A novel representational schema, PGschema, was developed that enables translation of phenotypic, genetic and their closely related information found in textual narratives to a well-defined data structure comprising phenotypic and genetic concepts from established ontologies along with modifiers and relationships. Evaluation for coverage of a selected set of entities showed that 90% of the information could be represented (95% confidence interval: 86-93%; n = 268). Moreover, PGschema can be expressed automatically in an XML format using natural language techniques to process the text. To our knowledge, we are providing the first evaluation of a translational schema for NLP that contains declarative knowledge about genes and their associated biomedical data (e.g. phenotypes). AVAILABILITY: http://zellig.cpmc.columbia.edu/PGschema

Abstracting and Indexing↗

Terminology model discovery using natural language processing and visualization techniques.

Medical terminologies are important for unambiguous encoding and exchange of clinical information. The traditional manual method of developing terminology models is time-consuming and limited in the number of phrases that a human developer can examine. In this paper, we present an automated method for developing medical terminology models based on natural language processing (NLP) and information visualization techniques. Surgical pathology reports were selected as the testing corpus for developing a pathology procedure terminology model. The use of a general NLP processor for the medical domain, MedLEE, provides an automated method for acquiring semantic structures from a free text corpus and sheds light on a new high-throughput method of medical terminology model development. The use of an information visualization technique supports the summarization and visualization of the large quantity of semantic structures generated from medical documents. We believe that a general method based on NLP and information visualization will facilitate the modeling of medical terminologies.

Automation↗

Genestrace: phenomic knowledge discovery via structured terminology.

The era of applied genomic medicine is quickly approaching accompanied by the increasing availability of detailed genetic information. Understanding the genetic etiology behind complex, multi-gene diseases remains an important challenge. In order to uncover the putative genetic etiology of complex diseases, we designed a method that explores the relationships between two major terminological and ontological resources: the Unified Medical Language System (UMLS) and the Gene Ontology (GO). The UMLS has a mainly clinical emphasis; Gene Ontology has become the standard for biological annotations of genes and gene products. Using statistical and semantic relationships within and between the two resources, we are able to infer relationships between disease concepts in the UMLS and gene products annotated using GO and its associated databases. We validated our inferences by comparing them to the known gene-disease relationships, as defined in the Online Mendelian Inheritance in Man's morbidmap (OMIM). The proof-of-concept methods presented here are unique in that they bypass the ambiguity of the direct extraction of gene or disease term from MEDLINE. Additionally, our methods provide direct links to clinically significant diseases through established terminologies or ontologies. The preliminary results presented here indicate the potential utility of exploiting the existing, manually curated relationships in biomedical resources as a tool for the discovery of potentially valuable new gene-disease relationships.

Computational Biology↗

Partitioning knowledge bases between advanced notification and clinical decision support systems.

Due to the varying rates of change of ephemeral administrative and enduring clinical knowledge in decision support systems (DSSs), the functional partition of knowledge base (KB) components can lead to more efficient and cost-effective system implementation and maintenance. Our prototype loosely couples a clinical event monitor developed by Columbia University Medical Center (CUMC) with a secure notification service proxy developed by IBM Research to form a novel and complex clinical event communication service.

Computer Security↗

Visualizing information across multidimensional post-genomic structured and textual databases.

MOTIVATION: Visualizing relationships among biological information to facilitate understanding is crucial to biological research during the post-genomic era. Although different systems have been developed to view gene-phenotype relationships for specific databases, very few have been designed specifically as a general flexible tool for visualizing multidimensional genotypic and phenotypic information together. Our goal is to develop a method for visualizing multidimensional genotypic and phenotypic information and a model that unifies different biological databases in order to present the integrated knowledge using a uniform interface. RESULTS: We developed a novel, flexible and generalizable visualization tool, called PhenoGenesviewer (PGviewer), which in this paper was used to display gene-phenotype relationships from a human-curated database (OMIM) and from an automatic method using a Natural Language Processing tool called BioMedLEE. Data obtained from multiple databases were first integrated into a uniform structure and then organized by PGviewer. PGviewer provides a flexible query interface that allows dynamic selection and ordering of any desired dimension in the databases. Based on users' queries, results can be visualized using hierarchical expandable trees that present views specified by users according to their research interests. We believe that this method, which allows users to dynamically organize and visualize multiple dimensions, is a potentially powerful and promising tool that should substantially facilitate biological research. AVAILABILITY: PhenogenesViewer as well as its support and tutorial are available at http://www.dbmi.columbia.edu/pgviewer/ CONTACT: Lussier@dbmi.columbia.edu.

Computer Graphics↗

A tool for abstracting relevant classes of concepts: the Common Ancestry Summarizer.

Controlled Medical Terminologies (CMTs) are indispensable tools for present medical information systems and medical informatics researches. Concept-oriented architecture and multiple-hierarchy have been accepted as two of the desiderata of contemporary CMTs. A common problem for informaticians working with class-based methods for terminologies is to find the Most Relevant Common Ancestor (MRCA) for a group of concepts, the common ancestor which has the closest semantic distance to all the given concepts. Finding an ancestor concept to summarize a group of concepts is required to map concepts from the knowledgebase to those of appropriate granularity in CMTs. However, manually exploring the hierarchical relationships and determining the MRCA are daunting tasks due to the massive size, the multiple hierarchies and in some case the cycles of the hierarchical networks of terminologies. In our study, we developed a web-based visualization tool, the Common Ancestry Summarizer (CAS), that graphically displays the hierarchical relations of multiple concepts within CMTs, and also assist users to identify the MRCA of these concepts by identifying the Lowest Common Ancestors (LCAs). The CAS has been shown useful in organizing classes for the controlled terms used in clinical guidelines (e.g. "the bottom-up method").

Algorithms↗

Automating terminological networks to link heterogeneous biomedical databases.

As cross-disciplinary research escalates, researchers are facing the challenge of linking disparate biomedical databases that have been developed without common indexes. Manually indexing these large-scale databases is laborious and often impractical. Solutions involving mediating terminologies have been proposed, but coordination of terms from the databases of interest to these mediating terminologies is also laborious, and regular synchronization between indexes is an additional problem. In this study we describe a novel method of linking heterogeneous databases using terminology networks constructed with automated mapping methods. Linkage was established between two disparate biomedical databases (SNOMED-CT and HDG), using two relevant intermediating databases (UMLS and OMIM). One gold standard of 514 distinct matches is used as proof-of-principle. In conclusion, as hypothesized, 1) Manually curated pathways provide high precision, but offer low recall, 2) the automated terminology pathways can significantly increase recall at acceptable precision. Taken together, our conclusion may suggest the combined manual and automated terminology networks could offer recall and precision in an incremental manner

Abstracting and Indexing↗

Mining OMIM for insight into complex diseases.

Understanding clinical phenotypes through their corresponding genotypes is one of the principal goals of genetic research. Though achieving this goal is relatively simple with single gene syndromes, more complex diseases often consist of varied clinical phenotypes that may be the result of interactions among multiple genetic loci. Microarray technology has brought the phenotype -genotype relationship to the molecular level, using differently behaving cancers, for example, as the basis for comparing patterns of gene expression. With this feasibility study, we attempted to use similar methods of analysis at the clinical level, in order to evaluate our hypothesis that the clustering of clinical phenotypes would provide information that would be useful in elucidating their underlying genotypes. Because of its breadth of content and detailed descriptions, we used OMIM as our source material for phenotypic and genetic information. After processing the source material, we then performed self-organizing map and hierarchical clustering analysis on representative diseases by phenotypic category. Through pre-determined queries over this analysis, we made two findings of potential clinical significance, one concerning diabetes and another concerning progressive neurologic diseases. Our methods provide a formal approach to analyzing phenotypes among diverse diseases, and may help indicate fruitful areas for further research into their underlying genetic causes.

Cluster Analysis↗

Putting data integration into practice: using biomedical terminologies to add structure to existing data sources.

A major purpose of biomedical terminologies is to provide uniform concept representation, allowing for improved methods of analysis of biomedical information. While this goal is being realized in bioinformatics, with the emergence of the Gene Ontology as a standard, there is still no real standard for the representation of clinical concepts. As discoveries in biology and clinical medicine move from parallel to intersecting paths, standardized representation will become more important. A large portion of significant data, however, is mainly represented as free text, upon which conducting computer-based inferencing is nearly impossible. In order to test our hypothesis that existing biomedical terminologies, specifically the UMLS Metathesaurus and SNOMED CT, could be used as templates to implement semantic and logical relationships over free text data that is important both clinically and biologically, we chose to analyze OMIM (Online Mendelian Inheritance in Man). After finding OMIM entries' conceptual equivalents in each respective terminology, we extracted the semantic relationships that were present and evaluated a subset of them for semantic, logical, and biological legitimacy. Our study reveals the possibility of putting the knowledge present in biomedical terminologies to its intended use, with potentially clinically significant consequences.

Databases, Genetic↗

Adapting current Arden Syntax knowledge for an object oriented event monitor.

Arden Syntax for Medical Logic Module (MLM)1 was designed for writing and sharing task-specific health knowledge in 1989. Several researchers have developed frameworks to improve the sharability and adaptability of Arden Syntax MLMs, an issue known as "curly braces" problem. Karadimas et al proposed an Arden Syntax MLM-based decision support system that uses an object oriented model and the dynamic linking features of the Java platform.2 Peleg et al proposed creating a Guideline Expression Language (GEL) based on Arden Syntax's logic grammar.3 The New York Presbyterian Hospital (NYPH) has a collection of about 200 MLMs. In a process of adapting the current MLMs for an object-oriented event monitor, we identified two problems that may influence the "curly braces" one: (1) the query expressions within the curly braces of Arden Syntax used in our institution are cryptic to the physicians, institutional dependent and written ineffectively (unpublished results), and (2) the events are coded individually within a curly braces, resulting sometimes in a large number of events - up to 200.

Decision Support Systems, Clinical↗

A "systematics" tool for medical terminologies.

Finding the hierarchical relations amongst multiple terms within medical terminologies that support multiple parents to a term is a common task, especially for trainees and knowledge engineers implementing or maintaining medical logic modules or guidelines. Examples of such terminologies include the UMLS and the Medical Entity Dictionary (MED). In addition, the task of identifying and discriminating amongst some common ancestors to a list of terms is a recurrent theme. This is also a common concern in the science of classification (systematics). Some nearest common ancestors have distinct valuable properties for classification and simplification of lists. Although there exist some visualized navigating and editing tools for the UMLS and the MED, they are browsers that show a large number of unrelated and irrelevant relationships to the task at hand. While algorithms have been well studied in computer science to solve such a problem over semantic networks and trees, to our knowledge, they have not been used with a visualization tool in biomedicine. We developed a visualized tool that graphically displays the hierarchical relations of multiple terms, and helps identifying the nearest common ancestors of these terms.

Classification↗

A knowledge framework for computational molecular-disease relationships in cancer.

Biomedical knowledge is growing at an exponential rate, with new discoveries being published across a range of information sources. A coded, fully-computable, and integrated approach to this information could increase the efficiency of its use, through improved retrieval as well as the eventual ability to apply decision support tools to the knowledge base. Though multiple knowledge bases (KBs) and databases (DBs) concerning gene-disease relationships exist, few present the information in a coded, easily computable form. Focusing on molecular-disease relationships in cancer (gene-disease and protein-disease), we evaluated articles in major biomedical journals, in order to develop both the framework for a knowledge model as well as evaluation criteria. We then used these criteria to evaluate major KBs, DBs, and terminologies. We discovered that although both the high-level as well as the specific molecular-disease relationships present in our test set were mapped in many of the databases, they generally were not applied together in a coded form. We propose a rationale behind a model mediated schema for the integration of these resources.

Artificial Intelligence↗

Web-based tailoring and its effect on self-efficacy: results from the MI-HEART randomized controlled trial.

This paper reports the effects of a tailored Web-based delivery system on self-efficacy as it relates to a patients' response to acute myocardial information (AMI) symptoms. The data reported are from MI-HEART, a randomized trial examining ways in which a clinical information system can favorably influence the appropriateness and rapidity of decision-making in patients suffering from symptoms of acute myocardial infarction. Participants were randomized into one of three groups: tailored Web-based, non-tailored Web-based and non-tailored paper based. A theoretically based behavioral-cognitive model was used to identify key variables upon which to tailor education material. A key variable in the model is self-efficacy, operationalized with a three-dimensional scaling. Results show trends in improved self-efficacy scores for all groups at 1-month follow up, with sustained significant increases in baseline to 3-month scores only in the tailored Web-based group. One possible explanation could be related to "hit-count", which was significantly higher in the tailored group. This study is a first step in quantifying the contribution of Web-based tailoring over non-tailoring in changing key determinants of patient delay to AMI symptoms.

Algorithms↗

An integrative model for in-silico clinical-genomics discovery science.

Human Genome discovery research has set the pace for Post-Genomic Discovery Research. While post-genomic fields focused at the molecular level are intensively pursued, little effort is being deployed in the later stages of molecular medicine discovery research, such as clinical-genomics. The objective of this study is to demonstrate the relevance and significance of integrating mainstream clinical informatics decision support systems to current bioinformatics genomic discovery science. This paper is a feasibility study of an original model enabling novel "in-silico" clinical-genomic discovery science and that demonstrates its feasibility. This model is designed to mediate queries among clinical and genomic knowledge bases with relevant bioinformatic analytic tools (e.g. gene clustering). Briefly, trait-disease-gene relationships were successfully illustrated using QMR, OMIM, SNOMED-RT, GeneCluster and TreeView. The analyses were visualized as two-dimensional dendrograms of clinical observations clustered around genes. To our knowledge, this is the first study using knowledge bases of clinical decision support systems for genomic discovery. Although this study is a proof of principle, it provides a framework for the development of clinical decision-support-system driven, high-throughput clinical-genomic technologies which could potentially unveil significant high-level functions of genes.

Artificial Intelligence↗

Extended attributes of event monitor systems for criteria-based notification modalities.

The efficacy of event monitors (EMs) at reducing morbidity and mortality of certain clinical conditions (CCs) is well established. In addition, studies have shown that user inverted exclamation mark s preferences on the modality of notification are correlated to the type of reminder or alert. Nonetheless, few institutions have implemented large scale automated monitoring of a considerable number of distinct CCs, and to our knowledge, none of these sizable projects also offer user-customizable communication modalities (CMs) over all monitored conditions. As both the numbers of CMs and CCs increase, the complexity of customizing user preferences amplifies following a geometric progression. This paper demonstrates an automated approach, based on generic notification attributes (NAs) and notification criteria (NC), which significantly simplifies the management and personalization of the CMs for institutions where the manual assignment of a CM for every alert is forbidding. The methods by which these NAs were developed, their significance for existing CCs and their implementation using the Arden Syntax and Guideline interchange format (GLIF) are described. The proposed Criteria-Based Notification is shown to improve two facets of the management of event monitors: 1) the assignment of CMs becomes independent from clinical conditions, de-facto removing institution-specific CMs from the knowledge bases of the event monitors and inserting CC-specific and institution-independent NAs, thus increasing their reusability and sharability; 2) knowledge-based independent NAs facilitate both institution-level management and user-level preference configuration.

Decision Making, Computer-Assisted↗

Improving the human readability of Arden Syntax medical logic modules using a concept-oriented terminology and object-oriented programming expressions.

Medical logic modules are a procedural representation for sharing task-specific knowledge for decision support systems. Based on the premise that clinicians may perceive object-oriented expressions as easier to read than procedural rules in Arden Syntax-based medical logic modules, we developed a method for improving the readability of medical logic modules. Two approaches were applied: exploiting the concept-oriented features of the Medical Entities Dictionary and building an executable Java program to replace Arden Syntax procedural expressions. The usability evaluation showed that 66% of participants successfully mapped all Arden Syntax rules to Java methods. These findings suggest that these approaches can play an essential role in the creation of human readable medical logic modules and can potentially increase the number of clinical experts who are able to participate in the creation of medical logic modules. Although our approaches are broadly applicable, we specifically discuss the relevance to concept-oriented nursing terminologies and automated processing of task-specific nursing knowledge.

Attitude of Health Personnel↗