Search PubMed⌕ Search

Biomedical subjects

Carol Friedman

Publications and source records attributed to Carol Friedman.

At least 37 records · Page 2Linked to original sources

Extending a medical language processing system to the functional status domain.

The World Health Organization's International Classification of Functioning, Disability, and Health (ICF) provides a common framework for describing functional status information (FSI) in health records. Given the expense of manual coding, we are investigating the use of natural language processing (NLP) for automated FSI coding. We used an existing NLP system that was originally designed to encode clinical information. The system's lexicon and coding table were modified and preprocessing and postprocessing programs were created, allowing for automated assignment of selected ICF codes.

Activities of Daily Living↗

Natural language processing in the molecular imaging domain.

Molecular imaging represents the intersection between imaging and genomic sciences. There has been a surge in research literature and information in both sciences. Information contained within molecular imaging literature could be used to 1) link to genomic and imaging information resources and 2) to organize and index images. This research focuses on the adaptation, evaluation, and application of BioMedLEE, a natural language processing system (NLP), in the automated extraction of information from molecular imaging abstracts.

Cell Line↗

Visualizing information across multidimensional post-genomic structured and textual databases.

MOTIVATION: Visualizing relationships among biological information to facilitate understanding is crucial to biological research during the post-genomic era. Although different systems have been developed to view gene-phenotype relationships for specific databases, very few have been designed specifically as a general flexible tool for visualizing multidimensional genotypic and phenotypic information together. Our goal is to develop a method for visualizing multidimensional genotypic and phenotypic information and a model that unifies different biological databases in order to present the integrated knowledge using a uniform interface. RESULTS: We developed a novel, flexible and generalizable visualization tool, called PhenoGenesviewer (PGviewer), which in this paper was used to display gene-phenotype relationships from a human-curated database (OMIM) and from an automatic method using a Natural Language Processing tool called BioMedLEE. Data obtained from multiple databases were first integrated into a uniform structure and then organized by PGviewer. PGviewer provides a flexible query interface that allows dynamic selection and ordering of any desired dimension in the databases. Based on users' queries, results can be visualized using hierarchical expandable trees that present views specified by users according to their research interests. We believe that this method, which allows users to dynamically organize and visualize multiple dimensions, is a potentially powerful and promising tool that should substantially facilitate biological research. AVAILABILITY: PhenogenesViewer as well as its support and tutorial are available at http://www.dbmi.columbia.edu/pgviewer/ CONTACT: Lussier@dbmi.columbia.edu.

Computer Graphics↗

Gene name ambiguity of eukaryotic nomenclatures.

MOTIVATION: With more and more scientific literature published online, the effective management and reuse of this knowledge has become problematic. Natural language processing (NLP) may be a potential solution by extracting, structuring and organizing biomedical information in online literature in a timely manner. One essential task is to recognize and identify genomic entities in text. 'Recognition' can be accomplished using pattern matching and machine learning. But for 'identification' these techniques are not adequate. In order to identify genomic entities, NLP needs a comprehensive resource that specifies and classifies genomic entities as they occur in text and that associates them with normalized terms and also unique identifiers so that the extracted entities are well defined. Online organism databases are an excellent resource to create such a lexical resource. However, gene name ambiguity is a serious problem because it affects the appropriate identification of gene entities. In this paper, we explore the extent of the problem and suggest ways to address it. RESULTS: We obtained gene information from 21 organisms and quantified naming ambiguities within species, across species, with English words and with medical terms. When the case (of letters) was retained, official symbols displayed negligible intra-species ambiguity (0.02%) and modest ambiguities with general English words (0.57%) and medical terms (1.01%). In contrast, the across-species ambiguity was high (14.20%). The inclusion of gene synonyms increased intra-species ambiguity substantially and full names contributed greatly to gene-medical-term ambiguity. A comprehensive lexical resource that covers gene information for the 21 organisms was then created and used to identify gene names by using a straightforward string matching program to process 45,000 abstracts associated with the mouse model organism while ignoring case and gene names that were also English words. We found that 85.1% of correctly retrieved mouse genes were ambiguous with other gene names. When gene names that were also English words were included, 233% additional 'gene' instances were retrieved, most of which were false positives. We also found that authors prefer to use synonyms (74.7%) to official symbols (17.7%) or full names (7.6%) in their publications. CONTACT: lifeng.chen@dbmi.columbia.edu

Abstracting and Indexing↗

Automated encoding of clinical documents based on natural language processing.

OBJECTIVE: The aim of this study was to develop a method based on natural language processing (NLP) that automatically maps an entire clinical document to codes with modifiers and to quantitatively evaluate the method. METHODS: An existing NLP system, MedLEE, was adapted to automatically generate codes. The method involves matching of structured output generated by MedLEE consisting of findings and modifiers to obtain the most specific code. Recall and precision applied to Unified Medical Language System (UMLS) coding were evaluated in two separate studies. Recall was measured using a test set of 150 randomly selected sentences, which were processed using MedLEE. Results were compared with a reference standard determined manually by seven experts. Precision was measured using a second test set of 150 randomly selected sentences from which UMLS codes were automatically generated by the method and then validated by experts. RESULTS: Recall of the system for UMLS coding of all terms was .77 (95% CI.72-.81), and for coding terms that had corresponding UMLS codes recall was .83 (.79-.87). Recall of the system for extracting all terms was .84 (.81-.88). Recall of the experts ranged from .69 to .91 for extracting terms. The precision of the system was .89 (.87-.91), and precision of the experts ranged from .61 to .91. CONCLUSION: Extraction of relevant clinical information and UMLS coding were accomplished using a method based on NLP. The method appeared to be comparable to or better than six experts. The advantage of the method is that it maps text to codes along with other related information, rendering the coded output suitable for effective retrieval.

Abstracting and Indexing↗

Cancer mortality surveillance--United States, 1990-2000.

PROBLEM/CONDITION: Cancer is the second leading cause of death in the United States and is expected to become the leading cause of death within the next decade. Considerable variation exists in cancer mortality between the sexes and among different racial/ethnic populations and geographic locations. The description of mortality data by state, sex, and race/ethnicity is essential for cancer-control researchers to target areas of need and develop programs that reduce the burden of cancer. REPORTING PERIOD COVERED: 1990-2000. DESCRIPTION OF SYSTEM: Mortality data from CDC were used to calculate death rates and trends, categorized by state, sex, and race/ethnicity. Trend analyses for 1990-2000 are presented for all cancer sites combined and for the four leading cancers causing death (lung/bronchus, colorectal, prostate, and breast) categorized by state, sex, and race/ethnicity. Death rates per 100,000 population for the 10 primary cancer sites with the highest age-adjusted rates are also presented for each state and the District of Columbia by sex. For males, the 10 primary sites include lung/bronchus, prostate, colon/rectum, pancreas, leukemia, non-Hodgkin lymphoma, liver/intrahepatic bile duct, esophagus, stomach, and urinary bladder. For females, the 10 primary sites include lung/bronchus, breast, colon/rectum, pancreas, ovary, non-Hodgkin lymphoma, leukemia, brain/other nervous system, uterine corpus, and myeloma. RESULTS: For 1990-2000, cancer mortality decreased among the majority of racial/ethnic populations and geographic locations in the United States. Statistically significant decreases in mortality among all races combined occurred with lung and bronchus cancer among men (--1.7%/year); colorectal cancer among men and women (--2.0%/year and--1.7%/year, respectively); prostate cancer (--2.6%/year); and female breast cancer (--2.3%/year). For 1990-2000, cancer mortality remained stable among American Indian/Alaska Native populations. Statistically significant increases in lung and bronchus cancer mortality occurred among women of all racial/ethnic backgrounds, except among Asian/Pacific Islanders. INTERPRETATION: Although cancer remains the second leading cause of death in the United States, the overall declining trend in cancer mortality demonstrates considerable progress in cancer prevention, early detection, and treatment. PUBLIC HEALTH ACTION: More effective tobacco-cessation programs are necessary to reduce lung and bronchus cancer mortality among women and sustain the decrease in lung and bronchus cancer mortality among men. Additional programs that deter smoking initiation among adolescents are essential to ensure future decreases in lung and bronchus cancer mortality. Continued research in primary prevention, screening methods, and therapeutics is needed to further reduce disparities and improve quality of life and survival among all populations.

Adult↗

A multi-aspect comparison study of supervised word sense disambiguation.

OBJECTIVE: The aim of this study was to investigate relations among different aspects in supervised word sense disambiguation (WSD; supervised machine learning for disambiguating the sense of a term in a context) and compare supervised WSD in the biomedical domain with that in the general English domain. METHODS: The study involves three data sets (a biomedical abbreviation data set, a general biomedical term data set, and a general English data set). The authors implemented three machine-learning algorithms, including (1) naïve Bayes (NBL) and decision lists (TDLL), (2) their adaptation of decision lists (ODLL), and (3) their mixed supervised learning (MSL). There were six feature representations (various combinations of collocations, bag of words, oriented bag of words, etc.) and five window sizes (2, 4, 6, 8, and 10). RESULTS: Supervised WSD is suitable only when there are enough sense-tagged instances with at least a few dozens of instances for each sense. Collocations combined with neighboring words are appropriate selections for the context. For terms with unrelated biomedical senses, a large window size such as the whole paragraph should be used, while for general English words a moderate window size between 4 and 10 should be used. The performance of the authors' implementation of decision list classifiers for abbreviations was better than that of traditional decision list classifiers. However, the opposite held for the other two sets. Also, the authors' mixed supervised learning was stable and generally better than others for all sets. CONCLUSION: From this study, it was found that different aspects of supervised WSD depend on each other. The experiment method presented in the study can be used to select the best supervised WSD classifier for each ambiguous term.

Abbreviations as Topic↗

Probabilistic inference of molecular networks from noisy data sources.

Information on molecular networks, such as networks of interacting proteins, comes from diverse sources that contain remarkable differences in distribution and quantity of errors. Here, we introduce a probabilistic model useful for predicting protein interactions from heterogeneous data sources. The model describes stochastic generation of protein-protein interaction networks with real-world properties, as well as generation of two heterogeneous sources of protein-interaction information: research results automatically extracted from the literature and yeast two-hybrid experiments. Based on the domain composition of proteins, we use the model to predict protein interactions for pairs of proteins for which no experimental data are available. We further explore the prediction limits, given experimental data that cover only part of the underlying protein networks. This approach can be extended naturally to include other types of biological data sources.

Algorithms↗

GeneWays: a system for extracting, analyzing, visualizing, and integrating molecular pathway data.

The immense growth in the volume of research literature and experimental data in the field of molecular biology calls for efficient automatic methods to capture and store information. In recent years, several groups have worked on specific problems in this area, such as automated selection of articles pertinent to molecular biology, or automated extraction of information using natural-language processing, information visualization, and generation of specialized knowledge bases for molecular biology. GeneWays is an integrated system that combines several such subtasks. It analyzes interactions between molecular substances, drawing on multiple sources of information to infer a consensus view of molecular networks. GeneWays is designed as an open platform, allowing researchers to query, review, and critique stored information.

Artificial Intelligence↗

Assessing the burden of disease among an employed population: implications for employer-sponsored prevention programs.

Escalating healthcare costs have led employers to identify ways to assess the actual burden of disease among their employees. One such measure is the use of disability-adjusted life-years (DALYs). DALYs were calculated for the General Motors (GM) population for 1994 through 1998 using data from GM's Mortality Registry, published life tables, and age- and sex-specific disease incidence and disability data from the U.S. Burden of Disease Study. Chronic diseases accounted for 45% (245,844 of 540,450) of total DALYs lost. Ischemic heart disease, stroke, lung cancer, and chronic obstructive pulmonary disease led the list for both men and women and accounted for 39% and 31%, respectively, of the top 10 DALYs lost. Disease burden among employees could be reduced through targeted interventions aimed at the risk factors associated with the leading causes of DALYs.

Cost of Illness↗

A comparison of semantic categories of the ISO reference terminology models for nursing and the MedLEE natural language processing system.

Natural language processing (NLP) systems have demonstrated utility in parsing narrative texts for purposes such as surveillance and decision support. However, there has been little work related to NLP of nursing narratives. The purpose of this study was to compare the semantic categories of a NLP system (Medical Language Extraction and Encoding [MedLEE] system) with the semantic domains, categories, and attributes of the International Standards Organization(ISO) reference terminology models for nursing diagnoses and nursing actions. All but two MedLEE diagnosis and procedure-related semantic categories mapped to ISO models. In some instances, we found exact correspondence between the semantic structures of MedLEE and the ISO models. In other situations (e.g. aspects of site or location), the ISO model was not as granular as MedLEE. For clinical procedure and non-invasive examination, two ISO nursing action model components (action and target) were required to represent the MedLEE semantic category. The ISO model requires additional specification of selected semantic categories for the abstract semantic domains in order to achieve the objective of using NLP to parse and encode data from nursing narratives. Our analysis also suggests areas for extension of MedLEE.

Natural Language Processing↗

Facilitating cancer research using natural language processing of pathology reports.

Many ongoing clinical research projects, such as projects involving studies associated with cancer, involve manual capture of information in surgical pathology reports so that the information can be used to determine the eligibility of recruited patients for the study and to provide other information, such as cancer prognosis. Natural language processing (NLP) systems offer an alternative to automated coding, but pathology reports have certain features that are difficult for NLP systems. This paper describes how a preprocessor was integrated with an existing NLP system (MedLEE) in order to reduce modification to the NLP system and to improve performance. The work was done in conjunction with an ongoing clinical research project that assesses disparities and risks of developing breast cancer for minority women. An evaluation of the system was performed using manually coded data from the research project's database as a gold standard. The evaluation outcome showed that the extended NLP system had a sensitivity of 90.6% and a precision of 91.6%. Results indicated that this system performed satisfactorily for capturing information for the cancer research project.

Biomedical Research↗

CliniViewer: a tool for viewing electronic medical records based on natural language processing and XML.

With the evolving use of computers in healthcare, the electronic medical record (EMR) is becoming more and more popular. A tool is needed that would enable physicians to accurately and efficiently access clinical information in multiple medical records associated with a particular patient. Both natural language processing (NLP) and the eXtensible Markup Language (XML) have been used in the clinical domain for capturing, representing, and utilizing clinical information and both have shown great potential. In this paper, we demonstrate another use of XML and NLP through CliniViewer, a tool that organizes and presents the clinical information in multiple records. We also describe the flexibility and capability provided when combining XML and NLP to summarize, navigate, and conceptualize structured information. The tool has been fully implemented and tested using patients with multiple discharge summaries.

Humans↗

Extracting phenotypic information from the literature via natural language processing.

In recent years, the amount of biomedical knowledge has been increasing exponentially. Several Natural Language Processing (NLP) systems have been developed to help researchers extract, encode and organize new information automatically from textual literature or narrative reports. Some of these systems focus on extracting biological entities or molecular interactions while others retrieve and encode clinical information. To exploit gene functions in the post-genome era, it is necessary to extract phenotypic information automatically from the literature as well. However, few NLP projects have focused on this. We present the development of a system called BioMedLEE that extracts a broad variety of phenotypic information from the biomedical literature. The system was developed by adapting MedLEE, an existing clinical information extraction NLP engine. A feasibility evaluation study of BioMedLEE was performed using 300 randomly chosen journal titles. Results showed that experts achieved an average precision rate of 65.4%, (95%CI: [58.0%, 72.8%]) and a recall rate of 73.0%, (95%CI: [66.2%, 80.0%]). BioMedLEE had 64.0% precision and 77.1% recall respectively, according to expert agreements.

Databases as Topic↗

The influence of year-end bonuses on colorectal cancer screening.

OBJECTIVE: To estimate the effect of physician bonus eligibility on colorectal cancer (CRC) screening, controlling for patient and primary care physician characteristics. STUDY DESIGN: Retrospective study using managed care plan claims data from 2000 and 2001. METHODS: Data on 50-year-old commercially insured patients in a managed care health plan were linked to enrollment and provider files. The data included information on 6749 patients (3058 in 2000 and 3691 in 2001). Multivariate logistic regression models were used to assess the association between CRC screening receipt and physician bonus eligibility. RESULTS: From 2000 to 2001, CRC screening use increased from 23.4% to 26.4% (P < .01). Results from the multivariate logistic regression analysis revealed that the probability that a patient received a CRC screening was approximately 3 percentage points higher in the bonus year, 2001 (P < .01). CONCLUSIONS: Bonuses targeted at individual physicians were associated with increased use of CRC screening tests. However, more research is needed to examine the effect of performance-based incentives on resource use and the quality of medical care. Specifically, there is a need to determine whether explicit financial incentives are effective in reducing racial disparities in the quality of patient care. This has particular relevance for CRC screening given that black patients are less likely to be screened, they have higher CRC incidence and mortality rates compared with other racial groups, and screening has been shown to be more cost effective in this population.

Colorectal Neoplasms↗

Effect of the frequency of delivery of reminders and an influenza tool kit on increasing influenza vaccination rates among adults with high-risk conditions.

OBJECTIVE: To evaluate the incremental effect of a second client reminder postcard or an influenza tool kit targeted toward employers on increasing influenza vaccination rates among adults age < 65 years at high risk for complications from influenza illness. METHODS: In this demonstration study, enrollees of 3 managed care organizations (n = 8881) were randomized at the employer level into 4 arms: 1 postcard, 2 postcards, 1 postcard + tool kit, and 2 postcards + tool kit. The postcards and tool kits were mailed during the fall of 2001, and their effect on influenza vaccination rates was assessed through a survey. RESULTS: Compared with a single postcard, 2 postcards increased vaccination rates by 4 percentage points (adjusted relative risk = 1.05; P < .05) among persons aged 50 to 64 years but did not have any effect among younger adults. Older adults had a greater burden of disease and reported more favorable knowledge and attitudes toward the influenza vaccine. The influenza tool kit did not appear to have any incremental effect on vaccination rates. CONCLUSIONS: Our findings underscore the necessity of evaluating the effectiveness of interventions in different population subgroups and of identifying factors that modify the effectiveness of interventions. Rigorous assessment of intervention effectiveness in managed care settings will enable decision makers to optimize use of scarce healthcare dollars for improving the health and well-being of enrollees.

Adolescent↗

A vocabulary development and visualization tool based on natural language processing and the mining of textual patient reports.

Medical terminologies are critical for automated healthcare systems. Some terminologies, such as the UMLS and SNOMED are comprehensive, whereas others specialize in limited domains (i.e., BIRADS) or are developed for specific applications. An important feature of a terminology is comprehensive coverage of relevant clinical terms and ease of use by users, which include computerized applications. We have developed a method for facilitating vocabulary development and maintenance that is based on utilization of natural language processing to mine large collections of clinical reports in order to obtain information on terminology as expressed by physicians. Once the reports are processed and the terms structured and collected into an XML representational schema, it is possible to determine information about terms, such as frequency of occurrence, compositionality, relations to other terms (such as modifiers), and correspondence to a controlled vocabulary. This paper describes the method and discusses how it can be used as a tool to help vocabulary builders navigate through the terms physicians use, visualize their relations to other terms via a flexible viewer, and determine their correspondence to a controlled vocabulary.

Algorithms↗

Mining terminological knowledge in large biomedical corpora.

Terminological knowledge of the biomedical domain is important for natural language processing (NLP) and information retrieval (IR) applications, and a number of terminological knowledge sources, such as LocusLink, GeneBank, and the UMLS, already exist. However, because of the tremendous amount of research activity in the field, new terms and symbols are continually being created, many of which are published in the literature, but are not available in any of the other resources. Therefore, effective mining of the literature for new terminology is critical for furthering NLP and IR applications. Abbreviations are widely used in the biomedical domain, and the understanding of abbreviations requires a terminological knowledge base that consists of abbreviations with their associated senses. In previous work, several methods have been developed for automatic construction of abbreviation knowledge bases from parenthetical expressions. However, these methods pair abbreviations and their expansions based on manually crafted patterns or rules. In this paper, we propose an automatic method, which is not based on patterns or rules but is based on the use of collocations, to extract a set of related terms from parenthetical expressions including abbreviations associated with their expansions and other types of related terms such as synonyms, or hyponyms etc. Our method is based on the observation that terms associated with parenthetical expressions i) are usually related, and ii) are often collocations because they tend to co-occur more often than expected by chance. Our method was applied to the collection of MEDLINE abstracts. The method and the results were evaluated using two collections: Berman's handcrafted abbreviation list and the LocusLink collection.

Abstracting and Indexing↗