Search PubMed⌕ Search

Biomedical subjects

Carol Friedman

Publications and source records attributed to Carol Friedman.

At least 19 recordsLinked to original sources

Colorectal cancer in U.S. adults younger than 50 years of age, 1998-2001.

BACKGROUND: Colorectal cancer (CRC) incidence rates are increasing among persons younger than 50 years of age, a population routinely not screened unless an individual has a high risk of CRC. This population-based study focuses primarily on describing the CRC burden for persons in this age group. METHODS: The data used for this study were derived from the National Program of Cancer Registries (NPCR) and Surveillance, Epidemiology, and End Results (SEER) surveillance systems. Age-adjusted incidence rates, rate ratios, and their corresponding 95% confidence intervals were calculated. RESULTS: CRC is ranked among the top 10 cancers occurring in males and females aged 20-49 years regardless of race. Persons younger than 50 years were more likely to present with less localized and more distant disease than do older adults. Among younger adults, age-adjusted incidence rates for poorly differentiated cancers were twice as high as rates for well-differentiated cancers. Incidence rates for poorly differentiated cancers were 60% higher than that for well-differentiated cancers diagnosed in older adults. Rates were significantly higher for blacks and significantly lower for Asians/Pacific Islanders when compared with that for whites for the most demographic and tumor characteristics examined. CONCLUSIONS: This study confirms the findings of previous population-based studies suggesting that younger patients present with more advanced disease than do older patients. This study also identifies racial and ethnic disparities in CRC incidence in this population. These findings suggest the need for additional studies to understand the behavior and etiology of CRC in blacks.

Adolescent↗

Bio-Ontology and text: bridging the modeling gap.

MOTIVATION: Natural language processing (NLP) techniques are increasingly being used in biology to automate the capture of new biological discoveries in text, which are being reported at a rapid rate. Yet, information represented in NLP data structures is classically very different from information organized with ontologies as found in model organisms or genetic databases. To facilitate the computational reuse and integration of information buried in unstructured text with that of genetic databases, we propose and evaluate a translational schema that represents a comprehensive set of phenotypic and genetic entities, as well as their closely related biomedical entities and relations as expressed in natural language. In addition, the schema connects different scales of biological information, and provides mappings from the textual information to existing ontologies, which are essential in biology for integration, organization, dissemination and knowledge management of heterogeneous phenotypic information. A common comprehensive representation for otherwise heterogeneous phenotypic and genetic datasets, such as the one proposed, is critical for advancing systems biology because it enables acquisition and reuse of unprecedented volumes of diverse types of knowledge and information from text. RESULTS: A novel representational schema, PGschema, was developed that enables translation of phenotypic, genetic and their closely related information found in textual narratives to a well-defined data structure comprising phenotypic and genetic concepts from established ontologies along with modifiers and relationships. Evaluation for coverage of a selected set of entities showed that 90% of the information could be represented (95% confidence interval: 86-93%; n = 268). Moreover, PGschema can be expressed automatically in an XML format using natural language techniques to process the text. To our knowledge, we are providing the first evaluation of a translational schema for NLP that contains declarative knowledge about genes and their associated biomedical data (e.g. phenotypes). AVAILABILITY: http://zellig.cpmc.columbia.edu/PGschema

Abstracting and Indexing↗

Machine learning and word sense disambiguation in the biomedical domain: design and evaluation issues.

BACKGROUND: Word sense disambiguation (WSD) is critical in the biomedical domain for improving the precision of natural language processing (NLP), text mining, and information retrieval systems because ambiguous words negatively impact accurate access to literature containing biomolecular entities, such as genes, proteins, cells, diseases, and other important entities. Automated techniques have been developed that address the WSD problem for a number of text processing situations, but the problem is still a challenging one. Supervised WSD machine learning (ML) methods have been applied in the biomedical domain and have shown promising results, but the results typically incorporate a number of confounding factors, and it is problematic to truly understand the effectiveness and generalizability of the methods because these factors interact with each other and affect the final results. Thus, there is a need to explicitly address the factors and to systematically quantify their effects on performance. RESULTS: Experiments were designed to measure the effect of "sample size" (i.e. size of the datasets), "sense distribution" (i.e. the distribution of the different meanings of the ambiguous word) and "degree of difficulty" (i.e. the measure of the distances between the meanings of the senses of an ambiguous word) on the performance of WSD classifiers. Support Vector Machine (SVM) classifiers were applied to an automatically generated data set containing four ambiguous biomedical abbreviations: BPD, BSA, PCA, and RSV, which were chosen because of varying degrees of differences in their respective senses. Results showed that: 1) increasing the sample size generally reduced the error rate, but this was limited mainly to well-separated senses (i.e. cases where the distances between the senses were large); in difficult cases an unusually large increase in sample size was needed to increase performance slightly, which was impractical, 2) the sense distribution did not have an effect on performance when the senses were separable, 3) when there was a majority sense of over 90%, the WSD classifier was not better than use of the simple majority sense, 4) error rates were proportional to the similarity of senses, and 5) there was no statistical difference between results when using a 5-fold or 10-fold cross-validation method. Other issues that impact performance are also enumerated. CONCLUSION: Several different independent aspects affect performance when using ML techniques for WSD. We found that combining them into one single result obscures understanding of the underlying methods. Although we studied only four abbreviations, we utilized a well-established statistical method that guarantees the results are likely to be generalizable for abbreviations with similar characteristics. The results of our experiments show that in order to understand the performance of these ML methods it is critical that papers report on the baseline performance, the distribution and sample size of the senses in the datasets, and the standard deviation or confidence intervals. In addition, papers should also characterize the difficulty of the WSD task, the WSD situations addressed and not addressed, as well as the ML methods and features used. This should lead to an improved understanding of the generalizablility and the limitations of the methodology.

Algorithms↗

Quantitative assessment of dictionary-based protein named entity tagging.

OBJECTIVE: Natural language processing (NLP) approaches have been explored to manage and mine information recorded in biological literature. A critical step for biological literature mining is biological named entity tagging (BNET) that identifies names mentioned in text and normalizes them with entries in biological databases. The aim of this study was to provide quantitative assessment of the complexity of BNET on protein entities through BioThesaurus, a thesaurus of gene/protein names for UniProt knowledgebase (UniProtKB) entries that was acquired using online resources. METHODS: We evaluated the complexity through several perspectives: ambiguity (i.e., the number of genes/proteins represented by one name), synonymy (i.e., the number of names associated with the same gene/protein), and coverage (i.e., the percentage of gene/protein names in text included in the thesaurus). We also normalized names in BioThesaurus and measures were obtained twice, once before normalization and once after. RESULTS: The current version of BioThesaurus has over 2.6 million names or 2.1 million normalized names covering more than 1.8 million UniProtKB entries. The average synonymy is 3.53 (2.86 after normalization), ambiguity is 2.31 before normalization and 2.32 after, while the coverage is 94.0% based on the BioCreAtive data set comprising MEDLINE abstracts containing genes/proteins. CONCLUSION: The study indicated that names for genes/proteins are highly ambiguous and there are usually multiple names for the same gene or protein. It also demonstrated that most gene/protein names appearing in text can be found in BioThesaurus.

Dictionaries as Topic↗

Human and automated coding of rehabilitation discharge summaries according to the International Classification of Functioning, Disability, and Health.

OBJECTIVE: The International Classification of Functioning, Disability, and Health (ICF) is designed to provide a common language and framework for describing health and health-related states. The goal of this research was to investigate human and automated coding of functional status information using the ICF framework. DESIGN: The authors extended an existing natural language processing (NLP) system to encode rehabilitation discharge summaries according to the ICF. MEASUREMENTS: The authors conducted a formal evaluation, comparing the coding performed by expert coders, non-expert coders, and the NLP system. RESULTS: Automated coding can be used to assign codes using the ICF, with results similar to those obtained by human coders, at least for the selection of ICF code and assignment of the performance qualifier. Coders achieved high agreement on ICF code assignment. CONCLUSION: This research is a key next step in the development of the ICF as a sensitive and universal classification of functional status information. It is worthwhile to continue to investigate automated ICF coding.

Activities of Daily Living↗

Terminology model discovery using natural language processing and visualization techniques.

Medical terminologies are important for unambiguous encoding and exchange of clinical information. The traditional manual method of developing terminology models is time-consuming and limited in the number of phrases that a human developer can examine. In this paper, we present an automated method for developing medical terminology models based on natural language processing (NLP) and information visualization techniques. Surgical pathology reports were selected as the testing corpus for developing a pathology procedure terminology model. The use of a general NLP processor for the medical domain, MedLEE, provides an automated method for acquiring semantic structures from a free text corpus and sheds light on a new high-throughput method of medical terminology model development. The use of an information visualization technique supports the summarization and visualization of the large quantity of semantic structures generated from medical documents. We believe that a general method based on NLP and information visualization will facilitate the modeling of medical terminologies.

Automation↗

Annual report to the nation on the status of cancer, 1975-2002, featuring population-based trends in cancer treatment.

BACKGROUND: The American Cancer Society (ACS), the Centers for Disease Control and Prevention (CDC), the National Cancer Institute (NCI), and the North American Association of Central Cancer Registries (NAACCR) collaborate annually to provide information on cancer rates and trends in the United States. This year's report updates statistics on the 15 most common cancers in the five major racial/ethnic populations in the United States for 1992-2002 and features population-based trends in cancer treatment. METHODS: The NCI, the CDC, and the NAACCR provided information on cancer cases, and the CDC provided information on cancer deaths. Reported incidence and death rates were age-adjusted to the 2000 U.S. standard population, annual percent change in rates for fixed intervals was estimated by linear regression, and annual percent change in trends was estimated with joinpoint regression analysis. Population-based treatment data were derived from the Surveillance, Epidemiology, and End Results (SEER) Program registries, SEER-Medicare linked databases, and NCI Patterns of Care/Quality of Care studies. RESULTS: Among men, the incidence rates for all cancer sites combined were stable from 1995 through 2002. Among women, the incidence rates increased by 0.3% annually from 1987 through 2002. Death rates in men and women combined decreased by 1.1% annually from 1993 through 2002 for all cancer sites combined and also for many of the 15 most common cancers. Among women, lung cancer death rates increased from 1995 through 2002, but lung cancer incidence rates stabilized from 1998 through 2002. Although results of cancer treatment studies suggest that much of contemporary cancer treatment for selected cancers is consistent with evidence-based guidelines, they also point to geographic, racial, economic, and age-related disparities in cancer treatment. CONCLUSIONS: Cancer death rates for all cancer sites combined and for many common cancers have declined at the same time as the dissemination of guideline-based treatment into the community has increased, although this progress is not shared equally across all racial and ethnic populations. Data from population-based cancer registries, supplemented by linkage with administrative databases, are an important resource for monitoring the quality of cancer treatment. Use of this cancer surveillance system, along with new developments in medical informatics and electronic medical records, may facilitate monitoring of the translation of basic science and clinical advances to cancer prevention, detection, and uniformly high quality of care in all areas and populations of the United States.

Age Distribution↗

Qualitative assessment of the International Classification of Functioning, Disability, and Health with respect to the desiderata for controlled medical vocabularies.

BACKGROUND: The International Classification of Functioning, Disability, and Health (ICF), a classification system published in 2001 by the World Health Organization (WHO), provides a common language and framework for describing functional status information (FSI) in health records. METHODS: Informed by ongoing research in coding FSI in patient records, this paper qualitatively assesses the ICF framework with respect to the desiderata for controlled medical vocabularies, an enumerated a list of desirable qualities for controlled medical vocabularies proposed by Cimino [J.J. Cimino, Desiderata for controlled medical vocabularies in the twenty-first century, Meth. Inform. Med. 37 (1998) 394-403]. RESULTS: The ICF satisfies 5 of the 12 desiderata. Five points were not satisfied and two points could not be evaluated. CONCLUSION: The ICF is a rich source of relevant terms, concepts, and relationships, but it was not developed in consideration of requirements for formal terminologies. Therefore, it could serve as a base from which to develop a formal terminology of functioning and disability. This assessment is a key next step in the development of the ICF as a sensitive, universal measure of functional status.

Decision Support Techniques↗

Investigation of postoperative allograft-associated infections in patients who underwent musculoskeletal allograft implantation.

BACKGROUND: The rate at which allografts are used in surgical procedures has doubled in the United States during the past decade. In 2002, one outpatient surgical center (SC-X) identified a cluster of surgical site infections (SSIs) after anterior cruciate ligament reconstructive surgery (ACLRS). Therefore, we conducted an investigation to determine the extent of the outbreak and to identify risk factors. METHODS: Our investigation included retrospective cohort and observational studies. A case patient was defined as any patient who acquired a SSI after undergoing ACLRS at SC-X between February 2000 and June 2002 (the study period). Data collected included demographic characteristics, clinical information, and graft details, such as processing method (i.e., aseptic or sterile). RESULTS: Of 331 patients who underwent ACLRS during the study period, 11 (3.3%) met the case definition. All infections occurred at the tibial fixation site of the graft and involved 8 different microorganisms; the median time to a positive culture result was 55 days after ACLRS. The infection rate for patients who received aseptically processed allografts was 4.4% (11 of 250 patients), compared with 0% (0 of 81) for patients who received autografts or sterile allografts (P=.07). Use of a supplementary staple for tibial fixation, compared with other fixation methods that did not involve such staples, increased the risk of infection 10-fold in univariate analysis (relative risk [RR], 10.0; 95% confidence interval [CI], 3.0-32.9) and 9-fold when controlling for tissue processing method (RR, 9.0; 95% CI, 2.8-28.8). CONCLUSIONS: The use of sterile allograft tissue appears to be associated with a significant reduction in the risk of postoperative infection, particularly in the presence of adjunctive fixation. Larger clinical studies are necessary to confirm this observation.

Adolescent↗

ISO reference terminology models for nursing: applicability for natural language processing of nursing narratives.

Natural language processing (NLP) systems have demonstrated utility in parsing narrative texts for purposes such as surveillance and decision support. However, there has been little work related to NLP of nursing narratives. The purpose of this study was to compare the semantic categories of a NLP system (Medical Language Extraction and Encoding [MedLEE] system) with the semantic domains, categories, and attributes of the International Standards Organization (ISO) reference terminology models for nursing diagnoses and nursing actions. All but two MedLEE diagnosis and procedure-related semantic categories mapped to ISO models. In some instances, we found exact correspondence between the semantic structures of MedLEE and the ISO models. In other situations (e.g. aspects of Site or Location), the ISO model was not as granular as MedLEE. For clinical procedure and non-invasive examination, two ISO nursing action model components (Action and Target) mapped to a single MedLEE semantic category. The ISO models are applicable to NLP of nursing narratives. However, the ISO models require additional specification of selected semantic categories for the abstract semantic domains in order to achieve the objective of using NLP to parse and encode data from nursing narratives. Our analysis also suggests areas for extension of MedLEE particularly in regard to represent nursing actions.

Diagnosis, Computer-Assisted↗

Extracting information on pneumonia in infants using natural language processing of radiology reports.

Natural language processing (NLP) is critical for improvement of the healthcare process because it can encode clinical data in patient documents. Many clinical applications such as decision support require coded data to function appropriately. However, in order to be applicable for healthcare, performance must be adequate. A valuable automated application is the detection of infectious diseases, such as surveillance of pneumonia in newborns (e.g., neonates) because the disease produces significant rates of morbidity and mortality, and manual surveillance is challenging. Studies have demonstrated that automated surveillance using NLP is a useful adjunct to manual surveillance and an effective tool for infection control practitioners. This paper presents a study evaluating the feasibility of an NLP-based monitoring system to screen for healthcare-associated pneumonia in neonates. We estimated sensitivity, specificity, and positive predictive value by comparing results with clinicians' judgments. Sensitivity was 71% and specificity was 99%. Our results demonstrated that the automated method was feasible.

Database Management Systems↗

Use of computerized surveillance to detect nosocomial pneumonia in neonatal intensive care unit patients.

BACKGROUND: Pneumonia surveillance is difficult and time-consuming. The definition is complicated, and there are many opportunities for subjectivity in determining infection status. OBJECTIVE: To compare traditional infection control professional (ICP) surveillance for pneumonia among neonatal intensive care unit (NICU) patients with computerized surveillance of chest x-ray reports using an automated detection system based on a natural language processor. METHODS: This system evaluated chest x-rays from 2 NICUs over a 2-year period. It flagged x-rays indicative of pneumonia according to rules derived from the National Nosocomial Infection Surveillance System definition as applied to radiology reports. Data from the automated system were compared with pneumonia data collected prospectively by an ICP. RESULTS: Sensitivity of the computerized surveillance in NICU 1 was 71%, and specificity was 99.8%. The positive predictive value was 7.9%, and the negative predictive value (NPV) was >99%. Data from NICU 2 were incomplete. CONCLUSIONS: Computer-assisted surveillance has the potential to decrease ICP workload and make pneumonia surveillance feasible. The high NPV means the system can safely screen out many chest x-rays of noninfected patients. However, all data must be available to the computer system and must be analyzed the same way for results to be comparable.

Computers↗

Does health insurance coverage of office visits influence colorectal cancer testing?

OBJECTIVE: To assess the effect of differing health insurance coverage of physician office visits on the use of colorectal cancer (CRC) tests among an employed and insured population. METHOD: Cohort study of persons ages 50 to 64 years enrolled in fee-for-service (FFS) or preferred provider organization (PPO) health plans, where FFS plan enrollees bear disproportionate share of office visit coverage, for the period 1995 through 1999. RESULTS: Compared with FFS plans, enrollees in PPO plans were significantly more likely to obtain CRC tests [adjusted relative risk (RR(a)), 1.27; 95% confidence intervals (CI), 1.21-1.24]. The association was more pronounced among hourly individuals (RR(a), 1.43; 95% CI, 1.41-1.45) than among salaried individuals (RR(a), 1.09; 95% CI, 1.05-1.10), consistent with a greater differential in office visit coverage among the hourly group. CONCLUSIONS: Disproportionate cost-sharing seems to have a negative effect on the use of CRC tests most likely by discouraging nonacute care physician office visits.

Cohort Studies↗

System architecture for temporal information extraction, representation and reasoning in clinical narrative reports.

Exploring temporal information in narrative Electronic Medical Records (EMRs) is essential and challenging. We propose an architecture for an integrated approach to process temporal information in clinical narrative reports. The goal is to initiate and build a foundation that supports applications which assist healthcare practice and research by including the ability to determine the time of clinical events (e.g., past vs. present). Key components include: (1) an annotation schema for temporal expressions and the development of an associated tagger; (2) a natural language processing (NLP) system for encoding and extracting medical events and associating them with formalized temporal data; (3) a post-processor, with a knowledge-based subsystem to help discover implicit information, that resolves temporal expressions and deals with issues such as granularity and vagueness; and (4) a reasoning mechanism which models clinical reports as Simple Temporal Problems (STPs).

Humans↗

Extending a medical language processing system to the functional status domain.

The World Health Organization's International Classification of Functioning, Disability, and Health (ICF) provides a common framework for describing functional status information (FSI) in health records. Given the expense of manual coding, we are investigating the use of natural language processing (NLP) for automated FSI coding. We used an existing NLP system that was originally designed to encode clinical information. The system's lexicon and coding table were modified and preprocessing and postprocessing programs were created, allowing for automated assignment of selected ICF codes.

Activities of Daily Living↗

Natural language processing in the molecular imaging domain.

Molecular imaging represents the intersection between imaging and genomic sciences. There has been a surge in research literature and information in both sciences. Information contained within molecular imaging literature could be used to 1) link to genomic and imaging information resources and 2) to organize and index images. This research focuses on the adaptation, evaluation, and application of BioMedLEE, a natural language processing system (NLP), in the automated extraction of information from molecular imaging abstracts.

Cell Line↗