Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Information Retrieval Systems”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Automated classification of encounter notes in a computer based medical record.

Harvard Community Health Plan is exploring emerging information technologies for means to use the text portion of its 25 year old computerized medical record system. The Center for Intelligent Information Retrieval is developing systems to answer the question: to what extent can automated information systems replace manual chart review of encounter notes? INQUERY, a probabilistic inference net information retrieval system, and FIGLEAF, an inductive decision tree text classifier are applied to the problem of classifying electronic encounter notes to identify acute exacerbations in pediatric asthmatics. Both systems achieve average precisions of greater than 80%, with a new enhancement to INQUERY's relevance feedback, the top performer. Refinement of the systems and plans for their integration are discussed.

Asthma↗

SRS--an indexing and retrieval tool for flat file data libraries.

SRS (Sequence Retrieval System) is an information indexing and retrieval system designed for libraries with a flat file format such as the EMBL nucleotide sequence databank, the SwissProt protein sequence databank or the Prosite library of protein subsequence consensus patterns. SRS supports the data structure of these libraries by providing special indices for implementing lists of subentities (e.g. feature tables) or hierarchically structured data-fields (e.g. taxonomic classification). A language (ODD) has been designed for the convenient specification of library format and organization, representation of individual data-fields within the system (design of indices) and structuring other data needed during retrieval. This ensures flexibility required for coping with different library formats, which are subject to continuous change. Queries and inspection of retrieved entries can be performed from a user interface with pull-down menus and windows. SRS supports various input and output formats but is particularly well adapted to the GCG programs.

Abstracting and Indexing↗

Possibilities for retrieving temporal information in MEDDOK retrieval system.

This paper describes shortly some aspects of the MEDDOK retrieval language especially the methods for making vertical retrieval. In addition the question is investigated how this retrieval can be done without scanning and searching most of the data pool. Another subject of the paper is the description of the underlying data structure.

Computers↗

Current status of the evaluation of information retrieval.

This is the second in the series of the articles on an application of the systems analytic approach to evaluation of information retrieval (IR). In the previous article a historical overview of IR was presented and existing terminological problems associated with IR were identified and discussed. In the presented article the current status of IR evaluation is summarized, and different evaluation approaches are discussed. The Cranfield evaluation model and the most often used relevance-based measures of recall and precision are explained, and their problems are presented. Possible evaluation alternatives to the Cranfield model are discussed, and the case for a systems analytic approach to IR is summarized.

Abstracting and Indexing↗

Textpresso: an ontology-based information retrieval and extraction system for biological literature.

We have developed Textpresso, a new text-mining system for scientific literature whose capabilities go far beyond those of a simple keyword search engine. Textpresso's two major elements are a collection of the full text of scientific articles split into individual sentences, and the implementation of categories of terms for which a database of articles and individual sentences can be searched. The categories are classes of biological concepts (e.g., gene, allele, cell or cell group, phenotype, etc.) and classes that relate two objects (e.g., association, regulation, etc.) or describe one (e.g., biological process, etc.). Together they form a catalog of types of objects and concepts called an ontology. After this ontology is populated with terms, the whole corpus of articles and abstracts is marked up to identify terms of these categories. The current ontology comprises 33 categories of terms. A search engine enables the user to search for one or a combination of these tags and/or keywords within a sentence or document, and as the ontology allows word meaning to be queried, it is possible to formulate semantic queries. Full text access increases recall of biological data types from 45% to 95%. Extraction of particular biological facts, such as gene-gene interactions, can be accelerated significantly by ontologies, with Textpresso automatically performing nearly as well as expert curators to identify sentences; in searches for two uniquely named genes and an interaction term, the ontology confers a 3-fold increase of search efficiency. Textpresso currently focuses on Caenorhabditis elegans literature, with 3,800 full text articles and 16,000 abstracts. The lexicon of the ontology contains 14,500 entries, each of which includes all versions of a specific word or phrase, and it includes all categories of the Gene Ontology database. Textpresso is a useful curation tool, as well as search engine for researchers, and can readily be extended to other organism-specific corpora of text. Textpresso can be accessed at http://www.textpresso.org or via WormBase at http://www.wormbase.org.

Abstracting and Indexing↗

Integrating query of relational and textual data in clinical databases: a case study.

OBJECTIVES: The authors designed and implemented a clinical data mart composed of an integrated information retrieval (IR) and relational database management system (RDBMS). DESIGN: Using commodity software, which supports interactive, attribute-centric text and relational searches, the mart houses 2.8 million documents that span a five-year period and supports basic IR features such as Boolean searches, stemming, and proximity and fuzzy searching. MEASUREMENTS: Results are relevance-ranked using either "total documents per patient" or "report type weighting." RESULTS: Non-curated medical text has a significant degree of malformation with respect to spelling and punctuation, which creates difficulties for text indexing and searching. Presently, the IR facilities of RDBMS packages lack the features necessary to handle such malformed text adequately. CONCLUSION: A robust IR+RDBMS system can be developed, but it requires integrating RDBMSs with third-party IR software. RDBMS vendors need to make their IR offerings more accessible to non-programmers.

Abstracting and Indexing↗

Information retrieval for pathology information systems.

In a medical information system there is a serious need for assessing patient information in various ways. There is also a need for summarizing information in simple statistical tabular form. In this part we present and evaluate different techniques for query formulation, such as: QBE (Query-by-Example), SQL (Structured Query Language), GCL (Graphical Query Language). We also present a query processor for use in a pathology information system. This software system incorporates: a) A semi graphical tree structure interface (SGTSI) for query formulation, and b) Summary tables handling.

Adult↗

Full-text document storage and retrieval in a clinical information system.

The overall design of the CIS at CPMC is heavily influenced by the decision support component. The type of automated decision support being implemented dictates the need for highly structured or coded data. The value of decision support systems has been well documented. The current reliance on free-text documents is natural and a rewarding first step to a more valuable mix of coded and free text. While the health care provider might find the textual comments of the various reports extremely useful, the capability of an automated system to vigilantly review every data element for trends and anomalies is becoming invaluable in today's ever more complex health care delivery environment. Other approaches such as optical imaging systems would facilitate human decision support, but do not supply data in a format that can be processed by automated decision support systems. The developers of the CIS at CPMC believe that data are most valuable when available for both human and automated decision support.

Clinical Medicine↗

[Development of a provisional information-retrieval descriptor language for "Roentgenology and Medical Radiology" for use in the Medinform system].

A problem information retrieval thesaurus consisting of descriptor and nondescriptor articles in alphabetic order was developed. It was conformed to and conjugated with the thesaurus in the field. Examples of descriptor and nondescriptor articles of the revised problem thesaurus were cited. The entire medical roentgenoradiology in the thesaurus in the field was represented by 20 descriptors. The problem thesaurus comprised approximately 1000 descriptors. The authors emphasized the role of the information retrieval thesaurus in which descriptors were renewed from an array of quasidescriptors every 3-5 yrs. The problem thesaurus is an element of the information retrieval language of the Medinform system intended for information support of specialists in roentgenoradiology.

Information Systems↗

The Arabidopsis Information Resource (TAIR): a comprehensive database and web-based information retrieval, analysis, and visualization system for a model plant.

Arabidopsis thaliana, a small annual plant belonging to the mustard family, is the subject of study by an estimated 7000 researchers around the world. In addition to the large body of genetic, physiological and biochemical data gathered for this plant, it will be the first higher plant genome to be completely sequenced, with completion expected at the end of the year 2000. The sequencing effort has been coordinated by an international collaboration, the Arabidopsis Genome Initiative (AGI). The rationale for intensive investigation of Arabidopsis is that it is an excellent model for higher plants. In order to maximize use of the knowledge gained about this plant, there is a need for a comprehensive database and information retrieval and analysis system that will provide user-friendly access to Arabidopsis information. This paper describes the initial steps we have taken toward realizing these goals in a project called The Arabidopsis Information Resource (TAIR) (www.arabidopsis.org).

Arabidopsis↗

BioSYNTHESIS: bridging the information gap.

BioSYNTHESIS is a prototype intelligent retrieval system under development as part of the IAIMS project at Georgetown University. The aim is to create an integrated system that can retrieve information located on disparate computer systems. The project work has been divided in two phases: BioSYNTHESIS I, development of a single menu to access various databases which reside on different computers; and BioSYNTHESIS II, development of a search component that facilitates complex searching for the user. BioSYNTHESIS II will accept a user's query and conduct a search for appropriate information in the IAIMS databases at Georgetown. For information not available at Georgetown, such as full text, it will access selected remote systems and translate the search query as appropriate for the target system. The search through various computer systems and different databases with unique storage and retrieval structures will be transparent to the user. BioSYNTHESIS I is complete and available to users. The design work for BioSYNTHESIS II is under development and will continue as a multiyear technical research effort of the proposed Georgetown IAIMS implementation project.

Computer Communication Networks↗

The design and structure of clinical research information systems. Implications for data retrieval and statistical analyses.

Data management software designed to support clinical data bases typically provides the user with the ability to "enter" and "retrieve" information according to simple user-specified criteria. In the medical research environment, such data base management systems can be self-limiting unless the user has carefully structured the data base schema to be consistent with subsequent statistical procedures used for analysis. For statistical purposes, the data base schema must be configured such that the dependent and independent variables are structurally situated to facilitate the use of statistical application programs. Furthermore, the analysis of time-oriented, prospective studies often requires the data base to be "relational." This may be inconsistent with data collection procedures that result in "hierarchical" schemata. Methodology for ensuring compatibility between the data base schema and subsequent statistical analyses is presented using examples derived from a multicenter clinical trial of diabetes and an observational data bank approach to disease surveillance in rheumatology.

Computers↗

The granularity of medical narratives and its effect on the speed and completeness of information retrieval.

OBJECTIVE: Using electronic rather than paper-based record systems improves clinicians' information retrieval from patient narratives. However, few studies address how data should be organized for this purpose. Information retrieval from clinical narratives containing free text involves two steps: searching for a labeled segment and reading its content. The authors hypothesized that physicians can retrieve information better when clinical narratives are divided into many small, labeled segments ("high granularity"). DESIGN: The study tested the ability of 24 internists and 12 residents at a teaching hospital to retrieve information from an electronic medical record--in terms of speed and completeness--when using different granularities of clinical narratives. Participants solved, without time pressure, predefined problems concerning three voluminous, inpatient case records. To mitigate confounding factors, participants were randomly allocated to a sequence that was balanced by patient case and learning effect. RESULTS: Compared with retrieval from undivided notes, information retrieval from problem-partitioned notes was 22 percent faster (statistically significant), whereas retrieval from notes divided into organ systems was only 11 percent faster (not statistically significant). Subdividing segments beyond organ systems was 13 percent slower (statistically significant) than not subdividing. Granularity of medical narratives affected the speed but not the completeness of information retrieval. CONCLUSION: Dividing voluminous free-text clinical narratives into labeled segments makes patient-related information retrieval easier. However, too much subdivision slows retrieval. Study results suggest that a coarser granularity is required for optimal information retrieval than for structured data entry. Validation of these conclusions in real-life clinical practice is recommended.

Cross-Over Studies↗