Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “information extraction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

A reliability study for evaluating information extraction from radiology reports.

GOAL: To assess the reliability of a reference standard for an information extraction task. SETTING: Twenty-four physician raters from two sites and two specialties judged whether clinical conditions were present based on reading chest radiograph reports. METHODS: Variance components, generalizability (reliability) coefficients, and the number of expert raters needed to generate a reliable reference standard were estimated. RESULTS: Per-rater reliability averaged across conditions was 0.80 (95% CI, 0.79-0.81). Reliability for the nine individual conditions varied from 0.67 to 0.97, with central line presence and pneumothorax the most reliable, and pleural effusion (excluding CHF) and pneumonia the least reliable. One to two raters were needed to achieve a reliability of 0.70, and six raters, on average, were required to achieve a reliability of 0.95. This was far more reliable than a previously published per-rater reliability of 0.19 for a more complex task. Differences between sites were attributable to changes to the condition definitions. CONCLUSION: In these evaluations, physician raters were able to judge very reliably the presence of clinical conditions based on text reports. Once the reliability of a specific rater is confirmed, it would be possible for that rater to create a reference standard reliable enough to assess aggregate measures on a system. Six raters would be needed to create a reference standard sufficient to assess a system on a case-by-case basis. These results should help evaluators design future information extraction studies for natural language processors and other knowledge-based systems.

Evaluation Studies as Topic↗

A knowledge-based approach to information extraction from surgical pathology reports.

We describe the development of a prototype system for knowledge-based information extraction from surgical pathology reports. The current system includes abstract problem solving methods and a frame-based knowledge representation of body parts, procedures, diseases, and findings for prostate and breast cases. The system currently extracts the organ, procedure, and diagnoses, and sets an agenda of goals for further processing. A potential advantage of this approach is the ability to increase specificity of information extraction.

Artificial Intelligence↗

Information extraction from full text scientific articles: where are the keywords?

BACKGROUND: To date, many of the methods for information extraction of biological information from scientific articles are restricted to the abstract of the article. However, full text articles in electronic version, which offer larger sources of data, are currently available. Several questions arise as to whether the effort of scanning full text articles is worthy, or whether the information that can be extracted from the different sections of an article can be relevant. RESULTS: In this work we addressed those questions showing that the keyword content of the different sections of a standard scientific article (abstract, introduction, methods, results, and discussion) is very heterogeneous. CONCLUSIONS: Although the abstract contains the best ratio of keywords per total of words, other sections of the article may be a better source of biologically relevant data.

Anatomy↗

Information extraction for enhanced access to disease outbreak reports.

Document search is generally based on individual terms in the document. However, for collections within limited domains it is possible to provide more powerful access tools. This paper describes a system designed for collections of reports of infectious disease outbreaks. The system, Proteus-BIO, automatically creates a table of outbreaks, with each table entry linked to the document describing that outbreak; this makes it possible to use database operations such as selection and sorting to find relevant documents. Proteus-BIO consists of a Web crawler which gathers relevant documents; an information extraction engine which converts the individual outbreak events to a tabular database; and a database browser which provides access to the events and, through them, to the documents. The information extraction engine uses sets of patterns and word classes to extract the information about each event. Preparing these patterns and word classes has been a time-consuming manual operation in the past, but automated discovery tools now make this task significantly easier. A small study comparing the effectiveness of the tabular index with conventional Web search tools demonstrated that users can find substantially more documents in a given time period with Proteus-BIO.

Abstracting and Indexing↗

Radio frequency photonic filter for highly resolved and ultrafast information extraction.

We present a method and devices for highly resolved carrier and information extraction of optically modulated radar signals. The extraction is done by passing the optical beam through a monitoring path that constitutes a finite impulse response filter. Replications of the monitoring signal realize the required spectral scan of the filter. Despite the fact that the filter configuration is fixed, each replication experiences different spectral filtering. The radar carrier is detected by observing the energy fluctuations in a low-rate output detector. The RF information is extracted by positioning a low-rate tunable filter at the detected carrier frequency.

Journal Article↗

Constructing biological knowledge bases by extracting information from text sources.

Recently, there has been much effort in making databases for molecular biology more accessible and interoperable. However, information in text form, such as MEDLINE records, remains a greatly underutilized source of biological information. We have begun a research effort aimed at automatically mapping information from text sources into structured representations, such as knowledge bases. Our approach to this task is to use machine-learning methods to induce routines for extracting facts from text. We describe two learning methods that we have applied to this task--a statistical text classification method, and a relational learning method--and our initial experiments in learning such information-extraction routines. We also present an approach to decreasing the cost of learning information-extraction routines by learning from "weakly" labeled training data.

Artificial Intelligence↗

Toward information extraction: identifying protein names from biological papers.

To solve the mystery of the life phenomenon, we must clarify when genes are expressed and how their products interact with each other. But since the amount of continuously updated knowledge on these interactions is massive and is only available in the form of published articles, an intelligent information extraction (IE) system is needed. To extract these information directly from articles, the system must firstly identify the material names. However, medical and biological documents often include proper nouns newly made by the authors, and conventional methods based on domain specific dictionaries cannot detect such unknown words or coinages. In this study, we propose a new method of extracting material names, PROPER, using surface clue on character strings. It extracts material names in the sentence with 94.70% precision and 98.84% recall, regardless of whether it is already known or newly defined.

Information Storage and Retrieval↗

How can information extraction ease formalizing treatment processes in clinical practice guidelines? A method and its evaluation.

OBJECTIVE: Formalizing clinical practice guidelines (CPGs) for a subsequent computer-supported processing is a challenging, but burdensome and time-consuming task. Existing methods and tools to support this task demand detailed medical knowledge, knowledge about the formal representations, and a manual modeling. Furthermore, formalized guideline documents mostly fall far short in terms of readability and understandability for the human domain modeler. METHODS AND MATERIAL: We propose a new multi-step approach using information extraction methods to support the human modeler by both automating parts of the modeling process and making the modeling process traceable and comprehensible. This paper addresses the first steps to obtain a representation containing processes which is independent of the final guideline representation language. RESULTS: We have developed and evaluated several heuristics without the need to apply natural language understanding and implemented them in a framework to apply them to several guidelines from the medical subject of otolaryngology. Findings in the evaluation indicate that using semi-automatic, step-wise information extraction methods are a valuable instrument to formalize CPGs. CONCLUSION: Our evaluation shows that a heuristic-based approach can achieve good results, especially for guidelines with a major portion of semi-structured text. It can be applied to guidelines irrespective to the final guideline representation format.

Artificial Intelligence↗

Picture perception: effects of luminance on available information and information-extraction rate.

In each of four experiments, complex visual stimuli--pictures and digit arrays--were remembered better when shown at high luminance than when shown at low luminance. Why does this occur? Two possibilities were considered: first that lowering luminance reduces the amount of available information in the stimulus, and second that lowering luminance reduces the rate at which the information is extracted from the stimulus. Evidence was found for both possibilities. When stimuli were presented at durations short enough to permit only a single eye fixation, luminance affected only the rate at which information is extracted: decreasing luminance by a factor of 100 caused information to be extracted more slowly by a factor that ranged, over experiments, from 1.4 to 2.0. When pictures were presented at durations long enough to permit multiple fixations, however, luminance affected the total amount of extractable information. In a fifth experiment, converging evidence was sought for the proposition that within the first eye fixation on a picture, luminance affects the rate of information extraction. If this proposition is correct and, in addition, the first eye fixation lasts until some criterion amount of information is extracted, then fixation duration should increase with decreasing luminance. This prediction was confirmed.

Attention↗

Information extraction from Korean radiology reports mingled two language.

This study presents overall of Information Extraction (IE) for SNUH (Seoul National University Hospital) radiology reports coexisted Korean and English using Concept Node (CN) which is a case frame as extraction rule. The following steps are performed: design conceptual model by terminology exploration based on lexical analysis, create a CN definition based on syntactic relationship pattern and implement automatic IE system using CN. Main purposes is to investigate whether syntactic and semantic analysis technique using extraction rule (CN) is effective for typical Korean medical text in mixed two different languages.

Humans↗

A nonparametric approach to extract information from interspike interval data.

In this work we develop an approach to extracting information from neural spike trains. Using the expectation-maximization (EM) algorithm, interspike interval data from experiments and simulations are fitted by mixtures of distributions, including Gamma, inverse Gaussian, log-normal, and the distribution of the interspike intervals of the leaky integrate-and-fire model. In terms of the Kolmogorov-Smirnov test for goodness-of-fit, our approach is proved successful (P>0.05) in fitting benchmark data for which a classical parametric approach has been shown to fail before. In addition, we present a novel method to fit mixture models to censored data, and discuss two examples of the application of such a method, which correspond to the case of multiple-trial and multielectrode array data. A MATLAB implementation of the algorithm is available for download from .

Action Potentials↗

BioIE: extracting informative sentences from the biomedical literature.

SUMMARY: BioIE is a rule-based system that extracts informative sentences relating to protein families, their structures, functions and diseases from the biomedical literaturE. Based on manual definition of templates and rules, it aims at precise sentence extraction rather than wide recall. After uploading source text or retrieving abstracts from MEDLINE, users can extract sentences based on predefined or user-defined template categories. BioIE also provides a brief insight into the syntactic and semantic context of the source-text by looking at word, N-gram and MeSH-term distributions. Important Applications of BioIE are in, for example, annotation of microarray data and of protein databases. AVAILABILITY: http://umber.sbs.man.ac.uk/dbbrowser/bioie/

Database Management Systems↗

Overview of BioCreAtIvE: critical assessment of information extraction for biology.

BACKGROUND: The goal of the first BioCreAtIvE challenge (Critical Assessment of Information Extraction in Biology) was to provide a set of common evaluation tasks to assess the state of the art for text mining applied to biological problems. The results were presented in a workshop held in Granada, Spain March 28-31, 2004. The articles collected in this BMC Bioinformatics supplement entitled "A critical assessment of text mining methods in molecular biology" describe the BioCreAtIvE tasks, systems, results and their independent evaluation. RESULTS: BioCreAtIvE focused on two tasks. The first dealt with extraction of gene or protein names from text, and their mapping into standardized gene identifiers for three model organism databases (fly, mouse, yeast). The second task addressed issues of functional annotation, requiring systems to identify specific text passages that supported Gene Ontology annotations for specific proteins, given full text articles. CONCLUSION: The first BioCreAtIvE assessment achieved a high level of international participation (27 groups from 10 countries). The assessment provided state-of-the-art performance results for a basic task (gene name finding and normalization), where the best systems achieved a balanced 80% precision / recall or better, which potentially makes them suitable for real applications in biology. The results for the advanced task (functional annotation from free text) were significantly lower, demonstrating the current limitations of text-mining approaches where knowledge extrapolation and interpretation are required. In addition, an important contribution of BioCreAtIvE has been the creation and release of training and test data sets for both tasks. There are 22 articles in this special issue, including six that provide analyses of results or data quality for the data sets, including a novel inter-annotator consistency assessment for the test set used in task 2.

Computational Biology↗

Synthetic Analysis for Extracting Information on Soil Salinity Using Remote Sensing and GIS: A Case Study of Yanggao Basin in China

/ This paper reports the experience of extracting information on the salinity of soil and offers a method of synthetic analysis. The experimental areas for analysis are located in Yanggao Basin, Shanxi Province, China. The types of soil are mainly meadow soil and salinized meadow soil. The method of synthetic analysis of salinity uses a geographic information system (GIS) as a tool, building a basic saltwater analysis model of saline soil and adjusting the result with expert experience after computer processing. The method of feature extraction has been used for remotely sensed data. An optimum combination of features has been determined and, after comparing several combinations in the Yanggao region, an improved result has been obtained after Kauth-Thomas (K-T) transformation. For precise quantitative analysis of the salinization, not only Thematic Mapper (TM) remote sensing data, but also two forms of non-remote-sensing data are needed: depth of groundwater and mineralization rate of groundwater according to the theory of genesis of soil. For the analysis of synthetic compounded multisources, a generalized Bayes classification is used after overlay, matching, and related coefficients have been determined. On the premise that various information sources are independent, global membership functions with probability are used to combine various pieces of information in order to apply them directly to the pixels and classifications of soil salinity. The experiment indicates that this analytical method is sound because of the increased speed of processing and its simplicity and improved precision of classification of salinity. Finally, it is necessary to examine and adjust the factors using expert intelligence. The experiment shows that synthetic analysis using the geographic information system can raise the precision of quantitative analysis of salinity, which has advantages for environmental monitoring and management.KEY WORDS: Salinity; Remote sensing; Thematic Mapper; Geographic information system; Classification

Journal Article↗

Event-related brain potentials as indices of information extraction and response priming.

Measures of overt response and of the event-related brain potential (ERP) were used to investigate the processing of a priming stimulus varying in its information content. Subjects were shown sequences of 2 letters that served as a priming and an imperative stimulus. In 3 randomly interspersed conditions the imperative stimulus had a 0.80, 0.50, or 0.20 probability of physically matching the priming letter. The different probability conditions were signaled by the position of a dot flanking the priming letter. Reaction time and accuracy data indicated that the subjects primed their responses as a function of the information conveyed by the priming stimulus. The amplitude and latency of the P300 to the priming stimulus were sensitive to the amount of information conveyed by the priming stimulus and the duration of the processing required. The readiness potential in the foreperiod was lateralized as a function of the priming stimulus. Furthermore, the larger the amplitude of the P300 to the priming stimulus, the larger the lateralization of the readiness potential, indicating that information extraction, indexed by the P300, was related to response priming, indexed by the readiness potential. The results indicate that ERP measures make manifest covert aspects of the priming process occurring in the foreperiod.

Adolescent↗

Zone analysis in biology articles as a basis for information extraction.

In the field of biomedicine, an overwhelming amount of experimental data has become available as a result of the high throughput of research in this domain. The amount of results reported has now grown beyond the limits of what can be managed by manual means. This makes it increasingly difficult for the researchers in this area to keep up with the latest developments. Information extraction (IE) in the biological domain aims to provide an effective automatic means to dynamically manage the information contained in archived journal articles and abstract collections and thus help researchers in their work. However, while considerable advances have been made in certain areas of IE, pinpointing and organizing factual information (such as experimental results) remains a challenge. In this paper we propose tackling this task by incorporating into IE information about rhetorical zones, i.e. classification of spans of text in terms of argumentation and intellectual attribution. As the first step towards this goal, we introduce a scheme for annotating biological texts for rhetorical zones and provide a qualitative and quantitative analysis of the data annotated according to this scheme. We also discuss our preliminary research on automatic zone analysis, and its incorporation into our IE framework.

Abstracting and Indexing↗

Bioinformatics methods for the analysis of expression arrays: data clustering and information extraction.

Expression arrays facilitate the monitoring of changes in the expression patterns of large collections of genes. The analysis of expression array data has become a computationally-intensive task that requires the development of bioinformatics technology for a number of key stages in the process, such as image analysis, database storage, gene clustering and information extraction. Here, we review the current trends in each of these areas, with particular emphasis on the development of the related technology being carried out within our groups.

Abstracting and Indexing↗

Identification of "pathologs" (disease-related genes) from the RIKEN mouse cDNA dataset using human curation plus FACTS, a new biological information extraction system.

BACKGROUND: A major goal in the post-genomic era is to identify and characterise disease susceptibility genes and to apply this knowledge to disease prevention and treatment. Rodents and humans have remarkably similar genomes and share closely related biochemical, physiological and pathological pathways. In this work we utilised the latest information on the mouse transcriptome as revealed by the RIKEN FANTOM2 project to identify novel human disease-related candidate genes. We define a new term "patholog" to mean a homolog of a human disease-related gene encoding a product (transcript, anti-sense or protein) potentially relevant to disease. Rather than just focus on Mendelian inheritance, we applied the analysis to all potential pathologs regardless of their inheritance pattern. RESULTS: Bioinformatic analysis and human curation of 60,770 RIKEN full-length mouse cDNA clones produced 2,578 sequences that showed similarity (70-85% identity) to known human-disease genes. Using a newly developed biological information extraction and annotation tool (FACTS) in parallel with human expert analysis of 17,051 MEDLINE scientific abstracts we identified 182 novel potential pathologs. Of these, 36 were identified by computational tools only, 49 by human expert analysis only and 97 by both methods. These pathologs were related to neoplastic (53%), hereditary (24%), immunological (5%), cardio-vascular (4%), or other (14%), disorders. CONCLUSIONS: Large scale genome projects continue to produce a vast amount of data with potential application to the study of human disease. For this potential to be realised we need intelligent strategies for data categorisation and the ability to link sequence data with relevant literature. This paper demonstrates the power of combining human expert annotation with FACTS, a newly developed bioinformatics tool, to identify novel pathologs from within large-scale mouse transcript datasets.

Animals↗