Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “information extraction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Protein structures and information extraction from biological texts: the PASTA system.

MOTIVATION: The rapid increase in volume of protein structure literature means useful information may be hidden or lost in the published literature and the process of finding relevant material, sometimes the rate-determining factor in new research, may be arduous and slow. RESULTS: We describe the Protein Active Site Template Acquisition (PASTA) system, which addresses these problems by performing automatic extraction of information relating to the roles of specific amino acid residues in protein molecules from online scientific articles and abstracts. Both the terminology recognition and extraction capabilities of the system have been extensively evaluated against manually annotated data and the results compare favourably with state-of-the-art results obtained in less challenging domains. PASTA is the first information extraction (IE) system developed for the protein structure domain and one of the most thoroughly evaluated IE system operating on biological scientific text to date. AVAILABILITY: PASTA makes its extraction results available via a browser-based front end: http://www.dcs.shef.ac.uk/nlp/pasta/. The evaluation resources (manually annotated corpora) are also available through the website: http://www.dcs.shef.ac.uk/nlp/pasta/results.html.

Abstracting and Indexing↗

Bioie: retargetable information extraction and ontological annotation of biological interactions from the literature.

The need for extracting general biological interactions of arbitrary types from the rapidly growing volume of the biomedical literature is drawing increased attention, while the need for this much diversity also requires both a robust treatment of complex linguistic phenomena and a method to consistently characterize the results. We present a biomedical information extraction system, BioIE, to address both of these needs by utilizing a full-fledged English grammar formalism, or a combinatory categorial grammar, and by annotating the results with the terms of Gene Ontology, which provides a common and controlled vocabulary. BioIE deals with complex linguistic phenomena such as coordination, relative structures, acronyms, appositive structures, and anaphoric expressions. In order to deal with real-world syntactic variations of ontological terms, BioIE utilizes the syntactic dependencies between words in sentences as well, based on the observation that the component words in an ontological term usually appear in a sentence with known patterns of syntactic dependencies.

Abstracting and Indexing↗

Adult age differences in the rate of information extraction during visual search.

We investigated adult age differences in the function relating accuracy to speed during visual-search performance. Twenty-four young adults between 17 and 26 years of age, and 24 older adults between 59 and 75 years of age, participated. The level of accuracy at which performance was first improved by increasing RT was higher for older adults than for young adults. Independently of this age difference in accuracy, however, a significant age-related slowing was present in the rate of information extraction during visual-search performance.

Adolescent↗

Extracting information from two-dimensional electrophoresis gels by partial least squares regression.

Two-dimensional gel electrophoresis (2-DE) produces large amounts of data and extraction of relevant information from these data demands a cautious and time consuming process of spot pattern matching between gels. The classical approach of data analysis is to detect protein markers that appear or disappear depending on the experimental conditions. Such biomarkers are found by comparing the relative volumes of individual spots in the individual gels. Multivariate statistical analysis and modelling of 2-DE data for comparison and classification is an alternative approach utilising the combination of all proteins/spots in the gels. In the present study it is demonstrated how information can be extracted by multivariate data analysis. The strategy is based on partial least squares regression followed by variable selection to find proteins that individually or in combination with other proteins vary informatively in relation to the experimental conditions. Finding of such coherent protein patterns leads to identification of potential relations between the involved proteins, and will be useful for focusing further investigation of proteins that relate to the chosen experimental conditions.

Electrophoresis, Gel, Two-Dimensional↗

Disambiguation data: extracting information from anonymized sources.

Privacy protection is an important consideration when releasing medical databases to the research community. We show that while recent advances in anonymization algorithms provide increased levels of protection, it is still possible to calculate approximations to the original data set. In some cases, one can even uniquely reconstruct entries in a table before anonymization. In this paper, we demonstrate how knowledge of an anonymization algorithm based on ambiguating data cell entries can be used to undo the anonymization process. We investigate the effect of this algorithm and its reversal on data sets of varying sizes and distributions. It is shown that by using a computationally complex disambiguation process, information on individuals can be extracted from an anonymized data set.

Adult↗

Annoyance perception of sound and information extraction.

The judgment of annoyance of distorted speech differs radically for different language groups. The results show that those who do comprehend a spoken language, base their annoyance-judgments on the informational content extracted while those who do not base it on the perceptual characteristics of meaningless sound (particularly loudness). A series of distorted German speech sounds were presented to two subject groups consisting of native Swedish and English speakers, and the results were compared with earlier results from groups of native German and Polish subjects. The 50 stimuli were generated from the very same speech signal distorted in two principle ways, either with repeated silent gaps or superimposed noise impulses. The perceived annoyance of the distorted speech was judged both by category scaling for all subject groups, and as a control for "ceiling" effects, also by magnitude estimation for the Swedish and the English subjects. There is a pronounced tendency for German subjects to judge the German speech distorted with silent gaps as more annoying than that distorted with superimposed noise impulses. In contrast, the Swedish, English, and Polish subjects judged the two German-speech distortions in reversed order with regard to annoyance. Thus for noncomprehending listeners, noise-distorted speech is more annoying but for comprehending listeners it is speech distorted by gaps. This means that impaired communication intrusiveness rather than loudness predominates in annoyance judgments from comprehending listeners.

Adult↗

System architecture for temporal information extraction, representation and reasoning in clinical narrative reports.

Exploring temporal information in narrative Electronic Medical Records (EMRs) is essential and challenging. We propose an architecture for an integrated approach to process temporal information in clinical narrative reports. The goal is to initiate and build a foundation that supports applications which assist healthcare practice and research by including the ability to determine the time of clinical events (e.g., past vs. present). Key components include: (1) an annotation schema for temporal expressions and the development of an associated tagger; (2) a natural language processing (NLP) system for encoding and extracting medical events and associating them with formalized temporal data; (3) a post-processor, with a knowledge-based subsystem to help discover implicit information, that resolves temporal expressions and deals with issues such as granularity and vagueness; and (4) a reasoning mechanism which models clinical reports as Simple Temporal Problems (STPs).

Humans↗

An information extraction and representation system for rapid review of the biomedical literature.

With the rapid expansion of scientific research, the ability to effectively find or integrate new domain knowledge in the sciences is proving increasingly difficult. The development of methods and tools for assisting researchers to effectively ex-tract problem-oriented knowledge from heterogeneous and massive information sources, and for using this knowledge in problem-solving is one of the most fundamental research directions for the information and computer sciences today. There is a need for new tools to support more precise identification of relevant research articles and provide visual clues regarding relationships among the document sets. We present the Telemakus system in which aggregated citation information and extracted research findings are displayed in a schema-based document surrogate and an interactive mapping tool provides graphical displays of research inter-relationships from documents across a domain. This system is an innovative approach to creating useful and precise document surrogates and may re-conceptualize the way we currently represent, retrieve, and assimilate research findings from the published literature.

Biomedical Research↗

Image processing, diagnostic information extraction and quantitative assessment in pathology.

As we enter the information age we hold strong beliefs in the benefits of digital technology applied to pathology: numerical representation offers objectivity. Digital knowledge may indeed lead to significant information discovery, and, processing systems might be designed to allow a true evolution of capabilities. Questions arise whether the methodology underlying quantitative analysis provides the information that we need and whether it is appropriate for some of the problems encountered in diagnostic and prognostic histopathology. While one certainly would not dispute the value of statistical procedures, the clinical needs call for individual patient targeted prognosis.

Diagnostic Imaging↗

Automated extraction of information on protein-protein interactions from the biological literature.

MOTIVATION: To understand biological process, we must clarify how proteins interact with each other. However, since information about protein-protein interactions still exists primarily in the scientific literature, it is not accessible in a computer-readable format. Efficient processing of large amounts of interactions therefore needs an intelligent information extraction method. Our aim is to develop an efficient method for extracting information on protein-protein interaction from scientific literature. RESULTS: We present a method for extracting information on protein-protein interactions from the scientific literature. This method, which employs only a protein name dictionary, surface clues on word patterns and simple part-of-speech rules, achieved high recall and precision rates for yeast (recall = 86.8% and precision = 94.3%) and Escherichia coli (recall = 82.5% and precision = 93.5%). The result of extraction suggests that our method should be applicable to any species for which a protein name dictionary is constructed. AVAILABILITY: The program is available on request from the authors.

Electronic Data Processing↗

Time course of linguistic information extraction from consecutive words during eye fixations in reading.

Sequential attention shift models of reading predict that an attended (typically fixated) word must be recognized before useful linguistic information can be obtained from the following (parafoveal) word. These models also predict that linguistic information is obtained from a parafoveal word immediately prior to a saccade toward it. To test these assumptions, sentences were constructed with a critical pretarget-target word sequence, and the temporal availability of the (parafoveal) target preview was manipulated while the pretarget word was fixated. Target viewing effects, examined as a function of prior target visibility, revealed that extraction of linguistic target information began 70-140 ms after the onset of pretarget viewing. Critically, acquisition of useful linguistic information from a target was not confined to the ending period of pretarget viewing. These results favor theoretical conceptions in which there is some temporal overlap in the linguistic processing of a fixated and parafoveally visible word during reading.

Attention↗

Temporal properties of information extraction in reading studied by a text-mask replacement technique.

A text was replaced with a visual mask for a fixed duration every time the reader made a saccade. The threshold duration was measured when the mask was presented at the beginning of a saccade or a fixation, or at a certain delay after the onset of a fixation. The effect of the mask on reading time, as well as on the subjective legibility of the text, was also investigated. It was shown that saccadic suppression exists in reading; subjects are not affected by a mask as long as it is inserted during the saccade. The visual sensitivity is recovered only partially at the initial part of fixation and is recovered fully approximately 70 msec after the beginning of the fixation. The visual information can be extracted during a later part of the fixation period as efficiently as or even more efficiently than during the early part of the fixation.

Adult↗

Extracting information masked by the chaotic signal of a time-delay system.

We further develop the method proposed by Bezruchko et al. [Phys. Rev. E 64, 056216 (2001)] for the estimation of the parameters of time-delay systems from time series. Using this method we demonstrate a possibility of message extraction for a communication system with nonlinear mixing of information signal and chaotic signal of the time-delay system. The message extraction procedure is illustrated using both numerical and experimental data and different kinds of information signals.

Journal Article↗

[Comparison of 2 possibilities for extracting information in the analysis of (multi-dimensional) contingency tables. Demonstrated with a medical-sociologic example].

In several papers the author has stressed the importance of qualitative characters in medical studies. From this follows the necessity of analysing more-dimensional contingency tables. In the first part of this paper the most important results are repeated. Two procedures may be distinguished: 1. A three-step procedure, called Elementary More-Dimensional Contingency Analysis (EMCTA) with and without model fitting 2. A combined Residual and Contrast Analysis (RA/CA). By the aid of a medical example it is shown, in the second part of the paper, that it is possible to extract rather easily all information about interesting dependencies be the aid of RA/CA in the case of one response variable.

Adult↗

BioRAT: extracting biological information from full-length papers.

MOTIVATION: Converting the vast quantity of free-format text found in journals into a concise, structured format makes the researcher's quest for information easier. Recently, several information extraction systems have been developed that attempt to simplify the retrieval and analysis of biological and medical data. Most of this work has used the abstract alone, owing to the convenience of access and the quality of data. Abstracts are generally available through central collections with easy direct access (e.g. PubMed). The full-text papers contain more information, but are distributed across many locations (e.g. publishers' web sites, journal web sites and local repositories), making access more difficult. In this paper, we present BioRAT, a new information extraction (IE) tool, specifically designed to perform biomedical IE, and which is able to locate and analyse both abstracts and full-length papers. BioRAT is a Biological Research Assistant for Text mining, and incorporates a document search ability with domain-specific IE. RESULTS: We show first, that BioRAT performs as well as existing systems, when applied to abstracts; and second, that significantly more information is available to BioRAT through the full-length papers than via the abstracts alone. Typically, less than half of the available information is extracted from the abstract, with the majority coming from the body of each paper. Overall, BioRAT recalled 20.31% of the target facts from the abstracts with 55.07% precision, and achieved 43.6% recall with 51.25% precision on full-length papers.

Abstracting and Indexing↗