Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “information extraction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Observing interactions of abusive and nonabusive dyads: information extracted by accurate and inaccurate judges.

This study describes the reported observations of protective services workers who varied in the accuracy with which they could discriminate abusive from nonabusive mother-child interactions. Participants viewed videotapes of two abusive and two nonabusive dyads. They noted behaviors that suggested an abusive history and behaviors that suggested a nonabusive history for each dyad. Comparisons of more and less accurate judges indicated that accurate judges observed more evidence for abuse when the dyad actually was abusive. The informational advantage of experts was not limited to detecting interactional difficulties in abusive dyads; they also reported observations of more favorable behaviors of nonabusive dyads. Highly accurate judges were more likely to observe communication patterns, task-oriented behavior, and other-directed behaviors than were less accurate judges. These findings suggest that specific information-processing factors may vary with clinical expertise.

Child↗

Implementation and evaluation of a negation tagger in a pipeline-based system for information extract from pathology reports.

We have developed a pipeline-based system for automated annotation of Surgical Pathology Reports with UMLS terms that builds on GATE--an open-source architecture for language engineering. The system includes a module for detecting and annotating negated concepts, which implements the NegEx algorithm--an algorithm originally described for use in discharge summaries and radiology reports. We describe the implementation of the system, and early evaluation of the Negation Tagger. Our results are encouraging. In the key Final Diagnosis section, with almost no modification of the algorithm or phrase lists, the system performs with precision of 0.84 and recall of 0.80 against a gold-standard corpus of negation annotations, created by modified Delphi technique by a panel of pathologists. Further work will focus on refining the Negation Tagger and UMLS Tagger and adding additional processing resources for annotating free-text pathology reports.

Algorithms↗

Information extraction from sound for medical telemonitoring.

Today, the growth of the aging population in Europe needs an increasing number of health care professionals and facilities for aged persons. Medical telemonitoring at home (and, more generally, telemedicine) improves the patient's comfort and reduces hospitalization costs. Using sound surveillance as an alternative solution to video telemonitoring, this paper deals with the detection and classification of alarming sounds in a noisy environment. The proposed sound analysis system can detect distress or everyday sounds everywhere in the monitored apartment, and is connected to classical medical telemonitoring sensors through a data fusion process. The sound analysis system is divided in two stages: sound detection and classification. The first analysis stage (sound detection) must extract significant sounds from a continuous signal flow. A new detection algorithm based on discrete wavelet transform is proposed in this paper, which leads to accurate results when applied to nonstationary signals (such as impulsive sounds). The algorithm presented in this paper was evaluated in a noisy environment and is favorably compared to the state of the art algorithms in the field. The second stage of the system is sound classification, which uses a statistical approach to identify unknown sounds. A statistical study was done to find out the most discriminant acoustical parameters in the input of the classification module. New wavelet based parameters, better adapted to noise, are proposed in this paper. The telemonitoring system validation is presented through various real and simulated test sets. The global sound based system leads to a 3% missed alarm rate and could be fused with other medical sensors to improve performance.

Activities of Daily Living↗

Optimal frequency ranges for extracting information on autonomic activity from the heart rate spectrogram.

Heart rate variability spectrum analysis provides useful quantitative indices of neural control of the SA node. This method is attractive both for its simplicity and the lack of invasive instrumentation, particularly for human investigation. The differing spectral characteristics of parasympathetic and sympathetic control of heart rate allows separate measurement. However, there are widely varying opinions as to the appropriate frequency bands to represent these two inputs. We compared the heart rate variability spectra of 16 humans in supine and upright positions. Adequate measures of parasympathetic or sympathetic activity change should correlate respectively inversely or directly with heart rate change. Frequently used spectral measures of sympathetic activation did not correlate with heart rate changes. With optimization of frequency bands, we found that restricting the sympathetic band to frequencies below 0.1 Hz and above 0.05 Hz (0.055 to either 0.086-0.098 Hz), and dividing by total spectral amplitude 0.004-0.5 Hz (to account for parasympathetic fluctuations within the sympathetic band) produced the best results. The parasympathetic band was best from 0.1 Hz to a frequency greater than that of the respiratory sinus arrhythmia. The optimization method detailed here is easily applied to circumstances other than active orthostasis, and should provide a means of empirically determining useful frequency limits.

Adolescent↗

Extracting information on pneumonia in infants using natural language processing of radiology reports.

Natural language processing (NLP) is critical for improvement of the healthcare process because it can encode clinical data in patient documents. Many clinical applications such as decision support require coded data to function appropriately. However, in order to be applicable for healthcare, performance must be adequate. A valuable automated application is the detection of infectious diseases, such as surveillance of pneumonia in newborns (e.g., neonates) because the disease produces significant rates of morbidity and mortality, and manual surveillance is challenging. Studies have demonstrated that automated surveillance using NLP is a useful adjunct to manual surveillance and an effective tool for infection control practitioners. This paper presents a study evaluating the feasibility of an NLP-based monitoring system to screen for healthcare-associated pneumonia in neonates. We estimated sensitivity, specificity, and positive predictive value by comparing results with clinicians' judgments. Sensitivity was 71% and specificity was 99%. Our results demonstrated that the automated method was feasible.

Database Management Systems↗

Extracting information on folding from the amino acid sequence: accurate predictions for protein regions with preferred conformation in the absence of tertiary interactions.

A recently developed procedure to predict backbone structure from the amino acid sequence [Rooman, M., Kocher, J. P., & Wodak, S. (1991) J. Mol. Biol, 221, 961-979] is fine tuned to identify protein segments, of length 5-15 residues, that adopt well-defined conformations in the absence of tertiary interactions. These segments are obtained by requiring that their predicted lowest energy structures have a sizable energy gap relative to other computed conformations. Applying this procedure to 69 proteins of known structure, we find that regions with largest energy gaps--those having highly preferred conformations--are also the most accurately predicted ones. On the basis of previous findings that such regions correlate well with sites that become structured early during folding, our approach provides the means of identifying such sites in proteins without prior knowledge of the tertiary structure. Furthermore, when predictions are performed so as to ignore the influence of residues flanking each segment along the sequence, a situation akin to excising the considered peptide from the rest of the chain, they offer the possibility of identifying protein segments liable to adopt well-defined conformations on their own. The described approach should have useful applications in experimental and theoretical investigations of protein folding and stability, and aid in designing peptide drugs and vaccines.

Amino Acid Sequence↗

Extracting information on folding from the amino acid sequence: consensus regions with preferred conformation in homologous proteins.

It is investigated whether protein segments predicted to have a well-defined conformational preference in the absence of tertiary interactions are conserved in families of homologous proteins. The prediction method follows the procedures of Rooman, M., Kocher, J.-P., and Wodak, S. (preceding paper in this issue). It uses a knowledge-based force field that incorporates only local interactions along the sequence and identifies segments whose lowest energy structure displays a sizable energy gap relative to other computed conformations. In 13 of the protein families and subfamilies considered that are sufficiently homologous to have similar 3D structures, at least one region is consistently predicted as having the same preferred conformation in virtually all family members. These regions are between 4 and 26 residues long. They are often located at chain ends and correspond primarily to segments of secondary structure heavily involved in interactions with the rest of the protein, suggesting that they could act as nuclei around which other parts of the structure would assemble. Experimental data on early folding intermediates or on protein fragments with appreciable structure in aqueous solution are available for more than half of the protein families. Comparison of our results with these data is quite favorable. They reveal that each of the experimentally identified early formed, or independently stable, substructures harbors at least one of the segments consistently predicted as having a preferred conformation by our procedure. The implications of our findings for the conservation of folding pathways in homologous proteins are discussed.

Adenylate Kinase↗

Extracting information from the power spectrum of synaptic noise.

In cortical neurons, synaptic "noise" is caused by the nearly random release of thousands of synapses. Few methods are presently available to analyze synaptic noise and deduce properties of the underlying synaptic inputs. We focus here on the power spectral density (PSD) of several models of synaptic noise. We examine different classes of analytically solvable kinetic models for synaptic currents, such as the "delta kinetic models," which use Dirac delta functions to represent the activation of the ion channel. We first show that, for this class of kinetic models, one can obtain an analytic expression for the PSD of the total synaptic conductance and derive equivalent stochastic models with only a few variables. This yields a method for constraining models of synaptic currents by analyzing voltage-clamp recordings of synaptic noise. Second, we show that a similar approach can be followed for the PSD of the the membrane potential (Vm) through an effective-leak approximation. Third, we show that this approach is also valid for inputs distributed in dendrites. In this case, the frequency scaling of the Vm PSD is preserved, suggesting that this approach may be applied to intracellular recordings of real neurons. In conclusion, using simple mathematical tools, we show that Vm recordings can be used to constrain kinetic models of synaptic currents, as well as to estimate equivalent stochastic models. This approach, therefore, provides a direct link between intracellular recordings in vivo and the design of models consistent with the dynamics and spectral structure of synaptic noise.

Animals↗

Serial processing is consistent with the time course of linguistic information extraction from consecutive words during eye fixations in reading: a response to Inhoff, Eiter, and Radach (2005).

A. W. Inhoff, B. M. Eiter, and R. Radach reported the results of 2 experiments that they claimed were problematic for serial attention models of eye movements in reading (such as the E-Z Reader model). In this reply, the authors demonstrate via argumentation and simulations that their data pose no serious problem for the E-Z Reader model or serial attention models in general.

Attention↗

Extracting information from cDNA arrays.

High-density DNA arrays allow measurements of gene expression levels (messenger RNA abundance) for thousands of genes simultaneously. We analyze arrays with spotted cDNA used in monitoring of expression profiles. A dilution series of a mouse liver probe is deployed to quantify the reproducibility of expression measurements. Saturation effects limit the accessible signal range at high intensities. Additive noise and outshining from neighboring spots dominate at low intensities. For repeated measurements on the same filter and filter-to-filter comparisons correlation coefficients of 0.98 are found. Next we consider the clustering of gene expression time series from stimulated human fibroblasts which aims at finding co-regulated genes. We analyze how preprocessing, the distance measure, and the clustering algorithm affect the resulting clusters. Finally we discuss algorithms for the identification of transcription factor binding sites from clusters of co-regulated genes. (c) 2001 American Institute of Physics.

Journal Article↗

Source extraction information from air quality data monitored in an Argentinean steel mill.

A statistical analysis of a series of ambient air concentrations of suspended particulate matter (SPM) and NO2 is presented. Measurements were taken at four sites that belong to an Argentinean steel mill and in another site located in its vicinity. The air pollutants were measured during a three-week exploratory sampling. The monitoring sites were selected on the basis of relevant characteristics of the emission sources and the corresponding climatological statistics of the last decade. Suspended particulate matter with aerodynamic diameter of less than 10 microm (PM10) and NO2 were continuously measured at only one site, while 1-hr samples of NO2 and 24-hr samples of total SPM and SO2 were collected at the other sites. The registered concentrations show that SPM was the pollutant of major concern. A first estimate about the nature of the contribution of the different sources of particles and NO2 present in the area was obtained through the statistical analysis of measured concentration data coupled with prevalent meteorological variables.

Air Pollutants, Occupational↗