Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Information Retrieval Systems”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Words, concepts, or both: optimal indexing units for automated information retrieval.

What is the best way to represent the content of documents in an information retrieval system? This study compares the retrieval effectiveness of five different methods for automated (machine-assigned) indexing using three test collections. The consistently best methods are those that use indexing based on the words that occur in the available text of each document. Methods used to map text into concepts from a controlled vocabulary showed no advantage over the word-based methods. This study also looked at an approach to relevance feedback which showed benefit for both word-based and concept-based methods.

Abstracting and Indexing↗

Rain and windchill as factors in the occurrence of pneumonia in sheep.

A computerised information retrieval system of abattoir pathology and meteorological data has been used to investigate the effect of prevailing weather conditions on the occurrence of pleurisy and pneumonia in the sheep population of Northern Ireland. Significant correlation coefficients were found between the percentage condemnations due to pleurisy and pneumonia in sheep and rainfall, windspeed, temperature and humidity. The most significant correlation was found with windspeed. The paper describes the calculation of a new meteorological variable, the rain/windchill factor. Very highly significant correlation coefficients were found between the percentage lung condemnations in sheep and the rain/windchill factor prevailing during the same month and both one and two months previously. The paper discusses the practical implications of these findings for sheep production and highlights the desirability of protecting sheep from adverse climatic conditions during the winter months.

Abattoirs↗

A randomized controlled trial of automated term composition.

OBJECTIVE: To compare the ability of an Automated Term Composition (ATC) algorithm with non-compositional mappings to provide coverage (exact mappings to a controlled vocabulary) for a randomly selected set of free text entries which were entered as headings to the Impression section of the clinical notes system at the Mayo Foundation. We also compare the results of four evaluators to determine the inter-observer variability and the variance between term sets, with respect to the accuracy of the mappings and the reliability of the failure analysis. METHODS: From a corpus of approximately 1,000,000 unique terms entered into the Impression/Report/Plan section of the clinical notes system in the calendar year 1997, we randomly selected 1,000 terms. We then further randomized these 1,000 terms into two groups of 500 (Sets A and B). We constructed two copies of the same term matching interface, one without ATC (alpha) and one with ATC (beta). We took four expert Indexers and assigned them to one of the following tasks. The first reviewer (R1) compared set A using the alpha program and then set B using the beta program (R1(Aalpha + Bbeta)). The second compared set A using the alpha program and then set B using the alpha program (R2(A + B) alpha). The third compared set B using the beta program and then set A using the beta program (R3(B + A) beta). The fourth compared set A using the beta program and then set B using the alpha program (R4(Abeta + Balpha)). RESULTS: The program with Automated Term Composition mapped 540 out of the 1,000 Concepts correctly (54.0%). The same program without ATC mapped only 276 out of the 1,000 Concepts correctly (27.6%). Therefore the program with ATC was significantly more effective at matching concepts in our problem lists than the same search engine without ATC (p < 0.0001; McNemar Method). These figures result from the comparison of the alpha program with the beta program by reviewers one and four. Failure analysis showed that with the alpha version 425 out of the 724 mismatches were because a base concept was missing from the retrieval set (58.7%) and 299 mismatches were from missing qualifiers or modifiers or both (41.3%). In the beta version of the program (with ATC) 340 out of the 460 mismatches were secondary to there being a missing base concept in the retrieval set (73.9%) and only 120 mismatches due to missing modifiers and or qualifiers (26.1%). CONCLUSIONS: Automated term composition provided significantly better coverage of a randomly chosen set of patient problems, diagnosed at the Mayo Clinic during the 1997 calendar year, when compared with the same information retrieval system without ATC. We believe that these results speak further to the excellent content coverage provided by the UMLS metathesaurus. These authors believe that increased structure, normalization of UMLS content and semantics, and better tools to make use of the currently available content such as automated term composition, are what is needed to leverage the production of commercially viable tools that provide access to controlled vocabularies for medicine.

Abstracting and Indexing↗

Profile of an inpatient population with a history of illicit drug use.

Using a data retrieval system, information was obtained from the records of all patients discharged over a three-year period with a primary or secondary diagnosis of illicit drug use. From 1987-1989, the number of patients and number of hospital days for patients with such diagnoses increased steadily. Seventy four percent of inpatients identified as users of illicit drugs were less than 35 years of age; this percentage was greatly influenced by the percentage (44-48%) of obstetrical patients in the population. Hospital charges for the major source of payment for the identified group exceeded $5 million in 1989. Medicaid was the major source of payment for the hospital care; the average length of hospital stay for that portion of the patients increased about 1.4 days over the three-year period. The number of hospitalized patients identified as users of illicit drugs increased over the period of study, and the population depended heavily on government sources for payment of inpatient services.

Adult↗

Ozone in the urban southeastern United States.

Ozone measurements (daily maximum values) from the Aerometric Information Retrieval System database are analyzed for selected sites, during 1980 to 1988, in southeastern USA. Frequency distributions, for most sites during most years, show a typical bell-shaped curve with the higher frequency around the yearly daily maximum ozone mean of about 100 to about 110 microg m(-3) (50-55 ppbv). Abnormal years in ozone concentration may skew the distribution as the mean shifts. A correlation of daily maximum ozone concentrations above 140 microg m(-3) (70 ppbv) between sites shows a division between the sites in the northern protion of the region and those in the southern portion of the region. Variations in ozone levels are well correlated over distances of several hundred kilometers, suggesting that high values are associated with synoptic scale episodes. An ozone exposure analysis also shows higher ozone exposures (250-300 ppm days) in the northerly sites as compared to the southerly sites (150-170 ppm days).

Journal Article↗

Cell wall-associated enzymes in fungi.

This review compiles and discusses previous reports on the identity of wall-associated enzymes (WAEs) in fungi and addresses critically the widely different terminologies used in the literature to specify the type of bonding of WAEs to other entities of the cell wall compartment, the extracellular matrix (ECM). A facile and rapid fractionation protocol for catalytically active WAEs is presented, which uses crude cell walls as the experimental material, a variety of test enzymes (including representatives of polysaccharide synthases and hydrolases, phosphatases, gamma-glutamyltransferases, pyridine-nucleotide dehydrogenases and phenol-oxidising enzymes) and a combination of simple hydrophilic and hydrophobic extractants. The protocol provides four fully operationally defined classes of WAEs, with constituent members of each class displaying the same basic type of physicochemical interaction with binding partners in situ. The routine application of the protocol to different species and cell types could yield easily accessible data useful for building-up a general objective information retrieval system of WAEs, suitable as an heuristic basis both for the unravelling of the role and for the biotechnological potentialities of WAEs. A detailed account is given of the function played in the ECM by WAEs in the metabolism of chitin (chitin synthase, chitinase and beta-N-acetylhexosaminidase) and of phenols (tyrosinase).

Cell Wall↗

Nursing informatics. Issues for critical care medicine.

The demands of today's health care arena have forced the issue of automation and computerization. Nursing, as the major stakeholder in the collecting, managing, processing, transforming, and communicating of information regarding the patient, has developed a new approach to these tasks. Nursing informatics, which is the application of computer science and information science, is being used to manage and process the data, information, and knowledge necessary in the discipline. Although still in its infancy, nursing informatics has started to have a major effect on health care information gathering and clinical practice despite the multiple barriers to its advancement. Critical care is a data-rich environmental that can benefit from better management and processing of the data derived from the critically ill patient. Nursing and medical informatics joining together to organize the data, coupled with the introduction of good DSS and the addition of information retrieval systems at the bedside and the on-line medical record, will have a positive effect on the critical care environment and on the critical care patient outcomes.

Critical Care↗

An exploratory look at hydrocarbon data from the Photochemical Assessment Monitoring Stations network.

This paper describes some characteristics of speciated nonmethane organic compound (NMOC) data collected in 1994 at five Photochemical Assessment Monitoring Stations (PAMS) and archived in the U.S. Environmental Protection Agency's Aerometric Information Retrieval System (AIRS). Topics include data completeness, distribution of individual NMOCs in concentration categories relative to minimum detectable levels, percentage of total NMOC associated with the sum of the 55 PAMS target compounds, and use of scatterplots to diagnose chromatographic misidentification of compounds. This is an early examination of a database that is expanding rapidly, and the insights presented here may be useful to both the producers and future users of the data for establishing consistency and quality control.

Air Pollutants↗

Defining the photochemical contribution to particulate matter in urban areas using time-series analysis.

The objective of this project is to demonstrate how the ambient air measurement record can be used to define the relationship between O3 (as a surrogate for photochemistry) and secondary particulate matter (PM) in urban air. The approach used is to develop a time-series transfer-function model describing the daily PM10 (PM with less than 10 microm aerodynamic diameter) concentration as a function of lagged PM and current and lagged O3, NO or NO2, CO, and SO2. Approximately 3 years of daily average PM10, daily maximum 8-hr average O3 and CO, daily 24-hr average SO2 and NO2, and daily 6:00 a.m.-9:00 a.m. average NO from the Aerometric Information Retrieval System (AIRS) air quality subsystem are used for this analysis. Urban areas modeled are Chicago, IL; Los Angeles, CA; Phoenix, AZ; Philadelphia, PA; Sacramento, CA; and Detroit, MI. Time-series analysis identified significant autocorrelation in the O3, PM10, NO, NO2, CO, and SO2 series. Cross correlations between PM10 (dependent variable) and gaseous pollutants (independent variables) show that all of the gases are significantly correlated with PM10 and that O3 is also significantly correlated lagged up to two previous days. Once a transfer-function model of current PM10 is defined for an urban location, the effect of an O3-control strategy on PM concentrations is estimated by calculating daily PM10 concentrations with reduced O3 concentrations. Forecasted summertime PM10 reductions resulting from a 5 percent decrease in ambient O3 range from 1.2 microg/m3 (3.03%) in Chicago to 3.9 microg/m3 (7.65%) in Phoenix.

Air Pollutants↗

Spatial variability of PM2.5 in urban areas in the United States.

Data from the U.S. Environmental Protection Agency's Aerometric Information Retrieval System (now known as the Air Quality System) database for 1999 and 2000 have been used to characterize the spatial variability of concentrations of particulate matter with aerodynamic diameter < or = 2.5 microg (PM2.5) in 27 urban areas across the United States. Different measures were used to quantify the degree of uniformity of PM2.5 concentrations in the urban areas characterized. It was observed that PM2.5 concentrations varied to differing degrees in the urban areas examined. Analyses of several urban areas in the Southeast indicated high correlations between site pairs and spatial uniformity in concentration fields. Considerable spatial variation was found in other regions, especially in the West. Even within urban areas in which all site pairs were highly correlated, a variable degree of heterogeneity in PM2.5 concentrations was found. Thus, even though concentrations at pairs of sites were highly correlated, their concentrations were not necessarily the same. These findings indicate that the potential for exposure misclassification errors in time-series epidemiologic studies exists.

Air Pollutants↗

Assessing source characteristics of PM2.5 in the eastern United States using positive matrix factorization.

Fine aerosol (PM2.5) measurements obtained from the first year of operation of the nationwide network of PM2.5 monitors were studied with the factor analysis technique of positive matrix factorization (PMF). PM2.5 mass concentration data were extracted from the Atmospheric Information Retrieval System (AIRS) database of the U.S. Environmental Protection Agency (EPA). PMF was applied to measurements at more than 350 monitoring locations in the eastern half of the United States. Data consisted of PM2.5 24-hr averaged concentrations measured every third day from April through December 1999. The PMF model suggested six factors representing source influences to the PM2.5 mass concentrations at measurement sites. Factor 5, covering much of the Appalachian states, exhibited significant seasonal behavior.

Air Pollutants↗

PIR-ALN: a database of protein sequence alignments.

MOTIVATION: The Protein Information Resource (PIR) maintains a database of annotated and curated alignments in order to visually represent interrelationships among sequences in the PIR-International Protein Sequence Database, to spread and standardize protein names, features and keywords among members of a family or superfamily, and to aid us in classifying sequences, in identifying conserved regions, and in defining new homology domains. RESULTS: Release 22.0, (December 1998), of the PIR-ALN database contains a total of 3806 alignments, including 1303 superfamily, 2131 family and 372 homology domain alignments. This is an appropriate dataset to develop and extract patterns, test profiles, train neural networks or build Hidden Markov Models (HMMs). These alignments can be used to standardize and spread annotation to newer members by homology, as well as to understand the modular architecture of multidomain proteins. PIR-ALN includes 529 alignments that can be used to develop patterns not represented in PROSITE, Blocks, PRINTS and Pfam databases. The ATLAS information retrieval system can be used to browse and query the PIR-ALN alignments. AVAILABILITY: PIR-ALN is currently being distributed as a single ASCII text file along with the title, member, species, superfamily and keyword indexes. The quarterly and weekly updates can be accessed via the WWW at pir.georgetown.edu. The quarterly updates can also be obtained by anonymous FTP from the PIR FTP site at NBRF.Georgetown.edu, directory [ANONYMOUS.PIR.ALIGNMENT].

Amino Acid Sequence↗

The comparative toxicogenomics database: a cross-species resource for building chemical-gene interaction networks.

Chemicals in the environment play a critical role in the etiology of many human diseases. Despite their prevalence, the molecular mechanisms of action and the effects of chemicals on susceptibility to disease are not well understood. To promote understanding of these mechanisms, the Comparative Toxicogenomics Database (CTD; http://ctd.mdibl.org/) presents scientifically reviewed and curated information on chemicals, relevant genes and proteins, and their interactions in vertebrates and invertebrates. CTD integrates sequence, reference, species, microarray, and general toxicology information to provide a unique centralized resource for toxicogenomic research. The database also provides visualization capabilities that enable cross-species comparisons of gene and protein sequences. These comparisons will facilitate understanding of structure-function correlations and the genetic basis of susceptibility. Manual curation and integration of cross-species chemical-gene and chemical-protein interactions from the literature are now underway. These data will provide information for building complex interaction networks. New CTD features include (1) cross-species gene, rather than sequence, query and visualization capabilities; (2) integrated cross-links to microarray data from chemicals, genes, and sequences in CTD; (3) a reference set related to chemical-gene and protein interactions identified by an information retrieval system; and (4) a "Chemicals in the News" initiative that provides links from CTD chemicals to environmental health articles from the popular press. Here we describe these new features and our novel cross-species curation of chemical-gene and chemical-protein interactions.

Animals↗

Effect of ambient air pollution on pulmonary exacerbations and lung function in cystic fibrosis.

Information concerning the impact of environmental factors on cystic fibrosis (CF) is limited. We conducted a cohort study to assess the impact of air pollutants in CF. The study included patients over the age of 6 years enrolled in the Cystic Fibrosis Foundation National Patient Registry in 1999 and 2000. Exposure was assessed by linking air pollution values from the Aerometric Information Retrieval System with the patients' home zip code. After adjusting for confounders, a 10 microg/m(3) rise in particulate matter (both with a median aerodynamic diameter of 10 microm (PM(10)) or less and with an aerodynamic diameter of 2.5 microm or less (PM(2.5)) was associated with an 8% (95% confidence interval [CI], 2-15%) and 21% (95% CI, 7-33%) increase in the odds of two or more exacerbations, respectively; a 10-ppb rise in ozone was associated with a 10% (95% CI, 3-17%) increase in odds of two or more exacerbations. For every increase in PM(2.5) of 10 microg/m(3), there was an associated fall in FEV(1) of 24 ml (7-40) (95% CI) after adjusting for confounders. PM(2.5)'s association with mortality did not achieve statistical significance (adjusted RR = 1.32 per 10 microg/m(3) 0.91-1.93; 95% CI). Annual average exposures to particulate air pollution was associated with an increased risk of pulmonary exacerbations and a decline in lung function, suggesting a role of environmental exposures on prognosis in CF.

Adolescent↗

The effect of air pollution on inner-city children with asthma.

The effect of daily ambient air pollution was examined within a cohort of 846 asthmatic children residing in eight urban areas of the USA, using data from the National Cooperative Inner-City Asthma Study. Daily air pollution concentrations were extracted from the Aerometric Information Retrieval System database from the Environment Protection Agency in the USA. Mixed linear models and generalized estimating equation models were used to evaluate the effects of several air pollutants (ozone, sulphur dioxide (SO2), nitrogen dioxide (NO2) and particles with a 50% cut-off aerodynamic diameter of 10 microm (PM10) on peak expiratory flow rate (PEFR) and symptoms in 846 children with a history of asthma (ages 4-9 yrs). None of the pollutants were associated with evening PEFR or symptom reports. Only ozone was associated with declines in morning % PEFR (0.59% decline (95% confidence interval (CI) 0.13-1.05%) per interquartile range (IQR) increase in 5-day average ozone). In single pollutant models, each pollutant was associated with an increased incidence of morning symptoms: (odds ratio (OR)=1.16 (95% CI 1.02-1.30) per IQR increase in 4-day average ozone, OR=1.32 (95% CI 1.03-1.70) per IQR increase in 2-day average SO2, OR=1.48 (95% CI 1.02-2.16) per IQR increase in 6-day average NO2 and OR=1.26 (95% CI 1.0-1.59) per IQR increase in 2-day average PM10. This longitudinal analysis supports previous time-series findings that at levels below current USA air-quality standards, summer-air pollution is significantly related to symptoms and decreased pulmonary function among children with asthma.

Air Pollutants↗

Machine learning and word sense disambiguation in the biomedical domain: design and evaluation issues.

BACKGROUND: Word sense disambiguation (WSD) is critical in the biomedical domain for improving the precision of natural language processing (NLP), text mining, and information retrieval systems because ambiguous words negatively impact accurate access to literature containing biomolecular entities, such as genes, proteins, cells, diseases, and other important entities. Automated techniques have been developed that address the WSD problem for a number of text processing situations, but the problem is still a challenging one. Supervised WSD machine learning (ML) methods have been applied in the biomedical domain and have shown promising results, but the results typically incorporate a number of confounding factors, and it is problematic to truly understand the effectiveness and generalizability of the methods because these factors interact with each other and affect the final results. Thus, there is a need to explicitly address the factors and to systematically quantify their effects on performance. RESULTS: Experiments were designed to measure the effect of "sample size" (i.e. size of the datasets), "sense distribution" (i.e. the distribution of the different meanings of the ambiguous word) and "degree of difficulty" (i.e. the measure of the distances between the meanings of the senses of an ambiguous word) on the performance of WSD classifiers. Support Vector Machine (SVM) classifiers were applied to an automatically generated data set containing four ambiguous biomedical abbreviations: BPD, BSA, PCA, and RSV, which were chosen because of varying degrees of differences in their respective senses. Results showed that: 1) increasing the sample size generally reduced the error rate, but this was limited mainly to well-separated senses (i.e. cases where the distances between the senses were large); in difficult cases an unusually large increase in sample size was needed to increase performance slightly, which was impractical, 2) the sense distribution did not have an effect on performance when the senses were separable, 3) when there was a majority sense of over 90%, the WSD classifier was not better than use of the simple majority sense, 4) error rates were proportional to the similarity of senses, and 5) there was no statistical difference between results when using a 5-fold or 10-fold cross-validation method. Other issues that impact performance are also enumerated. CONCLUSION: Several different independent aspects affect performance when using ML techniques for WSD. We found that combining them into one single result obscures understanding of the underlying methods. Although we studied only four abbreviations, we utilized a well-established statistical method that guarantees the results are likely to be generalizable for abbreviations with similar characteristics. The results of our experiments show that in order to understand the performance of these ML methods it is critical that papers report on the baseline performance, the distribution and sample size of the senses in the datasets, and the standard deviation or confidence intervals. In addition, papers should also characterize the difficulty of the WSD task, the WSD situations addressed and not addressed, as well as the ML methods and features used. This should lead to an improved understanding of the generalizablility and the limitations of the methodology.

Algorithms↗

Associations between air pollution and mortality in Phoenix, 1995-1997.

We evaluated the association between mortality outcomes in elderly individuals and particulate matter (PM) of varying aerodynamic diameters (in micrometers) [PM(10), PM(2.5), and PM(CF )(PM(10) minus PM(2.5))], and selected particulate and gaseous phase pollutants in Phoenix, Arizona, using 3 years of daily data (1995-1997). Although source apportionment and epidemiologic methods have been previously combined to investigate the effects of air pollution on mortality, this is the first study to use detailed PM composition data in a time-series analysis of mortality. Phoenix is in the arid Southwest and has approximately 1 million residents (9. 7% of the residents are > 65 years of age). PM data were obtained from the U.S. Environmental Protection Agency (EPA) National Exposure Research Laboratory Platform in central Phoenix. We obtained gaseous pollutant data, specifically carbon monoxide, nitrogen dioxide, ozone, and sulfur dioxide data, from the EPA Aerometric Information Retrieval System Database. We used Poisson regression analysis to evaluate the associations between air pollution and nonaccidental mortality and cardiovascular mortality. Total mortality was significantly associated with CO and NO(2) (p < 0.05) and weakly associated with SO(2), PM(10), and PM(CF) (p < 0. 10). Cardiovascular mortality was significantly associated with CO, NO(2), SO(2), PM(2.5), PM(10), PM(CF) (p < 0.05), and elemental carbon. Factor analysis revealed that both combustion-related pollutants and secondary aerosols (sulfates) were associated with cardiovascular mortality.

Aged↗

Information technologies for clinical toxicology in Russia.

OBJECTIVE: To describe Poison Information in Russia. RESULTS: The Moscow Toxicology Information and Advisory Center was created in 1993 as an institution of the Russian Federation Ministry of Health and Medical Industry. The Toxicology Information and Advisory Center is the first in a network of over 20 toxicology information centers to be created in different regions of Russia by 1998. At present the Toxicology Information and Advisory Center serves over 20 million people in the Moscow region with episodic inquiries from other areas. A prototype national bank of clinical and toxicological data on acute chemical poisoning and an information retrieval system POISON have been created. Work is underway to create computerized systems for data analysis of telephone inquiries on diagnosis and treatment protocols of acute poisoning.

Computer Communication Networks↗