Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Information Retrieval Systems”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Predicting Lexical Relations between Biomedical Terms: towards a Multilingual Morphosemantics-based System.

This paper addresses the issue of how semantic information can be automatically assigned to compound terms, i.e. both a definition and a set of semantic relations. This issue is particularly crucial when elaborating multilingual databases and when developing cross-language information retrieval systems. The paper shows how morpho-semantics can contribute in the constitution of multilingual lexical networks in biomedical corpora. It presents a system capable of labelling terms with morphologically related words, i.e. providing them with a definition, and grouping them according to synonymy, hyponymy and proximity relations. The approach requires the interaction of three techniques: (1) a la morphosemantic parser, (2) a multilingual table defining basic relations between word roots, and (3) a set of language-independant rules to draw up the list of related terms. This approach has been fully implemented for French, on an about 29,000 terms biomedical lexicon, resulting to more than 3,000 lexical families.

Language↗

Database of protein sequence alignments: PIR-ALN.

The Protein Information Resource (PIR) has been maintaining a database of curated protein sequence alignments since 1991. The collection includes superfamily, family and homology domain alignments. CLUSTAL V/W is used to generate multiple sequence alignments and ALNED, an interactive alignment editor, is used to check and correct them. The database has helped in classifying sequences, in defining new homology domains, and in spreading and standardizing protein names, features and keywords among members of a family or superfamily. The ATLAS information retrieval system can be used to browse and query the PIR-ALN alignments. The quarterly and weekly updates can be accessed via the WWW at http://www-nbrf. georgetown.edu/pir/

Databases, Factual↗

Interpreting natural language queries using the UMLS.

This paper describes AQUA (A QUery Analyzer), the natural language front end of a prototype information retrieval system. AQUA translates a user's natural language query into a representation in the Conceptual Graph formalism. The graph is then used by subsequent components to search various resources such as databases of the medical literature. The focus of the parsing method is on semantics rather than syntax, with semantic restrictions being provided by the UMLS Semantic Net. The intent of the approach is to provide a method that can be emulated easily in applications that require simple natural language interfaces.

Humans↗

Multi-media orientation and education programs for MEDLARS used the the Medical Information Center (MIC), Stockholm.

The Mediatron is a modified stereo-tape recorder which is designed to carry out simultaneous recording of audio-commentaries, trigger pulses for photographic slides, and digital signals from a computerized information-retrieval system. Two MEDLARS programs were produced: an orientation program and a self-instructional package. Technical procedures and experiences are briefly discussed.

Humans↗

Free-text medical document retrieval via phrase-based vector space model.

Many information retrieval systems are based on vector space model (VSM) that represents a document as a vector of index terms. Concepts have been proposed to replace word stems as the index terms to improve retrieval accuracy. However, past research revealed that such systems did not outperform the traditional stem-based systems. Incorporating conceptual similarity derived from knowledge sources should have the potential to improve retrieval accuracy. Yet the incompleteness of the knowledge source precludes significant improvement. To remedy this problem, we propose to represent documents using phrases. A phrase consists of multiple concepts and word stems. The similarity between two phrases is jointly determined by their conceptual similarity and their common word stems. The document similarity can in turn be derived from phrase similarities. Using OHSUMED as a test collection and UMLS as the knowledge source, our experiment results reveal that phrase-based VSM yields a 16% increase of retrieval accuracy compared to the stem-based model.

Abstracting and Indexing↗

A performance and failure analysis of SAPHIRE with a MEDLINE test collection.

OBJECTIVE: Assess the performance of the SAPHIRE automated information retrieval system. DESIGN: Comparative study of automated and human searching of a MEDLINE test collection. MEASUREMENTS: Recall and precision of SAPHIRE were compared with those attributes of novice physicians, expert physicians, and librarians for a test collection of 75 queries and 2,334 citations. Failure analysis assessed the efficacy of the Metathesaurus as a concept vocabulary; the reasons for retrieval of nonrelevant articles and nonretrieval of relevant articles; and the effect of changing the weighting formula for relevance ranking of retrieved articles. RESULTS: Recall and precision of SAPHIRE were comparable to those of both physician groups, but less than those of librarians. CONCLUSION: The current version of the Metathesaurus, as utilized by SAPHIRE, was unable to represent the conceptual content of one-fourth of physician-generated MEDLINE queries. The most likely cause for retrieval of nonrelevant articles was the presence of some or all of the search terms in the article, with frequencies high enough to lead to retrieval. The most likely cause for nonretrieval of relevant articles was the absence of the actual terms from the query, with synonyms or hierarchically related terms present instead. There were significant variations in performance when SAPHIRE's concept-weighing formulas were modified.

Abstracting and Indexing↗

Wide-area network connecting a hospital drug informatics center with a university.

A wide-area network (WAN) connecting a new drug informatics center in a university-affiliated hospital with the university's campus-based computer network is described. In 1994 a pharmacy school developed a drug informatics center in an affiliated hospital. The center was originally designed around a local-area network (LAN) to be located at the hospital and planned to provide clients with easy access to typical productivity software and various electronic information resources. Only occasional modem connections to the university network were envisioned. However, large price increases in information retrieval systems and decreases in the cost of a frame relay connection (T1 line) to the campus network led to the installation of a WAN when the drug informatics center was established. Technical, political, and legal problems were overcome, and the connection was made. The WAN gave faculty and students at the hospital access to many of the university's computing and Internet resources. In addition, the faculty and students have access to various files and programs available only on the drug informatics center's file server at the affiliated hospital. It cost about $6500 to install all WAN equipment and maintain the frame relay for the first year, or a third of what would have been necessary for information retrieval software had a separate LAN been established at the hospital. A WAN connecting a drug informatics center and a university's computer network gave the center access to more electronic information resources at lower cost than would have been possible with a separate LAN.

Computer Communication Networks↗

Information retrieval: an overview of system characteristics.

The paper gives an overview of characteristics of information retrieval (IR) systems. The characteristics are identified from the descriptions of 23 IR systems. Four IR models are discussed: the Boolean model, the vector model, the probabilistic model and the connectionistic model. Twelve other characteristics of IR models are identified: search intermediary, domain knowledge, relevance feedback, natural language interface, graphical query language, conceptual queries, full-text IR, field searching, fuzzy queries, hypertext integration, machine learning, and ranked output. Finally, the relevance of IR systems for the World Wide Web is established.

Algorithms↗

Histone Sequence Database: new histone fold family members.

Searches of the major public protein databases with core and linker chicken and human histone sequences have resulted in the compilation of an annotated set of histone protein sequences. In addition, new database searches with two distinct motif search algorithms have identified several members of the histone fold family, including human DRAP1 and yeast CSE4. Database resources include information on conflicts between similar sequence entries in different source databases, multiple sequence alignments, links to the Entrez integrated information retrieval system, structures for histone and histone fold proteins, and the ability to visualize structural data through Cn3D. The database currently contains >1000 protein sequences, which are searchable by protein type, accession number, organism name, or any other free text appearing in the definition line of the entry. All sequences and alignments in this database are available through the World Wide Web at http://www.nhgri.nih. gov/DIR/GTB/HISTONES or http://www.ncbi.nlm.nih. gov/Baxevani/HISTONES

Amino Acid Sequence↗

RIPS: a UNIX-based reference information program for scientists.

A set of programs is described which implement a personal reference management and information retrieval system on a UNIX-based minicomputer. The system operates in a multiuser configuration with a host of user-friendly utilities that assist entry of reference material, its retrieval, and formatted printing for associated tasks. A search command language was developed without restriction in keyword vocabulary, number of keywords, or level of parenthetical expression nesting. The system is readily transported, and by design is applicable to any academic specialty.

Books↗

"Bag of words" is not enough for strength of evidence classification.

Incorporation of evidence from clinical research requires critical appraisal of its quality. Information retrieval systems can facilitate clinicians' judgments by automatically labeling retrieved citations with their strength of evidence categories. Preliminary results of such a text classification experiment involving MEDLINE citations show that a "bag of words" approach is insufficient for accurate classification.

Artificial Intelligence↗

Use of an index to reflect the aggregate burden of long-term exposure to criteria air pollutants in the United States.

Air pollution control in the United States for five common pollutants--particulate matter, ground-level ozone, sulfur dioxide, nitrogen dioxide, and carbon monoxide--is based partly on the attainment of ambient air quality standards that represent a level of air pollution regarded as safe. Regulatory and health agencies often focus on whether standards for short periods are attained; the number of days that standards are exceeded is used to track progress. Efforts to explain air pollution to the public often incorporate an air quality index that represents daily concentrations of pollutants. While effects of short-term exposures have been emphasized, research shows that long-term exposures to lower concentrations of air pollutants can also result in adverse health effects. We developed an aggregate index that represents long-term exposure to these pollutants, using 1995 monitoring data for metropolitan areas obtained from the U.S. Environmental Protection Agency's Aerometric Information Retrieval System. We compared the ranking of metropolitan areas under the proposed aggregate index with the ranking of areas by the number of days that short-term standards were exceeded. The geographic areas with the highest burden of long-term exposures are not, in all cases, the same as those with the most days that exceeded a short-term standard. We believe that an aggregate index of long-term air pollution offers an informative addition to the principal approaches currently used to describe air pollution exposures; further work on an aggregate index representing long-term exposure to air pollutants is warranted.

Air Pollutants↗

CDD: a Conserved Domain Database for protein classification.

The Conserved Domain Database (CDD) is the protein classification component of NCBI's Entrez query and retrieval system. CDD is linked to other Entrez databases such as Proteins, Taxonomy and PubMed, and can be accessed at http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=cdd. CD-Search, which is available at http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi, is a fast, interactive tool to identify conserved domains in new protein sequences. CD-Search results for protein sequences in Entrez are pre-computed to provide links between proteins and domain models, and computational annotation visible upon request. Protein-protein queries submitted to NCBI's BLAST search service at http://www.ncbi.nlm.nih.gov/BLAST are scanned for the presence of conserved domains by default. While CDD started out as essentially a mirror of publicly available domain alignment collections, such as SMART, Pfam and COG, we have continued an effort to update, and in some cases replace these models with domain hierarchies curated at the NCBI. Here, we report on the progress of the curation effort and associated improvements in the functionality of the CDD information retrieval system.

Amino Acid Sequence↗

What factors are associated with the integration of evidence retrieval technology into routine general practice settings?

BACKGROUND: Information retrieval systems have the potential to improve patient care but little is known about the variables which influence clinicians' uptake and use of systems in routine work. AIM: To determine which factors influenced use of an online evidence retrieval system. DESIGN OF STUDY: Computer logs and pre- and post-system survey analysis of a 4-week clinical trial of the Quick Clinical online evidence system involving 227 general practitioners across Australia. RESULTS: Online evidence use was not linked to general practice training or clinical experience but female clinicians conducted more searches than their male counterparts (mean use=14.38 searches, S.D.=11.68 versus mean use=8.50 searches, S.D.=9.99; t=2.67, d.f.=157, P=0.008). Practice characteristics such as hours worked, type and geographic location of clinic were not associated with search activity. Information seeking was also not related to participants' perceived information needs, computer skills, training nor Internet connection speed. Clinicians who reported direct improvements in patient care as a result of system use had significantly higher rates of system use than other users (mean use=12.55 searches, S.D.=13.18 versus mean use=8.15 searches, S.D.=9.18; t=2.322, d.f.=154 P=0.022). Comparison of participants' views pre- and post- the trial, showed that post-trial clinicians expressed more positive views about searching for information during a consultation (chi(2)=27.40, d.f.=4, P< or =0.001) and a significantly greater number reported seeking information between consultations as a result of having access to an online evidence system in their consulting rooms (chi(2)=9.818, d.f.=2, P=0.010). CONCLUSION: Clinicians' use of an online evidence system was directly related to their reported experiences of improvements in patient care. Post-trial clinicians positively changed their views about having time to search for information and pursued more questions during clinic hours.

Adult↗

ASTER: an integration of the AQUIRE data base and the QSAR system for use in ecological risk assessments.

Ecological risk assessments are used by the US Environmental Protection Agency (US EPA) and other governmental agencies to assist in determining the probability and magnitude of deleterious effects of hazardous chemicals on plants and animals. These assessments are important steps in formulating regulatory decisions. The completion of an ecological risk assessment requires the gathering of ecotoxicological hazard and environmental exposure information. This information is evaluated in the risk characterization section to assist in making the final risk assessment. ASTER (ASsessment Tools for the Evaluation of Risk) was designed by the US EPA Environmental Research Laboratory-Duluth (ERL-D) to assist regulators in producing assessments. ASTER is an integration of the ACQUIRE (AQUatic toxicity Information REtrieval system) and QSAR (Quantitative Structure Activity Relationships) systems. ACQUIRE is a data base of aquatic toxicity tests and QSAR is comprised of a data base of measured physicochemical properties, and various QSAR models that estimate physicochemical and ecotoxicological endpoints. ASTER will be available to international governmental agencies through the US EPA National Computing Center.

Animals↗

The SAPHIRE server: a new algorithm and implementation.

SAPHIRE is an experimental information retrieval system implemented to test new approaches to automated indexing and retrieval of medical documents. Due to limitations in its original concept-matching algorithm, a modified algorithm has been implemented which allows greater flexibility in partial matching and different word order within concepts. With the concomitant growth in client-server applications and the Internet in general, the new algorithm has been implemented as a server that can be accessed via other applications on the Internet.

Abstracting and Indexing↗