Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dictionary”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Developing national health information in Australia.

Two significant developments in the past two years have given impetus to development of health information in Australia. In March 1993, the former National Minimum Data Set was revised and published as the National Health Data Dictionary. Second, establishment of an agreement in June 1993, between the Commonwealth and State/Territory government health authorities, the Australian Bureau of Statistics, and the Australian Institute of Health and Welfare initiated a process of working cooperatively to develop national health information. Australia, like many other countries, suffers from inconsistent health data definitions, lack of timely data, poor data quality, gaps in data coverage, and barriers to accessing the data. The National Health Information Agreement [1] came into effect on June 1, 1993 and seeks to provide a national framework and processes to improve national health information, that is, information on health of the population; determinants of the population's health; provision and utilization of health promotion and disease prevention programs and health services including: outcomes and outputs, resource use and costs, access by and distribution to population groups; relationships between these elements; and the language necessary to facilitate provision of services and collection of national health information. The major implementation mechanism of the Agreement is a rolling three-year National Health Information Work Program of national health information activities. The activities range from development work on standard hospital charts of accounts, on health outcome measures, and on new collections such as outpatients to improved definitions and the enhancement of existing collections such as mental health and vital statistics. The Work Program is published annually. A first priority is to improve the data collections available. This is being achieved through the setting of national data definitions and standards. The Agreement recognizes the National Health Data Dictionary (NHDD) as the authoritative set of national definitions and is a significant initiative aimed at improving Australia's health information. The dictionary is the repository of the agreed common language, use of the definitions facilities the description and comparison of health and health services nationally [2]. The National Health Data Dictionary currently covers institutionally provided health care, the national health labor force, and is expanding to cover other major areas, including outpatient services, community care, and mental health. The NHDD is reviewed and maintained by the National Health Data Committee and the overall coordination of definition development projects and publication is undertaken by the Institute. The placement of an agreed definition in the NHDD does not automatically mean that it has a place in a national data collection. The use of the dictionary definition will allow comparison by and between service providers. In order for a data item to be eligible for inclusion in a national minimum data set, the definition of that item must be contained in the NHDD. During the first three months of 1995, the Australian Institute of Health and Welfare will conduct a national project to develop a model for the health system in Australia. The model will provide a common vocabulary and information architecture in order to facilitate better quality health information, and consequently better health for Australians. It is expected that the development of the model will bring several benefits including facilitating the more rapid and accurate assembly of appropriate clinical information to support improved customer service and outcomes, provide a mechanism for achieving better quality information, reduce the costs of data collection; provide enabling mechanisms for the integration of systems via data standards and reduce the costs of acquiring information systems through reduced development and tailoring costs for suppliers. (abst

Australia↗

Recognition of human genes by stochastic parsing.

A gene finding system, GeneDecoder, based on a parsing technique using a stochastic grammar and dictionary of genetic words is introduced. The structure of human genes are expressed by a stochastic grammar and a dictionary, whose components are the genetic words consisting of genetic phonemes, built as hidden Markov models (HMMs). The HMMs represent the nucleotide acid bases, the codons, and the amino acids. The genetic words in the dictionary are described by the sequence of these HMMs and represent exons, introns, intergenic regions, tRNA regions and signals in DNA sequences. The statistics between these regions are expressed by the grammar, which is a stochastic network of the genetic words. Using the same kind of technique of speech recognition by HMMs with a word dictionary and a grammar, the stochastic network of genetic words enables the motif dictionary to be used during the parsing of the DNA sequences. At the same time, stochastic features of donor/acceptor sites, information of the di-codon statistics, and other important features are integrated into stochastic scores during the parsing. As a result, while the system parses DNA sequences and finds the exon/intron structures, the protein motifs are automatically annotated in the regions. It helps to identify the functions of the genes and reduces the cost of homology search for each hypothetical coding regions. This method is different from simply using the information of homology search. This method uses the information of the motif patterns during the parsing process, but searching the motif patterns after/before finding the coding regions cannot directly affect the parsing process itself. Experimental results have shown that this method reasonably finds and annotates the motifs in the exons in the DNA sequence of human.

Amino Acid Sequence↗

Organizing the present, looking to the future: an online knowledge repository to facilitate collaboration.

BACKGROUND: Comprehensive data available in the Canadian province of Manitoba since 1970 have aided study of the interaction between population health, health care utilization, and structural features of the health care system. Given a complex linked database and many ongoing projects, better organization of available epidemiological, institutional, and technical information was needed. OBJECTIVE: The Manitoba Centre for Health Policy and Evaluation wished to develop a knowledge repository to handle data, document research Methods, and facilitate both internal communication and collaboration with other sites. METHODS: This evolving knowledge repository consists of both public and internal (restricted access) pages on the World Wide Web (WWW). Information can be accessed using an indexed logical format or queried to allow entry at user-defined points. The main topics are: Concept Dictionary, Research Definitions, Meta-Index, and Glossary. The Concept Dictionary operationalizes concepts used in health research using administrative data, outlining the creation of complex variables. Research Definitions specify the codes for common surgical procedures, tests, and diagnoses. The Meta-Index organizes concepts and definitions according to the Medical Sub-Heading (MeSH) system developed by the National Library of Medicine. The Glossary facilitates navigation through the research terms and abbreviations in the knowledge repository. An Education Resources heading presents a web-based graduate course using substantial amounts of material in the Concept Dictionary, a lecture in the Epidemiology Supercourse, and material for Manitoba's Regional Health Authorities. Confidential information (including Data Dictionaries) is available on the Centre's internal website. RESULTS: Use of the public pages has increased dramatically since January 1998, with almost 6,000 page hits from 250 different hosts in May 1999. More recently, the number of page hits has averaged around 4,000 per month, while the number of unique hosts has climbed to around 400. CONCLUSIONS: This knowledge repository promotes standardization and increases efficiency by placing concepts and associated programming in the Centre's collective memory. Collaboration and project management are facilitated.

Databases as Topic↗

STAR/mmCIF: an ontology for macromolecular structure.

MOTIVATION: Crystallographers were motivated 10 years ago to develop a simple and consistent data representation for the exchange and archiving of data associated with the crystallographic experiment and the final structure. As this process evolved (and the data grew at near exponential rates) came the recognition that this representation should also facilitate the automated management of the data and, with the aid of additional software for verification and validation, provide improved consistency and accuracy and hence improved scientific inquiry. This realization led to a new Dictionary Definition Language (DDL) and an extensive dictionary based on this DDL for describing macromolecular structure. In broad terms this could be considered an ontology. An important feature in the development of the ontology was the endorsement and ongoing maintenance and support of the International Union of Crystallography (IUCr). While the description of macromolecular structure and the x-ray crystallographic experiment used to derive it represent explicit data, the ontology is extensible and applicable to other less well-characterized data domains. RESULTS: Details of the DDL, the dictionaries that have been developed, and software for reading and using this ontology are presented. AVAILABILITY: Extensive documentation, software tools and the DDL and dictionaries are available from http://ndbserver.rutgers.edu/mmcif and associated mirror sites. CONTACT: Bourne: bourne@sdsc.edu and Westbrook:jwest@rcsp.rutgers.edu

Crystallography, X-Ray↗

PDBML: the representation of archival macromolecular structure data in XML.

SUMMARY: The Protein Data Bank (PDB) has recently released versions of the PDB Exchange dictionary and the PDB archival data files in XML format collectively named PDBML. The automated generation of these XML files is driven by the data dictionary infrastructure in use at the PDB. The correspondences between the PDB dictionary and the XML schema metadata are described as well as the XML representations of PDB dictionaries and data files.

Amino Acid Sequence↗

Image decomposition via the combination of sparse representations and a variational approach.

The separation of image content into semantic parts plays a vital role in applications such as compression, enhancement, restoration, and more. In recent years, several pioneering works suggested such a separation be based on variational formulation and others using independent component analysis and sparsity. This paper presents a novel method for separating images into texture and piecewise smooth (cartoon) parts, exploiting both the variational and the sparsity mechanisms. The method combines the basis pursuit denoising (BPDN) algorithm and the total-variation (TV) regularization scheme. The basic idea presented in this paper is the use of two appropriate dictionaries, one for the representation of textures and the other for the natural scene parts assumed to be piecewise smooth. Both dictionaries are chosen such that they lead to sparse representations over one type of image-content (either texture or piecewise smooth). The use of the BPDN with the two amalgamed dictionaries leads to the desired separation, along with noise removal as a by-product. As the need to choose proper dictionaries is generally hard, a TV regularization is employed to better direct the separation process and reduce ringing artifacts. We present a highly efficient numerical scheme to solve the combined optimization problem posed by our model and to show several experimental results that validate the algorithm's performance.

Algorithms↗

Unsupervised analysis of polyphonic music by sparse coding.

We investigate a data-driven approach to the analysis and transcription of polyphonic music, using a probabilistic model which is able to find sparse linear decompositions of a sequence of short-term Fourier spectra. The resulting system represents each input spectrum as a weighted sum of a small number of "atomic" spectra chosen from a larger dictionary; this dictionary is, in turn, learned from the data in such a way as to represent the given training set in an (information theoretically) efficient way. When exposed to examples of polyphonic music, most of the dictionary elements take on the spectral characteristics of individual notes in the music, so that the sparse decomposition can be used to identify the notes in a polyphonic mixture. Our approach differs from other methods of polyphonic analysis based on spectral decomposition by combining all of the following: (a) a formulation in terms of an explicitly given probabilistic model, in which the process estimating which notes are present corresponds naturally with the inference of latent variables in the model; (b) a particularly simple generative model, motivated by very general considerations about efficient coding, that makes very few assumptions about the musical origins of the signals being processed; and (c) the ability to learn a dictionary of atomic spectra (most of which converge to harmonic spectral profiles associated with specific notes) from polyphonic examples alone-no separate training on monophonic examples is required.

Journal Article↗

Unbiased high resolution method of EEG analysis in time-frequency space.

Matching Pursuit (MP)--a method of high-resolution signal analysis--is described in the context of other methods operating in time-frequency space. The method relies on an adaptive approximation of a signal by means of waveforms chosen from a very large and redundant dictionary of functions. The MP performance is illustrated by simulations and examples of sleep spindles and slow wave activity analysis. An improvement of the original procedure, relying on the introduction of stochastic dictionaries, is proposed. A comparison of the performance of dyadic and stochastic dictionaries is presented. MP with stochastic dictionaries is characterized by an unmatched resolution in time-frequency space; moreover it allows for parametric description of all (periodic and transient) signal features in the framework of the same formalism. Matching pursuit is especially suitable for analysis of non-stationary signals and is a unique tool for the investigation of dynamic changes of brain activity.

Brain↗

Standardized terminology for clinical trial protocols based on top-level ontological categories.

This paper describes a new method for the ontologically based standardization of concepts with regard to the quality assurance of clinical trial protocols. We developed a data dictionary for medical and trial-specific terms in which concepts and relations are defined context-dependently. The data dictionary is provided to different medical research networks by means of the software tool Onto-Builder via the internet. The data dictionary is based on domain-specific ontologies and the top-level ontology of GOL. The concepts and relations described in the data dictionary are represented in natural language, semi-formally or formally according to their use.

Clinical Trials as Topic↗

Categorization of free-text problem lists: an effective method of capturing clinical data.

Problem lists assist in organizing patient information in computer based medical records. However, in order to use problem lists for billing, research, decision support and standardization, a categorization of the problems entered is required. We describe the problem list component of our computerized patient record, the On-line Medical Record (OMR), which combines a free-text entry mechanism with a categorization scheme, using a dictionary containing 846 terms. All 118,040 problems entered during the system's six years of use have been analyzed, 477 clinicians have entered a mean +/- S.D. of 238 +/- 604 problems into 22,311 patient records. The average number of problems in each patient's file was 5.1 +/- 3.9. Comments were typed for 80,281 (68%) of the problems, ranging in length from 1 to 2456 characters, with a mean length of 98 +/- 110 characters. Half the problems were entered on the day of the encounter with the patient. Overall, 66% of all problems were categorized in relation to terms from the problem dictionary. Lexical analysis of all problem names showed that 80% could be mapped to Meta 1.4, Snomed 3.0 or a pre-release version of Read 3.0. We conclude that a problem list entry scheme combining free-text entry and optional categorization using a dictionary can result in a high proportion of problems being categorized as desired. Improvement of the system by elimination of unused dictionary terms and addition of 1000 terms identified by the lexical analysis is likely to result in even higher categorization rates.

Humans↗

Integrated clinical information system.

SIDOCI (Système Informatisé de DOnnées Cliniques Intégrées) is a Canadian joint venture introducing newly-operating paradigms into hospitals. The main goal of SIDOCI is to maintain the quality of care in todayUs tightening economy. SIDOCI is a fully integrated paperless patient-care system which automates and links all information about a patient. Data is available on-line and instantaneously to doctors, nurses, and support staff in the format that best suits their specific requirements. SIDOCI provides a factual and chronological summary of the patient's progress by drawing together clinical information provided by all professionals working with the patient, regardless of their discipline, level of experience, or physical location. It also allows for direct entry of the patient's information at the bedside. Laboratory results, progress notes, patient history and graphs are available instantaneously on screen, eliminating the need for physical file transfers. The system, incorporating a sophisticated clinical information database, an intuitive graphical user interface, and customized screens for each medical discipline, guides the user through standard procedures. Unlike most information systems created for the health care industry, SIDOCI is longitudinal, covering all aspects of the health care process through its link to various vertical systems already in place. A multidisciplinary team has created a clinical dictionary that provides the user with most of the information she would normally use: symptoms, signs, diagnoses, allergies, medications, interventions, etc. This information is structured and displayed in such a manner that health care professionals can document the clinical situation at the touch of a finger. The data is then encoded into the patient's file. Once encoded, the structured data is accessible for research, statistics, education, and quality assurance. This dictionary complies with national and international nomenclatures. It also contains personalized profiles: questionnaires based on the predetermined choices of the information most relevant to the specific user. The SIDOCI clinical dictionary also includes the hospital's suggested or mandatory interventions, clinical guidelines, and protocols. These clinical guidelines are customized at the hospital, service, and professional levels. Common interventions have been regrouped so that health professionals may apply the appropriate diagnostic, therapeutic, educational, or other intervention plans. The clinical dictionary also serves as a teaching and continuing education tool. The patient profile is a permanent record containing information on allergies, blood type, primary and secondary diagnoses, ongoing treatments, and prior hospitalizations. The problem list dealing with the current hospitalization includes symptoms, signs, and diagnoses. This standard clinical record facilitates communication between the services and provides a quick overview of the patient's history should emergency treatment be required. This health information system integrates Requests and Results, Progress Notes, and Analysis of the results. In addition, functions inherent to a patient's clinical cycle such as Administrative Management of episodes, Adaptation to physical and professional structures of the hospital, Messages between health professionals, and Electronic signature constitute the basis of SIDOCI. The most exciting aspect of this research project is its social impact: a more efficient health care system will improve the lives of all citizens. Moreover this applied research project involves the information industry and directly calls for the input of users such as doctors, nurses and hospital support staff.

Hospital Information Systems↗

Resolving abbreviations to their senses in Medline.

MOTIVATION: Biological literature contains many abbreviations with one particular sense in each document. However, most abbreviations do not have a unique sense across the literature. Furthermore, many documents do not contain the long forms of the abbreviations. Resolving an abbreviation in a document consists of retrieving its sense in use. Abbreviation resolution improves accuracy of document retrieval engines and of information extraction systems. RESULTS: We combine an automatic analysis of Medline abstracts and linguistic methods to build a dictionary of abbreviation/sense pairs. The dictionary is used for the resolution of abbreviations occurring with their long forms. Ambiguous global abbreviations are resolved using support vector machines that have been trained on the context of each instance of the abbreviation/sense pairs, previously extracted for the dictionary set-up. The system disambiguates abbreviations with a precision of 98.9% for a recall of 98.2% (98.5% accuracy). This performance is superior in comparison with previously reported research work. AVAILABILITY: The abbreviation resolution module is available at http://www.ebi.ac.uk/Rebholz/software.html.

Abbreviations as Topic↗

Development of a standardized language for case management among high-risk elderly.

Consistency and communication remain key barriers to tracking case management outcomes and developing the best practices. A dictionary of case management problems, goals, interventions, and outcomes was developed to support a prevention-oriented case management program targeted on elderly high-risk patients. Case management featured an annual screening questionnaire, appointment monitoring, disease education, self-management support, and ongoing care coordination. The dictionary resulted in a Standardized Language for Case Management (SLED). This has since been reviewed and modified on the basis of comments and recommendations from 5 leading case management organizations and is aligned with Standards of Practice for Case Management. The article provides a description of the standardized language terms, the rationale underlying the documentation, examples of how this dictionary of definitions can be incorporated into the daily practice of case management, and examples of some of the benefits to the field that can be achieved with the use of a common data-recording system.

Aged↗

A frequency-based technique to improve the spelling suggestion rank in medical queries.

OBJECTIVE: There is an abundance of health-related information online, and millions of consumers search for such information. Spell checking is of crucial importance in returning pertinent results, so the authors propose a technique for increasing the effectiveness of spell-checking tools used for health-related information retrieval. DESIGN: A sample of incorrectly spelled medical terms was submitted to two different spell-checking tools, and the resulting suggestions, derived under two different dictionary configurations, were re-sorted according to how frequently each term appeared in log data from a medical search engine. MEASUREMENTS: Univariable analysis was carried out to assess the effect of each factor (spell-checking tool, dictionary type, re-sort, or no re-sort) on the probability of success. The factors that were statistically significant in the univariable analysis were then used in multivariable analysis to evaluate the independent effect of each of the factors. RESULTS: The re-sorted suggestions proved to be significantly more accurate than the original list returned by the spell-checking tool. The odds of finding the correct suggestion in the number one rank were increased by 63% after re-sorting using the authors' method. This effect was independent of both the dictionary and the spell-checking tools that were used. CONCLUSION: Using knowledge about the frequency of a given word's occurrence in the medical domain can significantly improve spelling correction for medical queries.

Analysis of Variance↗

Generality of connotative meaning across methods and subjects.

A test was made of the generality of the connotative meaning scores of evaluation, activity, and potency in a dictionary of the 1000 most frequent words in English that has been used as the basis of a computer system to measure the expression of emotional tone. This was done by correlating those scores with the values for evaluation and activity given for words also found in two independently created dictionaries, one based upon adults' ratings and one upon ratings by children. In replicated findings reported by the authors of the dictionary for adults' responses support for the generality of evaluation scores was more strong than that for activity scores. For the ratings by children the generality of both evaluation and activity dimensions received strong support.

Adult↗

Evaluation of the DEFINDER system for fully automatic glossary construction.

In this paper we present a quantitative and qualitative evaluation of DEFINDER, a rule-based system that mines consumer-oriented full text articles in order to extract definitions and the terms they define. The quantitative evaluation shows that in terms of precision and recall as measured against human performance, DEFINDER obtained 87% and 75% respectively, thereby revealing the incompleteness of existing resources and the ability of DEFINDER to address these gaps. Our basis for comparison is definitions from on-line dictionaries, including the UMLS Metathesaurus. Qualitative evaluation shows that the definitions extracted by our system are ranked higher in terms of user-centered criteria of usability and readability than are definitions from on-line specialized dictionaries. The output of DEFINDER can be used to enhance these dictionaries. DEFINDER output is being incorporated in a system to clarify technical terms for non-specialist users in understandable non-technical language.

Dictionaries, Medical as Topic↗

SESAM: a relational database for structure and sequence of macromolecules.

A system is described that provides ways of integrating data on protein structure, sequence, and survey results, with molecular graphics and molecular mechanics software. Its major component is the relational database SESAM, presently implemented under the commercial package SYBASE. By design, the database allows full integration--within the same data organization--of raw data on protein structure, sequence, ligands, and heterogroups, obtained from the Brookhaven Protein Databank, with pure sequence information available from other databanks such as SWISS-PROT. It contains in addition higher level descriptions of structural and topological properties, as well as survey results, obtained by executing specialized computer programs. Aside from the very useful attribute of closely combining structural and nonstructural information, other important features distinguish it from analogous systems developed elsewhere. It includes a molecular dictionary with complete description of geometric properties and energy parameters used in modeling and conformational energy calculations. Using this dictionary, structural data are validated by checking for localized inconsistencies in atomic coordinates, atomic symbols, chirality definitions, and flagging errors and incomplete entries. Because of both the dictionary and the validation procedures, SESAM can be readily interfaced with conventional molecular graphics and mechanics software packages, or with other specialized application programs. With the aid of appropriate interfaces, data access is sufficiently fast for SESAM to be interrogated interactively. Prototypes of user interfaces, as well as an interface with the molecular graphics package BRUGEL, are described and the power of the system is illustrated in applications such as homology-based protein modeling, computer-aided protein design, protein structure predictions, analysis of local structure motifs, and of relationships between protein sequence and structure.

Amino Acid Sequence↗

Forms control and error detection procedures used at the Coordinating Center of the Multiple Risk Factor Intervention Trial (MRFIT).

Although methods used for data collection and quality assurance for large-scale clinical trials are important to critical reading of trial results and have been published, such reporting is the exception rather than the rule. In the MRFIT, systematic methods for processing large volumes of data over a long period of time were developed. The methods were designed to detect and control a variety of errors and to leave a complete audit trial of the processing of forms and corrections to forms. Many of these methods evolved and were refined during the course of the study as a result of trial and error. If one were to start over, the methods described herein would be modified. The field of data processing is evolving, and it is important for statistical and data processing staff of coordinating centers to recognize this and continually evaluate and update their methods. For example, the simultaneous entry and computer editing of forms is becoming more feasible with time. Also, more sophisticated intelligent data entry equipment is available for central use. Near the end of MRFIT, some data received at the Coordinating Center were entered and edited on a minicomputer. The parameter-driven edits described previously were performed at the time of data entry. Additional modifications to the content of the data dictionary for future studies are also being considered. The incorporation into the data dictionary of consistency checks (both deterministic and probabilistic) between fields on different forms would facilitate the specification of complex edit checks and would provide better documentation of the edit checks actually performed. Incorporating definitions of the numeric codes for each field would improve the documentation and facilitate reporting using statistical packages. Dedicated computer hardware should also be a major consideration of coordinating centers in future clinical trials. For MRFIT, a dedicated system was used from 1978 to the end of the trial. With the continued decline in hardware costs, dedicated systems can and should be considered, even for trials much smaller than MRFIT. We believe the system developed for processing data in the MRFIT has several advantages. It satisfies the requirements identified by Karrison or a system of data editing and control, it is largely self-documenting as a result of the data dictionary approach taken, and it is easily adaptable to other clinical studies.

Clinical Trials as Topic↗