Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dictionary”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Gene and protein nomenclature in public databases.

BACKGROUND: Frequently, several alternative names are in use for biological objects such as genes and proteins. Applications like manual literature search, automated text-mining, named entity identification, gene/protein annotation, and linking of knowledge from different information sources require the knowledge of all used names referring to a given gene or protein. Various organism-specific or general public databases aim at organizing knowledge about genes and proteins. These databases can be used for deriving gene and protein name dictionaries. So far, little is known about the differences between databases in terms of size, ambiguities and overlap. RESULTS: We compiled five gene and protein name dictionaries for each of the five model organisms (yeast, fly, mouse, rat, and human) from different organism-specific and general public databases. We analyzed the degree of ambiguity of gene and protein names within and between dictionaries, to a lexicon of common English words and domain-related non-gene terms, and we compared different data sources in terms of size of extracted dictionaries and overlap of synonyms between those. The study shows that the number of genes/proteins and synonyms covered in individual databases varies significantly for a given organism, and that the degree of ambiguity of synonyms varies significantly between different organisms. Furthermore, it shows that, despite considerable efforts of co-curation, the overlap of synonyms in different data sources is rather moderate and that the degree of ambiguity of gene names with common English words and domain-related non-gene terms varies depending on the considered organism. CONCLUSION: In conclusion, these results indicate that the combination of data contained in different databases allows the generation of gene and protein name dictionaries that contain significantly more used names than dictionaries obtained from individual data sources. Furthermore, curation of combined dictionaries considerably increases size and decreases ambiguity. The entries of the curated synonym dictionary are available for manual querying, editing, and PubMed- or Google-search via the ProThesaurus-wiki. For automated querying via custom software, we offer a web service and an exemplary client application.

Abstracting and Indexing↗

Towards semantic integration within an existing medical information system.

Talking about the problems of integration in medical information systems, the necessity to provide end users with a consistent and coherent view of patient's data, has been largely reported. In order to attempt this goal, systems need to perform semantic integration. We propose a pragmatic way to describe the semantics of the elements of a database, based on a bottom-up three steps process: 1. a back documentation of the elements of the system from their description contained in the data catalog of the database 2. a first semantic extension to transform a data catalog into a data dictionary 3. a second semantic extension to create a dictionary of the medical concepts from a data dictionary. This dictionary of concepts can be considered as the final result of "semantic integration". It contains a set of entities directly understandable by the end users. These entities are deduced or built from the elements collected and characterized in the data dictionary. This work reports the conceptual analysis and the implementation of such a data dictionary.

Database Management Systems↗

Building a controlled health vocabulary in Japanese.

OBJECTIVES: This study is aimed at developing a controlled clinical vocabulary for use in electronic patient record (EPR) systems. METHODS: In this paper, we propose a model for building the vocabulary. The model is composed of a Canonical Term Dictionary, an Atom Dictionary, a Composite Atom Dictionary, and an Index. Parsing and composing functions are included in this model. Canonical terms were extracted from reference terminologies. Atoms were extracted from the Canonical Term Dictionary and reduced to a set from which the Composite Atom Dictionary can be built. The index was built to link these two dictionaries. For testing the model, we compiled a sample vocabulary and applied the model to a SNOMED translation system (English to Japanese) and a term similarity estimation system. RESULTS: The sample vocabulary consisted of 15,600 atomic terms and 4,450 composite terms. 33,441 SNOMED terms were translated by the SNOMED translation system. The system gave adequate Japanese candidates in 56.3% of cases. The similarity estimation system found an average of 5.4 candidates when the equality ratio was over 50%. CONCLUSIONS: The trial applications produced good results. The model seems promising for building a standard clinical vocabulary system. This system can be applied in certain other Asian countries, such as China and Korea.

Algorithms↗

New parameters for the refinement of nucleic acid-containing structures.

Structures at atomic resolution (up to 1.0 A) which contain bases, sugars or the phosphodiester linkage, were selected from the Nucleic Acid Database or the Cambridge Structural Database to build a nucleic acid dictionary from X-ray refined structures. The dictionary consists of the average values for bond distances, bond angles and dihedral angles. The variance of the sample is used to provide information about the expected r.m.s. deviations of the refined parameters. A dictionary was constructed for refinement trials in X-PLOR. The dictionary includes RNA and DNA in C2'-endo and C3'-endo sugar pucker conformations, as well as values for the backbone dihedrals. Tests were performed on the dictionary using three structures: a B-DNA, a Z-DNA and a protein-DNA complex. During the course of refinement, all three structures showed significant improvements as measured by r.m.s. deviations and R factors when compared to the previous DNA dictionary.

Journal Article↗

Utilizing weakly controlled vocabulary for sentence segmentation in biomedical literature.

Since biomedical texts contain a wide variety of domain specific terms, building a large dictionary to perform term matching is of great relevance. However, due to the existence of null boundary between adjacent terms, this matching is not a trivial problem. Moreover, it is known that generative words cannot be comprehensively included in a dictionary because their possible variations are infinite. In this study, we report our approach to dictionary building and term matching in biomedical texts. Large amount of terms with/without part-of-speech (POS) and/or category information were gathered, and a completion program generated approximately 1.36 million term variants to avoid stemming problems when matching terms. The dictionary was stored in a relational database management system (RDBMS) for quick lookup, and used by a matching program. Since the matching operation is not restricted to a substring surrounded by space characters, we can avoid the problem of null boundaries. This feature is also useful for generative words. Experimental results on GENIA corpus are promising: nearly half of the possible terms were correctly recognized as a meaningful segment, and most of the remaining half could be correctly recognized by some post-processing process, like chunking and further decomposition. It should be remarked that although we have not used term cost, connectivity cost, or syntactic information, reasonable segmentation and dictionary lookup were performed in most cases.

Abstracting and Indexing↗

Less is more: towards an optimal universal description of protein folds.

MOTIVATION: Identification and characterization of protein structure regularities can reveal the mechanisms governing protein structure, function and evolution. Here we focus on an intermediate level of regularity. We have developed automated methods to systematically construct a dictionary of supersecondary structures that can be used as 'protein parts' to describe fold-sized structures. RESULTS: The dictionary was constructed by aligning representative structures of all known folds, clustering similar substructures and selecting the most descriptive substructures in a minimum description length fashion. We show that the dictionary is compact and descriptive, capable of describing a substantial fraction of all known protein folds. We performed simulations using independent sets of training and testing folds. Dictionaries generated using the training set had high coverage over the folds in the testing set, suggesting that dictionary entries reflect general features of protein structures and should be capable of describing novel protein folds.

Algorithms↗

Continuous speech recognition for clinicians.

The current generation of continuous speech recognition systems claims to offer high accuracy (greater than 95 percent) speech recognition at natural speech rates (150 words per minute) on low-cost (under $2000) platforms. This paper presents a state-of-the-technology summary, along with insights the authors have gained through testing one such product extensively and other products superficially. The authors have identified a number of issues that are important in managing accuracy and usability. First, for efficient recognition users must start with a dictionary containing the phonetic spellings of all words they anticipate using. The authors dictated 50 discharge summaries using one inexpensive internal medicine dictionary ($30) and found that they needed to add an additional 400 terms to get recognition rates of 98 percent. However, if they used either of two more expensive and extensive commercial medical vocabularies ($349 and $695), they did not need to add terms to get a 98 percent recognition rate. Second, users must speak clearly and continuously, distinctly pronouncing all syllables. Users must also correct errors as they occur, because accuracy improves with error correction by at least 5 percent over two weeks. Users may find it difficult to train the system to recognize certain terms, regardless of the amount of training, and appropriate substitutions must be created. For example, the authors had to substitute "twice a day" for "bid" when using the less expensive dictionary, but not when using the other two dictionaries. From trials they conducted in settings ranging from an emergency room to hospital wards and clinicians' offices, they learned that ambient noise has minimal effect. Finally, they found that a minimal "usable" hardware configuration (which keeps up with dictation) comprises a 300-MHz Pentium processor with 128 MB of RAM and a "speech quality" sound card (e.g., SoundBlaster, $99). Anything less powerful will result in the system lagging behind the speaking rate. The authors obtained 97 percent accuracy with just 30 minutes of training when using the latest edition of one of the speech recognition systems supplemented by a commercial medical dictionary. This technology has advanced considerably in recent years and is now a serious contender to replace some or all of the increasingly expensive alternative methods of dictation with human transcription.

Dictionaries, Medical as Topic↗

Code generation through annotation of macromolecular structure data.

The maintenance of software which uses a rapidly evolving data annotation scheme is time consuming and expensive. At the same time without current software the annotation scheme itself becomes limited and is less likely to be widely adopted. A solution to this problem has been developed for the macromolecular Crystallographic Information File (mmCIF) annotation scheme. The approach could be generalized for a variety of annotation schemes used or proposed for molecular biology data. mmCIF provides a highly structured and complete annotation for describing NMR and X-ray crystallographic data and the resulting macromolecular structures. This annotation is maintained in the mmCIF dictionary which currently contains over 3,200 terms. A major challenge is to maintain code for converting between mmCIF and Protein Data Bank (PDB) annotations while both continue to evolve. The solution has been to define a simple domain specific language (DSL) which is added to the extensive annotation already found in the mmCIF dictionary. The DSL calls specific mapping modules for each category of data item in the mmCIF dictionary. Adding or changing the mapping between PDB and mmCIF items of data is straightforward since data categories (and hence mapping modules) correspond to elements of macromolecular structure familiar to the experimentalist. Each time a change is made to the macromolecular annotation the appropriate change is made to the easily located and modifiable mapping modules. A code generator is then called which reads the mapping modules and creates a new executable for performing the data conversion. In this way code is easily kept current by individuals with limited programming skill, but who have an understanding of macromolecular structure and details of the annotation scheme. Most important, the conversion process becomes part of the global dictionary and is not open to a variety of interpretations by different research groups writing code based on dictionary contents. Details of the DSL and code generator are provided.

Crystallography, X-Ray↗

Automated coded ambulatory problem lists: evaluation of a vocabulary and a data entry tool.

BACKGROUND: Problem lists are fundamental to electronic medical records (EMRs). However, obtaining an appropriate problem list dictionary is difficult, and getting users to code their problems at the time of data entry can be challenging. OBJECTIVE: To develop a problem list dictionary and search algorithm for an EMR system and evaluate its use. METHODS: We developed a problem list dictionary and lookup tool and implemented it in several EMR systems. A sample of 10,000 problem entries was reviewed from each system to assess overall coding rates. We also performed a manual review of a subset of entries to determine the appropriateness of coded entries, and to assess the reasons other entries were left uncoded. RESULTS: The overall coding rate varied significantly between different EMR implementations (63-79%). Coded entries were virtually always appropriate (99%). The most frequent reasons for uncoded entries were due to user interface failures (44-45%), insufficient dictionary coverage (20-32%), and non-problem entries (10-12%). CONCLUSION: The problem list dictionary and search algorithm has achieved a good coding rate, but the rate is dependent on the specific user interface implementation. Problem coding is essential for providing clinical decision support, and improving usability should result in better coding rates.

Algorithms↗

Mass data massage: an automated data processing system used for NHEXAS, Arizona. National Human Exposure Assessment Survey.

Data entry and management are critical components of all large survey projects; data quality objectives must be met and data must be quickly and readily accessible. We developed a comprehensive system for data entry and management utilizing scannable forms with bubble fields and handwriting recognition. This 'Mass Data Massage' (MDM) system had three components: (1) form creation and database definition; (2) programming of data dictionaries for documentation and preliminary logic and range checks; and (3) data entry, management and documentation using the 'Mass Data Cleaning Program' (MDCP). Scannable forms were written in Teleform, where the data field definition, variable names and ranges were defined as the form was created. Completed forms were returned from the field, subjected to final field quality control (QC) checks, and transferred to the data management section. They were batched and coded as necessary. Once a batch of data was scanned and visually verified, the operator called up the menu for the MDCP. The MDCP had 31 program modules with 500-1200 lines of code each. The operator could select and run the appropriate dictionary on each data batch 'correcting' apparent errors in responses. This process was iterative until the data batch passed all dictionary checks. Proposed 'changes' were forwarded to the data coordinator (DC) for acceptance or rejection. After all errors had been resolved, each data batch was subjected to a 10% quality assurance (QA) check. The original data batch and associated file of applied changes were archived. Time expenditure using the scanning approach varied with the number of questions and the types of responses (handwritten or bubble fields). One-page forms took 42-60% of the time needed for hand entry; forms longer than 10 pages took 35-38% of the time. Use of faster machines will further speed the process. The main advantage of the system was the reduction of systematic errors. Scanning alone reduced errors found on 995 NHEXAS Baseline Questionnaires. Overall, the dictionary identified 0.55% errors on the scanned forms. Ten percent QC checks, performed on corrected batches ready for appendage to the master database, revealed an overall error rate of 0.02%. Similar checks on a laboratory form scanned from numeric handwriting detected 0.3% errors following dictionary application and 0.2% errors during the 10% QA check. This system was faster, more accurate, and more cost-effective than hand entry of data. A batch of data that took >1 week to process using the hand entry method was processed within 1 day using MDM. Human coding of specific answers and the final verification were the most time-consuming processes.

Arizona↗

MDB: a database system utilizing automatic construction of modules and STAR-derived universal language.

MOTIVATION: The value of information greatly increases if stored in databases. The objective was to construct a multi-purpose database system primarily designed to store and provide access to three-dimensional structures of biological molecules including theoretical models. RESULTS: A dictionary defining data format and structure for three-dimensional models of biological molecules (MDB dictionary) was developed. The dictionary was written using universal, standardized data description language. This language can be applied to describe data with no restrictions on their origin or type, including metadata. Thus both the data definitions (format) and database descriptions are created using the uniform language and processed with universal software. A database and data design technique that allowed use of dictionaries to automatically construct relational databases was developed. This technique was employed to construct the MDB database system. Data design developed and applied in the MDB project makes it possible to carry out data curation utilizing the database engine to identify errors. It also allows storage and query of data at different levels of consistency with the standard format specifications, i.e. both the correctly formatted data, and data that requires further curation. AVAILABILITY: The MDB dictionary is available at http://www.gwer.ch/proteinstructure/mdb and as part of the PDB resources at http://pdb.rutgers.edu/mmcif/.

Computational Biology↗

On optimizing syntactic pattern recognition using tries and AI-based heuristic-search strategies.

This paper deals with the problem of estimating, using enhanced artificial-intelligence (AI) techniques, a transmitted string X* by processing the corresponding string Y, which is a noisy version of X*. It is assumed that Y contains substitution, insertion, and deletion (SID) errors. The best estimate X+ of X* is defined as that element of a dictionary H that minimizes the generalized Levenshtein distance (GLD) D (X, Y) between X and Y, for all X epsilon H. In this paper, it is shown how to evaluate D (X, Y) for every X epsilon H simultaneously, when the edit distances are general and the maximum number of errors is not given a priori, and when H is stored as a trie. A new scheme called clustered beam search (CBS) is first introduced, which is a heuristic-based search approach that enhances the well-known beam-search (BS) techniques used in AI. The new scheme is then applied to the approximate string-matching problem when the dictionary is stored as a trie. The new technique is compared with the benchmark depth-first search (DFS) trie-based technique (with respect to time and accuracy) using large and small dictionaries. The results demonstrate a marked improvement of up to 75% with respect to the total number of operations needed on three benchmark dictionaries, while yielding an accuracy comparable to the optimal. Experiments are also done to show the benefits of the CBS over the BS when the search is done on the trie. The results also demonstrate a marked improvement (more than 91%) for large dictionaries.

Algorithms↗

Evaluation of controlled vocabulary resources for development of a consumer entry vocabulary for diabetes.

BACKGROUND: Digital information technology can facilitate informed decision making by individuals regarding their personal health care. The digital divide separates those who do and those who do not have access to or otherwise make use of digital information. To close the digital divide, health care communications research must address a fundamental issue, the consumer vocabulary problem: consumers of health care, at least those who are laypersons, are not always familiar with the professional vocabulary and concepts used by providers of health care and by providers of health care information, and, conversely, health care and health care information providers are not always familiar with the vocabulary and concepts used by consumers. One way to address this problem is to develop a consumer entry vocabulary for health care communications. OBJECTIVES: To evaluate the potential of controlled vocabulary resources for supporting the development of consumer entry vocabulary for diabetes. METHODS: We used folk medical terms from the Dictionary of American Regional English project to create extended versions of 3 controlled vocabulary resources: the Unified Medical Language System Metathesaurus, the Eurodicautom of the European Commission's Translation Service, and the European Commission Glossary of popular and technical medical terms. We extracted consumer terms from consumer-authored materials, and physician terms from physician-authored materials. We used our extended versions of the vocabulary resources to link diabetes-related terms used by health care consumers to synonymous, nearly-synonymous, or closely-related terms used by family physicians. We also examined whether retrieval of diabetes-related World Wide Web information sites maintained by nonprofit health care professional organizations, academic organizations, or governmental organizations can be improved by substituting a physician term for its related consumer term in the query. RESULTS: The Dictionary of American Regional English extension of the Metathesaurus provided coverage, either direct or indirect, of approximately 23% of the natural language consumer-term-physician-term pairs. The Dictionary of American Regional English extension of the Eurodicautom provided coverage for 16% of the term pairs. Both the Metathesaurus and the Eurodicautom indirectly related more terms than they directly related. A high percentage of covered term pairs, with more indirectly covered pairs than directly covered pairs, might be one way to make the most out of expensive controlled vocabulary resources. We compared retrieval of diabetes-related Web information sites using the physician terms to retrieval using related consumer terms We based the comparison on retrieval of sites maintained by non-profit healthcare professional organizations, academic organizations, or governmental organizations. The number of such sites in the first 20 Results from a search was increased by substituting a physician term for its related consumer term in the query. This suggests that the Dictionary of American Regional English extensions of the Metathesaurus and Eurodicautom may be used to provide useful links from natural language consumer terms to natural language physician terms. CONCLUSIONS: The Dictionary of American Regional English extensions of the Metathesaurus and Eurodicautom should be investigated further for support of consumer entry vocabulary for diabetes.

Community Participation↗

[The concept of nursing. A systematic analysis].

Nursing is a field which employs numerous people whose formation and professional roles are regulated by a wide set of rules and norms, and one which benefits innumerable patients. A paradox exists when, upon consulting various dictionaries, one of these the Royal Academy Dictionary (Diccionario de la Real Academia), the term nursing is defined as a physical space/set of nurses without any reference whatsoever to nurses' role in the health and educational systems. The definition we find in a dictionary should correspond to the concept the general public holds and reflect the true meaning of the profession. During a three year period, academic years 1994-96, on the first day of class, we asked each student to define or describe what he/she understood nursing to be. We consider that their responses, in a large sense, should correspond to the idea which the general public holds about nursing. An analysis of the content of these definitions allowed us to establish that there is a clear confrontation between what dictionaries state and what students, as a sample of society as a whole, think. Our results permit us to offer a definition for nursing which may be incorporated into new dictionary editions for the purpose of completing the already existing ones.

Attitude of Health Personnel↗

[Old English plant names from the linguistic and lexicographic viewpoint].

Roughly 1350 Old English plant names have come down to us; this is a relatively large number considering that the attested Old English vocabulary comprises ca. 24 000 words. The plant names are not only interesting for botanists, historians of medicine and many others, but also for philologists and linguists; among other aspects they can investigate their etymology, their morphology (including word-formation) and their meaning and motivation. Practically all Old English texts where plant names occur have been edited (including glosses and glossaries), the names have been listed in the Old English dictionaries, and some specific studies have been devoted to them. Nevertheless no comprehensive systematic analysis of their linguistic structure has been made. Ulrike Krischke is preparing such an analysis. A proper dictionary of the Old English plant names is also a desideratum, especially since the Old English dictionaries available and in progress normally do not deal with morphological and semantic aspects, and many do not provide etymological information. A plant-name dictionary concentrating on this information is being prepared by Hans Sauer and Ulrike Krischke. In our article here, we sketch the state of the art (ch. 1), we deal with some problems of the analysis of Old English plant names (ch. 2), e.g. the delimitation of the word-field plant names, the identification of the plants, errors and problematic spellings in the manuscripts. In ch. 3 we sketch the etymological structure according to chronological layers (Indo-European, Germanic, West-Germanic, Old English) as well as according to the distinction between native words and loan-words; in the latter category, we also mention loan-formations based on Latin models. In ch. 4 we survey the morphological aspects (simplex vs. complex words); among the complex nouns, compounds are by far the largest group (and among those, the noun + noun compounds), but there are also a few suffix formations. We also briefly present some morphological peculiarities, e.g. formations with blocked (unique) morphemes, the question of homonyms, cases of obscuration and of popular etymology. In ch. 5 we outline semantic structures, and in ch. 6, we introduce the structure of the proposed dictionary of the Old English plant names, also providing several specimen entries.

Animals↗