Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dictionary”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

A hospital-wide clinical findings dictionary based on an extension of the International Classification of Diseases (ICD).

The use of a controlled vocabulary set in a hospital-wide clinical information system is of crucial importance for many departmental database systems to communicate and exchange information. In the absence of an internationally recognized clinical controlled vocabulary set, a new extension of the International statistical Classification of Diseases (ICD) is proposed. It expands the scope of the standard ICD beyond diagnosis and procedures to clinical terminology. In addition, the common Clinical Findings Dictionary (CFD) further records the definition of clinical entities. The construction of the vocabulary set and the CFD is incremental and manual. Tools have been implemented to facilitate the tasks of defining/maintaining/publishing dictionary versions. The design of database applications in the integrated clinical information system is driven by the CFD which is part of the Medical Questionnaire Designer tool. Several integrated clinical database applications in the field of diabetes and neuro-surgery have been developed at the HUG.

Databases as Topic↗

Requirements and design aspects of a data model for a data dictionary in paediatric oncology.

German children suffering from cancer are mostly treated within the framework of multicentre clinical trials. An important task of conducting these trials is an extensive information and knowledge exchange, which has to be based on a standardised documentation. To support this effort, it is the aim of a nationwide project to define a standardised terminology that should be used by clinical trials for therapy documentation. In order to support terminology maintenance we are currently developing a data dictionary. In this paper we describe requirements and design aspects of the data model used for the data dictionary as first results of our research. We compare it with other terminology systems.

Child↗

Available sources of veterinary biographies useful in compiling contributions for an International Dictionary of Veterinary Biography(IDVB).

In the framework of the WAHVM-project to compile an International Dictionary of Veterinary Biography, two modern tools are described that may assist the contributors to this dictionary in finding references concerning the persons of their choice. The first is the World Biographical Index (WBI), an initiative of K. G. Saur Publishing, that can be consulted in print and online as well. The index covers more than 5000 sources of biographical information of all times and regions of the world. Its 7th online-edition (www.biblio.tu-bs/de/wbi) identifies the works where information can be found on the lives of 4551 veterinarians, for the greater part of the past. The editions are regularly expanded. The second tool is a CD-ROM, published by the WAHVM, with 13.000 references to books, periodical articles and dissertations on veterinary history. More than 40% of these references have biographical relevance. The index on personal names contains more than 2700 names. Both tools are complementary because the source materials differ largely.

Biographies as Topic↗

Integration of nursing assessment concepts into the medical entities dictionary using the LOINC semantic structure as a terminology model.

Recent investigations have tested the applicability of various terminology models for the representing nursing concepts including those related to nursing diagnoses, nursing interventions, and standardized nursing assessments as a prerequisite for building a reference terminology that supports the nursing domain. We used the semantic structure of Clinical LOINC (Logical Observations, Identifiers, Names, and Codes) as a reference terminology model to support the integration of standardized assessment terms from two nursing terminologies into the Medical Entities Dictionary (MED), the concept-oriented, metadata dictionary at New York Presbyterian Hospital. Although the LOINC semantic structure was used previously to represent laboratory terms in the MED, selected hierarchies and semantic slots required revisions in order to incorporate the nursing assessment concepts. This project was an initial step in integrating nursing assessment concepts into the MED in a manner consistent with evolving standards for reference terminology models. Moreover, the revisions provide the foundation for adding other types of standardized assessments to the MED.

Dictionaries, Medical as Topic↗

On the cultural unity of Europe and a European dictionary project.

There is a cultural unity of Europe, notwithstanding the Reformation. This unity must be 'consciously' recognized, in order to develop the concept of European togetherness. The paper will examine the common roots of the European vocabulary, submitting the project for the compilation of some dictionaries, to which European colleagues will be called to contribute. European languages must be studied no more as 'national' languages, but in the context of the European tradition, emphasizing what they have in common, and not what separates them. Hence the necessity of compiling dictionaries on a new basis, of which early, summary indications are given.

Anthropology, Cultural↗

Word segmentation processing: a way to exponentially extend medical dictionaries.

One of the most critical problems of automatic natural language processing (NLP) is the size of the medical lexicons. The set of compound medical words and the continual creation of new terms renders medical lexicons exhaustive beyond question. The structure of such dictionaries usually consists of two parts: 1) the morphological and sometimes syntactical information necessary to identify, on a grapheme level, a given word in a sentence, and 2) the part often devoted to conceptual knowledge associated with the recognized word. It is only when these two prerequisites are fulfilled that an attempt to understand the meaning of a whole expression is possible. The approach developed in this paper is a pragmatic way to rapidly increase the lexico-semantic part of medical dictionaries. We developed a semi-automatic tool, as a prototype to demonstrate the feasibility of this approach. This tool is able to translate almost any diagnosis expressed in French into its equivalent in the ICD-9CM coding scheme.

Dictionaries, Medical as Topic↗

Quality control methods for data entry in pathology using a computerized data management system based on an extended data dictionary.

In pathology, computerized data management systems have been used increasingly to facilitate a more efficient supply of information. Since data entry precedes data utilization, the reliability of the information stored strongly depends on the quality of data input. Despite its potential capability, most personal computer-based database software does not provide versatile and user-friendly data validation procedures. Therefore, we developed a data dictionary-driven data management system that enables the user to perform extensive validation routines without the need for hard programming. Using examples from an existing database for endometrial carcinomas, different types of data errors and their error traps are explained. It is pointed out that data type definitions, defaults, templates, or picture clauses are suitable means to avoid formal errors. Validations on data domains and ranges test whether data fall into a predefined scope. Relational checks control data validity within a context of different data items, whereas process routines provide automatic data computation, thereby circumventing user input. By exploiting the facilities of an extended data dictionary, a powerful tool is made available to secure various aspects of data integrity simultaneously with input. In this way, computerized data quality control can improve the efficiency and reliability of data management tasks in pathology.

Medical Informatics Computing↗

The HLA dictionary 2001: a summary of HLA-A, -B, -C, -DRB1/3/4/5, -DQB1 alleles and their association with serologically defined HLA-A, -B, -C, -DR, and -DQ antigens.

This report presents the serologic equivalents of 123 HLA-A, 272 HLA-B, and 155 HLA-DRB1 alleles. The equivalents cover over 64 percent of the presently identified HLA-A, -B, and -DRB1 alleles. The dictionary is an update of the one published in 1999 (Schreuder GMTh, Hurley CK, Marsh SGE, Lau M, Maiers M, Kollman C, Noreen H. The HLA dictionary 1999: a summary of HLA-A, -B, -C, -DRB1/3/4/5, -DQB1 alleles and their association with serologically defined HLA-A, -B, -C, -DR and -DQ antigens. Tissue Antigens 54:407, 1999) and also includes equivalents for HLA-C, DRB3, DRB4, DRB5, and DQB1 alleles. The data summarize information obtained by the WHO Nomenclature Committee for Factors of the HLA System, the International Cell Exchange (UCLA), the National Marrow Donor Program (NMDP), and individual laboratories. In addition, a listing is provided of alleles which are expressed as antigens with serologic reaction patterns that differ from the well-established HLA specificities. The equivalents provided will be useful in guiding searches for unrelated hematopoietic stem cell donors in which patients and/or potential donors are typed by either serology or DNA-based methods. These equivalents will also serve typing and matching procedures for organ transplant programs where HLA typings from donors and from recipients on waiting lists represent mixtures of serologic and molecular typings. The tables with HLA equivalents and a questionnaire for submission of serologic reaction patterns for poorly identified allelic products will also be available on the WMDA web page: www.worldmarrow.org.

Alleles↗

Optimally sparse representation in general (nonorthogonal) dictionaries via l minimization.

Given a dictionary D = {d(k)} of vectors d(k), we seek to represent a signal S as a linear combination S = summation operator(k) gamma(k)d(k), with scalar coefficients gamma(k). In particular, we aim for the sparsest representation possible. In general, this requires a combinatorial optimization process. Previous work considered the special case where D is an overcomplete system consisting of exactly two orthobases and has shown that, under a condition of mutual incoherence of the two bases, and assuming that S has a sufficiently sparse representation, this representation is unique and can be found by solving a convex optimization problem: specifically, minimizing the l(1) norm of the coefficients gamma. In this article, we obtain parallel results in a more general setting, where the dictionary D can arise from two or several bases, frames, or even less structured systems. We sketch three applications: separating linear features from planar ones in 3D data, noncooperative multiuser encoding, and identification of over-complete independent component models.

Journal Article↗

Building a dictionary for genomes: identification of presumptive regulatory sites by statistical analysis.

The availability of complete genome sequences and mRNA expression data for all genes creates new opportunities and challenges for identifying DNA sequence motifs that control gene expression. An algorithm, "MobyDick," is presented that decomposes a set of DNA sequences into the most probable dictionary of motifs or words. This method is applicable to any set of DNA sequences: for example, all upstream regions in a genome or all genes expressed under certain conditions. Identification of words is based on a probabilistic segmentation model in which the significance of longer words is deduced from the frequency of shorter ones of various lengths, eliminating the need for a separate set of reference data to define probabilities. We have built a dictionary with 1,200 words for the 6, 000 upstream regulatory regions in the yeast genome; the 500 most significant words (some with as few as 10 copies in all of the upstream regions) match 114 of 443 experimentally determined sites (a significance level of 18 standard deviations). When analyzing all of the genes up-regulated during sporulation as a group, we find many motifs in addition to the few previously identified by analyzing the subclusters individually to the expression subclusters. Applying MobyDick to the genes derepressed when the general repressor Tup1 is deleted, we find known as well as putative binding sites for its regulatory partners.

Algorithms↗

PNAD-CSS: a workbench for constructing a protein name abbreviation dictionary.

MOTIVATION: Since their initial development, integration and construction of databases for molecular-level data have progressed. Though biological molecules are related to each other and form a complex system, the information is stored in the vast archives of the literature or in diverse databases. There is no unified naming convention for biological object, and biological terms may be ambiguous or polysemic. This makes the integration and interaction of databases difficult. In order to eliminate these problems, machine-readable natural language resources appear to be quite promising. We have developed a workbench for protein name abbreviation dictionary (PNAD) building. RESULTS: We have developed PNAD Construction Support System (PNAD-CSS), which offers various convenient facilities to decrease the construction costs of a protein name abbreviation dictionary of which entries are collected from abstracts in biomedical papers. The system allows the users to concentrate on higher level interpretation by removing some troublesome tasks, e.g. management of abstracts, extracting protein names and their abbreviations, and so on. To extract a pair of protein names and abbreviations, we have developed a hybrid system composed of the PROPER System and the PNAD System. The PNAD System can extract the pairs from parenthetical-paraphrases involved in protein names, the PROPER System identified these paris, with 98.95% precision, 95.56% recall and 97.58% complete precision. AVAILABILITY: PROPER System is freely available from http://www.hgc.inc.u-tokyo.ac.jp/service/tooldoc /KeX/intro.html. The other software are also available on request. Contact the authors. CONTACT: mikio@ims.u-tokyo.ac.jp

Databases, Factual↗

Vocabulon: a dictionary model approach for reconstruction and localization of transcription factor binding sites.

MOTIVATION: Gene expression arrays enable measurements of transcription values for a large number or all genes in the genome. In order to better interpret these results and to use them to reconstruct transcription networks, information on location of binding sites for regulatory proteins in the entire genome is needed. In particular, this represents an open problem in Escherichia coli. RESULTS: We describe the first implementation of dictionary-style models to the study of transcription factors binding sites in an entire genome. Vocabulon's unique feature is that it can both reconstruct binding sites characterized by unknown motifs and impute locations of known binding sites in long sequences by simultaneous search. On one hand, the dictionary model specifies a probability for the entire sequence taking simultaneously into account all the possible binding sites. This greatly reduces the number of false positives. On the other hand, the possibility of refining motif description, as an increasing number of binding sites are identified, augments the sensitivity of the method. We illustrate these properties with examples in E.coli. The results of gene expression arrays are used both to guide the search and corroborate it.

Algorithms↗

Dictionary-driven prokaryotic gene finding.

Gene identification, also known as gene finding or gene recognition, is among the important problems of molecular biology that have been receiving increasing attention with the advent of large scale sequencing projects. Previous strategies for solving this problem can be categorized into essentially two schools of thought: one school employs sequence composition statistics, whereas the other relies on database similarity searches. In this paper, we propose a new gene identification scheme that combines the best characteristics from each of these two schools. In particular, our method determines gene candidates among the ORFs that can be identified in a given DNA strand through the use of the Bio-Dictionary, a database of patterns that covers essentially all of the currently available sample of the natural protein sequence space. Our approach relies entirely on the use of redundant patterns as the agents on which the presence or absence of genes is predicated and does not employ any additional evidence, e.g. ribosome-binding site signals. The Bio-Dictionary Gene Finder (BDGF), the algorithm's implementation, is a single computational engine able to handle the gene identification task across distinct archaeal and bacterial genomes. The engine exhibits performance that is characterized by simultaneous very high values of sensitivity and specificity, and a high percentage of correctly predicted start sites. Using a collection of patterns derived from an old (June 2000) release of the Swiss-Prot/TrEMBL database that contained 451 602 proteins and fragments, we demonstrate our method's generality and capabilities through an extensive analysis of 17 complete archaeal and bacterial genomes. Examples of previously unreported genes are also shown and discussed in detail.

Algorithms↗

WordSpy: identifying transcription factor binding motifs by building a dictionary and learning a grammar.

Transcription factor (TF) binding sites or motifs (TFBMs) are functional cis-regulatory DNA sequences that play an essential role in gene transcriptional regulation. Although many experimental and computational methods have been developed, finding TFBMs remains a challenging problem. We propose and develop a novel dictionary based motif finding algorithm, which we call WordSpy. One significant feature of WordSpy is the combination of a word counting method and a statistical model which consists of a dictionary of motifs and a grammar specifying their usage. The algorithm is suitable for genome-wide motif finding; it is capable of discovering hundreds of motifs from a large set of promoters in a single run. We further enhance WordSpy by applying gene expression information to separate true TFBMs from spurious ones, and by incorporating negative sequences to identify discriminative motifs. In addition, we also use randomly selected promoters from the genome to evaluate the significance of the discovered motifs. The output from WordSpy consists of an ordered list of putative motifs and a set of regulatory sequences with motif binding sites highlighted. The web server of WordSpy is available at http://cic.cs.wustl.edu/wordspy.

Algorithms↗

Mapping transcriptional responses to cellular perturbation dictionaries with RNA fingerprinting.

Single-cell perturbation dictionaries provide systematic measurements of how cells respond to genetic and chemical perturbations, and create the opportunity to assign causal interpretations to observational data. Here, we introduce RNA fingerprinting, a statistical framework that maps transcriptional responses from new experiments onto reference perturbation dictionaries. RNA fingerprinting learns denoised perturbation "fingerprints" from single-cell data, then probabilistically assigns query cells to one or more candidate perturbations while accounting for uncertainty. We benchmark our method across ground-truth datasets, demonstrating accurate assignments at single-cell resolution, scalability to genome-wide screens, and the ability to resolve combinatorial perturbations. We demonstrate its broad utility across diverse biological settings: identifying context-specific regulators of p53 under ribosomal stress, characterizing drug mechanisms of action and dose-dependent off-target effects, and uncovering cytokine-driven B cell heterogeneity during secondary influenza infection in vivo. Together, these results establish RNA fingerprinting as a versatile framework for interpreting single-cell datasets by linking cellular states to the underlying perturbations which generated them.

Journal Article↗

Image denoising via sparse and redundant representations over learned dictionaries.

We address the image denoising problem, where zero-mean white and homogeneous Gaussian additive noise is to be removed from a given image. The approach taken is based on sparse and redundant representations over trained dictionaries. Using the K-SVD algorithm, we obtain a dictionary that describes the image content effectively. Two training options are considered: using the corrupted image itself, or training on a corpus of high-quality image database. Since the K-SVD is limited in handling small image patches, we extend its deployment to arbitrary image sizes by defining a global image prior that forces sparsity over patches in every location in the image. We show how such Bayesian treatment leads to a simple and effective denoising algorithm. This leads to a state-of-the-art denoising performance, equivalent and sometimes surpassing recently published leading alternative denoising methods.

Algorithms↗

[Comments on Nigel Wiseman's A Practical Dictionary of Chinese Medicine: on the use of Western medical terms in English glossary of Chinese medicine].

Mr. Wiseman believes that Western medical terms chosen as equivalents of Chinese medical terms should be the words known to all speakers and not requiring any specialist knowledge or instrumentation to understand or identify, and strictly technical Western medical terms should be avoided regardless of their conceptual conformity to the Chinese terms. According to such criteria, many inappropriate Western medical terms are selected as English equivalents by the authors of the Dictionary, and on the other hand, many ready-made appropriate Western medical terms are replaced by loan English terms with the Chinese style of word formation. The experience obtained by translating Western medical terms into Chinese when Western medicine was first introduced to China should be helpful for developing English equivalents at present. However, the authors of the Dictionary adhere to their own opinions and reject others' experience. The English terms thus created do not reflect the genuine meaning of the Chinese terms, but make the English glossary in chaos. The so-called true face of traditional Chinese revealed by such terms is merely the Chinese custom of word formation and metaphoric rhetoric. In other words, traditional Chinese medicine is not regarded as a system of medicine but merely some Oriental folklore.

Medicine, Chinese Traditional↗

[Comments on "A practical dictionary of Chinese medicine" by Wiseman].

At least 24 Chinese-English dictionaries of Chinese Medicine have been published in China during the recent 24 years (1984-2003). This thesis comments on "A Practical Dictionary of Chinese Medicine" by Wiseman, agreeing on its establishing principles, sources and formation methods of the English system of Chinese medical terminology, and pointing out the defect. The author holds that study on the origin and development of TCM terms, standardization of Chinese medical terms in different layers, i.e. Chinese medical in classic, in commonly used modern TCM terms, and integrative medical texts, are prerequisites to the standardization of English translation of Chinese medical terms.

Book Reviews as Topic↗