Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dictionary”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Improving the performance of dictionary-based approaches in protein name recognition.

Dictionary-based protein name recognition is often a first step in extracting information from biomedical documents because it can provide ID information on recognized terms. However, dictionary-based approaches present two fundamental difficulties: (1) false recognition mainly caused by short names; (2) low recall due to spelling variations. In this paper, we tackle the former problem using machine learning to filter out false positives and present two alternative methods for alleviating the latter problem of spelling variations. The first is achieved by using approximate string searching, and the second by expanding the dictionary with a probabilistic variant generator, which we propose in this paper. Experimental results using the GENIA corpus revealed that filtering using a naive Bayes classifier greatly improved precision with only a slight loss of recall, resulting in 10.8% improvement in F-measure, and dictionary expansion with the variant generator gave further 1.6% improvement and achieved an F-measure of 66.6%.

Abstracting and Indexing↗

REFMAC5 dictionary: organization of prior chemical knowledge and guidelines for its use.

One of the most important aspects of macromolecular structure refinement is the use of prior chemical knowledge. Bond lengths, bond angles and other chemical properties are used in restrained refinement as subsidiary conditions. This contribution describes the organization and some aspects of the use of the flexible and human/machine-readable dictionary of prior chemical knowledge used by the maximum-likelihood macromolecular-refinement program REFMAC5. The dictionary stores information about monomers which represent the constitutive building blocks of biological macromolecules (amino acids, nucleic acids and saccharides) and about numerous organic/inorganic compounds commonly found in macromolecular crystallography. It also describes the modifications the building blocks undergo as a result of chemical reactions and the links required for polymer formation. More than 2000 monomer entries, 100 modification entries and 200 link entries are currently available. Algorithms and tools for updating and adding new entries to the dictionary have also been developed and are presented here. In many cases, the REFMAC5 dictionary allows entirely automatic generation of restraints within REFMAC5 refinement runs.

Chemical Phenomena↗

Building a protein name dictionary from full text: a machine learning term extraction approach.

BACKGROUND: The majority of information in the biological literature resides in full text articles, instead of abstracts. Yet, abstracts remain the focus of many publicly available literature data mining tools. Most literature mining tools rely on pre-existing lexicons of biological names, often extracted from curated gene or protein databases. This is a limitation, because such databases have low coverage of the many name variants which are used to refer to biological entities in the literature. RESULTS: We present an approach to recognize named entities in full text. The approach collects high frequency terms in an article, and uses support vector machines (SVM) to identify biological entity names. It is also computationally efficient and robust to noise commonly found in full text material. We use the method to create a protein name dictionary from a set of 80,528 full text articles. Only 8.3% of the names in this dictionary match SwissProt description lines. We assess the quality of the dictionary by studying its protein name recognition performance in full text. CONCLUSION: This dictionary term lookup method compares favourably to other published methods, supporting the significance of our direct extraction approach. The method is strong in recognizing name variants not found in SwissProt.

Abstracting and Indexing↗

Creating a medical English-Swedish dictionary using interactive word alignment.

BACKGROUND: This paper reports on a parallel collection of rubrics from the medical terminology systems ICD-10, ICF, MeSH, NCSP and KSH97-P and its use for semi-automatic creation of an English-Swedish dictionary of medical terminology. The methods presented are relevant for many other West European language pairs than English-Swedish. METHODS: The medical terminology systems were collected in electronic format in both English and Swedish and the rubrics were extracted in parallel language pairs. Initially, interactive word alignment was used to create training data from a sample. Then the training data were utilised in automatic word alignment in order to generate candidate term pairs. The last step was manual verification of the term pair candidates. RESULTS: A dictionary of 31,000 verified entries has been created in less than three man weeks, thus with considerably less time and effort needed compared to a manual approach, and without compromising quality. As a side effect of our work we found 40 different translation problems in the terminology systems and these results indicate the power of the method for finding inconsistencies in terminology translations. We also report on some factors that may contribute to making the process of dictionary creation with similar tools even more expedient. Finally, the contribution is discussed in relation to other ongoing efforts in constructing medical lexicons for non-English languages. CONCLUSION: In three man weeks we were able to produce a medical English-Swedish dictionary consisting of 31,000 entries and also found hidden translation errors in the utilized medical terminology systems.

Database Management Systems↗

Model-based semantic dictionaries for medical language understanding.

Semantic dictionaries are emerging as a major cornerstone towards achieving sound natural language understanding. Indeed, they constitute the main bridge between words and conceptual entities that reflect their meanings. Nowadays, more and more wide-coverage lexical dictionaries are electronically available in the public domain. However, associating a semantic content with lexical entries is not a straightforward task as it is subordinate to the existence of a fine-grained concept model of the treated domain. This paper presents the benefits and pitfalls in building and maintaining multilingual dictionaries, the semantics of which is directly established on an existing concept model. Concrete cases, handled through the GALEN-IN-USE project, illustrate the use of such semantic dictionaries for the analysis and generation of multilingual surgical procedures.

Dictionaries, Medical as Topic↗

Terminological reference of a knowledge-based system: the data dictionary.

The development of open and integrated knowledge bases makes new demands on the definition of the used terminology. The definition should be realized in a data dictionary separated from the knowledge base. Within the works done at a reference model of medical knowledge, a data dictionary has been developed and used in different applications: a term definition shell, a documentation tool and a knowledge base. The data dictionary includes that part of terminology, which is largely independent of a certain knowledge model. For that reason, the data dictionary can be used as a basis for integrating knowledge bases into information systems, for knowledge sharing and reuse and for modular development of knowledge-based systems.

Artificial Intelligence↗

Relational expressions in STAR file dictionaries.

The STAR File (J. Chem. Inf Comput. Sci. 1994, 34, 505-508) is used widely in structural chemistry for exchanging numerical and text information with scientific journals and databases. These exchanges are increasingly dependent on data dictionaries to facilitate automatic data validation and checking. Definitions in data dictionaries are constructed using attribute descriptors, and this paper describes a method attribute for specifying the relationships between data items as an executable script written in a new relational expression language called dREL. The addition of this attribute improves the precision and the semantic content of dictionaries by providing relational representations of data, as well as facilitating the direct evaluation of derivable data items. The capacity to evaluate derivative data directly from the combination of primitive data and dictionary expressions is expected to change future archival approaches. The design concepts of the relational expression language dREL parser, which are applicable to any discipline, are described.

Journal Article↗

Building dictionaries of 1D and 3D motifs by mining the Unaligned 1D sequences of 17 archaeal and bacterial genomes.

We have used the Teiresias algorithm to carry out unsupervised pattern discovery in a database containing the unaligned ORFs from the 17 publicly available complete archaeal and bacterial genomes and build a 1D dictionary of motifs. These motifs which we refer to as seqlets account for and cover 97.88% of this genomic input at the level of amino acid positions. Each of the seqlets in this 1D dictionary was located among the sequences in Release 38.0 of the Protein Data Bank and the structural fragments corresponding to each seqlet's instances were identified and aligned in three dimensions: those of the seqlets that resulted in RMSD errors below a pre-selected threshold of 2.5 Angstroms were entered in a 3D dictionary of structurally conserved seqlets. These two dictionaries can be thought of as cross-indices that facilitate the tackling of tasks such as automated functional annotation of genomic sequences, local homology identification, local structure characterization, comparative genomics, etc.

Algorithms↗

The HLA Dictionary 2004: a summary of HLA-A, -B, -C, -DRB1/3/4/5 and -DQB1 alleles and their association with serologically defined HLA-A, -B, -C, -DR and -DQ antigens.

This report presents serological equivalents of HLA-A, -B, -C, -DRB1, -DRB3, -DRB4, -DRB5 and -DQB1 alleles. The dictionary is an update of that published in 2001. The data summarize equivalents obtained by the World Health Organization Nomenclature Committee for Factors of the HLA System, the International Cell Exchange (UCLA), the National Marrow Donor Program (NMDP), recent publications and individual laboratories. This latest update of the dictionary is enhanced by the inclusion of results from studies performed during the 13th International Histocompatibility Workshop and from neural network analyses. A summary of the data as recommended serological equivalents is presented as expert assigned types. The tables include remarks for alleles, which are or may be expressed as antigens with serological reaction patterns that differ from the well-established HLA specificities. The equivalents provided will be useful in guiding searches for unrelated haematopoietic stem cell donors in which patients and/or potential donors are typed by either serology or DNA-based methods. The serological DNA equivalent dictionary will also aid in typing and matching procedures for organ transplant programmes whose waiting lists of potential donors and recipients comprise mixtures of serological and DNA-based typings. The tables with HLA equivalents and a questionnaire for submission of serological reaction patterns for poorly identified allelic products will be made available through the WMDA web page (http://www.worldmarrow.org) and, in the near future, also in a searchable form on the IMGT/HLA database.

Alleles↗

Comparison of American medical dictionaries.

Although American medical dictionaries are a valuable part of any medical library collection, the attributes of each of the four major dictionaries are often unknown and the reference material contained in each unused. The medical librarian should be aware of the differences and values of each dictionary and try to have at least one edition of each available to library users in order to maintain an adequate reference collection.

Book Selection↗

Dictionary building via unsupervised hierarchical motif discovery in the sequence space of natural proteins.

Using Teiresias, a pattern discovery method that identifies all motifs present in any given set of protein sequences without requiring alignment or explicit enumeration of the solution space, we have explored the GenPept sequence database and built a dictionary of all sequence patterns with two or more instances. The entries of this dictionary, henceforth named seqlets, cover 98.12% of all amino acid positions in the input database and in essence provide a comprehensive finite set of descriptors for protein sequence space. As such, seqlets can be effectively used to describe almost every naturally occurring protein. In fact, seqlets can be thought of as building blocks of protein molecules that are a necessary (but not sufficient) condition for function or family equivalence memberships. Thus, seqlets can either define conserved family signatures or cut across molecular families and previously undetected sequence signals deriving from functional convergence. Moreover, we show that seqlets also can capture structurally conserved motifs. The availability of a dictionary of seqlets that has been derived in such an unsupervised, hierarchical manner is generating new opportunities for addressing problems that range from reliable classification and the correlation of sequence fragments with functional categories to faster and sensitive engines for homology searches, evolutionary studies, and protein structure prediction.

Amino Acid Motifs↗

What's in a name? Comments on the dermatological dictionary by Ledier, Rosenblum, and Carter.

BACKGROUND: Any scientific discipline needs a sharply defined set of words for exact and reproducible communication. Surprisingly this has never been a strong point in dermatological science, especially as regards the living gross pathology of skin disease. OBJECTIVE: This article briefly reviews the four editions of a dermatological dictionary of words and phrases and gives some thoughts on their usefulness. CONCLUSION: My conclusions are twofold: we need a new dictionary; at the very least, we need a reprinting of the fourth edition of "A Dictionary of Dermatologic Terms" by Carter.

Dermatology↗

Using medical dictionaries to teach the critical evaluation of information sources.

Bibliographic instruction (BI) is the formal teaching of information skills by library professionals. It is argued that biomedical BI must involve the teaching, not only of information retrieval, but also of evaluative skills. A course-integrated BI session on the critical evaluation of health information sources is described. This session uses medical and nursing dictionaries as a subject for investigation. Students in small groups evaluate various dictionaries according to selected criteria. Then they rank the dictionaries and defend the ranking before the rest of the class. It appears to be an effective learning experience from which participants emerge with sharper evaluative skills.

Dictionaries, Medical as Topic↗

The ICF as a framework for national data: the introduction of ICF into Australian data dictionaries.

PURPOSE: A country's data may influence and inform its policy and services, if suitably designed. This paper describes how two related and interacting activities--work on disability concepts and classification as well as the preparation of national data dictionaries--have been carried out in Australia, with this purpose. METHOD: Three key ingredients were combined. A broadly based advisory group was established to ensure the use of disability concepts that are meaningful not only to policy makers but also to the Australian community. This group advised on two 'twin' activities: participation in the revision of the key international classification for disability, and specification of data elements for a national data dictionary according to international standards. RESULTS: National data elements were developed, based on the Beta-2 draft ICIDH-2, and accepted for use in Australian national data dictionaries on a trial basis. CONCLUSION: The purpose and process have been accepted as valuable, and there is interested anticipation of new Australian standard data elements based on the ICF.

Activities of Daily Living↗

SaRAD: a Simple and Robust Abbreviation Dictionary.

MOTIVATION: Due to recent interest in the use of textual material to augment traditional experiments it has become necessary to automatically cluster, classify and filter natural language information. RESULTS: The Simple and Robust Abbreviation Dictionary (SaRAD) provides an easy to implement, high performance tool for the construction of a biomedical symbol dictionary. The algorithms, applied to the MEDLINE document set, result in a high quality dictionary and toolset to disambiguate abbreviation symbols automatically.

Abbreviations as Topic↗

The HLA Dictionary 2004: a summary of HLA-A, -B, -C, -DRB1/3/4/5 and -DQB1 alleles and their association with serologically defined HLA-A, -B, -C, -DR and -DQ antigens.

This report presents serologic equivalents of human leucocyte antigen (HLA)-A, -B, -C, -DRB1, -DRB3, -DRB4, -DRB5 and -DQB1 alleles. The dictionary is an update of the one published in 2001. The data summarize equivalents obtained by the World Health Organization Nomenclature Committee for factors of the HLA System, the International Cell Exchange, the National Marrow Donor Program, recent publications and individual laboratories. This latest update of the dictionary is enhanced by the inclusion of results from studies performed during the 13th International Histocompatibility Workshop and from neural network analyses. A summary of the data as recommended serologic equivalents is presented as expert assigned types. The tables include remarks for alleles, which are or may be expressed as antigens with serologic reaction patterns that differ from the well-established HLA specificities. The equivalents provided will be useful in guiding searches for unrelated hematopoietic stem cell donors in which patients and/or potential donors are typed by either serology or DNA-based methods. The serological DNA equivalent dictionary will also aid in typing and matching procedures for organ transplant programs whose waiting lists of potential donors and recipients comprise of mixtures of serologic and DNA-based typings. The tables with HLA equivalents and a questionnaire for submission of serologic reaction patterns for poorly identified allelic products will be made available through the WMDA web page: www.worldmarrow.org. and in the near future also in a searchable form on the IMGT/HLA database.

Dictionaries, Medical as Topic↗

Gene/protein name recognition based on support vector machine using dictionary as features.

BACKGROUND: Automated information extraction from biomedical literature is important because a vast amount of biomedical literature has been published. Recognition of the biomedical named entities is the first step in information extraction. We developed an automated recognition system based on the SVM algorithm and evaluated it in Task 1.A of BioCreAtIvE, a competition for automated gene/protein name recognition. RESULTS: In the work presented here, our recognition system uses the feature set of the word, the part-of-speech (POS), the orthography, the prefix, the suffix, and the preceding class. We call these features "internal resource features", i.e., features that can be found in the training data. Additionally, we consider the features of matching against dictionaries to be external resource features. We investigated and evaluated the effect of these features as well as the effect of tuning the parameters of the SVM algorithm. We found that the dictionary matching features contributed slightly to the improvement in the performance of the f-score. We attribute this to the possibility that the dictionary matching features might overlap with other features in the current multiple feature setting. CONCLUSION: During SVM learning, each feature alone had a marginally positive effect on system performance. This supports the fact that the SVM algorithm is robust on the high dimensionality of the feature vector space and means that feature selection is not required.

Algorithms↗

Note regarding the word 'behavior' in glossaries of introductory textbooks, dictionaries, and encyclopedias devoted to psychology.

Glossaries of introductory textbooks in psychology, biology, and animal behavior were surveyed to find whether they induded the word 'behavior'. In addition to texts, encyclopedias and dictionaries devoted to the study of behavior were also surveyed. Of the 138 tests sampled across all three fields, only 38 (27%) included the term 'behavior' in their glossaries. Of the 15 encyclopedias and dictionaries surveyed, only 5 defined 'behavior'. To assess whether the term 'behavior' has disappeared from textbook glossaries or whether it has usually been absent, we sampled 23 introductory psychology texts written from 1886 to 1958. Only two texts contained glossaries, and the word 'behavior' was defined in both. An informal survey was conducted of students enrolled in introductory classes in psychology, biology, and animal behavior to provide data on the consistency of definitions. Students were asked to "define the word 'behavior'." Analysis indicated the definition was dependent upon the course. We suggest that future introductory textbook authors and editors of psychology-based dictionaries and encyclopedias include 'behavior' in their glossaries.

Adolescent↗