Search PubMed⌕ Search

PubMed · 16535893

Guidelines for developing a data dictionary.

Abstract

The source did not provide an abstract. Follow the original record for more information.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

AHIMA e-HIM Workgroup on EHR Data Content. 2006. Guidelines for developing a data dictionary.. https://pubmed.ncbi.nlm.nih.gov/16535893/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Quantitative assessment of dictionary-based protein named entity tagging.

OBJECTIVE: Natural language processing (NLP) approaches have been explored to manage and mine information recorded in biological literature. A critical step for biological literature mining is biological named entity tagging (BNET) that identifies names mentioned in text and normalizes them with entries in biological databases. The aim of this study was to provide quantitative assessment of the complexity of BNET on protein entities through BioThesaurus, a thesaurus of gene/protein names for UniProt knowledgebase (UniProtKB) entries that was acquired using online resources. METHODS: We evaluated the complexity through several perspectives: ambiguity (i.e., the number of genes/proteins represented by one name), synonymy (i.e., the number of names associated with the same gene/protein), and coverage (i.e., the percentage of gene/protein names in text included in the thesaurus). We also normalized names in BioThesaurus and measures were obtained twice, once before normalization and once after. RESULTS: The current version of BioThesaurus has over 2.6 million names or 2.1 million normalized names covering more than 1.8 million UniProtKB entries. The average synonymy is 3.53 (2.86 after normalization), ambiguity is 2.31 before normalization and 2.32 after, while the coverage is 94.0% based on the BioCreAtive data set comprising MEDLINE abstracts containing genes/proteins. CONCLUSION: The study indicated that names for genes/proteins are highly ambiguous and there are usually multiple names for the same gene or protein. It also demonstrated that most gene/protein names appearing in text can be found in BioThesaurus.

Dictionaries as Topic↗

Evaluation of the expressiveness of an ICNP-based nursing data dictionary in a computerized nursing record system.

This study evaluated the domain completeness and expressiveness issues of the International Classification for Nursing Practice-based (ICNP) nursing data dictionary (NDD) through its application in an enterprise electronic medical record (EMR) system as a standard vocabulary at a single tertiary hospital in Korea. Data from 2,262 inpatients obtained over a period of 9 weeks (May to July 2003) were extracted from the EMR system for analysis. Among the 530,218 data-input events, 401,190 (75.7%) were entered from the NDD, 20,550 (3.9%) used only free text, and 108,478 (20.4%) used a combination of coded data and free text. A content analysis of the free-text events showed that 80.3% of the expressions could be found in the NDD, whereas 10.9% were context-specific expressions such as direct quotations of patient complaints and responses, and references to the care plan or orders of physicians. A total of 7.8% of the expressions was used for a supplementary purpose such as adding a conjunction or end verb to make an expression appear as natural language. Only 1.0% of the expressions were identified as not being covered by the NDD. This evaluation study demonstrates that the ICNP-based NDD has sufficient power to cover most of the expressions used in a clinical nursing setting.

Dictionaries as Topic↗