Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dictionary”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

From EEG dependency multichannel matching pursuit to sparse topographic EEG decomposition.

In this work, we present a multichannel EEG decomposition model based on an adaptive topographic time-frequency approximation technique. It is an extension of the Matching Pursuit algorithm and called dependency multichannel matching pursuit (DMMP). It takes the physiologically explainable and statistically observable topographic dependencies between the channels into account, namely the spatial smoothness of neighboring electrodes that is implied by the electric leadfield. DMMP decomposes a multichannel signal as a weighted sum of atoms from a given dictionary where the single channels are represented from exactly the same subset of a complete dictionary. The decomposition is illustrated on topographical EEG data during different physiological conditions using a complete Gabor dictionary. Further the extension of the single-channel time-frequency distribution to a multichannel time-frequency distribution is given. This can be used for the visualization of the decomposition structure of multichannel EEG. A clustering procedure applied to the topographies, the vectors of the corresponding contribution of an atom to the signal in each channel produced by DMMP, leads to an extremely sparse topographic decomposition of the EEG.

Algorithms↗

Recognizing names in biomedical texts: a machine learning approach.

MOTIVATION: With an overwhelming amount of textual information in molecular biology and biomedicine, there is a need for effective and efficient literature mining and knowledge discovery that can help biologists to gather and make use of the knowledge encoded in text documents. In order to make organized and structured information available, automatically recognizing biomedical entity names becomes critical and is important for information retrieval, information extraction and automated knowledge acquisition. RESULTS: In this paper, we present a named entity recognition system in the biomedical domain, called PowerBioNE. In order to deal with the special phenomena of naming conventions in the biomedical domain, we propose various evidential features: (1) word formation pattern; (2) morphological pattern, such as prefix and suffix; (3) part-of-speech; (4) head noun trigger; (5) special verb trigger and (6) name alias feature. All the features are integrated effectively and efficiently through a hidden Markov model (HMM) and a HMM-based named entity recognizer. In addition, a k-Nearest Neighbor (k-NN) algorithm is proposed to resolve the data sparseness problem in our system. Finally, we present a pattern-based post-processing to automatically extract rules from the training data to deal with the cascaded entity name phenomenon. From our best knowledge, PowerBioNE is the first system which deals with the cascaded entity name phenomenon. Evaluation shows that our system achieves the F-measure of 66.6 and 62.2 on the 23 classes of GENIA V3.0 and V1.1, respectively. In particular, our system achieves the F-measure of 75.8 on the "protein" class of GENIA V3.0. For comparison, our system outperforms the best published result by 7.8 on GENIA V1.1, without help of any dictionaries. It also shows that our HMM and the k-NN algorithm outperform other models, such as back-off HMM, linear interpolated HMM, support vector machines, C4.5, C4.5 rules and RIPPER, by effectively capturing the local context dependency and resolving the data sparseness problem. Moreover, evaluation on GENIA V3.0 shows that the post-processing for the cascaded entity name phenomenon improves the F-measure by 3.9. Finally, error analysis shows that about half of the errors are caused by the strict annotation scheme and the annotation inconsistency in the GENIA corpus. This suggests that our system achieves an acceptable F-measure of 83.6 on the 23 classes of GENIA V3.0 and in particular 86.2 on the "protein" class, without help of any dictionaries. We think that a F-measure of 90 on the 23 classes of GENIA V3.0 and in particular 92 on the "protein" class, can be achieved through refining of the annotation scheme in the GENIA corpus, such as flexible annotation scheme and annotation consistency, and inclusion of a reasonable biomedical dictionary. AVAILABILITY: A demo system is available at http://textmining.i2r.a-star.edu.sg/NLS/demo.htm. Technology license is available upon the bilateral agreement.

Abstracting and Indexing↗

Spontaneous perilymphatic fistula: myth or fact.

Controversy exists surrounding the diagnosis of spontaneous perilymphatic fistula. In an effort to help resolve this controversy the author conducted a review of the literature as well as a review of 212 of his patients who underwent surgical exploration for suspected perilymphatic fistula. Interpretation of the literature reviewed was hampered by the lack of a uniformly accepted definition for the word spontaneous. Dorland's Medical Dictionary defines spontaneous as that which occurs without external influence. Webster's Dictionary, on the other hand, provides a much more confining definition of the word by stating that a spontaneous event is one that occurs or is produced by its own energy. Only 58 percent of the author's 212 patients had an antecedent history of an external event that may have precipitated the suspected perilymphatic fistula (trauma, flying, diving) while almost 41 percent recalled an antecedent event of internal origin (lifting, straining, sneezing, nose blowing). If one were to support the definition of spontaneous provided by Dorland's Medical Dictionary, then the 41 percent of patients who had no antecedent history of external event would have to be considered as having spontaneous perilymphatic fistula. If, on the other hand, one were to endorse the definition of spontaneous provided by Webster's then less than 2 percent of the author's patients would have to be considered as having spontaneous perilymphatic fistula.

Adolescent↗

Accuracy of a voice-to-text personal dictation system in the generation of radiology reports.

OBJECTIVE: Systems that convert spoken words directly into text have recently become available. The purpose of this study was to test the accuracy, on a word-for-word basis, of one such system for generating radiology reports. MATERIALS AND METHODS: The IBM Personal Dictation System (IPDS), with the optional add-on radiology vocabulary, was assessed (the system is now known as VoiceType Dictation). The system requires one to use discrete speech (i.e., with a momentary pause between words). Two hundred current radiology reports, including 100 consecutive chest radiographs, 50 consecutive thoracic CT scans, and 50 random sonography and angiography reports, were read to the system. Before testing, the IPDS had been used in the dictation of material related to chest radiology. All errors were noted on a word-for-word basis and categorized as follows: incorrect, partially incorrect (no effect on meaning), not in dictionary (word was then added), homophone, or formatting. Medical words (those thought to be relevant to the meaning of the report) were considered separately. Specific assessments of numbers and dates were made. Words added to the dictionary were reread after the 200 reports had been assessed. RESULTS: When all mistakes were considered, the accuracy was 0.99 for chest radiology and 0.96 for material not concerning chest radiology. When only relevant mistakes on medical words were considered, the accuracies were 0.98 and 0.96, respectively. Accuracy for numbers was 0.95 and for dates 0.97. Redictation of the 22 words previously not in the dictionary was 100% accurate. CONCLUSION: The IPDS is an accurate system for direct voice-to-text dictation of radiology reports that improves with continued use. The most important future enhancement of such systems will be to allow the more natural continuous or conversational speech style.

Humans↗

Continuity of care: some experiences and thoughts.

Continuity of health care is a goal to be achieved. Most are for it. Many claim to provide it. But how do we know we have it? What are the key features of continuity? While dictionaries do not define the phrase "continuity of health care," we do find definitions of "continuity." The Oxford English Dictionary, Second Edition, includes in its definitions: "the state or quality of being uninterrupted in sequence or succession, or in essence or idea; connectedness, coherence, unbroken..." Stedman's Medical Dictionary includes: "absence of interruption, a succession of parts intimately united..." These definitions stress an uninterrupted succession and include the concept that there needs to be a connection to the parts. Without that connection, continuity, in health care delivery or elsewhere, does not exist.

Continuity of Patient Care↗

Domain analysis and modeling to improve comparability of health statistics.

Health statistics is an essential element to improve the ability of managers of health institutions, healthcare researchers, policy makers, and health professionals to formulate appropriate course of reactions and to make decisions based on evidence. To ensure adequate health statistics, standards are of critical importance. A study on healthcare statistics domain analysis is underway in an effort to improve usability and comparability of health statistics. The ongoing study focuses on structuring the domain knowledge and making the knowledge explicit with a data element dictionary being the core. Supplemental to the dictionary are a domain term list, a terminology dictionary, and a data model to help organize the concepts constituting the health statistics domain.

Demography↗

The DairyCHAMP program: a computerised recording system for dairy herds.

The DairyCHAMP program is an animal health and management software program that helps daily animal management, herd performance monitoring and problem analysis. Data entry to the program uses a data dictionary and includes an error-checking system that ensures the consistency and appropriateness of data entered. DairyCHAMP performs health management functions, provides a convenient user interface, ensures uniform data across farms by using a standard data dictionary, can be fully integrated with decision-making software programs like DairyORACLE, and is flexible enough to be useful for many types of dairy facilities. Data are entered via a menu-based system. Animal events are organised around reproduction and lactation cycles and health records. Farm records include inventories for drug, feed and semen. Farm parameters can be established which customize the program for an individual farm. The database system is an integration of three schemas: the individual user's view, the community view and the storage system. The individual user's view must be easy to use, while the storage system must be compact enough to fit within the disc storage space on a microcomputer. This conflict requires a translation from one schema to another. The DairyCHAMP program accomplishes this through a coding system which assigns a code number to each event. The program can add synonyms to this event dictionary by assigning the same code number to the synonym the user chooses. The DairyCHAMP program provides access to the large amounts of data required to aid in daily animal management, allow performance monitoring and analyse problems. Its highly integrated system is efficient and easy to use and maintain.

Animal Husbandry↗

The concept of "template" assisted electronic medical record.

A new design for an electronic medical record with flexible template generation was developed. The "template" in this paper refers to one of the tools that calls a user's attention to entering the data items for each problem or to ordering tests when they are due. The template shows the patient's previous data and guides physicians to record the consistent description. Two kinds of medical data dictionaries are prepared. Data representations of signs, symptoms, laboratory tests, and other examinations are defined in the check items dictionary. The problem dictionary contains the possible patient problems with relation to check items and other kinds of related subjects. This system provides Problem Oriented Medical Record (POMR) and the graphical presentation of various patient data. The system was designed to establish a constant and integrated medical record.

Artificial Intelligence↗

[Information capacity of the nucleotide sequences and their fragments].

The problem of determining the information content of nucleotide sequences is discussed. Exact expressions for the reconstitution of higher-order frequency dictionaries from lower-order once were obtained by the maximum entropy method. In form, they are analogous to superpositional approximations known in statistical physics. The features of entropy characteristics of real nucleotide sequences are described that reliably distinguish them from random texts. Methods for comparing the information content of frequency dictionaries and assessing the residual uncertainty of the text at the known frequency dictionary are proposed.

Base Sequence↗

Protein names precisely peeled off free text.

MOTIVATION: Automatically identifying protein names from the scientific literature is a pre-requisite for the increasing demand in data-mining this wealth of information. Existing approaches are based on dictionaries, rules and machine-learning. Here, we introduced a novel system that combines a pre-processing dictionary- and rule-based filtering step with several separately trained support vector machines (SVMs) to identify protein names in the MEDLINE abstracts. RESULTS: Our new tagging-system NLProt is capable of extracting protein names with a precision (accuracy) of 75% at a recall (coverage) of 76% after training on a corpus, which was used before by other groups and contains 200 annotated abstracts. For our estimate of sustained performance, we considered partially identified names as false positives. One important issue frequently ignored in the literature is the redundancy in evaluation sets. We suggested some guidelines for removing overly inadequate overlaps between training and testing sets. Applying these new guidelines, our program appeared to significantly out-perform other methods tagging protein names. NLProt was so successful due to the SVM-building blocks that succeeded in utilizing the local context of protein names in the scientific literature. We challenge that our system may constitute the most general and precise method for tagging protein names. AVAILABILITY: http://cubic.bioc.columbia.edu/services/nlprot/

Abstracting and Indexing↗

Working more productively: tools for administrative data.

OBJECTIVE: This paper describes a web-based resource (http://www.umanitoba.ca/centres/mchp/concept/) that contains a series of tools for working with administrative data. This work in knowledge management represents an effort to document, find, and transfer concepts and techniques, both within the local research group and to a more broadly defined user community. Concepts and associated computer programs are made as "modular" as possible to facilitate easy transfer from one project to another. STUDY SETTING/DATA SOURCES: Tools to work with a registry, longitudinal administrative data, and special files (survey and clinical) from the Province of Manitoba, Canada in the 1990-2003 period. DATA COLLECTION: Literature review and analyses of web site utilization were used to generate the findings. PRINCIPAL FINDINGS: The Internet-based Concept Dictionary and SAS macros developed in Manitoba are being used in a growing number of research centers. Nearly 32,000 hits from more than 10,200 hosts in a recent month demonstrate broad interest in the Concept Dictionary. CONCLUSIONS: The tools, taken together, make up a knowledge repository and research production system that aid local work and have great potential internationally. Modular software provides considerable efficiency. The merging of documentation and researcher-to-researcher dissemination keeps costs manageable.

Databases as Topic↗

Multimedia search system for a textbook of urology in Japanese.

A search system for the textbook of urology has been developed and evaluated. The textbook is written in Japanese with 1.4 megabytes of text and has about 1000 pictures. The contents of the textbook can be seen with English or Japanese medical terms. The system contains a dictionary of Japanese medical terms. The dictionary has a network structure based on Japanese ideographic characters. Such a search system is useful for physicians because of its speed and the widely spread medical knowledge of many authors of the textbook.

Artificial Intelligence↗

An object-oriented approach for structuring the electronic medical record.

We implemented a framework for modelling the electronic medical record on top of an object-oriented model. Clinical patient data are structured in a uniform way through the use of a comprehensive data model. The meaning of the information elements is explicitly determined by a medical data dictionary. The data structures of both, medical record and data dictionary are implemented, using a semantically rich, object-oriented data model. We examined several possibilities for the graphical preparation of the inherently recursive data structures. Again, we use object-oriented frameworks for the implementation of flexible user interfaces to the electronic medical record with a consistent look-and-feel.

Data Collection↗

[Japanese medicines studied by Hepburn, an American missionary, in 1860s].

In the last days of the Tokugawa shogunate when Japan was opened its door to trade, J.C. Hepburn came to Japan, and he contributed to the modernization of Japan as a missionary, a doctor and an instructor. His great academic achievement was publishing a Japanese-English and English-Japanese dictionary. The dictionary contained words which were generally used in those days. It is said that it was way of living and culture in Japan as seen by an American. Words about medicine show the conditions of the business of medicine in those days. Until 1889 a pharmacist was not professionalized yet, nor was the separation of dispensary from medical practice effective. As most of doctors were Chinese herb doctors, there was not the word "MD." As Chinese doctors were the leading ones, lots of words about herb were mentioned. We can find only a few words about chemicals. The common people used patent medicine such as Daranisuke, Mankintan. Less than 100 years, sorts of medicine in Japan have changed to those of medicine in Western countries.

Dictionaries as Topic↗

Quantitative evaluation of English-Japanese machine translation of medical literature.

Although many machine-translation programs are currently available, few evaluation methods of such translation exist for any given application area. It is difficult to evaluate machine-translation systems objectively because the quality of a translation depends on the combination of three factors: the translation program, the dictionary, and the original document. In this study, we developed a quantitative evaluation method for assessing machine translation, which evaluates these three factors separately. We applied this method to the translation of English to Japanese for medical literature and the method proved to be a good indicator for further system improvement. Using this method we also discovered other important points for machine translation, such as the examination of target documents for the construction of a better application dictionary.

Computers↗

A conceptual graphs modeling of UMLS components.

The Unified Medical Language System (UMLS) of the U.S. National Library of Medicine is a complex collection of terms, concepts, and relationships derived from standard classifications. Potential applications would benefit from a high level representation of its components. This paper proposes a conceptual representation of both the Metathesaurus and the Semantic Network of the UMLS based on conceptual graphs. It shows that the addition of a dictionary of concepts to the UMLS knowledge base allows the capability to exploit it pertinently. This dictionary defines more precisely the core concepts and adds constraints on their use. Constraints are dedicated to guide an "intelligent" browsing of the UMLS knowledge sources.

Dictionaries as Topic↗

Automatic construction of knowledge base from biological papers.

We designed a system that acquires domain specific knowledge from human written biological papers, and we call this system IFBP (Information Finding from Biological Papers). IFBP is divided into three phases, Information Retrieval (IR), Information Extraction (IE) and Dictionary Construction (DC). We propose a query modification method using automatically constructed thesaurus for IR and a statistical keyword prediction method for IE. A dictionary of domain specific terms, which is one of the central knowledge sources for the task of knowledge acquisition, is also constructed automatically in the DC phase. IFBP is currently used for constructing the Transcription Factor DataBase (TFDB) and shows good performance. Since the model of knowledge base construction that is adopted into IFBP is carried out entirely automatically, this system can be easily ported across domains.

Algorithms↗

Look-up tables for protein solvent accessibility prediction and nearest neighbor effect analysis.

We developed dictionaries of two-, three-, and five-residue patterns in proteins and computed the average solvent accessibility of the central residues in their native proteins. These dictionaries serve as a look-up table for making subsequent predictions of solvent accessibility of amino acid residues. We find that predictions made in this way are very close to those made using more sophisticated methods of solvent accessibility prediction. We also analyzed the effect of immediate neighbors on the solvent accessibility of residues. This helps us in understanding how the same residue type may have different accessible surface areas in different proteins and in different positions of the same protein. We observe that certain residues have a tendency to increase or decrease the solvent accessibility of their neighboring residues in C- or N-terminal positions. Interestingly, the C-terminal and N-terminal neighbor residues are found to have asymmetric roles in modifying solvent accessibility of residues. As expected, similar neighbors enhance the hydrophobic or hydrophilic character of residues. Detailed look-up tables are provided on the web at www.netasa.org/look-up/.

Amino Acid Sequence↗