Search PubMed⌕ Search

Biomedical subjects

Yehoshua Perl

Publications and source records attributed to Yehoshua Perl.

15 recordsLinked to original sources

Analysis of a study of the users, uses, and future agenda of the UMLS.

OBJECTIVES: The UMLS constitutes the largest existing collection of medical terms. However, little has been published about the users and uses of the UMLS. This study sheds light on these issues. DESIGN: We designed a questionnaire consisting of 26 questions and distributed it to the UMLS user mailing list. Participants were assured complete confidentiality of their replies. To further encourage list members to respond, we promised to provide them with early results prior to publication. Sector analysis of the responses, according to employment organizations is used to obtain insights into some responses. RESULTS: We received 70 responses. The study confirms two intended uses of the UMLS: access to source terminologies (75%), and mapping among them (44%). However, most access is just to a few sources, led by SNOMED, MeSH, and ICD. Out of 119 reported purposes of use, terminology research (37), information retrieval (19), and terminology translation (14) lead. Four important observations are that the UMLS is widely used as a terminology (77%), even though it was not designed as one; many users (73%) want the NLM to mark concepts with multiple parents in an indented hierarchy and to derive a terminology from the UMLS (73%). Finally, auditing the UMLS is a top budget priority (35%) for users. CONCLUSIONS: The study reports many uses of the UMLS in a variety of subjects from terminology research to decision support and phenotyping. The study confirms that the UMLS is used to access its source terminologies and to map among them. Two primary concerns of the existing user base are auditing the UMLS and the design of a UMLS-based derived terminology.

Consumer Behavior↗

Auditing as part of the terminology design life cycle.

OBJECTIVE: To develop and test an auditing methodology for detecting errors in medical terminologies satisfying systematic inheritance. This methodology is based on various abstraction taxonomies that provide high-level views of a terminology and highlight potentially erroneous concepts. DESIGN: Our auditing methodology is based on dividing concepts of a terminology into smaller, more manageable units. First, we divide the terminology's concepts into areas according to their relationships/roles. Then each multi-rooted area is further divided into partial-areas (p-areas) that are singly-rooted. Each p-area contains a set of structurally and semantically uniform concepts. Two kinds of abstraction networks, called the area taxonomy and p-area taxonomy, are derived. These taxonomies form the basis for the auditing approach. Taxonomies tend to highlight potentially erroneous concepts in areas and p-areas. Human reviewers can focus their auditing efforts on the limited number of problematic concepts following two hypotheses on the probable concentration of errors. RESULTS: A sample of the area taxonomy and p-area taxonomy for the Biological Process (BP) hierarchy of the National Cancer Institute Thesaurus (NCIT) was derived from the application of our methodology to its concepts. These views led to the detection of a number of different kinds of errors that are reported, and to confirmation of the hypotheses on error concentration in this hierarchy. CONCLUSION: Our auditing methodology based on area and p-area taxonomies is an efficient tool for detecting errors in terminologies satisfying systematic inheritance of roles, and thus facilitates their maintenance. This methodology concentrates a domain expert's manual review on portions of the concepts with a high likelihood of errors.

Biology↗

Relationship structures and semantic type assignments of the UMLS Enriched Semantic Network.

OBJECTIVE: The Enriched Semantic Network (ESN) was introduced as an extension of the Unified Medical Language System (UMLS) Semantic Network (SN). Its multiple subsumption configuration and concomitant multiple inheritance make the ESN's relationship structures and semantic type assignments different from those of the SN. A technique for deriving the relationship structures of the ESN's semantic types and an automated technique for deriving the ESN's semantic type assignments from those of the SN are presented. DESIGN: The technique to derive the ESN's relationship structures finds all newly inherited relationships in the ESN. All such relationships are audited for semantic validity, and the blocking mechanism is used to block invalid relationships. The mapping technique to derive the ESN's semantic type assignments uses current SN semantic type assignments and preserves nonredundant categorizations, while preventing new redundant categorizations. RESULTS: Among the 426 newly inherited relationships, 326 are deemed valid. Seven blockings are applied to avoid inheritance of the 100 invalid relationships. Sixteen semantic types have different relationship structures in the ESN as compared to those in the SN. The mapping of semantic type assignments from the SN to the ESN avoids the generation of 26,950 redundant categorizations. The resulting ESN contains 138 semantic types, 149 IS-A links, 7,303 relationships, and 1,013,876 semantic type assignments. CONCLUSION: The ESN's multiple inheritance provides more complete relationship structures than in the SN. The ESN's semantic type assignments avoid the existing redundant categorizations appearing in the SN and prevent new ones that might arise due to multiple parents. Compared to the SN, the ESN provides a more accurate unifying semantic abstraction of the UMLS Metathesaurus.

Semantics↗

A lexical metaschema for the UMLS semantic network.

OBJECTIVE: A metaschema is a high-level abstraction network of the UMLS's semantic network (SN) obtained from a partition of the SN's collection of semantic types. Every metaschema has nodes, called meta-semantic types, each of which denotes a group of semantic types constituting a subject area of the SN. A new kind of metaschema, called the lexical metaschema, is derived from a lexical partition of the SN. The lexical metaschema is compared to previously derived metaschemas, e.g., the cohesive metaschema. DESIGN: A new lexical partitioning methodology is presented based on identical word-usage among the names of semantic types and the definitions of their respective children. The lexical metaschema is derived from the application of the methodology. We compare the constituent meta-semantic types and their underlying semantic-type groups with the previously derived cohesive metaschema. A similar comparison of the lexical partition and a published partition of the SN is also carried out. RESULTS: The lexical partition of the SN has 21 semantic-type groups, each of which represents a subject area. The lexical metaschema thus has 21 meta-semantic types, 19 meta-child-of hierarchical relationships, and 86 meta-relationships. Our comparison shows that 15 out of the 21 meta-semantic types in the lexical metaschema also appear in the cohesive metaschema, and 80 semantic types are covered by identical meta-semantic types or refinements between the two metaschemas. The comparison between the lexical partition and the semantic partition shows that they have very low similarity. CONCLUSION: The algorithmically derived lexical metaschema serves as an abstraction of the SN and provides views representing different subject areas. It compares favorably with the cohesive metaschema derived via the SN's relationship configuration.

Abstracting and Indexing↗

An expert study evaluating the UMLS lexical metaschema.

OBJECTIVE: A metaschema is an abstraction network of the UMLS's semantic network (SN) obtained from a connected partition of its collection of semantic types. A lexical metaschema was previously derived based on a lexical partition which partitioned the SN into semantic-type groups using identical word-usage among the names of semantic types and the definitions of their respective children. In this paper, a statistical analysis methodology is presented to evaluate the lexical metaschema based on a study involving a group of established UMLS experts. METHODS: In the study, each expert was asked to identify subject areas of the SN based on his or her understanding of the various semantic types. For this purpose, the expert scans the SN hierarchy top-down, identifying semantic types, which are important and different enough from their parent semantic types, as roots of their groups. From the response of each expert, an "expert metaschema" is constructed. The different experts' metaschemas can vary widely. So, additional metaschemas are obtained from aggregations of the experts' responses. Of special interest is the consensus metaschema which represents an aggregation of a simple majority of the experts' responses. Statistical analysis comparing the lexical metaschema with the experts' metaschemas and the consensus metaschema is presented. RESULTS: The analysis results shows that 17 out of the 21 meta-semantic types in the lexical metaschema also appear in the consensus metaschema (about 81%). There are 107 semantic types (about 79%) covered by identical meta-semantic types and refinements. The results show the high similarity between the two metaschemas. Furthermore, the statistical analysis shows that the lexical metaschema did not grossly underperform compared to the experts. CONCLUSION: Our study shows that the lexical metaschema provides a good approximation for a partition of meaningful subject areas in the SN, when compared to the consensus metaschema capturing the aggregation of a simple majority of the human experts' opinions.

Animals↗

An enriched unified medical language system semantic network with a multiple subsumption hierarchy.

OBJECTIVE: The Unified Medical Language System's (UMLS's) Semantic Network's (SN's) two-tree structure is restrictive because it does not allow a semantic type to be a specialization of several other semantic types. In this article, the SN is expanded into a multiple subsumption structure with a directed acyclic graph (DAG) IS-A hierarchy, allowing a semantic type to have multiple parents. New viable IS-A links are added as warranted. DESIGN: Two methodologies are presented to identify and add new viable IS-A links. The first methodology is based on imposing the characteristic of connectivity on a previously presented partition of the SN. Four transformations are provided to find viable IS-A links in the process of converting the partition's disconnected groups into connected ones. The second methodology identifies new IS-A links through a string matching process involving names and definitions of various semantic types in the SN. A domain expert is needed to review all the results to determine the validity of the new IS-A links. RESULTS: Nineteen new IS-A links are added to the SN, and four new semantic types are also created to support the multiple subsumption framework. The resulting network, called the Enriched Semantic Network (ESN), exhibits a DAG-structured hierarchy. A partition of the ESN containing 19 connected groups is also derived. CONCLUSION: The ESN is an expanded abstraction of the UMLS compared with the original SN. Its multiple subsumption hierarchy can accommodate semantic types with multiple parents. Its representation thus provides direct access to a broader range of subsumption knowledge.

Algorithms↗

Auditing concept categorizations in the UMLS.

The Unified Medical Language System (UMLS) integrates about 880,000 concepts from 100 biomedical terminologies. Each concept is categorized to at least one semantic type of the Semantic Network. During the integration, it is unavoidable that some categorization errors and inconsistencies will be introduced. In this paper, we present an auditing technique to find such errors and inconsistencies. Our technique is based on an expert reviewing the pure intersections of meta-semantic types of a metaschema, a compact abstract view of the UMLS Semantic Network. We use a divide and conquer approach, handling differently small pure intersections and medium to large pure intersections. By using this approach, we limit the number of concepts reviewed, for which we expect a high percentage of errors. We reviewed all concepts in 657 pure intersections containing one to 10 concepts. Various kinds of errors are identified and the analysis of the results are presented in the paper. Also, we checked the pure intersections containing more than 10 concepts for their semantic soundness, where the semantically suspicious pure intersections are presented in the paper and their concepts are reviewed.

Abstracting and Indexing↗

Designing metaschemas for the UMLS enriched semantic network.

The enriched semantic network (ESN) has previously been presented as an enhancement of the semantic network (SN) of the UMLS. The ESN's hierarchy is a DAG (Directed Acyclic Graph) structure allowing for multiple parents. The ESN is thus more complex than the SN and can be more difficult to view and comprehend. We have previously introduced the notion of a metaschema for the SN as a compact abstraction to support SN comprehension. We extend the definition of metaschema to make it applicable to a DAG classification hierarchy, such as the one exhibited by the ESN. We specify the requirements for and describe the general process of deriving such a metaschema. We derive two particular metaschemas of the ESN based on a pair of partitions. These two metaschemas and their underlying partitions are compared. Both metaschemas serve as compact representations of the ESN, allowing for convenient viewing of its hierarchy and easier comprehension.

Abstracting and Indexing↗

The cohesive metaschema: a higher-level abstraction of the UMLS Semantic Network.

The Unified Medical Language System (UMLS) joins together a group of established medical terminologies in a unified knowledge representation framework. Two major resources of the UMLS are its Metathesaurus, containing a large number of concepts, and the Semantic Network (SN), containing semantic types and forming an abstraction of the Metathesaurus. However, the SN itself is large and complex and may still be difficult to view and comprehend. Our structural partitioning technique partitions the SN into structurally uniform sets of semantic types based on the distribution of the relationships within the SN. An enhancement of the structural partition results in cohesive, singly rooted sets of semantic types. Each such set is named after its root which represents the common nature of the group. These sets of semantic types are represented by higher-level components called metasemantic types. A network, called a metaschema, which consists of the meta-semantic types connected by hierarchical and semantic relationships is obtained and provides an abstract view supporting orientation to the SN. The metaschema is utilized to audit the UMLS classifications. We present a set of graphical views of the SN based on the metaschema to help in user orientation to the SN. A study compares the cohesive metaschema to metaschemas derived semantically by UMLS experts.

Algorithms↗

Partitioning the UMLS semantic network.

The unified medical language system (UMLS) integrates many well-established biomedical terminologies. The UMLS semantic network (SN) can help orient users to the vast knowledge content of the UMLS Metathesaurus (META) via its abstract conceptual view. However, the SN itself is large and complex and may still be difficult to comprehend. Our technique partitions the SN into smaller meaningful units amenable to display on limited-sized computer screens. The basis for the partitioning is the distribution of the relationships within the SN. Three rules are applied to transform the original partition into a second more cohesive partition.

Algorithms↗

Evaluation and application of a semantic network partition.

Semantic networks (SNs) are excellent knowledge representation structures. However, large semantic networks are hard to comprehend. To overcome this difficulty, several methods of partitioning have been developed that rely on different mixes of structural and semantic methods. However, little has appeared in the literature concerning the question whether a partition of a semantic network creates subnetworks that agree with human insight. We address this issue by presenting a comparison between the results of an algorithmic partitioning method and a partition created by a group of experts. Subsequently, we show how a network partition can be used to generate various partial views of a semantic network, which facilitate user orientation. Examples from the Unified Medical Language System (UMLS) SN are used to demonstrate partial views.

Algorithms↗

Using the metaschema to audit UMLS classification errors.

The Unified Medical Language System integrates about 800,000 concepts from 99 biomedical terminologies. Each concept is assigned to at least one semantic type of the Semantic Network. During the integration, it is unavoidable that some classification errors and inconsistencies will be introduced. In this paper, we present an auditing technique to find such errors and inconsistencies. Our technique is based on an expert reviewing the pure intersections of meta-semantic types of the metaschema, a compact abstract view of the Semantic Network. Results regarding the pure intersections are reported. The analysis results for pure intersections with 1 to 6 concepts are presented. Various kinds of errors are identified.

Semantics↗

Auditing the UMLS for redundant classifications.

The UMLS's Semantic Network (SN) serves as a valuable abstraction for the underlying concept repository called the Metathesaurus (META). Specifically, the SN forms a classification layer for the META, with each of the META's constituent concepts assigned to one or more semantic types in the SN. The rule in the design of the SN is to have concepts explicitly assigned to the lowest possible semantic types in the SN's IS-A hierarchy. Implicit assignment to higher semantic types can be inferred via the IS-A relationships. However, in subsequent versions of the UMLS, unnecessary, simultaneous assignments to descendant and ancestor semantic types have been discovered (e.g., 8,622 in the UMLS 1998 version and 12,657 in the 2001 version). The assignment of concepts to such ancestor semantic types is called redundant classification. There is a need for an automated auditing tool that can identify all these redundant classifications. In this paper, an efficient algorithm for this auditing task is introduced. Details of its application to the current (2001) version of the UMLS are presented and the results are discussed.

Algorithms↗

Enriching the structure of the UMLS semantic network.

The Unified Medical Language System's (UMLS's) Semantic Network (SN)---consisting of a network of semantic types---has a two-tree structure, where each semantic type has at most one parent semantic type. This arrangement is restrictive because some semantic types are, by their definition, specializations of several parents. As a proposed enhancement to the SN, its semantic types have previously been partitioned into groups, each of which contains semantic types of some specific area. However, some groups of this proposed partition contain forest (i.e., multiple-tree) structures or even isolated semantic types. Both situations imply a disconnected internal structure. Connectivity is actually one way to assess the proposed "semantic validity" principle for partitions. It is a desired, although not required, property. In this paper, we introduce a methodology for identifying "missing" IS-A links and adding them to the SN. This process transforms the SN into a Directed Acyclic Graph (DAG) structure, with semantic types permitted to have multiple parents. A result of our methodology is the transformation of the proposed SN partition into groups satisfying the connectivity property.

Semantics↗