Search PubMed⌕ Search

PubMed · 15542017

Biomedical named entity recognition using two-phase model based on SVMs.

Abstract

Named entity (NE) recognition has become one of the most fundamental tasks in biomedical knowledge acquisition. In this paper, we present a two-phase named entity recognizer based on SVMs, which consists of a boundary identification phase and a semantic classification phase of named entities. When adapting SVMs to named entity recognition, the multi-class problem and the unbalanced class distribution problem become very serious in terms of training cost and performance. We try to solve these problems by separating the NE recognition task into two subtasks, where we use appropriate SVM classifiers and relevant features for each subtask. In addition, by employing a hierarchical classification method based on ontology, we effectively solve the multi-class problem concerning semantic classification. The experimental results on the GENIA corpus show that the proposed method is effective not only in reducing computational cost but also in improving performance. The F-score (beta=1) for the boundary identification is 74.8 and the F-score for the semantic classification is 66.7.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ki-Joong Lee, Young-Sook Hwang, Seonho Kim, Hae-Chang Rim. 2004. Biomedical named entity recognition using two-phase model based on SVMs.. https://doi.org/10.1016/j.jbi.2004.08.012

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Can electronic search engines optimize screening of search results in systematic reviews: an empirical study.

BACKGROUND: Most electronic search efforts directed at identifying primary studies for inclusion in systematic reviews rely on the optimal Boolean search features of search interfaces such as DIALOG and Ovid. Our objective is to test the ability of an Ultraseek search engine to rank MEDLINE records of the included studies of Cochrane reviews within the top half of all the records retrieved by the Boolean MEDLINE search used by the reviewers. METHODS: Collections were created using the MEDLINE bibliographic records of included and excluded studies listed in the review and all records retrieved by the MEDLINE search. Records were converted to individual HTML files. Collections of records were indexed and searched through a statistical search engine, Ultraseek, using review-specific search terms. Our data sources, systematic reviews published in the Cochrane library, were included if they reported using at least one phase of the Cochrane Highly Sensitive Search Strategy (HSSS), provided citations for both included and excluded studies and conducted a meta-analysis using a binary outcome measure. Reviews were selected if they yielded between 1000-6000 records when the MEDLINE search strategy was replicated. RESULTS: Nine Cochrane reviews were included. Included studies within the Cochrane reviews were found within the first 500 retrieved studies more often than would be expected by chance. Across all reviews, recall of included studies into the top 500 was 0.70. There was no statistically significant difference in ranking when comparing included studies with just the subset of excluded studies listed as excluded in the published review. CONCLUSION: The relevance ranking provided by the search engine was better than expected by chance and shows promise for the preliminary evaluation of large results from Boolean searches. A statistical search engine does not appear to be able to make fine discriminations concerning the relevance of bibliographic records that have been pre-screened by systematic reviewers.

Abstracting and Indexing↗

Searching for observational studies: what does citation tracking add to PubMed? A case study in depression and coronary heart disease.

BACKGROUND: PubMed is the most widely used method for searches of the medical literature, but fails to identify many relevant articles. Electronic citation tracking offers an alternative search method. METHODS: Articles investigating the role of depression in the aetiology and prognosis of coronary heart disease were sought through two methods: a) PubMed, and b) citation tracking where Science Citation Index was searched for all articles which cited ("forward citation tracking") or were cited by ("backward citation tracking") any of the articles in an index review. The number and quality of eligible articles identified by the two methods were compared. RESULTS: 50 articles that were not already included in the index review met our inclusion criteria; 11 were identified through Science Citation Index alone, 8 through PubMed alone, and 31 through both methods. Articles identified by Science Citation Index alone were published in higher impact factor journals, were larger and were less likely to show a positive association. CONCLUSION: Science Citation Index identified more eligible articles than PubMed, and these differed qualitatively. Failing to use citation tracking in a systematic review of observational studies may result in bias.

Abstracting and Indexing↗

Hubs of knowledge: using the functional link structure in Biozon to mine for biologically significant entities.

BACKGROUND: Existing biological databases support a variety of queries such as keyword or definition search. However, they do not provide any measure of relevance for the instances reported, and result sets are usually sorted arbitrarily. RESULTS: We describe a system that builds upon the complex infrastructure of the Biozon database and applies methods similar to those of Google to rank documents that match queries. We explore different prominence models and study the spectral properties of the corresponding data graphs. We evaluate the information content of principal and non-principal eigenspaces, and test various scoring functions which combine contributions from multiple eigenspaces. We also test the effect of similarity data and other variations which are unique to the biological knowledge domain on the quality of the results. Query result sets are assessed using a probabilistic approach that measures the significance of coherence between directly connected nodes in the data graph. This model allows us, for the first time, to compare different prominence models quantitatively and effectively and to observe unique trends. CONCLUSION: Our tests show that the ranked query results outperform unsorted results with respect to our significance measure and the top ranked entities are typically linked to many other biological entities. Our study resulted in a working ranking system of biological entities that was integrated into Biozon at http://biozon.org.

Abstracting and Indexing↗