Search PubMed⌕ Search

PubMed · 9929278

Building manageable rough set classifiers.

Abstract

An interesting aspect of techniques for data mining and knowledge discovery is their potential for generating hypotheses by discovering underlying relationships buried in the data. However, the set of possible hypotheses is often very large and the extracted models may become prohibitively complex. It is therefore typically desirable to only consider the "strongest" hypotheses, so that smaller models can be obtained that also retain good classificatory capabilities. This paper outlines how rule-based classifiers based on rough set theory and Boolean reasoning that are both small and perform well can be developed. Applied to a real-world medical dataset, the final models are shown to exhibit good performance using only a subset of the available information. Furthermore, the number of resulting rules is low and enables practical a posteriori inspection and interpretation of the models.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A Ohrn, L Ohno-Machado, T Rowland. 1998. Building manageable rough set classifiers.. https://pubmed.ncbi.nlm.nih.gov/9929278/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Development of an ecologic marine classification in the new zealand region.

We describe here the development of an ecosystem classification designed to underpin the conservation management of marine environments in the New Zealand region. The classification was defined using multivariate classification using explicit environmental layers chosen for their role in driving spatial variation in biologic patterns: depth, mean annual solar radiation, winter sea surface temperature, annual amplitude of sea surface temperature, spatial gradient of sea surface temperature, summer sea surface temperature anomaly, mean wave-induced orbital velocity at the seabed, tidal current velocity, and seabed slope. All variables were derived as gridded data layers at a resolution of 1 km. Variables were selected by assessing their degree of correlation with biologic distributions using separate data sets for demersal fish, benthic invertebrates, and chlorophyll-a. We developed a tuning procedure based on the Mantel test to refine the classification's discrimination of variation in biologic character. This was achieved by increasing the weighting of variables that play a dominant role and/or by transforming variables where this increased their correlation with biologic differences. We assessed the classification's ability to discriminate biologic variation using analysis of similarity. This indicated that the discrimination of biologic differences generally increased with increasing classification detail and varied for different taxonomic groups. Advantages of using a numeric approach compared with geographic-based (regionalisation) approaches include better representation of spatial patterns of variation and the ability to apply the classification at widely varying levels of detail. We expect this classification to provide a useful framework for a range of management applications, including providing frameworks for environmental monitoring and reporting and identifying representative areas for conservation.

Classification↗

SINEs of progress: Mobile element applications to molecular ecology.

Mobile elements represent a unique and under-utilized set of tools for molecular ecologists. They are essentially homoplasy-free characters with the ability to be genotyped in a simple and efficient manner. Interpretation of the data generated using mobile elements can be simple compared to other genetic markers. They exist in a wide variety of taxa and are useful over a wide selection of temporal ranges within those taxa. Furthermore, their mode of evolution instills them with another advantage over other types of multilocus genotype data: the ability to determine loci applicable to a range of time spans in the history of a taxon. In this review, I discuss the application of mobile element markers, especially short interspersed elements (SINEs), to phylogenetic and population data, with an emphasis on potential applications to molecular ecology.

Classification↗