Search PubMed⌕ Search

PubMed · 12477931

Soft and hard classification by reproducing kernel Hilbert space methods.

Abstract

Reproducing kernel Hilbert space (RKHS) methods provide a unified context for solving a wide variety of statistical modelling and function estimation problems. We consider two such problems: We are given a training set [yi, ti, i = 1, em leader, n], where yi is the response for the ith subject, and ti is a vector of attributes for this subject. The value of y(i) is a label that indicates which category it came from. For the first problem, we wish to build a model from the training set that assigns to each t in an attribute domain of interest an estimate of the probability pj(t) that a (future) subject with attribute vector t is in category j. The second problem is in some sense less ambitious; it is to build a model that assigns to each t a label, which classifies a future subject with that t into one of the categories or possibly "none of the above." The approach to the first of these two problems discussed here is a special case of what is known as penalized likelihood estimation. The approach to the second problem is known as the support vector machine. We also note some alternate but closely related approaches to the second problem. These approaches are all obtained as solutions to optimization problems in RKHS. Many other problems, in particular the solution of ill-posed inverse problems, can be obtained as solutions to optimization problems in RKHS and are mentioned in passing. We caution the reader that although a large literature exists in all of these topics, in this inaugural article we are selectively highlighting work of the author, former students, and other collaborators.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Grace Wahba. 2002-12-11. Soft and hard classification by reproducing kernel Hilbert space methods.. https://doi.org/10.1073/pnas.242574899

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Development of an ecologic marine classification in the new zealand region.

We describe here the development of an ecosystem classification designed to underpin the conservation management of marine environments in the New Zealand region. The classification was defined using multivariate classification using explicit environmental layers chosen for their role in driving spatial variation in biologic patterns: depth, mean annual solar radiation, winter sea surface temperature, annual amplitude of sea surface temperature, spatial gradient of sea surface temperature, summer sea surface temperature anomaly, mean wave-induced orbital velocity at the seabed, tidal current velocity, and seabed slope. All variables were derived as gridded data layers at a resolution of 1 km. Variables were selected by assessing their degree of correlation with biologic distributions using separate data sets for demersal fish, benthic invertebrates, and chlorophyll-a. We developed a tuning procedure based on the Mantel test to refine the classification's discrimination of variation in biologic character. This was achieved by increasing the weighting of variables that play a dominant role and/or by transforming variables where this increased their correlation with biologic differences. We assessed the classification's ability to discriminate biologic variation using analysis of similarity. This indicated that the discrimination of biologic differences generally increased with increasing classification detail and varied for different taxonomic groups. Advantages of using a numeric approach compared with geographic-based (regionalisation) approaches include better representation of spatial patterns of variation and the ability to apply the classification at widely varying levels of detail. We expect this classification to provide a useful framework for a range of management applications, including providing frameworks for environmental monitoring and reporting and identifying representative areas for conservation.

Classification↗

SINEs of progress: Mobile element applications to molecular ecology.

Mobile elements represent a unique and under-utilized set of tools for molecular ecologists. They are essentially homoplasy-free characters with the ability to be genotyped in a simple and efficient manner. Interpretation of the data generated using mobile elements can be simple compared to other genetic markers. They exist in a wide variety of taxa and are useful over a wide selection of temporal ranges within those taxa. Furthermore, their mode of evolution instills them with another advantage over other types of multilocus genotype data: the ability to determine loci applicable to a range of time spans in the history of a taxon. In this review, I discuss the application of mobile element markers, especially short interspersed elements (SINEs), to phylogenetic and population data, with an emphasis on potential applications to molecular ecology.

Classification↗