Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “classifier”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Differential gene expression as a potential classifier of 2-(4-amino-3-methylphenyl)-5-fluorobenzothiazole-sensitive and -insensitive cell lines.

2-(4-Amino-3-methylphenyl)-5-fluorobenzothiazole (5F-203) is a candidate antitumor agent empirically discovered with the aid of the National Cancer Institute (NCI) Anticancer Drug Screen. In an effort to determine whether basal expression of genes could be used to classify cell sensitivity to this agent, serial analysis of gene expression (SAGE) libraries were generated for three sensitive and two insensitive human tumor cell lines. When the SAGE tags expressed within these cell line libraries were compared and evaluated for differences, several genes seemed more highly expressed in 5F-203-sensitive cell lines than in the insensitive cell lines. Constitutive expressions of 15 genes identified by the analysis were then measured by quantitative reverse-transcription polymerase chain reaction in the 60 cell lines of the NCI Anticancer Drug Screen. This generated a pattern of relative basal gene expression across the cell lines and also confirmed the differential expression of SAGE-discovered genes within the initial set of five cell lines. Further analyses of these expression data in 60 cell lines suggested that a smaller subset of genes could be used to classify tumor cell sensitivity to 5F-203. In contrast, the same set of genes did not predict with equivalent precision sensitivity to unrelated active drugs, and another set of genes could not better classify the cell lines in terms of 5F-203 sensitivity. These results suggest that constitutive gene expression profiles, in which the genes are not necessarily known to be related to the mechanism of action of a given drug, may be viewed as a general tool to extend and improve the concept of a single predictive gene to groups of predictive genes.

Cell Line, Tumor↗

Altered enzyme-linked immunosorbent assay immunoglobulin M (IgM)/IgG optical density ratios can correctly classify all primary or secondary dengue virus infections 1 day after the onset of symptoms, when all of the viruses can be isolated.

We compared dengue virus (DV) isolation rates and tested whether acute primary (P) and acute/probable acute secondary (S/PS) DV infections could be correctly classified serologically when the patients' first serum (S1) samples were obtained 1 to 3 days after the onset of symptoms (AOS). DV envelope/membrane protein-specific immunoglobulin M (IgM) capture and IgG capture enzyme-linked immunosorbent assay (ELISA) titrations (1/log(10) 1.7 to 1 log(10) 6.6 dilutions) were performed on 100 paired S1 and S2 samples from suspected DV infections. The serologically confirmed S/PS infections were divided into six subgroups based on their different IgM and IgG responses. Because of their much greater dynamic ranges, IgG/IgM ELISA titer ratios were more accurate and reliable than IgM/IgG optical density (OD) ratios recorded at a single cutoff dilution for discriminating between P and S/PS infections. However, 62% of these patients' S1 samples were DV IgM and IgG titer negative ( or=2.60 and <2.60) discriminatory IgM/IgG OD (DOD) ratios on these S1 samples than those published previously to correctly classify the highest percentage of these P and S/PS infections. The DV isolation rate was highest (12/12; 100%) using IgG and IgM titer-negative S1 samples collected 1 day AOS, when 100% of them were correctly classified as P or S/PS infections using these higher DOD ratios.

Colombia↗

Total and occupationally active life expectancies in relation to social class and marital status in men classified as healthy at 20 in Finland.

STUDY OBJECTIVE: To study differences in total life expectancy and in occupationally active life expectancy in relation to social class and marital status in men classified as healthy as young adults. DESIGN: Historical cohort study. SETTING: Finland. PARTICIPANTS: Altogether 1662 men classified as completely healthy at the time of induction to military service (mean birth year 1923), who had been selected as referents for a study of former athletes. Mean follow up time was 46 years. MEASUREMENTS: Vital status was determined by follow up through local parish data up to 1990. Mortality data were obtained from the Cause of Death bureau of the Central Statistical Office of Finland. Occurrence of work disability was assessed from nationwide disability pension register data. Mean total life expectancy and mean occupationally active life expectancy (end points disability pension or death before age 65 years) were estimated. Social class was based on the major lifetime occupation, while marital status was classified as "never married" or "ever married" at the end of follow up. MAIN RESULTS: Mean total life expectancy was highest among executives and managers (73.2 (95% confidence interval (CI): 70.3, 76.1) years), next highest in clerical (white collar) workers (72.0 (70.0, 74.1) years), and lowest in unskilled blue collar workers (63.65 (61.1, 66.2) years). Skilled workers and farmers were intermediate. For the occupationally active life expectancy estimates, a similar gradient was observed: highest for executives (61.9 (60.7, 63.1) years) and lowest for the unskilled (52.2 (50.2, 54.2) years). The ratio of occupationally active life expectancy to total life expectancy was highest for executives (85%) and lowest for farmers (81%) and unskilled workers (82%). CONCLUSIONS: The social class gradient known to exist for mortality is also present for occupational disability. Social class and marital status differences in mortality are already evident in early adulthood and continue into old age. Those with the highest life expectancy also have the largest proportion of their life span free of occupationally incapacitating disability.

Adult↗

Comparison of multilayer neural network and Nearest Neighbor Classifiers for handwritten digit recognition.

The basic Nearest Neighbor Classifier (NNC) is often inefficient for classification in terms of memory space and computing time needed if all training samples are used as prototypes. These problems can be solved by reducing the number of prototypes using clustering algorithms and optimizing the prototypes using a special neural network model. In this paper, we compare the performance of the multilayer neural network and an Optimized Nearest Neighbor Classifier (ONNC) for handwritten digit recognition applications. We show that an ONNC can have the same recognition performance as an equivalent neural network classifier. The ONNC can be efficiently implemented using prototype and variable ranking, partial summation and distance triangular inequality based strategies. It requires the same memory space as, but less, training time and classification time than the neural network.

Computers↗

Empirical error-confidence curves for neural network and Gaussian classifiers.

"Error-Confidence" measures the probability that the proportion of errors made by a classifier will be within epsilon of EB, the optimal (Bayes) error. Probably Almost Bayes (PAB) theory attempts to quantify how this confidence increases with the number of training samples. We investigate the relationship empirically by comparing average error versus number of training patterns (m) for linear and neural network classifiers. On Gaussian problems, the resulting EC curves demonstrate that the PAB bounds are extremely conservative. Asymptotic statistics predicts a linear relationship between the logarithms of the average error and the number of training patterns. For low Bayes error rates we found excellent agreement between the prediction and the linear discriminant performance. At higher Bayes error rates we still found a linear relationship, but with a shallower slope than the predicted-1. When the underlying true model is a three-layer network, the EC curves show a greater dependence on classifier capacity, and the linear predictions no longer seem to hold.

Bayes Theorem↗

Protein structure and fold prediction using Tree-Augmented naïve Bayesian classifier.

Due to the large volume of protein sequence data, computational methods to determine the structure class and the fold class of a protein sequence have become essential. Several techniques based on sequence similarity, Neural Networks, Support Vector Machines (SVMs), etc. have been applied. Since most of these classifiers use binary classifiers for multi-classification, there may be (N) c2 classifiers required. This paper presents a framework using the Tree-Augmented Bayesian Networks (TAN) which performs multi-classification based on the theory of learning Bayesian Networks and using improved feature vector representation of (Ding et al., 2001). In order to enhance TAN's performance, pre-processing of data is done by feature discretization and post-processing is done by using Mean Probability Voting (MPV) scheme. The advantage of using Bayesian approach over other learning methods is that the network structure is intuitive. In addition, one can read off the TAN structure probabilities to determine the significance of each feature (say, hydrophobicity) for each class, which helps to further understand the complexity in protein structure. The experiments on the datasets used in three prominent recent works show that our approach is more accurate than other discriminative methods. The framework is implemented on the BAYESPROT web server and it is available at http://www-appn.comp.nus.edu.sg/~bioinfo/bayesprot/Default.htm. More detailed results are also available on the above website.

Algorithms↗

Attempts to physiologically classify human thenar motor units.

1. This study was designed to determine whether human thenar motor units can be classified into types by the same physiological criteria used for other mammalian limb motor units and to consider whether such classification is functionally relevant. 2. Contractile responses of 25 human thenar single motor units were examined when their motor axons were stimulated intraneurally at rates from 1 to 100 Hz and intermittently at 40 Hz in a conventional 2-min fatigue test. Twitch and tetanic forces were measured together with various indexes of contractile rate. 3. Twitch contraction times and subtetanic to maximum tetanic force ratios were both distributed continuously. "Sag" in tension was not evident in unfused force profiles. Thus these units could not be divided into fast and slow types by the use of traditional contractile rate criteria. 4. Most units were fatigue resistant, with force fatigue indexes (FI) ranging from 0.33 to 1.14. None could be classified as fatiguable (FI less than 0.25). Seven units (28%) fell into the fatigue-intermediate (FI = 0.25-0.75) category, whereas 18 units (72%) had FI greater than 0.75, i.e., they were fatigue-resistant units. However, these units could not be classified by conventional FI and contractile rate criteria, because fatigue-resistant and fatigue-intermediate units had similar contractile rates. 5. Additional FI were calculated to describe changes in contractile rate. During the fatigue test, units behaved in one of three ways, showing 1) little change in either force or rate; 2) contractile slowing during the contraction and relaxation phases, with little or no force loss; or 3) both force and rate reduction.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

Prediction of radiation sensitivity using a gene expression classifier.

The development of a successful radiation sensitivity predictive assay has been a major goal of radiation biology for several decades. We have developed a radiation classifier that predicts the inherent radiosensitivity of tumor cell lines as measured by survival fraction at 2 Gy (SF2), based on gene expression profiles obtained from the literature. Our classifier correctly predicts the SF2 value in 22 of 35 cell lines from the National Cancer Institute panel of 60, a result significantly different from chance (P = 0.0002). In our approach, we treat radiation sensitivity as a continuous variable, significance analysis of microarrays is used for gene selection, and a multivariate linear regression model is used for radiosensitivity prediction. The gene selection step identified three novel genes (RbAp48, RGS19, and R5PIA) of which expression values are correlated with radiation sensitivity. Gene expression was confirmed by quantitative real-time PCR. To biologically validate our classifier, we transfected RbAp48 into three cancer cell lines (HS-578T, MALME-3M, and MDA-MB-231). RbAp48 overexpression induced radiosensitization (1.5- to 2-fold) when compared with mock-transfected cell lines. Furthermore, we show that HS-578T-RbAp48 overexpressors have a higher proportion of cells in G2-M (27% versus 5%), the radiosensitive phase of the cell cycle. Finally, RbAp48 overexpression is correlated with dephosphorylation of Akt, suggesting that RbAp48 may be exerting its effect by antagonizing the Ras pathway. The implications of our findings are significant. We establish that radiation sensitivity can be predicted based on gene expression profiles and we introduce a genomic approach to the identification of novel molecular markers of radiation sensitivity.

Carrier Proteins↗

Exploration of the precision of classifying sudden cardiac death. Implications for the interpretation of clinical trials.

BACKGROUND: As cardiovascular clinical trials improve in sophistication and therapies target specific cardiac mechanisms of death, a more objective and precise system to identify specific cause of death is needed. Ideally, sudden cardiac death would describe patients dying of ventricular tachycardia and ventricular fibrillation. In this context, we explored the precision of current sudden death classification and implications for clinical trials. METHODS AND RESULTS: Deaths were analyzed in 834 patients who received an automatic implantable cardioverter-defibrillator (ICD). Three arrhythmia experts used a standard prospective classification system to classify deaths into accepted categories: sudden cardiac, nonsudden cardiac, and noncardiac. New aspects to this study included analysis of autopsy results and ICD interrogation for arrhythmias at the time of death. All of the patients receiving the ICD previously had documented sustained ventricular tachycardia/fibrillation or cardiac arrest. Of the 109 subsequent deaths in the 834-patient database, 17 (16%) were classified as sudden cardiac. Compared with the nonsudden cardiac and noncardiac categories, sudden cardiac death was more often identified in outpatients (59% versus 10%) and witnessed less often (41% versus 86%; both P < .001). The autopsy information contradicted and changed the clinical perception of a "sudden cardiac death" in 7 cases (myocardial infarction [n = 1], pulmonary embolism [n = 2], cerebral infarction [n = 1], ruptured thoracic [n = 1], and abdominal aortic aneurysms [n = 2]). Interpretable ICD interrogation was available in 53% of the deaths (47% unavailable: buried, programmed off, or other technical reasons). When evaluated, only 7 of 17 "sudden deaths" were associated with ICD discharges near the time of death. CONCLUSIONS: Even in a group of patients with an ICD, deaths classified as sudden cardiac frequently were not associated with ventricular tachycardia or ventricular fibrillation and were often noncardiac. It is possible to create a wide range of sudden cardiac death rates (more than fourfold) using the identical clinical database despite objective, prespecified criteria. Autopsy results frequently reveal noncardiac causes of clinical events simulating sudden cardiac death. ICD interrogation revealed that ICD discharges were often related to terminal arrhythmias incidental to the primary pathophysiological process leading to death.

Cause of Death↗

Bayesian framework for least-squares support vector machine classifiers, gaussian processes, and kernel Fisher discriminant analysis.

The Bayesian evidence framework has been successfully applied to the design of multilayer perceptrons (MLPs) in the work of MacKay. Nevertheless, the training of MLPs suffers from drawbacks like the nonconvex optimization problem and the choice of the number of hidden units. In support vector machines (SVMs) for classification, as introduced by Vapnik, a nonlinear decision boundary is obtained by mapping the input vector first in a nonlinear way to a high-dimensional kernel-induced feature space in which a linear large margin classifier is constructed. Practical expressions are formulated in the dual space in terms of the related kernel function, and the solution follows from a (convex) quadratic programming (QP) problem. In least-squares SVMs (LS-SVMs), the SVM problem formulation is modified by introducing a least-squares cost function and equality instead of inequality constraints, and the solution follows from a linear system in the dual space. Implicitly, the least-squares formulation corresponds to a regression formulation and is also related to kernel Fisher discriminant analysis. The least-squares regression formulation has advantages for deriving analytic expressions in a Bayesian evidence framework, in contrast to the classification formulations used, for example, in gaussian processes (GPs). The LS-SVM formulation has clear primal-dual interpretations, and without the bias term, one explicitly constructs a model that yields the same expressions as have been obtained with GPs for regression. In this article, the Bayesian evidence framework is combined with the LS-SVM classifier formulation. Starting from the feature space formulation, analytic expressions are obtained in the dual space on the different levels of Bayesian inference, while posterior class probabilities are obtained by marginalizing over the model parameters. Empirical results obtained on 10 public domain data sets show that the LS-SVM classifier designed within the Bayesian evidence framework consistently yields good generalization performances.

Artificial Intelligence↗

The diabolo classifier

We present a new classification architecture based on autoassociative neural networks that are used to learn discriminant models of each class. The proposed architecture has several interesting properties with respect to other model-based classifiers like nearest-neighbors or radial basis functions: it has a low computational complexity and uses a compact distributed representation of the models. The classifier is also well suited for the incorporation of a priori knowledge by means of a problem-specific distance measure. In particular, we will show that tangent distance (Simard, Le Cun, & Denker, 1993) can be used to achieve transformation invariance during learning and recognition. We demonstrate the application of this classifier to optical character recognition, where it has achieved state-of-the-art results on several reference databases. Relations to other models, in particular those based on principal component analysis, are also discussed.

Journal Article↗

Accuracy-based learning classifier systems: models, analysis and applications to classification tasks.

Recently, Learning Classifier Systems (LCS) and particularly XCS have arisen as promising methods for classification tasks and data mining. This paper investigates two models of accuracy-based learning classifier systems on different types of classification problems. Departing from XCS, we analyze the evolution of a complete action map as a knowledge representation. We propose an alternative, UCS, which evolves a best action map more efficiently. We also investigate how the fitness pressure guides the search towards accurate classifiers. While XCS bases fitness on a reinforcement learning scheme, UCS defines fitness from a supervised learning scheme. We find significant differences in how the fitness pressure leads towards accuracy, and suggest the use of a supervised approach specially for multi-class problems and problems with unbalanced classes. We also investigate the complexity factors which arise in each type of accuracy-based LCS. We provide a model on the learning complexity of LCS which is based on the representative examples given to the system. The results and observations are also extended to a set of real world classification problems, where accuracy-based LCS are shown to perform competitively with respect to other learning algorithms. The work presents an extended analysis of accuracy-based LCS, gives insight into the understanding of the LCS dynamics, and suggests open issues for further improvement of LCS on classification tasks.

Algorithms↗

Bounding the effect of noise in multiobjective learning classifier systems.

This paper analyzes the impact of using noisy data sets in Pittsburgh-style learning classifier systems. This study was done using a particular kind of learning classifier system based on multiobjective selection. Our goal was to characterize the behavior of this kind of algorithms when dealing with noisy domains. For this reason, we developed a theoretical model for predicting the minimal achievable error in noisy domains. Combining this theoretical model for crisp learners with graphical representations of the evolved hypotheses through multiobjective techniques, we are able to bound the behavior of a learning classifier system. This kind of modeling lets us identify relevant characteristics of the evolved hypotheses, such as overfitting conditions that lead to hypotheses that poorly generalize the concept to be learned.

Algorithms↗

Rule fitness and pathology in learning classifier systems.

It has long been known that in some relatively simple reinforcement learning tasks traditional strength-based classifier systems will adapt poorly and show poor generalisation. In contrast, the more recent accuracy-based XCS, appears both to adapt and generalise well. In this work, we attribute the difference to what we call strong over general and fit over general rules. We begin by developing a taxonomy of rule types and considering the conditions under which they may occur. In order to do so an extreme simplification of the classifier system is made, which forces us toward qualitative rather than quantitative analysis. We begin with the basics, considering definitions for correct and incorrect actions, and then correct, incorrect, and overgeneral rules for both strength and accuracy-based fitness. The concept of strong overgeneral rules, which we claim are the Achilles' heel of strength-based classifier systems, are then analysed. It is shown that strong overgenerals depend on what we call biases in the reward function (or, in sequential tasks, the value function). We distinguish between strong and fit overgeneral rules, and show that although strong overgenerals are fit in a strength-based system called SB-XCS, they are not in XCS. Next we show how to design fit overgeneral rules for XCS (but not SB-XCS), by introducing biases in the variance of the reward function, and thus that each system has its own weakness. Finally, we give some consideration to the prevalence of reward and variance function bias, and note that non-trivial sequential tasks have highly biased value functions.

Algorithms↗

The use of sputum cultures in the evaluation of immigrants classified as tuberculosis suspects.

Because of possible deficiencies in the evaluation, based on symptoms and chest roentgenogram review, of new immigrants classified during the visa application process as tuberculosis suspects, a prospective (cohort) and a retrospective (case control) study were done to test the usefulness of routinely obtaining sputum specimens for culture in that setting. In the prospective study, 249 consecutive classified immigrants who were considered on the basis of clinical and roentgenographic findings to have nonprogressive tuberculosis submitted at least two sputums for culture: 13 (5.2%) had at least one culture positive for M. tuberculosis. Immigrants younger than 50 yr of age and refugees from Kampuchea and Laos had a fivefold to tenfold elevated risk of having a positive sputum culture. The cost per case detected of obtaining and processing sputum cultures was estimated to be +1,996 to +2,994. In the case-control study, 37 classified immigrants evaluated from 1981 through 1986 who had sputum cultures positive for M. tuberculosis even though they fulfilled clinical and roentgenographic criteria for nonprogressive tuberculosis served as control subjects. Several demographic, clinical, and roentgenographic factors were associated with an increased risk of being culture-positive: age younger than 50 yr, a positive tuberculin test, report of a cough, and a cavitary lesion on chest roentgenogram. The history of prior receipt of antituberculosis drugs was associated with having a negative culture, including a marked dose-response effect.

Asia, Southeastern↗

Heidelberg retina tomograph measurements of the optic disc and parapapillary retina for detecting glaucoma analyzed by machine learning classifiers.

PURPOSE: To determine whether topographical measurements of the parapapillary region analyzed by machine learning classifiers can detect early to moderate glaucoma better than similarly processed measurements obtained within the disc margin and to improve methods for optimization of machine learning classifier feature selection. METHODS: One eye of each of 95 patients with early to moderate glaucomatous visual field damage and of each of 135 normal subjects older than 40 years participating in the longitudinal Diagnostic Innovations in Glaucoma Study (DIGS) were included. Heidelberg Retina Tomograph (HRT; Heidelberg Engineering, Dossenheim, Germany) mean height contour was measured in 36 equal sectors, both along the disc margin and in the parapapillary region (at a mean contour line radius of 1.7 mm). Each sector was evaluated individually and in combination with other sectors. Gaussian support vector machine (SVM) learning classifiers were used to interpret HRT sector measurements along the disc margin and in the parapapillary region, to differentiate between eyes with normal and glaucomatous visual fields and to compare the results with global and regional HRT parameter measurements. The area under the receiver operating characteristic (ROC) curve was used to measure diagnostic performance of the HRT parameters and to evaluate the cross-validation strategies and forward selection and backward elimination optimization techniques that were used to generate the reduced feature sets. RESULTS: The area under the ROC curve for mean height contour of the 36 sectors along the disc margin was larger than that for the mean height contour in the parapapillary region (0.97 and 0.85, respectively). Of the 36 individual sectors along the disc margin, those in the inferior region between 240 degrees and 300 degrees, had the largest area under the ROC curve (0.85-0.91). With SVM Gaussian techniques, the regional parameters showed the best ability to discriminate between normal eyes and eyes with glaucomatous visual field damage, followed by the global parameters, mean height contour measures along the disc margin, and mean height contour measures in the parapapillary region. The area under the ROC curve was 0.98, 0.94, 0.93, and 0.85, respectively. Cross-validation and optimization techniques demonstrated that good discrimination (99% of peak area under the ROC curve) can be obtained with a reduced number of HRT parameters. CONCLUSIONS: Mean height contour measurements along the disc margin discriminated between normal and glaucomatous eyes better than measurements obtained in the parapapillary region.

Area Under Curve↗

Development and comparison of automated classifiers for glaucoma diagnosis using Stratus optical coherence tomography.

PURPOSE: To develop and compare the ability of several automated classifiers to differentiate between normal and glaucomatous eyes based on the quantitative assessment of summary data reports from Stratus optical coherence tomography (OCT; Carl Zeiss Meditec Inc., Dublin, CA) in a Chinese population in Taiwan. METHODS: One randomly selected eye from each of 89 patients with glaucoma and each of 100 age- and sex-matched normal individuals were included in the study. Measurements of glaucoma variables (retinal nerve fiber layer thickness and optic nerve head analysis results) were obtained by Stratus OCT. With the Stratus OCT parameters used as input, receiver operative characteristic (ROC) curves were generated by three methods, to classify eyes as either glaucomatous or normal: linear discriminant analysis (LDA), Mahalanobis distance (MD), and artificial neural network (ANN). The area under the ROC curve was optimized by principal component analysis (PCA). Classification accuracy was determined by cross validation. RESULTS: The average visual field mean deviation was -0.7 +/- 0.6 dB in the normal group and -2.7 +/- 1.9 dB in the glaucoma group. The areas under the ROC curves were 0.824 (LDA), 0.849 (MD), 0.821 (ANN), 0.915 (LDA with PCA), 0.991 (MD with PCA), and 0.874 (ANN with PCA). CONCLUSIONS: With Stratus OCT parameters used as input, automated classifiers show promise for discriminating between glaucomatous and normal eyes. MD measured from multivariate data can predict the severity of glaucoma through the construction of a measurement space. After PCA, implementation results show that the Mahalanobis space created by MD surpasses LDA and ANN in diagnosing glaucoma.

Adult↗

A comparison of taxonomic systems for classifying homeless men.

The present study compared the relative merits of two taxonomic systems for classifying homeless men. One system classified homeless men based on their past history of psychiatric disability. The other system classified individuals on the basis of their current psychiatric impairment. Both classification systems displayed significant discriminating power using a set of predictor variables that included demographic variables, childhood happiness, current life satisfaction, social support, stressful life events, and history of homelessness. Based on the percentage of correct classifications the system based on current impairment was superior to the system based on past history.

Demography↗