Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Discriminant Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

A discriminant analysis extension to mixed models.

Discriminant analysis is commonly used to classify an observation into one of two (or more) populations on the basis of correlated measurements. Classical discriminant analysis approaches require complete data for all observations. Our extension enables the use of all available longitudinal data, regardless of completeness. Traditionally a linear discriminant function assumes a common unstructured covariance matrix for both populations, which may be taken from a multivariate model. Here, we can model the correlated measurements and use a structured covariance in the discriminant function. We illustrate cases in which the estimated covariance structure is either compound symmetric, heterogeneous compound symmetric or heterogeneous autoregressive. Thus a structured covariance is incorporated into the discrimination process in contrast to standard discriminant analysis methodology. Simulations are performed to obtain a true measure of the effect of structure on the error rate. In addition, the usual multivariate expected value structure is altered. The impact on the discrimination process is contrasted when using the multivariate and random-effects covariance structures and expected values. The random-effects covariance structure leads to an improvement in the error rate in small samples. To illustrate the procedure we consider repeated measurements data from a clinical trial comparing two active treatments; the goal is to determine if the treatment could be unblinded based on repeated anxiety score measurements.

Anti-Anxiety Agents↗

Implementation of linear and quadratic discriminant analysis incorporating costs of misclassification.

Discriminant analysis plays an important role in biological and medical research. In practice, standard linear and quadratic methods are often applied which assume equal costs of misclassification. However, there can be situations where misclassifications between certain groups may be more serious than between other groups. Such considerations can be taken into account by using classification methods which incorporate misclassification costs. The widely applied statistical packages BMDP, SAS, and SPSS do not offer the possibility of using unequal misclassification costs for discriminant analysis with more than two groups. In this paper a menu-driven, user-friendly PC program written in Borland Pascal is introduced which performs linear and quadratic discriminant analysis for g > or = 2 groups allowing for the incorporation of misclassification costs.

Bias↗

Beyond averaging. II. Single-trial classification of exogenous event-related potentials using stepwise discriminant analysis.

Using stepwise discriminant analysis (SWDA), single-trial event-related potentials (ERPs) were classified as to whether they were elicited by a checkerboard presented to the upper or lower visual half-field. Discriminant functions were computed on the basis of 'training sets' constructed of upper and lower half-field ERPs, and applied to 'test sets' of other ERPs elicited by the same stimuli. Individual-subject discriminant functions for data recorded at PZ classified the single ERPs in the test sets with a mean accuracy of 83.7% correct. The mean accuracy attained by individual-subject functions from the most discriminable scalp site for each subject was 87.8% correct, and that attained by an across-subjects function was 78.1% correct. Averaged ERPs showed the previously reported polarity reversal of corresponding exogenous components in the upper and lower half field wave forms. Moreover, the SWDA procedure chose ERP time points at the latencies of these exogenous components for discriminating the half-field ERPs. The results demonstrate that SWDA can accurately classify single ERPs in which the systematic variance is localized in exogenous components, having periods within the range of frequencies which typically compose the background EEG.

Adult↗

[Canonical discriminant analysis for hematological and serum biochemical changes during pregnancy period in squirrel monkeys (Saimiri sciurea)].

Hematological and serum biochemical data obtained from non-pregnant, pregnant and post-partum squirrel monkeys (Saimiri sciurea) were analyzed by canonical discriminant analysis (discriminant analysis with reduction of dimensionality). All animals were of wild origin and had been maintained under uniform environmental conditions at Tsukuba Primate Center for Medical Science, N.I.H., Japan. Months were standardized by the day of parturition. The calculated arithmetic means and standard deviations were listed for each item of measurement performed. Items detected statistically significant difference (p less than 0.01) between months were as follows: red blood cell count (RBC), mean corpuscular volume (MCV), hematocrit value (Ht), hemoglobin concentration (Hb), white blood cell count (WBC), albumin concentration (ALB), blood urea nitrogen concentration (BUN), total cholesterol concentration (T-CHO), triglyceride concentration (TG), alkaline phosphatase activity (ALP) and calcium concentration (Ca). Results of canonical discriminant analysis showed that the value of the first canonical variate (Z1) decreased from the early period of pregnancy to the middle period, and that the second canonical variate (Z2) decreased from the middle period of pregnancy to the end of pregnancy. The meaning of their changes were discussed.

Alkaline Phosphatase↗

Enhancing the usefulness of quadrant analysis in hospitals: a marriage with discriminant analysis.

This paper first looks at traditional quadrant analysis as used in contemporary research. Next, we show an extension to include multiple brands on the same quadrant chart--again consistent with current practice. Then, multiple discriminant analysis results are merged with the quadrant chart data. This makes the resultant chart even more actionable. While the example uses health care data, the methods outlined can be used in other segments of the service industry, as well as for more tangible products. The only limit is the researcher's imagination.

Consumer Behavior↗

Multivariate discriminant analysis of the electromyographic interference pattern: statistical approach to discrimination among controls, myopathies and neuropathies.

The stepwise linear discriminant analysis method is used to develop optimal combinations of features measured from the electromyographic interference pattern, with the aim of minimising the misclassification rate in controls while maximising the correct classification rates in patients with disease. This discriminant analysis among multiple groups leads to the determination of the optimal discriminating surface in a multivariable space and can also produce a severity of disease likelihood index. Applying these combinations of features to 186 studies performed in the biceps muscle, 81% of all studies are accurately classified as being normal, myopathic or neuropathic. An algorithm to perform this stepwise multigroup linear discriminant analysis is described.

Diagnosis, Differential↗

[Canonical discriminant analysis for pregnancy-related changes in hematological and serum biochemical values in tamarins (Saguinus spp].

Hematological and biochemical values obtained from 9 monkeys (Saguinus labiatus and S. mistax) during pre- and postpartum periods were analyzed by canonical discriminant analysis (discriminant analysis with reduction of dimensionality). All animals used were of wild origin and had been maintained under uniform environmental conditions at N. I. H., Japan. The items examined were as follows: white blood cell count (WBC), red blood cell count (RBC), hematocrit value (Ht), hemoglobin concentration (Hb), total protein concentration (TP), blood urea nitrogen (BUN), albumin concentration (ALB), albumin-globulin ratio (A/G), glutamic oxaloacetic transaminase activity (GGT), glutamic pyruvic transaminase activity (GPT), alkaline phosphatase activity (ALP) and total cholesterol concentration (CHO). The data obtained in the pre- and postpartum periods were divided into six chronological groups. The prepartum period was divided into Group I: weeks 15-10; Group II: weeks 9-7; Group III: weeks 6-4; and Group IV: weeks 3-0. The postpartum period was divided into Group V: weeks 0-4 and Group VI: weeks 5-7). In the later pregnancy period (Groups III and IV), significant decreases in RBC, Ht, Hb, TP and ALB, and a significant increase in CHO were observed. These values in the blood and serum continued after delivery (Groups V and VI). Results of canonical discriminant analysis showed that the value of the first canonical variate decreased according to the progress of pregnancy. The postpartum groups showed negative values. Although groups in the early

Analysis of Variance↗

Analytical methods to differentiate similar electroencephalographic spectra: neural network and discriminant analysis.

Differences in electroencephalographic (EEG) power spectra obtained under similar, but not identical, conditions may be difficult to discern using standard techniques. Statistical analysis may not be useful because of the large number of comparisons necessary. Visual recognition of differences also may be difficult. A new technique, neural network analysis, has been used successfully in other problems of pattern recognition and classification. We examined a number of methods of classifying similar EEG data: standard statistical analysis (analysis of variance), visual recognition, discriminant analysis, and neural network analysis. Twenty-nine volunteers received either thiopental (n = 9), midazolam (n = 10), or propofol (n = 10) in sedative doses in 3 different studies. These drugs produced very similar changes in the EEG power spectra. Except for beta 2 power during thiopental infusion, differences between drugs could not be detected using analysis of variance. Visual categorization was correct in 72% of the baseline EEGs, 70% of thiopental EEGs, 27% of propofol EEGs, and 46% of midazolam EEGs. A classification neural network (Learning Vector Quantization network) containing a Kohonen hidden layer was able to successfully classify 57 of 58 EEG samples (of 4 minutes' duration). Discriminant analysis had a similar rate of success. This level of performance was achieved by dividing the EEG power spectrum from 1 to 30 Hz into 15 2-Hz bandwidths. When the EEG power spectrum was divided into the "classical" frequency bandwidths (alpha, beta 1, beta 2, theta, delta), both neural network and discriminant analysis performance deteriorated. By training the network using only certain inputs we were able to identify drug-specific bandwidths that seemed to be important in correct classification. We conclude that propofol, thiopental, and midazolam produce different effects on the EEG and that both neural network and discriminant analysis are useful in identifying these differences. We also conclude that EEG spectra should be analyzed without using classical EEG bands (alpha, beta, etc.). Additionally, neural networks can be used to identify frequency bands that are "important" in specific drug effects on the EEG. Once a classification algorithm is obtained using either a neural network or discriminant analysis, it could be used as an on-line monitor to recognize drug-specific EEG patterns.

Adult↗

Exploratory investigation into mild brain injury and discriminant analysis with high frequency bands (32-64 Hz).

QEEG variables (five activation, two relationship variables, 19 locations and five bands up to 64 Hertz) were collected under eyes closed condition (under both 32 and 64 Hertz conditions) on 91 subjects, consisting of 32 mild brain-injured subjects (no loss of consciousness greater than 20 minutes) and 52 normals over the age of 14. An additional seven subjects who were unconscious greater than 20 minutes were available for analysis. Previous discriminant function analysis developed by Thatcher et al. was employed on the eyes closed 32 Hertz condition to ascertain its robustness for time periods greater than 1 year and for significant periods of unconsciousness. A separate discriminant for subjects was developed, employing only frontal high frequency coherence figures. The Thatcher discriminant could reliably (79%) identify all subjects up to 43 years post accident. The high frequency discriminant effectively identified 87% of the brain injured across all time periods (without significant loss of consciousness) and 100% of subjects within 1 year of accident. The combination of the discriminants resulted in a 100% accuracy rate for the 39 brain injured subjects for which discriminate values were available.

Adolescent↗

Subclass discriminant analysis.

Over the years, many Discriminant Analysis (DA) algorithms have been proposed for the study of high-dimensional data in a large variety of problems. Each of these algorithms is tuned to a specific type of data distribution (that which best models the problem at hand). Unfortunately, in most problems the form of each class pdf is a priori unknown, and the selection of the DA algorithm that best fits our data is done over trial-and-error. Ideally, one would like to have a single formulation which can be used for most distribution types. This can be achieved by approximating the underlying distribution of each class with a mixture of Gaussians. In this approach, the major problem to be addressed is that of determining the optimal number of Gaussians per class, i.e., the number of subclasses. In this paper, two criteria able to find the most convenient division of each class into a set of subclasses are derived. Extensive experimental results are shown using five databases. Comparisons are given against Linear Discriminant Analysis (LDA), Direct LDA (DLDA), Heteroscedastic LDA (HLDA), Nonparametric DA (NDA), and Kernel-Based LDA (K-LDA). We show that our method is always the best or comparable to the best.

Algorithms↗

A two-stage linear discriminant analysis via QR-decomposition.

Linear Discriminant Analysis (LDA) is a well-known method for feature extraction and dimension reduction. It has been used widely in many applications involving high-dimensional data, such as image and text classification. An intrinsic limitation of classical LDA is the so-called singularity problems; that is, it fails when all scatter matrices are singular. Many LDA extensions were proposed in the past to overcome the singularity problems. Among these extensions, PCA+LDA, a two-stage method, received relatively more attention. In PCA+LDA, the LDA stage is preceded by an intermediate dimension reduction stage using Principal Component Analysis (PCA). Most previous LDA extensions are computationally expensive, and not scalable, due to the use of Singular Value Decomposition or Generalized Singular Value Decomposition. In this paper, we propose a two-stage LDA method, namely LDA/QR, which aims to overcome the singularity problems of classical LDA, while achieving efficiency and scalability simultaneously. The key difference between LDA/QR and PCA+LDA lies in the first stage, where LDA/QR applies QR decomposition to a small matrix involving the class centroids, while PCA+LDA applies PCA to the total scatter matrix involving all training data points. We further justify the proposed algorithm by showing the relationship among LDA/QR and previous LDA methods. Extensive experiments on face images and text documents are presented to show the effectiveness of the proposed algorithm.

Algorithms↗

Natural discriminant analysis using interactive Potts models.

Natural discriminant analysis based on interactive Potts models is developed in this work. A generative model composed of piece-wise multivariate gaussian distributions is used to characterize the input space, exploring the embedded clustering and mixing structures and developing proper internal representations of input parameters. The maximization of a log-likelihood function measuring the fitness of all input parameters to the generative model, and the minimization of a design cost summing up square errors between posterior outputs and desired outputs constitutes a mathematical framework for discriminant analysis. We apply a hybrid of the mean-field annealing and the gradient-descent methods to the optimization of this framework and obtain multiple sets of interactive dynamics, which realize coupled Potts models for discriminant analysis. The new learning process is a whole process of component analysis, clustering analysis, and labeling analysis. Its major improvement compared to the radial basis function and the support vector machine is described by using some artificial examples and a real-world application to breast cancer diagnosis.

Journal Article↗

Multivariate analysis of microarray data by principal component discriminant analysis: prioritizing relevant transcripts linked to the degradation of different carbohydrates in Pseudomonas putida S12.

The value of the multivariate data analysis tools principal component analysis (PCA) and principal component discriminant analysis (PCDA) for prioritizing leads generated by microarrays was evaluated. To this end, Pseudomonas putida S12 was grown in independent triplicate fermentations on four different carbon sources, i.e. fructose, glucose, gluconate and succinate. RNA isolated from these samples was analysed in duplicate on an anonymous clone-based array to avoid bias during data analysis. The relevant transcripts were identified by analysing the loadings of the principal components (PC) and discriminants (D) in PCA and PCDA, respectively. Even more specifically, the relevant transcripts for a specific phenotype could also be ranked from the loadings under an angle (biplot) obtained after PCDA analysis. The leads identified in this way were compared with those identified using the commonly applied fold-difference and hierarchical clustering approaches. The different data analysis methods gave different results. The methods used were complementary and together resulted in a comprehensive picture of the processes important for the different carbon sources studied. For the more subtle, regulatory processes in a cell, the PCDA approach seemed to be the most effective. Except for glucose and gluconate dehydrogenase, all genes involved in the degradation of glucose, gluconate and fructose were identified. Moreover, the transcriptomics approach resulted in potential new insights into the physiology of the degradation of these carbon sources. Indications of iron limitation were observed with cells grown on glucose, gluconate or succinate but not with fructose-grown cells. Moreover, several cytochrome- or quinone-associated genes seemed to be specifically up- or downregulated, indicating that the composition of the electron-transport chain in P. putida S12 might change significantly in fructose-grown cells compared to glucose-, gluconate- or succinate-grown cells.

Carbohydrate Metabolism↗

Generalizing discriminant analysis using the generalized singular value decomposition.

Discriminant analysis has been used for decades to extract features that preserve class separability. It is commonly defined as an optimization problem involving covariance matrices that represent the scatter within and between clusters. The requirement that one of these matrices be nonsingular limits its application to data sets with certain relative dimensions. We examine a number of optimization criteria, and extend their applicability by using the generalized singular value decomposition to circumvent the nonsingularity requirement. The result is a generalization of discriminant analysis that can be applied even when the sample size is smaller than the dimension of the sample data. We use classification results from the reduced representation to compare the effectiveness of this approach with some alternatives, and conclude with a discussion of their relative merits.

Algorithms↗

Prediction of clinical outcome with microarray data: a partial least squares discriminant analysis (PLS-DA) approach.

Partial least squares discriminant analysis (PLS-DA) is a partial least squares regression of a set Y of binary variables describing the categories of a categorical variable on a set X of predictor variables. It is a compromise between the usual discriminant analysis and a discriminant analysis on the significant principal components of the predictor variables. This technique is specially suited to deal with a much larger number of predictors than observations and with multicollineality, two of the main problems encountered when analysing microarray expression data. We explore the performance of PLS-DA with published data from breast cancer (Perou et al. 2000). Several such analyses were carried out: (1) before vs after chemotherapy treatment, (2) estrogen receptor positive vs negative tumours, and (3) tumour classification. We found that the performance of PLS-DA was extremely satisfactory in all cases and that the discriminant cDNA clones often had a sound biological interpretation. We conclude that PLS-DA is a powerful yet simple tool for analysing microarray data.

Data Interpretation, Statistical↗

A comparison of two methods of discriminant analysis applied to binary data.

The results of applying classical linear discriminant analysis and kernel discriminant analysis to several real sets of multivariate binary data are presented. Classical discriminant analysis is intrinsically parametric and is usually presented as being well-suited to continuous variables; it is also well-known to be optimal when the (two) classes have normal distributions with identical covariance matrices. The kernel method, on the other hand, is nonparametric and, in the form used here, is ideally suited to binary data. The apparent error rates of the kernel method are found to be consistently less than those of the classical method. However, when the true error rates are estimated either by applying the classifiers to independent test sets, or by the leaving-one-out method from the design sets, no significant difference is discernible between the two types of classifier.

Enuresis↗

[Statistical evaluation of biochemical data by the method of discrimination analysis. Selection of the discriminant biochemical variables. Attempted biochemical discrimination of intrahepatic cholestasis, extrahepatic obstruction and liver cancer].

An original method of statistical treatment of biological data is proposed. It permits satisfactory biochemical classification of 322 patients divided up into 3 groups : intrahepatic cholestasis (235 patients), extrahepatic obstruction (44 patients) and carcinoma of the liver (43 patients). On the basis of 32 tests, it was possible to define discriminating areas permitting satisfactory diagnosis in 95 per cent of published cases. The reduction in the number of tests necessary for diagnosis was considered. The selection technic used was original to the extent that it dose not require, like most methods used today, the determination of better individual discriminators, but the establishment of a better discriminating subunit, obtained from the initial subunit composed of a group of variables. From the 32 parameters contained in the standard liver function tests, a search for a better discriminating subunit consisting of the best four tests, permitted the authors to select a group of 10 tests : bilirubin, alkaline phosphatase, 5-nucleotidase, Thymolturbidity, Cetavlon test, serum albumin, total LDH, TGP (ALAT), OCT, GLDH, of which the discriminating value remains very satisfactory.

Analysis of Variance↗

Discriminant analysis to evaluate clustering of gene expression data.

In this work we present a procedure that combines classical statistical methods to assess the confidence of gene clusters identified by hierarchical clustering of expression data. This approach was applied to a publicly released Drosophila metamorphosis data set [White et al., Science 286 (1999) 2179-2184]. We have been able to produce reliable classifications of gene groups and genes within the groups by applying unsupervised (cluster analysis), dimension reduction (principal component analysis) and supervised methods (linear discriminant analysis) in a sequential form. This procedure provides a means to select relevant information from microarray data, reducing the number of genes and clusters that require further biological analysis.

Animals↗