Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Principal component”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Principal components and further possibilities with the PANSS.

At the end of the last century, Hughlings-Jackson suggested that positive and negative syndromes should be kept apart in psychotic disorders. When the concepts of dementia praecox and schizophrenia were introduced by Kraepelin and Bleuler, emphasis was laid on the negative symptoms, regarded as fundamental. After the introduction of the "first rank symptoms" by Schneider emphasis switched to the positive symptoms in schizophrenia and these symptoms were included in most diagnostic criteria. In the 1980s Andreasen and Crow suggested a dichotomy into positive and negative syndromes in schizophrenia. Kay and co-workers introduced a Positive And Negative Syndrome Scale (PANSS) for schizophrenia. In the original studies satisfactory construct validity and inter-rater reliability were demonstrated. However, in studies outside the USA a high construct validity was found for the negative scales but not for the positive and general psychopathology scales. Furthermore, the inter-rater reliability of the negative scale was a problem. After introduction of the Structured Clinical Interview for the PANSS (SCI-PANSS) the inter-rater reliability increased for all three scales. In an early study Kay and Sevy found seven factors in a principal component analysis of the PANSS and suggested a four factor pyramidical model. Later principal component analyses by Lepine, Peralta et al. and Kawasaki et al. suggested that the four factor model was an oversimplification and Lindström and von Knorring suggested a five factor pyramidical model. A similar model was later suggested by Bell et al. after a reanalysis of the original series of Kay and Sevy.(ABSTRACT TRUNCATED AT 250 WORDS)

Affective Symptoms↗

Principal components analysis to summarize microarray experiments: application to sporulation time series.

A series of microarray experiments produces observations of differential expression for thousands of genes across multiple conditions. It is often not clear whether a set of experiments are measuring fundamentally different gene expression states or are measuring similar states created through different mechanisms. It is useful, therefore, to define a core set of independent features for the expression states that allow them to be compared directly. Principal components analysis (PCA) is a statistical technique for determining the key variables in a multidimensional data set that explain the differences in the observations, and can be used to simplify the analysis and visualization of multidimensional data sets. We show that application of PCA to expression data (where the experimental conditions are the variables, and the gene expression measurements are the observations) allows us to summarize the ways in which gene responses vary under different conditions. Examination of the components also provides insight into the underlying factors that are measured in the experiments. We applied PCA to the publicly released yeast sporulation data set (Chu et al. 1998). In that work, 7 different measurements of gene expression were made over time. PCA on the time-points suggests that much of the observed variability in the experiment can be summarized in just 2 components--i.e. 2 variables capture most of the information. These components appear to represent (1) overall induction level and (2) change in induction level over time. We also examined the clusters proposed in the original paper, and show how they are manifested in principal component space. Our results are available on the internet at http:¿www.smi.stanford.edu/project/helix/PCArray .

Computer Simulation↗

Mapping disease-susceptibility genes in admixed populations using interval principal component tests.

Family-based association approach for mapping disease-susceptibility genes of complex human diseases is a topical issue in genetic epidemiology. It is well known that admixture between genetically differentiated populations can result in high levels of linkage disequilibrium at loci separated far apart. This property has been capitalized upon to reduce the burden of genotyping in a genomewide association scan. The authors describe a new approach for admixture mapping--the "interval principal component test" (IPCT). The genome is divided into a multitude of non-overlapping "intervals" (with interval length of 10-20 cM) and the information of the markers in the same interval is integrated using the principal component analysis. Monte-Carlo simulation shows that an interval-by-interval scan using IPCT has much better performances than a conventional marker-by-marker scan using the transmission/disequilibrium test (TDT).

Computer Simulation↗

Principal component analysis to detect the similarity of distantly related proteins; its application to cytochromes c, c1 and f.

A new method has been developed for detecting the similarity between distantly related families of proteins. The amino acid sequences of each family of proteins are vertically aligned by a homologous alignment method and the physico-chemical properties of the amino acid residues at the corresponding site are evaluated simultaneously, by method of the principal component analysis. Taking into account the species diversity of each family of proteins, we assign the similar regions between the different families of proteins by the overlapping degree of the standard deviations around the mean values of the first principal component. To investigate the homologous relationship between the electron transport proteins in photosynthetic and O2 respiratory systems, this method has been applied to 70 species of mitochondrial cytochrome c, 4 species of cytochrome c1 and 7 species of cytochrome f. This analysis reveals that both cytochrome f and cytochrome c1 have large regions which are similar to those of cytochrome c. Assuming that these similar regions have the same stereochemical structures as those in cytochrome c, we can predict the outlines of the tertiary structures of cytochrome c1 and cytochrome f, respectively, each able to interact with its electron acceptor, cytochrome c and plastocyanin.

Amino Acid Sequence↗

Principal component analysis of chest wall movement in selected pathologies.

A method is presented for assessing a compact set of parameters characteristic of respiratory system functional status. 3D movements of points in the chest wall and the volumes of chest wall compartments (pulmonary rib cage, abdominal rib cage and abdomen) are considered. The co-ordinates of these points are measured using an opto-electronic system for 3D motion analysis. Principal component analysis is applied to these data. The behaviour of the eigenvectors of the covariance matrix of the 3D co-ordinates of the points on the chest wall shows close agreement with the pathology characteristics. The same is found for the percentage of total variance explained by the principal components of the volume variations. In this case, the higher values of variance percentage explained indicate independent motions (active or passive) in the degrees of freedom of the system identified by partitioning the total volume into compartments.

Adult↗

Two-dimensional correlation spectroscopy and principal component analysis studies of temperature-dependent IR spectra of cotton-cellulose.

The FTIR spectra were measured for raw Uplands Sicala-V2 cotton fibers over a temperature range of 40-325 degrees C to explore the temperature-dependent changes in the hydrogen bonds of cellulose. These cotton-cellulose spectra exhibited complicated patterns in the 3800-2800 cm(-1) region and thus were analyzed by both the exploratory principal component analysis (PCA) and two-dimensional (2-D) correlation spectroscopy methods. The exploratory PCA showed that the spectra separate into two groups on the basis of thermal degradation of the cotton-cellulose and the consequent breakage of intersheet H-bonds present in its structure. Frequency variables, which strongly contributed to each principal component highlighted in its loadings plot, were linked to the frequencies assigned to vibrations of the OH groups involved in different kinds of H-bonds, as well as to vibrations of the CH groups. Deeper insights into reorganization of the temperature-dependent hydrogen bonding were obtained by 2-D correlation spectroscopy. Synchronous and asynchronous spectra were analyzed in the temperature ranges of 40 to 150 and 250 to 320 degrees C, the ranges indicated by PCA. Detailed band assignments of the OH stretching region and changes in the patterns of the hydrogen bonding network of the cotton-cellulose were proposed with the aid of the 2-D correlation spectroscopy analysis. Below 150 degrees C, distinctly different bands assigned to the less stable Ialpha and the more stable Ibeta interchain H-bonds O-6-H-6...O-3' were observed at about 3230 and 3270 cm(-1), respectively. Evaporation of water entrapped in the cellulose network was examined by means of the band at about 3610 cm(-1). The cooperativity of hydrogen bonds, which play a key role in the cellulose conformation, was monitored by frequencies assigned to intrachain H-bonds. It was possible to separate the frequencies assigned to the O-2-H-2...O-6 and O-3-H-3...O-5 intrachain H-bonds into two separate ranges, the spread of which was controlled by the cooperativity effect. The temperature dependence of the asynchronous spectra indicated that the less stable O-3-H-3...O-5 bonds gave rise to an absorption extending from 3300 to 3384 cm(-1), while the more stable O-2-H-2...O-6 bonds were characterized by the absorption between 3400 and 3470 cm(-1). The final breaking of the inter- and intrachain H-bonds, which occurs at the higher temperatures, was monitored by the asynchronous peaks at 3533 and 3590 cm(-1), respectively. On the basis of both the exploratory PCA and 2-D correlation spectroscopy investigations, it was possible to extract well-defined wavenumber ranges assigned to different kinds of intra- and interchain hydrogen bonds, as well as to the free OH groups of the cotton-cellulose.

Algorithms↗

A principal components analysis of human odontometrics.

It has long been recognized that tooth crown diameters in hominoids are all positively intercorrelated one with another. This study reports on sex-specific correlation matrices derived from 2,650 individuals from the Solomon Islands, Melanesia. Mesiodistal and buccolingual diameters of all permanent teeth from one side are used, excluding third molars. Analysis discloses significant sex dimorphism in the strengths of the intercorrelations, with females being better integrated. Principal components analysis (PCA) provides an objective means of data reduction (shown here to be preferable to simple size summation methods) and decorrelation of the resulting linear combinations. Four components are extracted (with results being virtually identical in the two sexes) and arguments are put forth that varimax rotation to "a simpler solution" may be counterproductive. Before rotation, the four components are 1) overall size, 2) buccolingual widths contrasted with mesiodistal lengths, 3) anterior (I,C) contrasted with posterior (P,M) teeth, and 4) premolars contrasted with molars. Most of the explained (shared) variance (63%) extracted by PCA is in overall size of the dentition. There is a strong urge to view the results of these principal components analyses as reflective of biologically and genetically meaningful entities.

Female↗

Principal component analysis of pain-related cerebral potentials to mechanical and electrical stimulation in man.

Single trial event-related cerebral potentials (ERPs) in response to skin stimuli of various intensities and qualities in man were investigated in respect to their nociceptive information content. Electrical constant current stimuli (20 msec, 2 - 8 mA) and mechanical force controlled stimuli (20 msec, 0.8 - 3.2 N) were applied to the tip of the left middle finger. Four intensities of each stimulus quality were given, each intensity appearing 40 times in standardized randomized order. EEG segments (between 5 sec before and 500 msec after stimulus onset) were subjected to computer analysis. ERP wave form was shown to depend upon the amount of alpha waves in the prestimulus EEG. For analysis, only subjects with low power in the alpha band were selected. Principal component analysis was applied to all single trial ERPs measured using the variance-covariance matrix of association. Six principal components (PCs) were extracted accounting for about 90% of total variance. Five of the extracted PCs had well located loading maxima: PC1 (50 - 80 msec), PC4 (140 - 160 msec), PC3 (200 - 250 msec), PC4 (280 - 360 msec), PC5 (400 - 500 msec); PC6 appeared polyphasic. Analysis of variance of the mean PC scores revealed that one PC (PC1) discriminated between quality, and 4 PCs (PC1 - PC4) between quantity of stimulation. Eliminating effects of stimulus intensity resulted in two PCs (PC2, PC4) which distinguished exclusively between non-pain and pain. PCA applied to disjunctive subsets of ERPs, corresponding to the different experimental conditions, yielded practically identical sets of PCs, such that no specific ERP component emerged when pain was reported.

Adult↗

Methodological issues in determining the dimensionality of composite health measures using principal component analysis: case illustration and suggestions for practice.

During the early steps of the construction of composite health measures, principal component analysis (PCA) is commonly used to identify 'latent' factors that underlie observed variables and to determine the dimensionality of the instruments. The determination of the number of components to retain is critical to PCA: it markedly influences the factorial model identified and further conditions the validity of the constructed instrument. However, many researchers developing composite health measures seem to be unaware of the importance of this determination. The purposes of the paper are to illustrate (1) the variability of the factorial models obtained by using different published rules (n = 10) for determining the number of components to retain in PCA applied to two quality-of-life datasets, and (2) the value of a careful and diversified approach to the problem of the number of components to retain in PCA that we suggest, instead of the unsatisfactory 'rule-of-thumb' that many researchers use. This involves: (1) using robust rules (including parallel analysis and minimum average partial procedure) to generate a set of possible values for the number of components to retain, (2) repeating the analysis across samples, (3) comprehensively assessing the models obtained, and (4) considering complementary methods to PCA and especially confirmatory factor analysis.

Factor Analysis, Statistical↗

Distinguishing normal and abnormal tracheal breathing sounds by principal component analysis.

Expired and inspired tracheal breathing sounds (BS) were recorded from 10 normal subjects and 8 patients with respiratory diseases, including bronchial asthma, sarcoidosis, fibrosing lung disease, chronic bronchitis, and radiation pneumonitis. Frequency spectra were generated using Fast Fourier Transform (FFT), and we observed considerable differences between BS spectra of normal subjects and patients. The frequency of peak amplitude and mean frequency of the BS spectra of patients were significantly higher than those of normal subjects. Spectral features were extracted by dividing each spectra into equal frequency bands--each feature being the mean amplitude of each FFT element within a frequency band. We used Principal Component Analysis to compare spectral feature sets and found a clear separation between normal and abnormal tracheal BS for 10, 20, and 40 features/spectra. We conclude that Principal Component Analysis of BS could become a new method of diagnosing respiratory disease in an automated fashion.

Adult↗

The classification of solvents by combining classical QSPR methodology with principal component analysis.

The results of a quantitative structure-property relationship (QSPR) analysis of 127 different solvent scales and 774 solvents using the CODESSA PRO program are presented. QSPR models for each scale were constructed using only theoretical descriptors. The high quality of the models is reflected by the squared multiple correlation coefficients that range from 0.726 to 0.999; only 18 models have R2< 0.800. This enables direct theoretical calculation of predicted values for any scale and/or for any organic solvent, including those previously unmeasured. The molecular descriptors involved in the models are classified and discussed according to (i) the origin of their calculation (i.e., constitutional, geometric, charge-related, etc.) and (ii) the commonly accepted classification of physical interactions between the solute and solvent molecules in liquid (condensed) media. A reduced matrix 774 (solvents) x 100 (solvent scales) was selected for the principal component analysis (PCA) by taking into account only the solvent scales with more than 20 experimental data points. The first 5 principal components account for 75% of the total variance. The robustness of the PCA model obtained was validated by the comparison models development for restricted submatrices of data and with the results obtained for the full data set. The total variance accounted for by the first three PCs, for the submatrices with the same number of solvent scales but different numbers of solvents, varies from 68.2% to 59.0%. This demonstrates that the total variance described by the first 3 components is essentially stable as the number of solvents involved varies from 100 to 774. Subsequently, a matrix with 703 diverse solvents and 100 solvent scales was selected for the general classification of the solvents and scales according to the scores and loadings obtained from the PCA treatment. Classification of the theoretical molecular descriptors, derived from the chemical structure alone, according to their relevance to specific types of intermolecular interaction (cavity formation, electrostatic polarization, dispersion, and hydrogen bonding) in liquid media enables a more easily comprehensible physical interpretation of the QSPR of molecular properties in liquids and solutions. The reported QSPR models for solvent scales with theoretical molecular descriptors and the results of the PCA analysis are potentially of great practical importance, as they extend the applicability of correlations with empirical solvent scales to many previously unmeasured systems.

Journal Article↗

Principal components analysis competitive learning.

We present a new neural model that extends the classical competitive learning by performing a principal components analysis (PCA) at each neuron. This model represents an improvement with respect to known local PCA methods, because it is not needed to present the entire data set to the network on each computing step. This allows a fast execution while retaining the dimensionality-reduction properties of the PCA. Furthermore, every neuron is able to modify its behavior to adapt to the local dimensionality of the input distribution. Hence, our model has a dimensionality estimation capability. The experimental results we present show the dimensionality-reduction capabilities of the model with multisensor images.

Algorithms↗

Misallocation of variance in event-related potentials: simulation studies on the effects of test power, topography, and baseline-to-peak versus principal component quantifications.

Since Wood and McCarthy's simulation study (Electroenceph Clin Neurophysiol 1984;59:249-260), the use of principal component analysis (PCA) as a tool for the identification and quantification of event-related potentials (ERP) has been considered a challenge. Three relevant aspects have not been fully acknowledged in previous studies, however, and were therefore investigated in the present simulation study. Firstly, the impact of test power on the amount of variance misallocation was studied. Secondly, the impact of ERP component topography on variance misallocation was investigated. Thirdly, a systematic evaluation of variance misallocation in baseline-to-peak derived ERP measures was performed. Results based on an overall set of 2700 simulations indicate that: (a) variance misallocation is reduced to an almost acceptable level when an appropriate test power is simulated; (b) the overall amount of variance misallocation remains at an almost acceptable level when systematic topographic effects are simulated in combination with an appropriate test power; and (c) variance misallocation is in fact also a problem in baseline-to-peak measures. These findings confirm that, when used appropriately, PCA is a helpful and efficient tool for the identification and quantification of ERPs.

Algorithms↗

Exploratory studies of PM10 receptor and source profiling by GC/MS and principal component analysis of temporally and spatially resolved ambient samples.

For a recent exploratory study of particulate matter (PM) compositions, origins, and impacts in the El Paso/Juarez (Paso del Norte) airshed, the authors relied on solvent extraction (SX)-gas chromatography/mass spectrometry (GC/MS) procedures to characterize 24-hr quartz fiber (QF) filter samples obtained from nine spatially distributed high-volume (Hi-Vol) PM10 samplers as well as on thermal desorption (TD)-GC/MS methods to characterize 45 time-resolved (2-hr) filter samples obtained with modified 1-m3/hr PM10 samplers. Principal component analysis and related chemometric techniques were used for data reduction and data fusion as well as for multiway data correlation. A high degree of correspondence (R2 = 0.821) was found between the rapid TD-GC/MS method (which can be carried out on 2-hr filter slices containing only microgram amounts of sample) and conventional SX-GC/MS procedures. The four main source patterns of organic PM components observed in GC/MS profiles of both temporally and spatially resolved receptor samples obtained in the El Paso/Juarez border airshed during the study period are interpreted to represent (1) vehicular emissions plus resuspended urban dust; (2) biomass combustion; (3) native vegetation detritus and resuspended agricultural dust; and (4) waste burning. Moreover, principal component analysis of combined, variance-weighted, temporally resolved TD-GC/MS data and spatially resolved SX-GC/MS data was used to determine approximate source locations for specific PM components identified in time-resolved receptor sample profiles. The same approach can be used to determine approximate circadian concentration profiles of specific PM components identified in spatially resolved receptor sample profiles.

Agriculture↗

A weighted-principal component regression method for the identification of physiologic systems.

We introduce a system identification method based on weighted-principal component regression (WPCR). This approach aims to identify the dynamics in a linear time-invariant (LTI) model which may represent a resting physiologic system. It tackles the time-domain system identification problem by considering, asymptotically, frequency information inherent in the given data. By including in the model only dominant frequency components of the input signal(s), this method enables construction of candidate models that are specific to the data and facilitates a reduction in parameter estimation error when the signals are colored (as are most physiologic signals). Additionally, this method allows incorporation of preknowledge about the system through a weighting scheme. We present the method in the context of single-input and multi-input single-output systems operating in open-loop and closed-loop. In each scenario, we compare the WPCR method with conventional approaches and approaches that also build data-specific candidate models. Through both simulated and experimental data, we show that the WPCR method enables more accurate identification of the system impulse response function than the other methods when the input signal(s) is colored.

Algorithms↗

Description and critical appraisal of principal components analysis (PCA) methodology applied to pulsed-field gel electrophoresis profiles of methicillin-resistant Staphylococcus aureus isolates.

Principal components analysis (PCA) has been described for over 50 years; however, it is rarely applied to the analysis of epidemiological data. In this study PCA was critically appraised in its ability to reveal relationships between pulsed-field gel electrophoresis (PFGE) profiles of methicillin-resistant Staphylococcus aureus (MRSA) in comparison to the more commonly employed cluster analysis and representation by dendrograms. The PFGE type following SmaI chromosomal digest was determined for 44 multidrug-resistant hospital-acquired methicillin-resistant S. aureus (MR-HA-MRSA) isolates, two multidrug-resistant community-acquired MRSA (MR-CA-MRSA), 50 hospital-acquired MRSA (HA-MRSA) isolates (from the University Hospital Birmingham, NHS Trust, UK) and 34 community-acquired MRSA (CA-MRSA) isolates (from general practitioners in Birmingham, UK). Strain relatedness was determined using Dice band-matching with UPGMA clustering and PCA. The results indicated that PCA revealed relationships between MRSA strains, which were more strongly correlated with known epidemiology, most likely because, unlike cluster analysis, PCA does not have the constraint of generating a hierarchic classification. In addition, PCA provides the opportunity for further analysis to identify key polymorphic bands within complex genotypic profiles, which is not always possible with dendrograms. Here we provide a detailed description of a PCA method for the analysis of PFGE profiles to complement further the epidemiological study of infectious disease.

Cluster Analysis↗

A guide for applying principal-components analysis and confirmatory factor analysis to quantitative electroencephalogram data.

Principal-components analysis (PCA) has been used in quantitative electroencephalogram (qEEG) research to statistically reduce the dimensionality of the original qEEG measures to a smaller set of theoretically meaningful component variables. However, PCAs involving qEEG have frequently been performed with small sample sizes, producing solutions that are highly unstable. Moreover, solutions have not been independently confirmed using an independent sample and the more rigorous confirmatory factor analysis (CFA) procedure. This paper was intended to illustrate, by way of example, the process of applying PCA and CFA to qEEG data. Explicit decision rules pertaining to the application of PCA and CFA to qEEG are discussed. In the first of two experiments, PCAs were performed on qEEG measures collected from 102 healthy individuals as they performed an auditory continuous performance task. Component solutions were then validated in an independent sample of 106 healthy individuals using the CFA procedure. The results of this experiment confirmed the validity of an oblique, seven component solution. Measures of internal consistency and test-retest reliability for the seven component solution were high. These results support the use of qEEG data as a stable and valid measure of neurophysiological functioning. As measures of these neurophysiological processes are easily derived, they may prove useful in discriminating between and among clinical (neurological) and control populations. Future research directions are highlighted.

Adolescent↗

Principal component analysis of dissolution data with missing elements.

The use of principal component analysis (PCA) for incomplete dissolution data sets is examined. The PC space is constructed using a reference set and the test set is projected in that space. Several cases such as a reference set with missing data, an incomplete test set and both sets measured at different time points, are discussed using two examples: one simulation and one obtained from the pharmaceutical practice. From the many possibilities to deal with missing data, the expectation-maximization algorithm in combination with PCA was chosen. The influence on the similarity or f2 factor is examined too. The sampling with replacement or bootstrap technique, which can be used to obtain confidence limits, can also be used when missing data are present in one of the data sets.

Algorithms↗