Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Principal component”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Simultaneous determination of aniline and cyclohexylamine by principal component artificial neural networks.

A specterophotometric method for simultaneous determination of aniline and cyclohexylamine using principal component artificial neural networks is proposed. This method is based on the reactions involving aniline and/or cyclohexylamine, with bis(acetylacetoneethylendiamine)tributylphosphine cobalt(III) perchlorate as a complexing reagent. A nonionic surfactant, Triton X-100, was used for dissolving the complexes and intensifying the signals. The absorption data were based on the spectra registered in the range of 350 - 550 nm. An artificial neural network consisting of three layers of nodes was trained by applying a back-propagation learning rule. Sigmoid transfer functions were used in the hidden and output layers to facilitate nonlinear calibration. The predictive ability of artificial neural networks was examined for the determination of aniline and cyclohexylamine in synthetic mixtures.

Journal Article↗

A new approach for clinical biological assay comparison and standardization: application of principal component analysis to a multicenter study of twenty-one carcinoembryonic antigen immunoassay kits.

BACKGROUND: Principal component analysis (PCA) is a powerful mathematical method able to analyze data sets containing a large number of variables. To our knowledge, this method is applied here for the first time in the field of medical laboratory analysis. METHODS: PCA was used to evaluate the results of a blind comparative study of 21 carcinoembryonic antigen (CEA) reagent kits used to determine CEA concentration in a panel of sera from 80 patients. RESULTS: The mathematical technique first eliminated the variations attributable to the use of different calibrators. The PCA representation then gave a global view of the dispersion of the kits and allowed the identification of a main homogeneous group and of some discrepant kits. CONCLUSIONS: PCA applied to the in vitro diagnostic reagent field could contribute to the standardization process and improve the quality of medical laboratory analyses. A standardization method using a panel of patient sera is proposed.

Biomarkers, Tumor↗

Analysis of diffusion tensor magnetic resonance imaging data using principal component analysis.

An analysis method for diffusion tensor (DT) magnetic resonance imaging data is described, which, contrary to the standard method (multivariate fitting), does not require a specific functional model for diffusion-weighted (DW) signals. The method uses principal component analysis (PCA) under the assumption of a single fibre per pixel. PCA and the standard method were compared using simulations and human brain data. The two methods were equivalent in determining fibre orientation. PCA-derived fractional anisotropy and DT relative anisotropy had similar signal-to-noise ratio (SNR) and dependence on fibre shape. PCA-derived mean diffusivity had similar SNR to the respective DT scalar, and it depended on fibre anisotropy. Appropriate scaling of the PCA measures resulted in very good agreement between PCA and DT maps. In conclusion, the assumption of a specific functional model for DW signals is not necessary for characterization of anisotropic diffusion in a single fibre.

Adult↗

Recursive principal components analysis.

A recurrent linear network can be trained with Oja's constrained Hebbian learning rule. As a result, the network learns to represent the temporal context associated to its input sequence. The operation performed by the network is a generalization of Principal Components Analysis (PCA) to time-series, called Recursive PCA. The representations learned by the network are adapted to the temporal statistics of the input. Moreover, sequences stored in the network may be retrieved explicitly, in the reverse order of presentation, thus providing a straight-forward neural implementation of a logical stack.

Humans↗

Principal component analysis of the T wave and prediction of cardiovascular mortality in American Indians: the Strong Heart Study.

BACKGROUND: Increased QT interval dispersion (QTd) is a proposed ECG marker of vulnerability to ventricular arrhythmias and of cardiovascular (CV) mortality. However, principal component analysis (PCA) of the T-wave vector loop may more accurately represent repolarization abnormalities than QTd. METHODS AND RESULTS: Predictive values of QTd and PCA were assessed in 1839 American Indian participants in the first Strong Heart Study examination. T-wave loop morphology was quantified by the ratio of the second to first eigenvalues of the T-wave vector by PCA (PCA ratio); QTd was quantified as the difference between maximum and minimum QT intervals. After 3.7+/-0.9 years mean follow-up, there were 55 CV deaths. In univariate analyses, an increased PCA ratio predicted CV mortality in women (chi2=7.8, P=0.0053) and men (chi2=9.5, P=0.0021). In contrast, increased QTd was a significant predictor of CV mortality in women (chi2=30.6, P<0.0001) but not in men (chi2=2.0, P=NS). In multivariate Cox analyses controlling for risk factors and rate-corrected QT interval, the PCA ratio remained a significant predictor of CV mortality in women (chi2=4.0 P=0.043) and men (chi2=6.4, P=0.011); QTd was a significant predictor in women only (chi2=11.0, P=0.0009). PCA ratios >90th percentile (32% in women and 24.6% in men) identified women with a 3.68-fold increased risk of CV mortality (95% CI, 1.54 to 8.83) and men with a 2.77-fold increased risk (95% CI, 1.18 to 6.49). CONCLUSIONS: Abnormalities of repolarization measured by PCA of the T-wave loop predict CV death in men and women, supporting use of PCA for quantifying repolarization abnormalities.

Arizona↗

Characterization of contaminated soil and groundwater surrounding an illegal landfill (S. Giuliano, Venice, Italy) by principal component analysis and kriging.

The characterization of a hydrologically complex contaminated site bordering the lagoon of Venice (Italy) was undertaken by investigating soils and groundwaters affected by the chemical contaminants originated by the wastes dumped into an illegal landfill. Statistical tools such as principal components analysis and geostatistical techniques were applied to obtain the spatial distribution of chemical contaminants. Dissolved organic carbon (DOC), SO4(2-) and Cl- were used to trace the migration of the contaminants from the top soil to the underlying groundwaters. The chemical and hydrogeological available information was assembled to obtain the schematic of the conceptual model of the contaminated site capable to support the formulation of major exposure scenarios, which are also provided.

Environmental Monitoring↗

Principal component analysis of the conformational freedom within the EF-hand superfamily.

A database of nonredundant structures of EF-hand domains--i.e., pairs of helix-loop-helix motifs--has been assembled, and the six angles among the four helices re-determined. A principal component analysis of these angles allows us to use two such components (PC1 and PC2) to describe the system retaining 80% of the total variance. A PC2 against PC1 plot representation allows us to represent in a compact way the full range of structural diversity of EF-hand domains, their grouping into protein families, and the variation for each family upon calcium and peptide binding.

Amino Acid Motifs↗

NMR spectral quantitation by principal-component analysis. II. Determination of frequency and phase shifts.

This paper extends the use of principal-component analysis in spectral quantification to the estimation of frequency and phase shifts in a single resonant peak across a series of spectra. The estimated parameters can be used to correct the spectra accordingly, resulting in more accurate peak-area estimation. Further, the removal of the variations in phase and frequency cause by instrumental and experimental fluctuations makes it possible to determine more accurately the remaining variations, which bear biological significance. The procedure is demonstrated on simulated data, a 3D chemical-shift-imaging dataset acquired from a cylinder of inorganic phosphate (Pi), and a set of 736 31P NMR in vivo spectra taken from a kinetic study of rate muscle energetics. In all cases, the procedure rapidly and automatically identifies the frequency and phase shifts present in the individual spectra. In the kinetic study, the procedure is used twice, first to adjust the phase and frequency of a reference peak (phosphocreatine) and then to determine the individual frequencies of the Pi peak in each of the spectra which further can be used for estimation of pH changes during the experiment.

Computer Simulation↗

Standard errors of the principal component loadings for unstandardized and standardized variables.

The asymptotic standard errors of the estimates of the principal component loadings for standardized variables are derived under the assumption of multivariate normality. The standard errors are obtained for the usual unrotated case where a loading matrix is orthogonal and for the cases with orthogonally or obliquely rotated components. The corresponding standard errors for unstandardized variables and the asymptotic correlations among the estimators of the parameters for the unstandardized and standardized variables are simultaneously derived, together with the standard errors for the standardized variables. Results of a simulation illustrate the accuracy of the theoretical asymptotic standard errors and correlations.

Humans↗

GA-fisher: A new LDA-based face recognition algorithm with selection of principal components.

This paper addresses the dimension reduction problem in Fisherface for face recognition. When the number of training samples is less than the image dimension (total number of pixels), the within-class scatter matrix (Sw) in Linear Discriminant Analysis (LDA) is singular, and Principal Component Analysis (PCA) is suggested to employ in Fisherface for dimension reduction of Sw so that it becomes nonsingular. The popular method is to select the largest nonzero eigenvalues and the corresponding eigenvectors for LDA. To attenuate the illumination effect, some researchers suggested removing the three eigenvectors with the largest eigenvalues and the performance is improved. However, as far as we know, there is no systematic way to determine which eigenvalues should be used. Along this line, this paper proposes a theorem to interpret why PCA can be used in LDA and an automatic and systematic method to select the eigenvectors to be used in LDA using a Genetic Algorithm (GA). A GA-PCA is then developed. It is found that some small eigenvectors should also be used as part of the basis for dimension reduction. Using the GA-PCA to reduce the dimension, a GA-Fisher method is designed and developed. Comparing with the traditional Fisherface method, the proposed GA-Fisher offers two additional advantages. First, optimal bases for dimensionality reduction are derived from GA-PCA. Second, the computational efficiency of LDA is improved by adding a whitening procedure after dimension reduction. The Face Recognition Technology (FERET) and Carnegie Mellon University Pose, Illumination, and Expression (CMU PIE) databases are used for evaluation. Experimental results show that almost 5 % improvement compared with Fisherface can be obtained, and the results are encouraging.

Algorithms↗

An automated algorithm for the computation of brain volume change from sequential MRIs using an iterative principal component analysis and its evaluation for the assessment of whole-brain atrophy rates in patients with probable Alzheimer's disease.

This article introduces an automated method for the computation of changes in brain volume from sequential magnetic resonance images (MRIs) using an iterative principal component analysis (IPCA) and demonstrates its ability to characterize whole-brain atrophy rates in patients with Alzheimer's disease (AD). The IPCA considers the voxel intensity pairs from coregistered MRIs and identifies those pairs a sufficiently large distance away from the iteratively determined PCA major axis. Analyses of simulated and real MRI data support the underlying assumption of a linear relationship in paired voxel intensities, identify an outlier distance threshold that optimizes the trade-off between sensitivity and specificity in the detection of small volume changes while accounting for global intensity changes, and demonstrate an ability to detect changes as small as 0.04% of brain volume without confounding effects of between-scan shifts in voxel intensity. In eight patients with probable AD and eight age-matched normal control subjects, the IPCA was comparable to the established but partly manual digital subtraction (DS) method in characterizing annual rates of whole-brain atrophy: resulting rates were correlated (Spearman rank correlation = 0.94, P < 0.0005) and comparable in distinguishing probable AD from normal aging (IPCA-detected atrophy rates: 2.17 +/- 0.52% per year in the patients vs. 0.41 +/- 0.22% per year in the controls [Wilcoxon-Mann-Whitney test P = 7.8 x 10(-4)]; DS-detected atrophy rates: 3.51 +/- 1.31% per year in the patients vs. 0.48 +/- 0.29% per year in the controls [P = 7.8 x 10(-4)]). The IPCA could be used in tracking the progression of AD, evaluating the disease-modifying effects of putative treatments, and investigating the course of other normal and pathological changes in brain morphology.

Aged↗

Screening molecular associations with lipid membranes using natural abundance 13C cross-polarization magic-angle spinning NMR and principal component analysis.

We describe an NMR approach for detecting the interactions between phospholipid membranes and proteins, peptides, or small molecules. First, 1H-13C dipolar coupling profiles are obtained from hydrated lipid samples at natural isotope abundance using cross-polarization magic-angle spinning NMR methods. Principal component analysis of dipolar coupling profiles for synthetic lipid membranes in the presence of a range of biologically active additives reveals clusters that relate to different modes of interaction of the additives with the lipid bilayer. Finally, by representing profiles from multiple samples in the form of contour plots, it is possible to reveal statistically significant changes in dipolar couplings, which reflect perturbations in the lipid molecules at the membrane surface or within the hydrophobic interior.

Calcium-Binding Proteins↗

Effective dimensionality for principal component analysis of time series expression data.

Large-scale expression data are today measured for thousands of genes simultaneously. This development has been followed by an exploration of theoretical tools to get as much information out of these data as possible. Several groups have used principal component analysis (PCA) for this task. However, since this approach is data-driven, care must be taken in order not to analyze the noise instead of the data. As a strong warning towards uncritical use of the output from a PCA, we employ a newly developed procedure to judge the effective dimensionality of a specific data set. Although this data set is obtained during the development of rat central nervous system, our finding is a general property of noisy time series data. Based on knowledge of the noise-level for the data, we find that the effective number of dimensions that are meaningful to use in a PCA is much lower than what could be expected from the number of measurements. We attribute this fact both to effects of noise and the lack of independence of the expression levels. Finally, we explore the possibility to increase the dimensionality by performing more measurements within one time series, and conclude that this is not a fruitful approach.

Algorithms↗

Classification of high-speed gas chromatography-mass spectrometry data by principal component analysis coupled with piecewise alignment and feature selection.

A useful methodology is introduced for the analysis of data obtained via gas chromatography with mass spectrometry (GC-MS) utilizing a complete mass spectrum at each retention time interval in which a mass spectrum was collected. Principal component analysis (PCA) with preprocessing by both piecewise retention time alignment and analysis of variance (ANOVA) feature selection is applied to all mass channels collected. The methodology involves concatenating all concurrently measured individual m/z chromatograms from m/z 20 to 120 for each GC-MS separation into a row vector. All of the sample row vectors are incorporated into a matrix where each row is a sample vector. This matrix is piecewise aligned and reduced by ANOVA feature selection. Application of the preprocessing steps (retention time alignment and feature selection) to all mass channels collected during the chromatographic separation allows considerably more selective chemical information to be incorporated in the PCA classification, and is the primary novelty of the report. This methodology is objective and requires no knowledge of the specific analytes of interest, as in selective ion monitoring (SIM), and does not restrict the mass spectral data used, as in both SIM and total ion current (TIC) methods. Significantly, the methodology allows for the classification of data with low resolution in the chromatographic dimension because of the added selectivity from the complete mass spectral dimension. This allows for the successful classification of data over significantly decreased chromatographic separation times, since high-speed separations can be employed. The methodology is demonstrated through the analysis of a set of four differing gasoline samples that serve as model complex samples. For comparison, the gasoline samples are analyzed by GC-MS over both 10-min and 10-s separation times. The successfully classified 10-min GC-MS TIC data served as the benchmark analysis to compare to the 10-s data. When only alignment and feature selection was applied to the 10-s gasoline separations using GC-MS TIC data, PCA failed. PCA was successful for 10-s gasoline separations when the methodology was applied with all the m/z information. With ANOVA feature selection, chromatographic regions with Fisher ratios greater than 1500 were retained in a new matrix and subjected to PCA yielding successful classification for the 10-s separations.

Analysis of Variance↗

A study on water adsorption onto microcrystalline cellulose by near-infrared spectroscopy with two-dimensional correlation spectroscopy and principal component analysis.

Water adsorption onto microcrystalline cellulose (MCC) in the moisture content (M(c)) range of 0.2-13.4 wt % was investigated by near-infrared (NIR) spectroscopy. In order to distinguish heavily overlapping O-H stretching bands in the NIR region due to MCC and water, principal component analysis (PCA) and generalized two-dimensional correlation spectroscopy (2DCOS) were applied to the obtained spectra. The NIR spectra in four adsorption stages separated by PCA were analyzed by 2DCOS. For the low M(c) range of 0.2-3.1 wt %, a decrease in the free or weakly hydrogen-bonded (H-bonded) MCC OH band, increases in the H-bonded MCC OH bands, and increases in the adsorbed water OH bands are observed. These results suggest that the inter- and intrachain H-bonds of MCC are formed by monomeric water molecule adsorption. In the M(c) range of 3.8-7.1 wt %, spectral changes in the NIR spectra reveal that the aggregation of water molecules starts at the surface of MCC. For the high M(c) range of 8.1-13.4 wt %, the NIR results suggest that the formation of bulk water occurs. It is revealed from the present study that approximately 3-7 wt % of adsorbed water is responsible for the stabilization of the H-bond network in MCC at the cellulose-water surface.

Adsorption↗

T-RFLP combined with principal component analysis and 16S rRNA gene sequencing: an effective strategy for comparison of fecal microbiota in infants of different ages.

The fecal microbiota of two healthy Swedish infants was monitored over time by terminal restriction fragment length polymorphism (T-RFLP) analysis of amplified 16S rRNA genes. Principal component analysis (PCA) of the T-RFLP profiles revealed that the fecal flora in both infants was quite stable during breast-feeding and a major change occurred after weaning. The two infants had different sets of microbiota at all sampling time points. 16S rDNA clone libraries were constructed and the predominant terminal restriction fragments (T-RFs) were identified by comparing T-RFLP patterns in the fecal community with that of corresponding 16S rDNA clones. Sequence analysis indicated that the infants were initially colonized mostly by members of Enterobacteriaceae, Veillonella, Enterococcus, Streptococcus, Staphylococcus and Bacteroides. The members of Enterobacteriaceae and Bacteroides were predominant during breast-feeding in both infants. However, Enterobacteriaceae decreased while members of clostridia increased after weaning. T-RFLP in combination with PCA and 16S rRNA gene sequencing was shown to be an effective strategy for comparing fecal microbiota in infants and pointing out the major changes.

Base Sequence↗

Automatic detection of emboli in the TCD RF signal using principal component analysis.

The transcranial Doppler (TCD) radio-frequency (RF) signal can provide additional information on events recorded during ultrasonic monitoring. Embolic signals appear as uniform and predictable shapes within the RF signal, enabling pattern recognition and image processing techniques to be used for their automated detection. This paper uses principal component analysis (PCA) to characterise the typical variation in embolic signal shape, within the RF signal, using training sets of in vitro and in vivo data. PCA techniques are then utilised to discriminate between previously unseen embolic and artifact signals. Although the results of this study show that the algorithms described in this paper do not yet have the accuracy required for their use in a clinical setting, it does demonstrate that this novel technique has the potential to be developed further.

Algorithms↗

Principal components of the protein dynamical transition.

Proteins exhibit a solvent-driven dynamical transition at 180-220 K, manifested by nonlinearity in the temperature dependence of the average mean-square displacement. Here, molecular dynamics simulations of hydrated myoglobin show that the onset of the transition at approximately 180 K is characterized by the appearance of a single double-well principal component mode involving a global motion of two groups of helices. As the temperature is raised a few more quasiharmonic and multiminimum components successively appear. The results indicate an underlying simplicity in the protein dynamical transition.

Computer Simulation↗