Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Principal component”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Principal component analysis for selection of optimal SNP-sets that capture intragenic genetic variation.

Candidate gene association studies often utilize one single nucleotide polymorphism (SNP) for analysis, with an initial report typically not being replicated by subsequent studies. The failure to replicate may result from incomplete or poor identification of disease-related variants or haplotypes, possibly due to naive SNP selection. A method for identification of linkage disequilibrium (LD) groups and selection of SNPs that capture sufficient intra-genic genetic diversity is described. We assume all SNPs with minor allele frequency above a pre-determined frequency have been identified. Principal component analysis (PCA) is applied to evaluate multivariate SNP correlations to infer groups of SNPs in LD (LD-groups) and to establish an optimal set of group-tagging SNPs (gtSNPs) that provide the most comprehensive coverage of intra-genic diversity while minimizing the resources necessary to perform an informative association analysis. This PCA method differs from haplotype block (HB) and haplotype-tagging SNP (htSNP) methods, in that an LD-group of SNPs need not be a contiguous DNA fragment. Results of the PCA method compared well with existing htSNP methods while also providing advantages over those methods, including an indication of the optimal number of SNPs needed. Further, evaluation of the method over multiple replicates of simulated data indicated PCA to be a robust method for SNP selection. Our findings suggest that PCA may be a powerful tool for establishing an optimal SNP set that maximizes the amount of genetic variation captured for a candidate gene using a minimal number of SNPs.

Adult↗

A study of polycyclic aromatic hydrocarbons concentrations and source identifications by methods of diagnostic ratio and principal component analysis at Taichung chemical Harbor near Taiwan Strait.

Fine (PM(2.5)) and Coarse (PM(2.5-10)) particulates concentrations of ambient air particle-bound polycyclic aromatic hydrocarbons (PAHs) were measured simultaneously from February 2004 to January 2005 at the Taichung Harbor (TH) sampling site near Taiwan of central Taiwan. Particle-bound polycyclic aromatic hydrocarbons (PAHs) were collected on quartz filters, the collected sample used soxhlet analytical method extracted with a dichloromethane (DCM)/n-hexane mixture (50/50, v/v) for 24h, and then the extracts were subjected to gas chromatography-mass spectrometric (GC-MS) analysis. The results indicated that vehicle emissions, coal combustion, incomplete combustion and pyrolysis of fuel and oil burning were the main source of PAHs near Taiwan Strait of central Taiwan. Diagnostic ratio and principal component analysis (PCA) were also used to characterize and identify PAHs emission source in this study.

Air Pollutants↗

Differential gene expression profiles and identification of the genes relevant to clinicopathologic factors in colorectal cancer selected by cDNA array method in combination with principal component analysis.

The clinical outcome of patients with colorectal cancer frequently varies even if they are at the same clinicopathologic stage. Alternative superior tumor markers of colorectal cancer are needed for prediction of clinical outcome. To clarify the regulatory factors in colorectal cancers, we examined differential expression profiles using cDNA macroarray technique with surgically resected specimens obtained from the patients with colorectal cancer. The gene profiles by an average-linkage hierarchical clustering analysis were found to be almost separable into two groups: tumor group and normal mucosa group. The relationship between several clinicopathologic factors and cancer related genes were investigated by using statistical analyses including principal component analysis (PCA). c-myc-binding protein MM-1, and c-jun proto-oncogene were identified as possible markers of tumor histology and clinical prognosis and early growth response protein 1 (EGR1) was selected to play an important role in progression of clinical stage. We conclude that, with PCA method, we successfully selected the genes relevant to clinicopathologic factors using limited population of clinical samples.

Aged↗

Determination of the geographical origin of green coffee by principal component analysis of carbon, nitrogen and boron stable isotope ratios.

In this study we show that the continental origin of coffee can be inferred on the basis of coupling the isotope ratios of several elements determined in green beans. The combination of the isotopic fingerprints of carbon, nitrogen and boron, used as integrated proxies for environmental conditions and agricultural practices, allows discrimination among the three continental areas producing coffee (Africa, Asia and America). In these continents there are countries producing 'specialty coffees', highly rated on the market that are sometimes mislabeled further on along the export-sale chain or mixed with cheaper coffees produced in other regions. By means of principal component analysis we were successful in identifying the continental origin of 88% of the samples analyzed. An intra-continent discrimination has not been possible at this stage of the study, but is planned in future work. Nonetheless, the approach using stable isotope ratios seems quite promising, and future development of this research is also discussed.

Africa↗

Combining principal component and spectral analyses with the method of moments in studies of quantal transmission.

This chapter considers methods for measurements of postsynaptic responses and simple approaches to the estimation of parameters of quantal release in synapses of the central nervous system of vertebrates. The use of these methods is illustrated by the analysis of single-fibre and "minimal" monosynaptic postsynaptic potentials (PSPs) or currents (PSCs) recorded from neurons of the frog spinal cord and rat hippocampus. First, we briefly discuss traditional methods of the response measurements using peak amplitudes or areas, further focusing on a novel method based on multivariate statistical techniques of the principal component analysis (PCA). This approach provides typically better signal-to-noise ratios and is able to separate two or more response components, which can arise due to activation of more than one presynaptic fibre, axon collaterals, receptor subtypes or spatially separated transmission sites. Second, spectral analysis is introduced as the method of choice to verify whether the amplitude fluctuations of the postsynaptic responses have a quantal nature and to obtain estimations of the "basic" quantal parameters, i.e. the quantal size (Q) and mean quantal content (m), without introducing assumptions on release statistics. Third, we show how the method of moments could be applied in the framework of the Poisson and binomial models to estimate the basic quantal parameters and parameters p and n, which reflect the release probability and maximum number of quanta released (or the number of effective release sites), respectively. Fourth, we show that the analysis of the moments can also be instrumental to reveal non-uniformity of release probabilities and compare how several competing models of neurotransmitter release fit to multiple experimental data sets.

Animals↗

Finding haplotype tagging SNPs by use of principal components analysis.

The immense volume and rapid growth of human genomic data, especially single nucleotide polymorphisms (SNPs), present special challenges for both biomedical researchers and automatic algorithms. One such challenge is to select an optimal subset of SNPs, commonly referred as "haplotype tagging SNPs" (htSNPs), to capture most of the haplotype diversity of each haplotype block or gene-specific region. This information-reduction process facilitates cost-effective genotyping and, subsequently, genotype-phenotype association studies. It also has implications for assessing the risk of identifying research subjects on the basis of SNP information deposited in public domain databases. We have investigated methods for selecting htSNPs by use of principal components analysis (PCA). These methods first identify eigenSNPs and then map them to actual SNPs. We evaluated two mapping strategies, greedy discard and varimax rotation, by assessing the ability of the selected htSNPs to reconstruct genotypes of non-htSNPs. We also compared these methods with two other htSNP finders, one of which is PCA based. We applied these methods to three experimental data sets and found that the PCA-based methods tend to select the smallest set of htSNPs to achieve a 90% reconstruction precision.

Chromosome Mapping↗

Correlations and complementarities in data and methods through Principal Components Analysis (PCA) applied to the results of the SPIn-Eco Project.

This paper demonstrates how the results from different methods can be interpreted on the basis of a statistical approach that can help find new hints in the evaluation of sustainability at the territorial level. The SPIn-Eco Project for the Province of Siena (Italy) is an example of an environmental sustainability assessment of an area using methods that are suitable for a large system: Ecological Footprint, Greenhouse Gas Inventory, Extended Exergy Analysis, Emergy Evaluation, and Remote Sensing. The calculation of many indicators, derived from these methods, has prompted us to use a statistical method (Principal Components Analysis, PCA) to understand the degree of similarity/congruence of the indicators (here we have examined 26 of them) and the possibility of recognizing patterns or clusters in the description of the 36 municipalities that compose the Province of Siena. Among the results, unexpectedly, emergy flow and the Ecological Footprint resulted as being completely uncorrelated, apparently due to the importance that the non-renewable part of the emergy holds in the evaluation. The municipalities of the province are considerably spread out over the graphs, even though that of Siena is quite far from the rest along the first dimension. In addition, we were able to distinguish between more homogeneous districts (sets of municipalities), such as Val di Merse and Val d'Orcia, and very diverse ones, such as Val d'Elsa and Val di Chiana.

Ecosystem↗

[A principal component analysis of the AGGIR scale in demented elderly patients].

AGGIR grid is the national standardized instrument determining aimed at the dependency of old people in France living in institutions as well as in the community. Attribution of the governmental financial assistance APA (Allocation Personnalisée d'Autonomie) depends essentially on the classification of frail old people in 6 degrees of dependency (GIR1 to GIR6). The aim of the present study was to test the reliability of this grid to evaluate the degree of dependency in demented elderly people. Mild, moderate or severe demented patients were included in the study (n= 120). A factorial validation of the A GGIR grid was performed by principal components analysis (PCA). This analysis showed a 5-factor solution: factor 1 named the property factor (27 percent of the variance), factor 2 named the dynamic factor (21 percent),factor 3 named the cognitive factor (20 percent), factor 4 named the external mobility factor (11 percent) and factor 5 named the communication factor (11 percent). The result showed that the AGGIR grid takes physical dependency more into account than psychological and behavioral dependency. This result suggests a need for readjustment of the AGGIR grid for demented patients by adding new variables taking into account psychosocial and behavioral disorders.

Activities of Daily Living↗

Principal components of mania.

OBJECTIVE: An alternative to the categorical classification of psychiatric diseases is the dimensional study of the signs and symptoms of psychiatric syndromes. To date, there have been few reports about the dimensions of mania, and the existence of a depressive dimension in mania remains controversial. The aim of this study was to investigate the dimensions of manic disorder by using classical scales to study the signs and symptoms of affective disorders. METHODS: One-hundred and three consecutively admitted inpatients who met DSM IV criteria for bipolar disorder, manic or mixed were rated with the Young Mania Rating Scale (YMRS) and the Hamilton Depression Rating Scale (HDRS-21). A principal components factor analysis of the HDRS-21 and the YMRS was carried out. RESULTS: Factor analysis showed five independent and clinically interpretable factors corresponding to depression, dysphoria, hedonism, psychosis and activation. The distribution of factor scores on the depressive factor was bimodal, whereas it was unimodal on the dysphoric, hedonism and activation factors. Finally, the psychosis factor was not normally distributed. LIMITATIONS: Patients of the sample were all medicated inpatients. CONCLUSIONS: Mania seems to be composed of three core dimensions, i.e. hedonism, dysphoria and activation, and is frequently accompanied by a psychotic and a depressive factor. The existence of a depressive factor suggests that it is essential to evaluate depression during mania, and the distribution of the depressive factor supports the existence of two different states in mania.

Adolescent↗

[Application of principal component analysis (PCA) for the estimation of source of heavy metal contamination in marine sediments].

Concentrations of heavy metals and organic matter in the bottom sediments of Jiaozhou Bay were determined and the average enrichment factors (AEFs) were used simultaneously to evaluate the extent of metal enrichment-contamination. Results show that heavy metal contamination in this bay could be divided into three groups: negligible to low contamination (AEFs < 2), which is the case of Zn (AEF = 1.11), Pb (AEF = 1.15), Cr (AEF = 1.52), Mn (AEF = 0.80) and Fe (AEF = 0.45); moderate contamination (AEFs = 2 - 3), which is the case of Cu (AEF = 2.79) and Cd (AEF = 2.52); certain to severe contamination (AEFs > 3), As (AEF = 3.03) and Hg (AEF = 8.08) being included in this group. Principal component analysis (PCA) was applied to estimate the sources of heavy metal contamination. Results that the first three components accounted for 52.61%, 17.37% and 15.60% of the total variance respectively exhibited that industrial wastewater, degradation of organic matter and erosion of rocks were the main sources of heavy metal contamination. The Q-analysis of PCA indicated that 14 stations could be divided into five groups. This result not only reflected the pollution characteristic of surface sediments, but also provided fundamental evidences for the putative analysis that industrial discharge is the main source of heavy metal contamination in Jiaozhou Bay.

Geologic Sediments↗

Quantitation of resonances in biological 31P NMR spectra via principal component analysis: potential and limitations.

This paper examines the potential and limitations of peak area quantitation of biological NMR spectra using principal component analysis (PCA), including its requirement for prior knowledge. The principles of the method are presented without in-depth mathematical treatment. PCA is illustrated for simulated data, 31P NMR spectra obtained consecutively over 1-2.5 days from perfused Rat-2 cells metabolizing the choline analogue phosphoniumcholine (Chop) and in vivo proton-decoupled, NOE-enhanced, three-dimensional CSI localized 31P NMR spectra of the liver of healthy volunteers. The results show that PCA can be used to quantitate strongly overlapping peaks without prior knowledge of the peak shapes or positions and to reconstruct spectra with significantly reduced noise variance. Two major limitations of PCA are presented: (1) PCA cannot separate peaks whose intensities are well correlated; (2) PCA is sensitive to differences in chemical shift and line-width of peaks between spectra. The discussion focuses on what knowledge of the biological and spectroscopic features of the samples and the principles of PCA is necessary for peak area quantitation via PCA.

Humans↗

Quantification of benzodiazepine-induced topographic EEG changes by a computerized wave form recognition method: application of a principal component analysis.

Topographic EEG changes with medazepam and diazepam in normals were analyzed by the computerized wave form recognition method. A principal component (PC) analysis, using such EEG elements as wave percent-time (8 bands) and average amplitude (7 bands), resulted in a considerable reduction of variables (4 PCs). In O1, because of high positive loadings in the average amplitude in all bands and a decrease in the mean score with either drug, PC-1 represents a component which reacts in the form of diminution of average amplitude as a whole. In Fp1, C3 and O1, PC-2, with a bipolarity of alpha 2 versus beta 1 and beta 2 in the wave percent-time in loading profile, could be a component showing characteristic changes common to the 2 benzodiazepines. In C3, because of a significant difference in the mean score between the 2 drugs, PC-4 might be a between-drug difference component in which diazepam (medazepam) increases (decreases) slow activity. The relationship between the score at PC-4 in C3 and daytime sleepiness may signify that the slow components are associated with sedation. Based on the correlation at PC-2 in Fp1, a marked increase in beta 1 and beta 2 components (responder) rather means less sleepiness, and relative preservation of alpha 1 and alpha 2 (non-responder) more sleepiness.

Adult↗

Development and preliminary validation of a Greek-language outpatient satisfaction questionnaire with principal components and multi-trait analyses.

BACKGROUND: In the recent years there is a growing interest in Greece concerning the measurement of the satisfaction of patients who are visiting the outpatient clinics of National Health System (NHS) general acute hospitals. The aim of this study is therefore to develop a patient satisfaction questionnaire and provide its preliminary validation. METHODS: A questionnaire in Greek has been developed by literature review, researchers' on the spot observation and interviews. Pretesting has been followed by telephone surveys in two short-term general NHS hospitals in Macedonia, Greece. A proportional stratified random sample of 285 subjects and a second random sample of 100 outpatients, drawn on March 2004, have been employed for the analysis. These have resulted in scale creation via Principal Components Analysis and psychometric testing for internal consistency, test-retest and interrater reliability as well as construct validity. RESULTS: Four summated scales have emerged regarding the pure outpatient component of the patients' visits, namely medical examination, hospital environment, comfort and appointment time. Cronbach's alpha coefficients and Pearson, Spearman and intraclass correlations indicate a high degree of scale reliability and validity. Two other scales--lab appointment time and lab experience--capture the apparently distinct yet complementary visitor experience related to the radiographic and laboratory tests. Psychometric tests are equally promising, however, some discriminant validity differences lack statistical significance. CONCLUSION: The instrument appears to be reliable and valid regarding the pure outpatient experience, whereas more research employing larger samples is required in order to establish the apparent psychometric properties of the complementary radiographic and laboratory-testing process, which is only relevant to about 25% of the subjects analysed here.

Data Collection↗

Facial expression of children receiving immunizations: a principal components analysis of the child facial coding system.

OBJECTIVE: To identify the structure of facial reaction to procedural pain and to determine the subset of facial actions that best describe the response. DESIGN: Observational. SETTING: Five rural and five urban physicians' offices. PATIENTS: One hundred twenty-three children aged 4 to 5 years undergoing routine diphtheria, pertussis, tetanus, and polio immunization. OUTCOME MEASURES: The Child Facial Coding System, comprising 13 discrete facial actions, was used to code each second of five 10-second phases from videotape: baseline, preneedle, needle, postneedle, and posthandling. Parents and a technician provided visual analog scale ratings of children's pain. Children provided a self-report using a Faces Pain Scale, and parents and nurses rated the children's pain and anxiety using visual analog scales. RESULTS: A "pain face" similar to that reported in adults emerged with the onset of pain. Principal component analyses revealed the frequency and intensity of facial action during the needle phase could be represented by components reflecting pain sensation, a "brave face," and the children's expectations for pain. Children's Faces Pain Scale and adult visual analog scale ratings were best predicted by components reflecting pain sensation and expectations of high pain. CONCLUSIONS: These results provide a preliminary indication that the Child Facial Coding System can be reduced to components that reflect several aspects of children's acute pain experience and predict self-reports and observer reports of children's pain.

Acute Disease↗

Mining gene expression data by interpreting principal components.

BACKGROUND: There are many methods for analyzing microarray data that group together genes having similar patterns of expression over all conditions tested. However, in many instances the biologically important goal is to identify relatively small sets of genes that share coherent expression across only some conditions, rather than all or most conditions as required in traditional clustering; e.g. genes that are highly up-regulated and/or down-regulated similarly across only a subset of conditions. Equally important is the need to learn which conditions are the decisive ones in forming such gene sets of interest, and how they relate to diverse conditional covariates, such as disease diagnosis or prognosis. RESULTS: We present a method for automatically identifying such candidate sets of biologically relevant genes using a combination of principal components analysis and information theoretic metrics. To enable easy use of our methods, we have developed a data analysis package that facilitates visualization and subsequent data mining of the independent sources of significant variation present in gene microarray expression datasets (or in any other similarly structured high-dimensional dataset). We applied these tools to two public datasets, and highlight sets of genes most affected by specific subsets of conditions (e.g. tissues, treatments, samples, etc.). Statistically significant associations for highlighted gene sets were shown via global analysis for Gene Ontology term enrichment. Together with covariate associations, the tool provides a basis for building testable hypotheses about the biological or experimental causes of observed variation. CONCLUSION: We provide an unsupervised data mining technique for diverse microarray expression datasets that is distinct from major methods now in routine use. In test uses, this method, based on publicly available gene annotations, appears to identify numerous sets of biologically relevant genes. It has proven especially valuable in instances where there are many diverse conditions (10's to hundreds of different tissues or cell types), a situation in which many clustering and ordering algorithms become problematic. This approach also shows promise in other topic domains such as multi-spectral imaging datasets.

Algorithms↗

New semen quality scores developed by principal component analysis of semen characteristics.

The purpose of this study was to determine whether semen characteristics can be reduced to 2 semen quality (SQ) scores and whether these new scores can help the clinician in assessing the reproductive outcome. A cross-sectional sample of 250 patients seeking infertility treatment were analyzed for semen characteristics. In addition, 177 male-factor patients (prostatitis with infection, n = 40; varicocele, n = 77; varicocele with infections, n = 11; and vasectomy reversal, n = 43) were also assessed. Sperm motion kinetics were measured by computer-assisted semen analysis (CASA) (concentration, percent motility, curvilinear velocity [VCL], straight-line velocity [VSL], average path velocity [VAP], linearity [LIN], and amplitude of lateral head displacement [ALH]). Sperm morphology was assessed by both World Health Organization (WHO) guidelines and Tygerberg strict criteria. The principal component analysis model was used to construct an SQ score and a relative semen quality (RQ) score. A separate set of 25 normal donors was included as controls to determine normal ranges of the semen scores. Among the patient samples, SQ and RQ scores (median and 25% and 75% interquartile values) were 89.9, 25.1, and 130.4 and 106.1, 45.2, and 165.9, respectively. The SQ score for the varicocele and varicocele with infection groups was comparable (78.6 +/- 17.4 and 84.8 +/- 20.6) but significantly different from the control (100 +/- 10, P <.001 and.03). Vasectomy reversal patients had an SQ score of 78.2 plus or minus 16.8 that was significantly lower than controls (P <.001). The correlation among semen characteristics allows for the efficient combining of semen measures. The composite scores can summarize overall SQ and quantity. Both SQ and RQ scores provide meaningful information on the quality of semen specimens for the clinician.

Humans↗

Principal components analysis for the visualisation of multidimensional chemical data acquired by scanning Raman microspectroscopy.

Raman microspectroscopy is ideally suited to surface analysis as it allows detailed chemical information to be acquired from surfaces at a relatively high spatial resolution (typically 1 microm). Using a motorised sample table or probe, it is possible to raster scan a surface to obtain spatially resolved chemical information. Visualisation of the acquired data is a problem, however, as the spectrum acquired at each point can contain several hundred individual intensity measurements. Existing visualisation methods are limited to plotting each scanned point with an intensity determined from the measured intensity at a single wavenumber, or the similarly between the point's spectrum and a reference spectrum. Such methods are wasteful as a lot of acquired information is discarded, and results are prone to misinterpretation due to background variance and instrumental noise. In this paper we introduce a new method that uses principal components analysis (PCA) to reduce the spectrum at each point to three factors that are then used to define the red, green and blue components of the corresponding point on a false colour map. To increase the effective resolution, interpolation is used to approximate the colours corresponding to points between those actually scanned. To demonstrate the technique, the internal surface of a beverage can, contaminated with a 40 microm diameter carbonised oven impurity, consisting mainly of sp2- and sp3-hybridised saturated carbon bonds, has been used as a case study.

Beverages↗

Expression profiles of mouse dendritic cell sarcoma are similar to those of hematopoietic stem cells or progenitors by clustering and principal component analyses.

We isolated and screened two tumor cell clones DD1 and DG6 with different capacity of metastasis from the same parent cell line, a mouse dendritic cell (DC) sarcoma, using limited dilution method. The genome-wide expressions of DD1 and DG6 cells were detected by Affymetrix's MOE-430A microarray. The expression profiles related with mouse DC development were downloaded from GEO at NCBI and ArrayExpress at EBI database. In order to compare the expression of DC sarcoma and DC developmental arrays which was performed by MG-U74av2, we had screened the best matched probesets between MOE-430A and MG-U74av2 according to the probe identities from Affymetrix technical annotation. After the normalization of 11 housekeeping genes across the 34 arrays (2 DC sarcoma and 32 DC developmental arrays), all these expression profiles were analyzed by the methods of hierarchical clustering, principal component analysis, nearest-neighborhood, and self-organizing maps. The results indicate that expression profiles of DC sarcoma are closer to those of the DC progenitors and hematopoietic stem cells from bone marrow compared with the sorted DCs from spleen. The results support the hypothesis that cancers (tumors or sarcomas) arise from stem cells. It is suggested that the DC sarcomas are more similar to the DC progenitors and hematopoietic stem cells than the relative mature DCs in gene expressions on the large-scale.

Animals↗