Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dimensionality Reduction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Multilocus analysis of hypertension: a hierarchical approach.

While hypertension is a complex disease with a well-documented genetic component, genetic studies often fail to replicate findings. One possibility for such inconsistency is that the underlying genetics of hypertension is not based on single genes of major effect, but on interactions among genes. To test this hypothesis, we studied both single locus and multilocus effects, using a case-control design of subjects from Ghana. Thirteen polymorphisms in eight candidate genes were studied. Each candidate gene has been shown to play a physiological role in blood pressure regulation and affects one of four pathways that modulate blood pressure: vasoconstriction (angiotensinogen, angiotensin converting enzyme - ACE, angiotensin II receptor), nitric oxide (NO) dependent and NO independent vasodilation pathways and sodium balance (G protein-coupled receptor kinase, GRK4). We evaluated single site allelic and genotypic associations, multilocus genotype equilibrium and multilocus genotype associations, using multifactor dimensionality reduction (MDR). For MDR, we performed systematic reanalysis of the data to address the role of various physiological pathways. We found no significant single site associations, but the hypertensive class deviated significantly from genotype equilibrium in more than 25% of all multilocus comparisons (2,162 of 8,178), whereas the normotensive class rarely did (11 of 8,178). The MDR analysis identified a two-locus model including ACE and GRK4 that successfully predicted blood pressure phenotype 70.5% of the time. Thus, our data indicate epistatic interactions play a major role in hypertension susceptibility. Our data also support a model where multiple pathways need to be affected in order to predispose to hypertension.

Alleles↗

MDR and PRP: a comparison of methods for high-order genotype-phenotype associations.

Complex diseases such as cardiovascular disease are likely due to the effects of high-order interactions among multiple genes and demographic factors. Therefore, in order to understand their underlying biological mechanisms, we need to consider simultaneously the effects of genotypes across multiple loci. Statistical methods such as multifactor dimensionality reduction (MDR), the combinatorial partitioning method (CPM), recursive partitioning (RP), and patterning and recursive partitioning (PRP) are designed to uncover complex relationships without relying on a specific model for the interaction, and are therefore well-suited to this data setting. However, the theoretical overlap among these methods and their relative merits have not been well characterized. In this paper we demonstrate mathematically that MDR is a special case of RP in which (1) patterns are used as predictors (PRP), (2) tree growth is restricted to a single split, and (3) misclassification error is used as the measure of impurity. Both approaches are applied to a case-control study assessing the effect of eleven single nucleotide polymorphisms on coronary artery calcification in people at risk for cardiovascular disease.

Cardiovascular Diseases↗

Data-mining methods as useful tools for predicting individual drug response: application to CYP2D6 data.

OBJECTIVES: Selecting a maximally informative subset of polymorphisms to predict a clinical outcome, such as drug response, requires appropriate search methods due to the increased dimensionality associated with looking at multiple genotypes. In this study, we investigated the ability of several pattern recognition methods to identify the most informative markers in the CYP2D6 gene for the prediction of CYP2D6 metabolizer status. METHODS: Four data-mining tools were explored: decision trees, random forests, artificial neural networks, and the multifactor dimensionality reduction (MDR) method. Marker selection was performed separately in eight population samples of different ethnic origin to evaluate to what extent the most informative markers differ across ethnic groups. RESULTS: Our results show that the number of polymorphisms required to predict CYP2D6 metabolic phenotype with a high accuracy can be dramatically reduced owing to the strong haplotype block structure observed at CYP2D6. MDR and neural networks provided nearly identical results and performed the best. CONCLUSION: Data-mining methods, such as MDR and neural networks, appear as promising tools to improve the efficiency of genotyping tests in pharmacogenetics with the ultimate goal of pre-screening patients for individual therapy selection with minimum genotyping effort.

Cytochrome P-450 CYP2D6↗

Epistatic and pleiotropic effects of polymorphisms in the fibrinogen and coagulation factor XIII genes on plasma fibrinogen concentration, fibrin gel structure and risk of myocardial infarction.

An intricate interplay between the genes encoding fibrinogen gamma (FGG), alpha (FGA) and beta (FGB), coagulation factor XIII (F13A1) and interleukin 6 (IL6) and environmental factors is likely to influence plasma fibrinogen concentration, fibrin clot structure and risk of myocardial infarction (MI). In the present study, the potential contribution of SNPs harboured in the fibrinogen, IL6 and F13A1 genes to these biochemical and clinical phenotypes was examined. A database and biobank based on 387 survivors of a first MI and population-based controls were used. Sixty controls were selected according to FGG 9340T > C [rs1049636] genotype for studies on fibrin clot structure using the liquid permeation method. The multifactor dimensionality reduction method was used for interaction analyses. We here report that the FGA 2224G > A [rs2070011] SNP (9.2%), plasma fibrinogen concentration (13.1%) and age (8.1%) appeared as independent determinants of fibrin gel porosity. The FGA 2224G > A SNP modulated the relation between plasma fibrinogen concentration and fibrin clot porosity. The FGG-FGA*4 haplotype, composed of the minor FGG 9340C and FGA 2224A alleles, had similar effects, supporting its reported protective role in relation to MI. Significant epistasis on plasma fibrinogen concentration was detected between the FGA 2224G > A and F13A1 Val34Leu [rs5985] SNPs (p < 0.001). The FGG 9340T > C and FGB 1038G > A [rs1800791] SNPs appeared to interact on MI risk, explaining the association of FGG-FGB haplotypes with MI in the absence of effects of individual SNPs. Thus, epistatic and pleiotropic effects of polymorphisms contribute to the variation in plasma fibrinogen concentration, fibrin clot structure and risk of MI.

Environment↗

Renin-angiotensin system gene polymorphisms and atrial fibrillation.

BACKGROUND: The activated local atrial renin-angiotensin system (RAS) has been reported to play an important role in the pathogenesis of atrial fibrillation (AF). We hypothesized that RAS genes might be among the susceptibility genes of nonfamilial structural AF and conducted a genetic case-control study to demonstrate this. METHODS AND RESULTS: A total of 250 patients with documented nonfamilial structural AF and 250 controls were selected. The controls were matched to cases on a 1-to-1 basis with regard to age, gender, presence of left ventricular dysfunction, and presence of significant valvular heart disease. The ACE gene insertion/deletion polymorphism, the T174M, M235T, G-6A, A-20C, G-152A, and G-217A polymorphisms of the angiotensinogen gene, and the A1166C polymorphism of the angiotensin II type I receptor gene were genotyped. In multilocus haplotype analysis, the angiotensinogen gene haplotype profile was significantly different between cases and controls (chi2=62.5, P=0.0002). In single-locus analysis, M235T, G-6A, and G-217A were significantly associated with AF. Frequencies of the M235, G-6, and G-217 alleles were significantly higher in cases than in controls (P=0.000, 0.005, and 0.002, respectively). The odds ratios for AF were 2.5 (95% CI 1.7 to 3.3) with M235/M235 plus M235/T235 genotype, 3.3 (95% CI 1.3 to 10.0) with G-6/G-6 genotype, and 2.0 (95% CI 1.3 to 2.5) with G-217/G-217 genotype. Furthermore, significant gene-gene interactions were detected by the multifactor-dimensionality reduction method and multilocus linkage disequilibrium tests. CONCLUSIONS: This study demonstrates the association of RAS gene polymorphisms with nonfamilial structural AF and may provide the rationale for clinical trials to investigate the use of ACE inhibitor or angiotensin II antagonist in the treatment of structural AF.

Aged↗

What causes a neuron to spike?

The computation performed by a neuron can be formulated as a combination of dimensional reduction in stimulus space and the nonlinearity inherent in a spiking output. White noise stimulus and reverse correlation (the spike-triggered average and spike-triggered covariance) are often used in experimental neuroscience to "ask" neurons which dimensions in stimulus space they are sensitive to and to characterize the nonlinearity of the response. In this article, we apply reverse correlation to the simplest model neuron with temporal dynamics-the leaky integrate-and-fire model-and find that for even this simple case, standard techniques do not recover the known neural computation. To overcome this, we develop novel reverse-correlation techniques by selectively analyzing only "isolated" spikes and taking explicit account of the extended silences that precede these isolated spikes. We discuss the implications of our methods to the characterization of neural adaptation. Although these methods are developed in the context of the leaky integrate-and-fire model, our findings are relevant for the analysis of spike trains from real neurons.

Action Potentials↗

Mixtures of probabilistic principal component analyzers.

Principal component analysis (PCA) is one of the most popular techniques for processing, compressing, and visualizing data, although its effectiveness is limited by its global linearity. While nonlinear variants of PCA have been proposed, an alternative paradigm is to capture data complexity by a combination of local linear PCA projections. However, conventional PCA does not correspond to a probability density, and so there is no unique way to combine PCA models. Therefore, previous attempts to formulate mixture models for PCA have been ad hoc to some extent. In this article, PCA is formulated within a maximum likelihood framework, based on a specific form of gaussian latent variable model. This leads to a well-defined mixture model for probabilistic principal component analyzers, whose parameters can be determined using an expectation-maximization algorithm. We discuss the advantages of this model in the context of clustering, density modeling, and local dimensionality reduction, and we demonstrate its application to image compression and handwritten digit recognition.

Algorithms↗

Self-organization as an iterative kernel smoothing process.

Kohonen's self-organizing map, when described in a batch processing mode, can be interpreted as a statistical kernel smoothing problem. The batch SOM algorithm consists of two steps. First, the training data are partitioned according to the Voronoi regions of the map unit locations. Second, the units are updated by taking weighted centroids of the data falling into the Voronoi regions, with the weighing function given by the neighborhood. Then, the neighborhood width is decreased and steps 1, 2 are repeated. The second step can be interpreted as a statistical kernel smoothing problem where the neighborhood function corresponds to the kernel and neighborhood width corresponds to kernel span. To determine the new unit locations, kernel smoothing is applied to the centroids of the Voronoi regions in the topological space. This interpretation leads to some new insights concerning the role of the neighborhood and dimensionality reduction. It also strengthens the algorithm's connection with the Principal Curve algorithm. A generalized self-organizing algorithm is proposed, where the kernel smoothing step is replaced with an arbitrary nonparametric regression method.

Algorithms↗

DNA microarray data and contextual analysis of correlation graphs.

BACKGROUND: DNA microarrays are used to produce large sets of expression measurements from which specific biological information is sought. Their analysis requires efficient and reliable algorithms for dimensional reduction, classification and annotation. RESULTS: We study networks of co-expressed genes obtained from DNA microarray experiments. The mathematical concept of curvature on graphs is used to group genes or samples into clusters to which relevant gene or sample annotations are automatically assigned. Application to publicly available yeast and human lymphoma data demonstrates the reliability of the method in spite of its simplicity, especially with respect to the small number of parameters involved. CONCLUSIONS: We provide a method for automatically determining relevant gene clusters among the many genes monitored with microarrays. The automatic annotations and the graphical interface improve the readability of the data. A C++ implementation, called Trixy, is available from http://tagc.univ-mrs.fr/bioinformatics/trixy.html.

Algorithms↗

Feature selection and nearest centroid classification for protein mass spectrometry.

BACKGROUND: The use of mass spectrometry as a proteomics tool is poised to revolutionize early disease diagnosis and biomarker identification. Unfortunately, before standard supervised classification algorithms can be employed, the "curse of dimensionality" needs to be solved. Due to the sheer amount of information contained within the mass spectra, most standard machine learning techniques cannot be directly applied. Instead, feature selection techniques are used to first reduce the dimensionality of the input space and thus enable the subsequent use of classification algorithms. This paper examines feature selection techniques for proteomic mass spectrometry. RESULTS: This study examines the performance of the nearest centroid classifier coupled with the following feature selection algorithms. Student-t test, Kolmogorov-Smirnov test, and the P-test are univariate statistics used for filter-based feature ranking. From the wrapper approaches we tested sequential forward selection and a modified version of sequential backward selection. Embedded approaches included shrunken nearest centroid and a novel version of boosting based feature selection we developed. In addition, we tested several dimensionality reduction approaches, namely principal component analysis and principal component analysis coupled with linear discriminant analysis. To fairly assess each algorithm, evaluation was done using stratified cross validation with an internal leave-one-out cross-validation loop for automated feature selection. Comprehensive experiments, conducted on five popular cancer data sets, revealed that the less advocated sequential forward selection and boosted feature selection algorithms produce the most consistent results across all data sets. In contrast, the state-of-the-art performance reported on isolated data sets for several of the studied algorithms, does not hold across all data sets. CONCLUSION: This study tested a number of popular feature selection methods using the nearest centroid classifier and found that several reportedly state-of-the-art algorithms in fact perform rather poorly when tested via stratified cross-validation. The revealed inconsistencies provide clear evidence that algorithm evaluation should be performed on several data sets using a consistent (i.e., non-randomized, stratified) cross-validation procedure in order for the conclusions to be statistically sound.

Algorithms↗

A novel data mining method to identify assay-specific signatures in functional genomic studies.

BACKGROUND: The highly dimensional data produced by functional genomic (FG) studies makes it difficult to visualize relationships between gene products and experimental conditions (i.e., assays). Although dimensionality reduction methods such as principal component analysis (PCA) have been very useful, their application to identify assay-specific signatures has been limited by the lack of appropriate methodologies. This article proposes a new and powerful PCA-based method for the identification of assay-specific gene signatures in FG studies. RESULTS: The proposed method (PM) is unique for several reasons. First, it is the only one, to our knowledge, that uses gene contribution, a product of the loading and expression level, to obtain assay signatures. The PM develops and exploits two types of assay-specific contribution plots, which are new to the application of PCA in the FG area. The first type plots the assay-specific gene contribution against the given order of the genes and reveals variations in distribution between assay-specific gene signatures as well as outliers within assay groups indicating the degree of importance of the most dominant genes. The second type plots the contribution of each gene in ascending or descending order against a constantly increasing index. This type of plots reveals assay-specific gene signatures defined by the inflection points in the curve. In addition, sharp regions within the signature define the genes that contribute the most to the signature. We proposed and used the curvature as an appropriate metric to characterize these sharp regions, thus identifying the subset of genes contributing the most to the signature. Finally, the PM uses the full dataset to determine the final gene signature, thus eliminating the chance of gene exclusion by poor screening in earlier steps. The strengths of the PM are demonstrated using a simulation study, and two studies of real DNA microarray data--a study of classification of human tissue samples and a study of E. coli cultures with different medium formulations. CONCLUSION: We have developed a PCA-based method that effectively identifies assay-specific signatures in ranked groups of genes from the full data set in a more efficient and simplistic procedure than current approaches. Although this work demonstrates the ability of the PM to identify assay-specific signatures in DNA microarray experiments, this approach could be useful in areas such as proteomics and metabolomics.

Algorithms↗

The challenge for genetic epidemiologists: how to analyze large numbers of SNPs in relation to complex diseases.

Genetic epidemiologists have taken the challenge to identify genetic polymorphisms involved in the development of diseases. Many have collected data on large numbers of genetic markers but are not familiar with available methods to assess their association with complex diseases. Statistical methods have been developed for analyzing the relation between large numbers of genetic and environmental predictors to disease or disease-related variables in genetic association studies. In this commentary we discuss logistic regression analysis, neural networks, including the parameter decreasing method (PDM) and genetic programming optimized neural networks (GPNN) and several non-parametric methods, which include the set association approach, combinatorial partitioning method (CPM), restricted partitioning method (RPM), multifactor dimensionality reduction (MDR) method and the random forests approach. The relative strengths and weaknesses of these methods are highlighted. Logistic regression and neural networks can handle only a limited number of predictor variables, depending on the number of observations in the dataset. Therefore, they are less useful than the non-parametric methods to approach association studies with large numbers of predictor variables. GPNN on the other hand may be a useful approach to select and model important predictors, but its performance to select the important effects in the presence of large numbers of predictors needs to be examined. Both the set association approach and random forests approach are able to handle a large number of predictors and are useful in reducing these predictors to a subset of predictors with an important contribution to disease. The combinatorial methods give more insight in combination patterns for sets of genetic and/or environmental predictor variables that may be related to the outcome variable. As the non-parametric methods have different strengths and weaknesses we conclude that to approach genetic association studies using the case-control design, the application of a combination of several methods, including the set association approach, MDR and the random forests approach, will likely be a useful strategy to find the important genes and interaction patterns involved in complex diseases.

Editorial↗

A role for CETP TaqIB polymorphism in determining susceptibility to atrial fibrillation: a nested case control study.

BACKGROUND: Studies investigating the genetic and environmental characteristics of atrial fibrillation (AF) may provide new insights in the complex development of AF. We aimed to investigate the association between several environmental factors and loci of candidate genes, which might be related to the presence of AF. METHODS: A nested case-control study within the PREVEND cohort was conducted. Standard 12 lead electrocardiograms were recorded and AF was defined according to Minnesota codes. For every case, an age and gender matched control was selected from the same population (n = 194). In addition to logistic regression analyses, the multifactor-dimensionality reduction (MDR) method and interaction entropy graphs were used for the evaluation of gene-gene and gene-environment interactions. Polymorphisms in genes from the Renin-angiotensin, Bradykinin and CETP systems were included. RESULTS: Subjects with AF had a higher prevalence of electrocardiographic left ventricular hypertrophy, ischemic heart disease, hypertension, renal dysfunction, elevated levels of C-reactive protein (CRP) and increased urinary albumin excretion as compared to controls. The polymorphisms of the Renin-angiotensin system and Bradykinin gene did not show a significant association with AF (p > 0.05). The TaqIB polymorphism of the CETP gene was significantly associated with the presence of AF (p < 0.05). Using the MDR method, the best genotype-phenotype models included the combination of micro- or macroalbuminuria and CETP TaqIB polymorphism, CRP >3 mg/L and CETP TaqIB polymorphism, renal dysfunction and the CETP TaqIB polymorphism, and ischemic heart disease and CETP TaqIB polymorphism (1000 fold permutation testing, P < 0.05). Interaction entropy graph showed that the combination of albuminuria and CETP TaqIB polymorphism removed the most entropy. CONCLUSION: CETP TaqIB polymorphism is significantly associated with the presence of AF in the context of micro- or macroalbuminuria, elevated C-reactive protein, renal dysfunction, and ischemic heart disease.

Atrial Fibrillation↗

P-value based visualization of codon usage data.

Two important and not yet solved problems in bacterial genome research are the identification of horizontally transferred genes and the prediction of gene expression levels. Both problems can be addressed by multivariate analysis of codon usage data. In particular dimensionality reduction methods for visualization of multivariate data have shown to be effective tools for codon usage analysis. We here propose a multidimensional scaling approach using a novel similarity measure for codon usage tables. Our probabilistic similarity measure is based on P-values derived from the well-known chi-square test for comparison of two distributions. Experimental results on four microbial genomes indicate that the new method is well-suited for the analysis of horizontal gene transfer and translational selection. As compared with the widely-used correspondence analysis, our method did not suffer from outlier sensitivity and showed a better clustering of putative alien genes in most cases.

Journal Article↗

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis↗

Interaction between toll-like receptor 4 polymorphism and abdominal obesity on ovarian cancer risk in Chinese women.

OBJECTIVES: the aim of this study was to evaluate the impact of TLR4 gene single nucleotide polymorphisms (SNPs) and additional TLR4 gene SNP- SNP and SNP- abdominal obesity (AO) interaction on ovarian cancer (OC) risk. METHODS: Generalized multifactor dimensionality reduction method were utilized to identify the most informative interactions between four SNPs in the TLR4 gene and abdominal obesity. Logistic regression was employed to investigate the association between 4 SNPs within TLR4 gene and OC risk, and additional SNP- SNP and gene- AO interaction on OC risk, ORs (95%CI) were calculated. RESULTS: The analysis of logistic regression indicated a markedly elevated risk of OC in individuals carrying either the rs4986790-G or rs11536889-C alleles in the TLR4 gene compared to those with the standard genetic variations, adjusted ORs (95%CI) were 1.61 (1.28-1.96) and 1.48 (1.09-1.91). GMDR analysis indicated a significant two-locus model (p&#x2009;=&#x2009;0.018) involving rs4986790 and rs11536889, and a significant two-locus model (p&#x2009;=&#x2009;0.001) involving rs4986790 and AO. Participants with rs4986790- AG/GG and rs11536889GC/ CC genotype has the highest OC risk, compared to participants with rs4986790-AA and rs11536889-GG genotype, OR (95%CI)&#x2009;=&#x2009;2.58 (1.46-3.71), and abdominal obese participants with rs4986790- AG/GG genotype have the highest OC risk, compared to non- abdominal obese participants with rs4986790-AA genotype, OR (95%CI)&#x2009;=&#x2009;3.17 (1.78-4.58). CONCLUSIONS: The findings suggested that TLR4 gene rs4986790 and rs11536889 polymorphisms were associated with increased OC risk. Significant interaction also existed between rs4986790 and AO, which means that the WC levels may influence the impact of rs4986790 on OC risk.

Adult↗

ABO exon polymorphisms are related to ischemic stroke in a Chinese Han population.

BACKGROUND: Recent research have underscored the relation of ABO blood group system to cerebrovascular disorders predisposition. The present investigation endeavors to delve into the relationship between ABO polymorphisms and ischemic stroke (IS) risk. METHODS: A cohort of 646 IS patients and 649 matched healthy controls was recruited. Genotyping of five SNPs within ABO were conducted by Agena MassARRAY platform. Logistic regression models were employed to estimate odds ratios (ORs) and 95% confidence intervals (CIs). Additionally, SNP-SNP interaction was assessed by multifactor dimensionality reduction (MDR) method. Furthermore, Analysis of Variance (ANOVA) was utilized to explore the association between genotypes and blood lipid profiles. RESULTS: The study identified an elevated IS risk associated with rs8176740 and rs8176720 in the overall population. Notably, ABO rs8176720 emerged as the most informative single-locus model for IS susceptibility. These variants were related to an elevated IS risk, specifically in female subjects, the subgroup aged&#x2009;>&#x2009;64 years, non-smokers, drinkers or non-drinkers. Moreover, rs8176749 and rs8176745 were associated with red blood cell count levels and total bilirubin levels. CONCLUSION: This study firstly demonstrated the association of ABO rs8176740 and rs8176720 with IS incidence, which increased the understanding regarding the effect of ABO on IS pathogenesis.

Aged↗