Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dimensionality Reduction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Robust and comprehensive analysis of 20 osteoporosis candidate genes by very high-density single-nucleotide polymorphism screen among 405 white nuclear families identified significant association and gene-gene interaction.

UNLABELLED: Many "novel" osteoporosis candidate genes have been proposed in recent years. To advance our knowledge of their roles in osteoporosis, we screened 20 such genes using a set of high-density SNPs in a large family-based study. Our efforts led to the prioritization of those osteoporosis genes and the detection of gene-gene interactions. INTRODUCTION: We performed large-scale family-based association analyses of 20 novel osteoporosis candidate genes using 277 single nucleotide polymorphisms (SNPs) for the quantitative trait BMD variation and the qualitative trait osteoporosis (OP) at three clinically important skeletal sites: spine, hip, and ultradistal radius (UD). MATERIALS AND METHODS: One thousand eight hundred seventy-three subjects from 405 white nuclear families were genotyped and analyzed with an average density of one SNP per 4 kb across the 20 genes. We conducted association analyses by SNP- and haplotype-based family-based association test (FBAT) and performed gene-gene interaction analyses using multianalytic approaches such as multifactor-dimensionality reduction (MDR) and conditional logistic regression. RESULTS AND CONCLUSIONS: We detected four genes (DBP, LRP5, CYP17, and RANK) that showed highly suggestive associations (10,000-permutation derived empirical global p < or = 0.01) with spine BMD/OP; four genes (CYP19, RANK, RANKL, and CYP17) highly suggestive for hip BMD/OP; and four genes (CYP19, BMP2, RANK, and TNFR2) highly suggestive for UD BMD/OP. The associations between BMP2 with UD BMD and those between RANK with OP at the spine, hip, and UD also met the experiment-wide stringent criterion (empirical global p < or = 0.0007). Sex-stratified analyses further showed that some of the significant associations in the total sample were driven by either male or female subjects. In addition, we identified and validated a two-locus gene-gene interaction model involving GCR and ESR2, for which prior biological evidence exists. Our results suggested the prioritization of osteoporosis candidate genes from among the many proposed in recent years and revealed the significant gene-gene interaction effects influencing osteoporosis risk.

Adult↗

PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clustering.

The development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency.

Single-Cell Analysis↗

Single-nucleotide polymorphisms for diagnosis of salt-sensitive hypertension.

BACKGROUND: Salt-sensitive (SS) hypertension affects >30 million Americans and is often associated with low plasma renin activity. We tested the diagnostic validity of several candidate genes for SS and low-renin hypertension. METHODS: In Japanese patients with newly diagnosed, untreated hypertension (n = 184), we studied polymorphisms in 10 genes, including G protein-coupled receptor kinase type 4 (GRK4), some variations of which are associated with hypertension and impair D1 receptor (D1R)-inhibited renal sodium transport. We used the multifactor dimensionality reduction method to determine the genotype associated with salt sensitivity (> or =10% increase in blood pressure with high sodium intake) or low renin. To determine whether the GRK4 genotype is associated with impaired D1R function, we tested the natriuretic effect of docarpamine, a dopamine prodrug, in normotensive individuals with or without GRK4 polymorphisms (n = 18). RESULTS: A genetic model based on GRK4 R65L, GRK4 A142V, and GRK4 A486V was 94.4% predictive of SS hypertension, whereas the single-locus model with only GRK4 A142V was 78.4% predictive, and a 2-locus model of GRK4 A142V and CYP11B2 C-344T was 77.8% predictive of low-renin hypertension. Sodium excretion was inversely related to the number of GRK4 variants in hypertensive persons, and the natriuretic response to dopaminergic stimulation was impaired in normotensive persons having > or =3 GRK4 gene variants. CONCLUSIONS: GRK4 gene variants are associated with SS and low-renin hypertension. However, the genetic model predicting SS hypertension is different from the model for low renin, suggesting genetic differences in these 2 phenotypes. Like low-renin testing, screening for GRK4 variants may be a useful diagnostic adjunct for detection of SS hypertension.

Asian People↗

Myeloid landscape of BRAF-mutant papillary thyroid cancer and thyroiditis.

Papillary thyroid cancer (PTC) is less aggressive when associated with lymphocytic thyroiditis (LT), even in the presence of oncogenic BRAF, including smaller tumours, less lymph node involvement and reduced extrathyroidal extension. To investigate possible immune mechanisms underlying this association, we compared the tumour microenvironment of PTC-BRAF with LT and that without LT using single-cell RNA sequencing (scRNA-seq). Single-cell libraries were generated from fresh and fixed tumour samples with post-dissociation viability >70% using the 10x Genomics Chromium Platform and sequenced on an Illumina NovaSeq 6000. We analysed scRNA-seq data from 11 PTC-BRAF tumours: four with LT (one publicly available sample) and seven without LT. Downstream analyses included quality control, batch correction, dimensionality reduction, and differential gene expression analysis. We found that neutrophils were the predominant myeloid cell type in PTCs without LT. Thyrocytes without LT showed significant expression of the neutrophil recruitment chemokine ECRG4. In the absence of LT, neutrophils expressed oncogenic genes with poor clinical outcomes. In contrast, thyrocytes from tumours with LT showed increased expression of MHC-II antigen presentation, consistent with effective immune surveillance. Thyrocytes and macrophages in the presence of LT showed enrichment of interferon gamma response pathways. Our data suggest that LT in thyroid cancer is associated with enhanced antigen presentation and fewer features of pro-tumourigenic innate immune activity. These results identify previously under-recognised innate immune cell population and associated transcriptomic features, which suggest new mechanisms to target immune treatments in PTC refractory to other therapies.

Humans↗

A hybrid neural network system for prediction and recognition of promoter regions in human genome.

This paper proposes a high specificity and sensitivity algorithm called PromPredictor for recognizing promoter regions in the human genome. PromPredictor extracts compositional features and CpG islands information from genomic sequence, feeding these features as input for a hybrid neural network system (HNN) and then applies the HNN for prediction. It combines a novel promoter recognition model, coding theory, feature selection and dimensionality reduction with machine learning algorithm. Evaluation on Human chromosome 22 was approximately 66% in sensitivity and approximately 48% in specificity. Comparison with two other systems revealed that our method had superior sensitivity and specificity in predicting promoter regions. PromPredictor is written in MATLAB and requires Matlab to run. PromPredictor is freely available at http://www.whtelecom.com/Prompredictor.htm.

Computational Biology↗

Machine learning for detecting gene-gene interactions: a review.

Complex interactions among genes and environmental factors are known to play a role in common human disease aetiology. There is a growing body of evidence to suggest that complex interactions are 'the norm' and, rather than amounting to a small perturbation to classical Mendelian genetics, interactions may be the predominant effect. Traditional statistical methods are not well suited for detecting such interactions, especially when the data are high dimensional (many attributes or independent variables) or when interactions occur between more than two polymorphisms. In this review, we discuss machine-learning models and algorithms for identifying and characterising susceptibility genes in common, complex, multifactorial human diseases. We focus on the following machine-learning methods that have been used to detect gene-gene interactions: neural networks, cellular automata, random forests, and multifactor dimensionality reduction. We conclude with some ideas about how these methods and others can be integrated into a comprehensive and flexible framework for data mining and knowledge discovery in human genetics.

Algorithms↗

Recent Advances in Multi-Omics of Systemic Lupus Erythematosus.

This comprehensive narrative review examines recent advances in multi-omics research for Systemic Lupus Erythematosus (SLE), emphasizing integrated approaches over single-omics studies. The review critically evaluates technological advancements, methodological innovations, and clinical applications while identifying current limitations and future research directions. We conducted a comprehensive narrative review following SANRA guidelines, searching PubMed, Web of Science, Scopus, and Embase, covering publications from January 2018 to June 2025. The review focuses on studies integrating two or more omics layers in SLE research, with emphasis on computational methods, biomarker validation, and clinical applications. Multi-omics integration has revealed critical insights into SLE pathogenesis, including immune cell heterogeneity, gene-environment interactions, and metabolic dysregulation. However, significant challenges remain in data integration methodologies, small sample sizes, and biomarker reproducibility. Current computational approaches include early integration (concatenation), intermediate integration (joint dimensionality reduction), and late integration (ensemble methods). While multi-omics approaches offer unprecedented insights into SLE complexity, standardized integration protocols and robust validation frameworks are urgently needed. Small sample sizes and heterogeneity issues limit reproducibility, particularly affecting biomarker discovery and clinical translation. Multi-omics integration represents a paradigm shift toward precision medicine in SLE, but realizing this potential requires addressing current methodological limitations, standardizing validation processes, and developing robust computational frameworks for reliable clinical applications.

Humans↗

Selection of molecular descriptors with artificial intelligence for the understanding of HIV-1 protease peptidomimetic inhibitors-activity.

Quantitative Structure Activity Relationship (QSAR) techniques are used routinely by computational chemists in drug discovery and development to analyze datasets of compounds. Quantitative numerical methods like Partial Least Squares (PLS) and Artificial Neural Networks (ANN) have been used on QSAR to establish correlations between molecular properties and bioactivity. However, ANN may be advantageous over PLS because it considers the interrelations of the modeled variables. This study focused on the HIV-1 Protease (HIV-1 Pr) inhibitors belonging to the peptidomimetic class of compounds. The main objective was to select molecular descriptors with the best predictive value for antiviral potency (Ki). PLS and ANN were used to predict Ki activity of HIV-1 Pr inhibitors and the results were compared. To address the issue of dimensionality reduction, Genetic Algorithms (GA) were used for variable selection and their performance was compared against that of ANN. Finally, the structure of the optimum ANN achieving the highest Pearson's-R coefficient was determined. On the basis of Pearson's-R, PLS and ANN were compared to determine which exhibits maximum performance. Training and validation of models was performed on 15 random split sets of the master dataset consisted of 231 compounds. For each compound 192 molecular descriptors were considered. The molecular structure and constant of inhibition (Ki) were selected from the NIAID database. Study findings suggested that non-covalent interactions such as hydrophobicity, shape and hydrogen bonding describe well the antiviral activity of the HIV-1 Pr compounds. The significance of lipophilicity and relationship to HIV-1 associated hyperlipidemia and lipodystrophy syndrome warrant further investigation.

Algorithms↗

A first QSAR model for galectin-3 glycomimetic inhibitors based on 3D docked structures.

This study presents the first QSAR model for Galectin-3 glycomimetic inhibitors based on docked structures to the carbohydrate recognition domain (CRD). Quantitative numerical methods such as PLS (Partial Least Squares) and ANN (Artificial Neural Networks) have been used and compared on QSAR models to establish correlations between molecular properties and binding affinity values (Kd). Training and validation of QSAR predictive models was performed on a master dataset consisting of 136 compounds. The molecular structures and binding affinities (Kd) (136 compounds) were obtained from the literature. To address the issue of dimensionality reduction, molecular descriptors were selected with PLS contingency approach, ANN, PCA (Principal Component Analysis) and GA (Genetic Algorithms) for the best predictive Galectin-3 binding affinity (Kd). Final sets comprising 56, 31 and 35 descriptors were obtained with PLS, PCA and ANN, respectively. The objective of this prototype QSAR model is to serve as a first guideline for the design of novel and potent Gal-3 selective inhibitors with emphasis on modification at both C-3' and at O-3 positions.

Biomimetic Materials↗

Identification of critical heterodimer protein interface parameters by multi-dimensional scaling in euclidian space.

Protein subunit dimers are either homodimers (consisting of identical polypeptides) or heterodimers (consisting of different polypeptides). Protein dimers are involved in several cellular processes and an understanding of their molecular principle in complexations (subunit-subunit interaction) is essential. This is generally studied using 3D structures of homodimers and heterodimers determined by X-ray crystallography. However, the current knowledge on subunit interaction is limited due to lack of sufficient 3D dimer structures. It is our interest to study heterodimers using 3D structures to identify interaction parameters that would help in the development of a model to predict heterodimer interaction sites just from protein sequences. The efficiency of such models depends on the weighted contribution of numerous parameters characterizing heterodimer interfaces. Therefore, we studied the salient features of 111 interface parameters in 65 heterodimer structures. In this study, we applied multi-dimensional scaling for dimensionality reduction on these parameters to select the most critical ones that best characterize heterodimer interfaces. The significance of these parameters in subunit interaction is discussed.

Computational Biology↗

Bioinformatics approaches for detecting gene-gene and gene-environment interactions in studies of human disease.

Neurological and mental disorders occur often, with approximately 450 million people suffering from them worldwide. Like most other common diseases, neurological disorders are hypothesized to be highly complex, with interactions among genes and risk factors playing a major role in the process. In recent years it has become obvious that for common diseases there may be more complex interactions among genes with and without strong independent main effects. These effects are more difficult to detect using traditional methodologies. In this manuscript the author introduces the concept of epistasis and the challenges associated with detecting it. Next, she briefly mentions a number of bioinformatics approaches that have been developed to deal with this issue. Multifactor dimensionality reduction is a methodology developed specifically to deal with the challenge of detecting interaction effects in the absence of statistically detectable main effects in studies of common disorders, such as Alzheimer disease or brain cancer. Finally, the author describes the future directions for this technique and related methodologies.

Computational Biology↗

MKMC enables reference-free transcriptomic analysis using k-mer representations.

Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals-including sex differences in killifish liver-and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex- and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.

MKMC↗

Information encoding and computation with spikes and bursts.

Neurons compute and communicate by transforming synaptic input patterns into output spike trains. The nature of this transformation depends crucially on the properties of voltage-gated conductances in neuronal membranes. These intrinsic membrane conductances can enable neurons to generate different spike patterns including brief, high-frequency bursts that are commonly observed in a variety of brain regions. Here we examine how the membrane conductances that generate bursts affect neural computation and encoding. We simulated a bursting neuron model driven by random current input signal and superposed noise. We consider two issues: the timing reliability of different spike patterns and the computation performed by the neuron. Statistical analysis of the simulated spike trains shows that the timing of bursts is much more precise than the timing of single spikes. Furthermore, the number of spikes per burst is highly robust to noise. Next we considered the computation performed by the neuron: how different features of the input current are mapped into specific output spike patterns. Dimensional reduction and statistical classification techniques were used to determine the stimulus features triggering different firing patterns. Our main result is that spikes, and bursts of different durations, code for different stimulus features, which can be quantified without a priori assumptions about those features. These findings lead us to propose that the biophysical mechanisms of spike generation enables individual neurons to encode different stimulus features into distinct spike patterns.

Computer Simulation↗

[Dorsal stabilization of thoracic and lumbar vertebral injuries].

The concept of angle-stable transpedicular screw-rod instrumentation, realized in the different models of internal spine fixators with intrinsic stability, allows secure stabilization of the most unstable fracture patterns, limited-segment fixation and three-dimensional reduction of the fragments. The canal diameter is improved by ligamentotaxis and, if necessary, by hemilaminectomy and fragment impaction. Late collapse of the upper disk space must be anticipated and may lead to some increase in kyphotic deformity. The situations are identified where an additional formal interbody fusion is recommended.

Fracture Fixation, Internal↗

Ideal discrimination of discrete clinical endpoints using multilocus genotypes.

Multifactor Dimensionality Reduction (MDR) is a method for the classification and prediction of discrete clinical endpoints using attributes constructed from multilocus genotype data. Empirical studies with both real and simulated data suggest that MDR has good power for detecting gene-gene interactions in the absence of independent main effects. The purpose of this study is to develop an objective, theory-driven approach to evaluate the strengths and limitations of MDR. To accomplish this goal, we borrow concepts from ideal observer analysis used in visual perception to evaluate the theoretical limits of classifying and predicting discrete clinical endpoints using multilocus genotype data. We conclude that MDR ideally discriminates between low risk and high risk subjects using attributes constructed from multilocus genotype data. We also how that the classification approach used once a multilocus attribute is constructed is similar to that of a naive Bayes classifier. This study provides a theoretical foundation for the continued development, evaluation, and application of the MDR as a data mining tool in the domain of statistical genetics and genetic epidemiology.

Animals↗

Dynamic monitoring system for full-scale wastewater treatment plants.

This paper proposes a new process monitoring method using dynamic independent component analysis (ICA), ICA is a recently developed technique to extract the hidden factors that underlie sets of measurements, whereas principal component analysis (PCA) is a dimensionality reduction technique in terms of capturing the variance of the data. Its goal is to find a linear representation of non-Gaussian data so that the components are statistically independent. PCA aims at finding PCs that are uncorrelated and are linear combinations of the observed variables, while ICA is designed to separate the ICs that are independent and constitute the observed variables. The dynamic ICA monitoring method is applying ICA to the augmenting matrix with time-lagged variables. The dynamic monitoring method was applied to detect and monitor disturbances in a full-scale biological wastewater treatment (WWTP), which is characterized by a variety of dynamic and non-Gaussian characteristics. The dynamic ICA method showed more powerful monitoring performance on a WWTP application than the dynamic PCA method since it can extract source signals which are independent of time and cross-correlation of variables.

Algorithms↗

Risk factor interactions and genetic effects associated with post-operative atrial fibrillation.

Postoperative Atrial Fibrillation (PoAF) is the most common arrhythmia after heart surgery, and continues to be a major cause of morbidity. Due to the complexity of this condition, many genes and/or environmental factors may play a role in susceptibility. Previous findings have shown several clinical and genetic risk factors for the development of PoAF. The goal of this study was to determine whether interactions among candidate genes and a variety of clinical factors are associated with PoAF. We applied the Multifactor Dimensionality Reduction (MDR) method to detect interactions in a sample of 940 adult subjects undergoing elective procedures of the heart or great vessels, requiring general anesthesia and sternotomy or thoracotomy, where 255 developed PoAF. We took a random sample of controls matched to the 255 AF cases for a total sample size of 510 individuals. MDR is a powerful statistical approach used to detect gene-gene or gene-environment interactions in the presence or absence of statistically detectable main effects in pharmacogenomics studies. We chose polymorphisms in three (IL-6, ACE, and ApoE) candidate genes, all previously implicated in PoAF risk, and a variety of environmental factors for analysis. We detected a single locus effect of IL-6 which is able to correctly predict disease status with 58.8% (p<0.001) accuracy. We also detected an interaction between history of AF and length of hospital stay that predicted disease status with 68.34% (p<0.001) accuracy. These findings demonstrate the utility of novel computational approaches for the detection of disease susceptibility genes. While each of these results looks interesting, they only explain part of PoAF susceptibility. It will be important to collect a larger set of candidate genes and environmental factors to better characterize the development of PoAF. Applying this approach, we were able to elucidate potential associations with postoperative atrial fibrillation.

Adult↗

[Reverberation of the hindlimb rudimentation on its innervation in squamate reptiles].

When the dimensional reduction of the hind limb begins, a first caudal displacement of the lombar part of the lombo-sacral plexus - which involves the loss of the first root of the sacral part -- appears with a threshold in the increase in the number of presacral vertebrae. This a first indication of the serpentiform tendancy. Others thresholds can conduct to produce the disappearance of the sacral vertebrae and sacral root. The qualitative reduction only concerns the terminal branches of the plexus and does not seem to be associated with the vertebral elongation. If a caudo-proximal reduction of the brachial plexus occurs early in the lepidosaurian line and exists in all the Squamata, even in the Iguana which have well developed limbs, it is not the same for the reduction of the lombo-sacral plexus which does not appear in these Iguana. At last, if the reduction modalities of the both plexus are often differents, their supposed displacements facilitate the extension of the intermediate vertebral region.

Animals↗