Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dimensionality Reduction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Complex gene-gene interactions in multiple sclerosis: a multifactorial approach reveals associations with inflammatory genes.

The complex inheritance involved in multiple sclerosis (MS) risk has been extensively investigated, but our understanding of MS genetics remains rudimentary. In this study, we explore 51 single nucleotide polymorphisms (SNPs) in 36 candidate genes from the inflammatory pathway and test for gene-gene interactions using complementary case-control, discordant sibling pair, and trio family study designs. We used a sample of 421 carefully diagnosed MS cases and 96 unrelated, healthy controls; discordant sibling pairs from 146 multiplex families; and 275 trio families. We used multifactor dimensionality reduction to explore gene-gene interactions. Based on our analyses, we have identified several statistically significant models including both main effect models and two-locus, three-locus, and four-locus epistasis models that predict MS disease risk with between approximately 61% and 85% accuracy. These results suggest that significant epistasis, or gene-gene interactions, may exist even in the absence of statistically significant individual main effects.

Case-Control Studies↗

Multidimensional analysis of the concentrations of 17 substances in the CSF of schizophrenics and controls.

The concentrations of 17 substances were determined in the cerebrospinal fluid (CSF) of 28 paranoid schizophrenic patients and 16 controls. Results were standardized and simultaneously evaluated through Multidimensional Scaling (MDS). The full data set can be considered as a cloud of points consisting of the 44 subjects in the 17-dimensional parameter space. MDS seeks a two-dimensional representation of this 17-dimensional cloud of points, while retaining as much as possible the distances between the subjects. The two-dimensional reduction of the 17 CSF parameters correctly separated 15 of 16 controls from the schizophrenic subjects. This indicates that a biological heterogeneity between schizophrenic and nonschizophrenic subjects can be detected by the simultaneous analysis of the CSF concentrations of substances related directly or indirectly to the neuronal activity in the brain.

Adult↗

Multilocus interactions at maternal tumor necrosis factor-alpha, tumor necrosis factor receptors, interleukin-6 and interleukin-6 receptor genes predict spontaneous preterm labor in European-American women.

OBJECTIVE: We hypothesize that genetic variations (single nucleotide polymorphisms-SNPs) in the tumor necrosis factor-alpha (TNF-alpha), TNF receptors (TNFRI and TNFRII), interleukin-6 (IL-6) and IL-6 receptor (IL-6R) genes predict high-risk status for spontaneous preterm birth (sPTB) in European-American women. In this study we examine the allelic and genotypic variations and the gene-gene interactions in the TNF-alpha, TNFRs, IL-6, and IL-6R genes in maternal DNA samples by using a case-control model. STUDY DESIGN: Maternal DNA from cases of sPTB after preterm labor (n = 101) and controls (normal term labor and delivery) (n = 321) were genotyped for SNPs in the TNF-alpha (6), TNFRI (6), TNFRII (7), IL-6 (5), and IL-6R (3) loci. SNPs were tested for both allele and genotype differences (cases vs controls) with the use of standard genetic epidemiologic methods. Multilocus interaction was assessed with multifactor dimensionality reduction analysis (MDR) to test all single and multilocus combinations for the ability to predict sPTB. RESULTS: Few significant allelic and genotypic associations were detected between cases and controls in maternal DNA. Single locus analysis documented independent association of SNPs at -7294 (allele and genotype) of TNFRI and 24660 (genotype) TNFRII loci with sPTB. MDR revealed a significant 3 locus model that includes SNPs -3448 of TNF-alpha, -7227 of IL-6, and 33314 of IL-6R. This interactive model allowed the successful prediction of pre- to low-risk genotypes is 3.50 (95% CI 2.52-4.87, P < .001). CONCLUSION: This is the first report to document a multilocus interaction in sPTB that predicts 65.2% of the cases in a European-American sample. Although putatively significant associations with sPTB were seen at a few single locus sites in TNFRI and TNFRII, they were not as predictive as the 3-locus model produced by MDR, suggesting the use of multilocus analyses in gene association studies of complex disease such as sPTB.

Adult↗

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans↗

Brain tumor classification based on long echo proton MRS signals.

There has been a growing research interest in brain tumor classification based on proton magnetic resonance spectroscopy (1H MRS) signals. Four research centers within the EU funded INTERPRET project have acquired a significant number of long echo 1H MRS signals for brain tumor classification. In this paper, we present an objective comparison of several classification techniques applied to the discrimination of four types of brain tumors: meningiomas, glioblastomas, astrocytomas grade II and metastases. Linear and non-linear classifiers are compared: linear discriminant analysis (LDA), support vector machines (SVM) and least squares SVM (LS-SVM) with a linear kernel as linear techniques and LS-SVM with a radial basis function (RBF) kernel as a non-linear technique. Kernel-based methods can perform well in processing high dimensional data. This motivates the inclusion of SVM and LS-SVM in this study. The analysis includes optimal input variable selection, (hyper-) parameter estimation, followed by performance evaluation. The classification performance is evaluated over 200 stratified random samplings of the dataset into training and test sets. Receiver operating characteristic (ROC) curve analysis measures the performance of binary classification, while for multiclass classification, we consider the accuracy as performance measure. Based on the complete magnitude spectra, automated binary classifiers are able to reach an area under the ROC curve (AUC) of more than 0.9 except for the hard case glioblastomas versus metastases. Although, based on the available long echo 1H MRS data, we did not find any statistically significant difference between the performances of LDA and the kernel-based methods, the latter have the strength that no dimensionality reduction is required to obtain such a high performance.

Artificial Intelligence↗

Case-based retrieval to support the treatment of end stage renal failure patients.

OBJECTIVE: In the present paper, we describe an application of case-based retrieval to the domain of end stage renal failure patients, treated with hemodialysis. MATERIALS AND METHODS: Defining a dialysis session as a case, retrieval of past similar cases has to operate both on static and on dynamic features, since most of the monitoring variables of a dialysis session are time series. Retrieval is then articulated as a two-step procedure: (1) classification, based on static features and (2) intra-class retrieval, in which dynamic features are considered. As regards step (2), we concentrate on a classical dimensionality reduction technique for time series allowing for efficient indexing, namely discrete Fourier transform (DFT). Thanks to specific index structures (i.e. k -d trees), range queries (on local feature similarity) can be efficiently performed on our case base, allowing the physician to examine the most similar stored dialysis sessions with respect to the current one. RESULTS: The retrieval tool has been positively tested on real patients' data, coming from the nephrology and dialysis unit of the Vigevano hospital, in Italy. CONCLUSIONS: The overall system can be seen as a means for supporting quality assessment of the hemodialysis service, providing a useful input from the knowledge management perspective.

Decision Support Systems, Clinical↗

Renin-angiotensin system gene polymorphisms and coronary artery disease in a large angiographic cohort: detection of high order gene-gene interaction.

There have been many reports regarding the association between renin-angiotensin system (RAS) gene polymorphisms and coronary artery disease (CAD) or acute myocardial infarction (AMI), but the results are inconsistent. In the present study, we used several new approaches with multilocus data to reappraise this issue in a large and relatively homogeneous Taiwanese population. A total of 1254 consecutive patients who underwent cardiac catheterization (735 with documented coronary artery disease and 519 without) between 1996 and 2003 were recruited. Angiotensin-converting enzyme gene insertion/deletion (I/D) polymorphism; T174M, M235T, G-6A, A-20C, G-152A and G-217A polymorphisms of the angiotensinogen gene; and A1166C polymorphism of the angiotensin II type I receptor gene were genotyped. In single-locus analyses, no locus was associated with CAD, history of AMI and three-vessel CAD, either with or without adjustment for conventional CAD risk factors. For multilocus analyses, we recreated a balanced population, with the controls individually matched to the cases regarding the conventional CAD risk factors. We found that the angiotensinogen gene haplotype profile was significantly different between the cases and controls (chi2=31.6, P=0.030) in haplotype analyses. Furthermore, significant three-locus (G-217A, M235T and I/D) gene-gene interactions were detected by multifactor-dimensionality reduction method (highest cross-validation consistency 10.0, lowest prediction error 40.56%, P=0.017) and many even higher order gene-gene interactions by multilocus genotype disequilibrium tests (16 genotype disequilibria exclusively found in the controls, all of which included at least two genes among AGT, ACE and AT1R genes). Our study is the first to demonstrate epistatic, high-order, gene-gene interactions between RAS gene polymorphisms and CAD. These results are compatible with the concept of multilocus and multi-gene effects in complex diseases that would be missed with conventional approaches.

Aged↗

Gene-gene interactions for asthma and plasma total IgE concentration in Chinese children.

BACKGROUND: Asthma is a complex disease resulting from interactions between multiple genes and environmental factors. Study of gene-gene interactions could provide insight into asthma pathophysiology. OBJECTIVE: We investigated the interaction among 12 different loci in 8 candidate genes and asthma and increased plasma total IgE concentrations in 240 Chinese asthmatic subjects and 140 control subjects. METHODS: Genotyping was performed by means of RFLP analysis. Multifactor dimensionality reduction and logistic regression were used to analyze gene-gene interactions. RESULTS: A significant interaction was found between R130Q in the IL-13 gene (IL13) and I50V in the IL-4 receptor alpha gene (IL4RA) on the risk of asthma, with a cross-validation consistency of 10 of 10 and a prediction error of 33.7% (P = .014). The odds ratio of the high-risk to low-risk group was 2.6 (95% CI, 1.4-5.0; P = .004). For increased plasma total IgE concentration, the best 2-locus model consisted of R130Q in IL13 and C-431T in the thymus and activation-regulated chemokine gene (TARC). This model showed a maximum cross-validation consistency of 10 and a minimum prediction error of 36.1% (P = .022). The odds ratio of the high-risk to low-risk group was 3.9 (95% CI, 2.0-7.7; P = .0001). Logistic regression revealed significant interactions between IL13 and IL4RA for asthma (P = .042) and IL13 and TARC for increased total IgE concentration (P = .012). CONCLUSIONS: Our data suggest significant interactions between IL13 and IL4RA for asthma and IL13 and TARC for increased plasma total IgE concentrations in Chinese children.

Adolescent↗

Filter learning: application to suppression of bony structures from chest radiographs.

A novel framework for image filtering based on regression is presented. Regression is a supervised technique from pattern recognition theory in which a mapping from a number of input variables (features) to a continuous output variable is learned from a set of examples from which both input and output are known. We apply regression on a pixel level. A new, substantially different, image is estimated from an input image by computing a number of filtered input images (feature images) and mapping these to the desired output for every pixel in the image. The essential difference between conventional image filters and the proposed regression filter is that the latter filter is learned from training data. The total scheme consists of preprocessing, feature computation, feature extraction by a novel dimensionality reduction scheme designed specifically for regression, regression by k-nearest neighbor averaging, and (optionally) iterative application of the algorithm. The framework is applied to estimate the bone and soft-tissue components from standard frontal chest radiographs. As training material, radiographs with known soft-tissue and bone components, obtained by dual energy imaging, are used. The results show that good correlation with the true soft-tissue images can be obtained and that the scheme can be applied to images from a different source with good results. We show that bone structures are effectively enhanced and suppressed and that in most soft-tissue images local contrast of ribs decreases more than contrast between pulmonary nodules and their surrounding, making them relatively more pronounced.

Absorptiometry, Photon↗

Genetic variation in the mitochondrial enzyme carbamyl-phosphate synthetase I predisposes children to increased pulmonary artery pressure following surgical repair of congenital heart defects: a validated genetic association study.

Increased pulmonary artery pressure (PAP) can complicate the postoperative care of children undergoing surgical repair of congenital heart defects. Endogenous NO regulates PAP and is derived from arginine supplied by the urea cycle. The rate-limiting step in the urea cycle is catalyzed by a mitochondrial enzyme, carbamoyl-phosphate synthetase I (CPSI). A well-characterized polymorphism in the gene encoding CPSI (T1405N) has previously been implicated in neonatal pulmonary hypertension. A consecutive modeling cohort of children (N=131) with congenital heart defects requiring surgery was prospectively evaluated to determine key factors associated with increased postoperative PAP, defined as a mean PAP>20 mmHg for at least 1h during the 48h following surgery measured by an indwelling pulmonary artery catheter. Multiple dimensionality reduction (MDR) was used to both internally validate observations and develop optimal two-variable through five-variable models that were tested prospectively in a validation cohort (N=41). Unconditional logistic regression analysis of the modeling cohort revealed that age (OR=0.92, p=0.01), CPSI T1405N genotype (AC vs. AA: OR=4.08, p=0.04, CC vs. AA: OR=5.96, p=0.01), and Down syndrome (OR=5.25, p=0.04) were independent predictors of this complex phenotype. MDR predicted that the best two-variable model consisted of age and CPSI T1405N genotype (p<0.001). This two-variable model correctly predicted 73% of the outcomes from the validation cohort. A five-variable model that added race, gender and Down's syndrome was not significantly better than the two-variable model. In conclusion, the CPSI T1405N genotype appears to be an important new factor in predicting susceptibility to increased PAP following surgical repair of congenital cardiac defects in children.

Carbamoyl-Phosphate Synthase (Ammonia)↗

Methylome Profiling of Cartilage Tumors: A Promising New Diagnostic Tool?

DNA methylation and copy number variation (CNV) profiling has emerged as a promising tool for the classification of bone and soft tissue tumors. We evaluated its utility in cartilage tumors, where distinguishing low-grade from high-grade conventional central chondrosarcomas (CSs) and atypical cartilaginous tumors (ACTs) from enchondromas (ECs) is a frequent diagnostic challenge, particularly on biopsy material. We analyzed 214 chondrogenic tumors, including ECs, ACTs, conventional CSs, dedifferentiated chondrosarcomas (DDCSs), and clear cell CSs, and determined their IDH1/2 mutation status. Unsupervised dimensionality reduction of genome-wide DNA methylation patterns revealed 4 clusters among IDH-mutant (MUT) tumors (IDH-MUT-1: mostly ECs and ACTs and some high-grade CSs; IDH-MUT-2: predominantly high-grade CSs; IDH-MUT-3: largely DDCSs; and IDH-MUT-SB: distinct skull base group with a markedly different methylation pattern) and 2 clusters among IDH-wild-type (WT) tumors (IDH-WT-1 and IDH-WT-2: both primarily high-grade CSs, with IDH-WT-2 showing higher tumor grade and more extensive CNVs). Clear cell CSs formed a separate cluster. The amount of CNVs, including loss of CDKN2A, increased with tumor grade, reflecting increased genomic instability during chondrosarcoma progression. Supervised classifiers trained separately, both on methylation and CNV data, and distinguished low-grade and high-grade cartilaginous tumors with area under the curve values of 0.87 to 0.97 and 85% to 90% accuracy. Furthermore, we tested whether DDCSs can be distinguished from metastatic carcinomas and other high-grade sarcomas of the bone. Across 246 reference samples, a supervised classifier achieved 97.2% accuracy (area under the curve, 99.8%) and correctly identified 30 of 32 DDCSs (93.8%). These results indicate that DNA methylation and CNV data analysis provide a valuable tool for distinguishing most low- and high-grade CSs, with additional utility also in differentiating DDCS from morphologic mimics.

cartilaginous tumors↗

Computer-aided diagnosis of carotid atherosclerosis based on ultrasound image statistics, laws' texture and neural networks.

Quantitative characterisation of carotid atherosclerosis and classification into symptomatic or asymptomatic is crucial in planning optimal treatment of atheromatous plaque. The computer-aided diagnosis (CAD) system described in this paper can analyse ultrasound (US) images of carotid artery and classify them into symptomatic or asymptomatic based on their echogenicity characteristics. The CAD system consists of three modules: a) the feature extraction module, where first-order statistical (FOS) features and Laws' texture energy can be estimated, b) the dimensionality reduction module, where the number of features can be reduced using analysis of variance (ANOVA), and c) the classifier module consisting of a neural network (NN) trained by a novel hybrid method based on genetic algorithms (GAs) along with the back propagation algorithm. The hybrid method is able to select the most robust features, to adjust automatically the NN architecture and to optimise the classification performance. The performance is measured by the accuracy, sensitivity, specificity and the area under the receiver-operating characteristic (ROC) curve. The CAD design and development is based on images from 54 symptomatic and 54 asymptomatic plaques. This study demonstrates the ability of a CAD system based on US image analysis and a hybrid trained NN to identify atheromatous plaques at high risk of stroke.

Algorithms↗

A robust hybrid between genetic algorithm and support vector machine for extracting an optimal feature gene subset.

Development of a robust and efficient approach for extracting useful information from microarray data continues to be a significant and challenging task. Microarray data are characterized by a high dimension, high signal-to-noise ratio, and high correlations between genes, but with a relatively small sample size. Current methods for dimensional reduction can further be improved for the scenario of the presence of a single (or a few) high influential gene(s) in which its effect in the feature subset would prohibit inclusion of other important genes. We have formalized a robust gene selection approach based on a hybrid between genetic algorithm and support vector machine. The major goal of this hybridization was to exploit fully their respective merits (e.g., robustness to the size of solution space and capability of handling a very large dimension of feature genes) for identification of key feature genes (or molecular signatures) for a complex biological phenotype. We have applied the approach to the microarray data of diffuse large B cell lymphoma to demonstrate its behaviors and properties for mining the high-dimension data of genome-wide gene expression profiles. The resulting classifier(s) (the optimal gene subset(s)) has achieved the highest accuracy (99%) for prediction of independent microarray samples in comparisons with marginal filters and a hybrid between genetic algorithm and K nearest neighbors.

Algorithms↗

XRCC1 R399Q polymorphism is associated with response to platinum-based neoadjuvant chemotherapy in bulky cervical cancer.

OBJECTIVES: The aim of the study was to assess whether the genetic polymorphisms were associated with the tumor response in patients treated with platinum-based neoadjuvant chemotherapy (NAC) for bulky cervical cancer. METHODS: We retrospectively reviewed the clinical data and recruited paraffin-embedded, formalin-fixed tissues of 36 patients with bulky cervical carcinoma. All patients underwent two or three cycles of platinum-based NAC followed by radical hysterectomy. We determined the genotypes of each single nucleotide polymorphism (SNP) of ERCC1, ERCC2, GGH, GSTP1, MTHFR, SLC19A1 and XRCC1 using single base primer extension assay. RESULTS: The response to platinum-based NAC was higher in patients with SNP of XRCC1 R399Q (P=0.015), and there was a significant increased chance of treatment response in women with the G/G genotype (OR 38.0; 95% CI 1.66-870.45; P=0.023). The probability of response was also higher in patients with SNP of SLC19A1 6318C/T (P=0.032). There were dose-dependent influence of the number of alleles on the response to platinum-based chemotherapy (chi2 test for linear-by-linear association; P=0.009 for XRCC1 R399Q and P=0.017 for SLC19A1 6318C/T, respectively). Moreover, the multifactor dimensionality reduction (MDR) analysis documented that the combinations of XRCC1 R399Q and GGH-401C/T genetic polymorphisms were significantly associated with response to chemotherapy (P<0.0001). CONCLUSIONS: Genetic polymorphism of XRCC1 R399Q is associated with response to platinum-based NAC in bulky cervical cancer, and MDR analysis documented association between gene-gene interaction of XRCC1 R399Q and treatment response.

Antineoplastic Agents↗

Genetically constrained metabolic flux analysis.

Significant progress has been made in using existing metabolic databases to estimate metabolic fluxes. Traditional metabolic flux analysis generally starts with a predetermined metabolic network. This approach has been employed successfully to analyze the behaviors of recombinant strains by manually adding or removing the corresponding pathway(s) in the metabolic map. The current work focuses on the development of a new framework that utilizes genomic and metabolic databases, including available genetic/regulatory network structures and gene chip expression data, to constrain metabolic flux analysis. The genetic network consisting of the sensing/regulatory circuits will activate or deactivate a specific set of genes in response to external stimulus. The activation and/or repression of this set of genes will result in different gene expression levels that will in turn change the structure of the metabolic map. Hence, the metabolic map will automatically "adapt" to the external stimulus as captured by the genetic network. This adaptation selects a subnetwork from the pool of feasible reactions and so performs what we term "environmentally driven dimensional reduction." The Escherichia coli oxygen and redox sensing/regulatory system, which controls the metabolic patterns connected to glycolysis and the TCA cycle, was used as a model system to illustrate the proposed approach.

Citric Acid Cycle↗

Dynamics of influenza A drift: the linear three-strain model.

We analyze an epidemiological model consisting of a linear chain of three cocirculating influenza A strains that provide hosts exposed to a given strain with partial immune cross-protection against other strains. In the extreme case where infection with the middle strain prevents further infections from the other two strains, we reduce the model to a six-dimensional kernel capable of showing self-sustaining oscillations at relatively high levels of cross-protection. Dimensional reduction has been accomplished by a transformation of variables that preserves the eigenvalue responsible for the transition from damped oscillations to limit cycle solutions.

Antigenic Variation↗

Picture recognition in animals and humans.

The question of object-picture recognition has received relatively little attention in both human and comparative psychology; a paradoxical situation given the important use of image technology (e. g. slides, digitised pictures) made by neuroscientists in their experimental investigation of visual cognition. The present review examines the relevant literature pertaining to the question of the correspondence between and/or equivalence of real objects and their pictorial representations in animals and humans. Two classes of reactions towards pictures will be considered in turn: acquired responses in picture recognition experiments and spontaneous responses to pictures of biologically relevant objects (e.g. prey or conspecifics). Our survey will lead to the conclusion that humans show evidence of picture recognition from an early age; this recognition is, however, facilitated by prior exposure to pictures. This same exposure or training effect appears also to be necessary in nonhuman primates as well as in other mammals and in birds. Other factors are also identified as playing a role in the acquired responses to pictures: familiarity with and nature of the stimulus objects, presence of motion in the image, etc. Spontaneous and adapted reactions to pictures are a wide phenomenon present in different phyla including invertebrates but in most instances, this phenomenon is more likely to express confusion between objects and pictures than discrimination and active correspondence between the two. Finally, given the nature of a picture (e.g. bi-dimensionality, reduction of cues related to depth), it is suggested that object-picture recognition be envisioned in various levels, with true equivalence being a limited case, rarely observed in the behaviour of animals and even humans.

Animals↗

Application of pattern recognition and feature extraction techniques to volatile constituent metabolic profiles obtained by capillary gas chromatography.

The applicability of threshold logic units, a form of nonparametric pattern recognition, to the processing of metabolic profile data obtained by high-efficiency glass capillary column gas chromatography has been investigated. The test data included profiles of the volatile constituents of urine from normal individuals and from individuals with diabetes mellitus. A feature extraction algorithm allowed for dimensionality reduction and indicated the constituents most important in the normal versus pathological distinction. With an optimum number of dimensions, a normal versus pathological prediction rate of 93.75% was achieved. Gas chromatography-mass spectrometry was utilized to identify important profile constituents.

Chromatography, Gas↗