Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dimensionality Reduction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Density functional theory for general hard-core lattice gases.

We put forward a general procedure to obtain an approximate free-energy density functional for any hard-core lattice gas, regardless of the shape of the particles, the underlying lattice, or the dimension of the system. The procedure is conceptually very simple and recovers effortlessly previous results for some particular systems. Also, the obtained density functionals belong to the class of fundamental measure functionals and, therefore, are always consistent through dimensional reduction. We discuss possible extensions of this method to account for attractive lattice models.

Journal Article↗

Study of CP(N-1) theta-vacua by cluster simulation of SU(N) quantum spin ladders.

D-theory provides an alternative lattice regularization of the 2D CP(N-1) quantum field theory in which continuous classical fields emerge from the dimensional reduction of discrete SU(N) quantum spins. Spin ladders consisting of n transversely coupled spin chains lead to a CP(N-1) model with a vacuum angle theta=npi. In D-theory no sign problem arises and an efficient cluster algorithm is used to investigate theta-vacuum effects. At theta=pi there is a first order phase transition with spontaneous breaking of charge conjugation symmetry for CP(N-1) models with N>2.

Journal Article↗

Backreaction in acoustic black holes.

The backreaction equations for the linearized quantum fluctuations in an acoustic black hole are given. The solution near the horizon, obtained within a dimensional reduction, indicates that acoustic black holes, unlike Schwarzschild ones, get cooler as they radiate phonons. They show remarkable analogies with near-extremal Reissner-Nordström black holes.

Journal Article↗

The spectral dimension of the universe is scale dependent.

We measure the spectral dimension of universes emerging from nonperturbative quantum gravity, defined through state sums of causal triangulated geometries. While four dimensional on large scales, the quantum universe appears two dimensional at short distances. We conclude that quantum gravity may be "self-renormalizing" at the Planck scale, by virtue of a mechanism of dynamical dimensional reduction.

Journal Article↗

Probability-based current dipole localization from biomagnetic fields.

Focal biomagnetic sources are described as pointlike current dipoles. The dipole parameters, position, and moment coordinates are commonly determined from biomagnetic data using iterative nonlinear optimization algorithms such as the Levenberg-Marquardt algorithm. However, even for single-dipole sources, mislocalizations can occur due to side minima of the cost function or due to a wrong choice of the start vector. This can be shown by introducing a cost function where the independent variables are only the position coordinates instead of position and moment coordinates. This dimensional reduction--which is also possible for multiple dipole sources--is achieved by calculating the cost function at each position with the position and data-dependent, optimum dipole moments. We call these dipoles with--in a least squares sense--optimum moments, locally optimal dipoles. The visualization of such a single-dipole cost function and of the iteration steps of the Levenberg-Marquardt algorithm show why mislocalizations cannot be avoided. Therefore, we propose an alternative noniterative localization algorithm for single-dipole sources without this drawback. It uses localization probabilities calculated by means of the locally optimal dipoles. Besides the determination of the dipole parameters, the proposed algorithm furnishes a reliable error for each localization. Its effectiveness is shown with simulated and real patient data.

Algorithms↗

Automated design of robust discriminant analysis classifier for foot pressure lesions using kinematic data.

In the recent years, the use of motion tracking systems for acquisition of functional biomechanical gait data, has received increasing interest due to the richness and accuracy of the measured kinematic information. However, costs frequently restrict the number of subjects employed, and this makes the dimensionality of the collected data far higher than the available samples. This paper applies discriminant analysis algorithms to the classification of patients with different types of foot lesions, in order to establish an association between foot motion and lesion formation. With primary attention to small sample size situations, we compare different types of Bayesian classifiers and evaluate their performance with various dimensionality reduction techniques for feature extraction, as well as search methods for selection of raw kinematic variables. Finally, we propose a novel integrated method which fine-tunes the classifier parameters and selects the most relevant kinematic variables simultaneously. Performance comparisons are using robust resampling techniques such as Bootstrap 632+ and k-fold cross-validation. Results from experimentations with lesion subjects suffering from pathological plantar hyperkeratosis, show that the proposed method can lead to approximately 96% correct classification rates with less than 10% of the original features.

Adult↗

A hybrid SEM algorithm for high-dimensional unsupervised learning using a finite generalized Dirichlet mixture.

This paper applies a robust statistical scheme to the problem of unsupervised learning of high-dimensional data. We develop, analyze, and apply a new finite mixture model based on a generalization of the Dirichlet distribution. The generalized Dirichlet distribution has a more general covariance structure than the Dirichlet distribution and offers high flexibility and ease of use for the approximation of both symmetric and asymmetric distributions. We show that the mathematical properties of this distribution allow high-dimensional modeling without requiring dimensionality reduction and, thus, without a loss of information. This makes the generalized Dirichlet distribution more practical and useful. We propose a hybrid stochastic expectation maximization algorithm (HSEM) to estimate the parameters of the generalized Dirichlet mixture. The algorithm is called stochastic because it contains a step in which the data elements are assigned randomly to components in order to avoid convergence to a saddle point. The adjective "hybrid" is justified by the introduction of a Newton-Raphson step. Moreover, the HSEM algorithm autonomously selects the number of components by the introduction of an agglomerative term. The performance of our method is tested by the classification of several pattern-recognition data sets. The generalized Dirichlet mixture is also applied to the problems of image restoration, image object recognition and texture image database summarization for efficient retrieval. For the texture image summarization problem, results are reported for the Vistex texture image database from the MIT Media Lab.

Algorithms↗

A computer-aided diagnostic system to characterize CT focal liver lesions: design and optimization of a neural network classifier.

In this paper, a computer-aided diagnostic (CAD) system for the classification of hepatic lesions from computed tomography (CT) images is presented. Regions of interest (ROIs) taken from nonenhanced CT images of normal liver, hepatic cysts, hemangiomas, and hepatocellular carcinomas have been used as input to the system. The proposed system consists of two modules: the feature extraction and the classification modules. The feature extraction module calculates the average gray level and 48 texture characteristics, which are derived from the spatial gray-level co-occurrence matrices, obtained from the ROIs. The classifier module consists of three sequentially placed feed-forward neural networks (NNs). The first NN classifies into normal or pathological liver regions. The pathological liver regions are characterized by the second NN as cyst or "other disease." The third NN classifies "other disease" into hemangioma or hepatocellular carcinoma. Three feature selection techniques have been applied to each individual NN: the sequential forward selection, the sequential floating forward selection, and a genetic algorithm for feature selection. The comparative study of the above dimensionality reduction methods shows that genetic algorithms result in lower dimension feature vectors and improved classification performance.

Algorithms↗

Probabilistic independent component analysis for functional magnetic resonance imaging.

We present an integrated approach to probabilistic independent component analysis (ICA) for functional MRI (FMRI) data that allows for nonsquare mixing in the presence of Gaussian noise. In order to avoid overfitting, we employ objective estimation of the amount of Gaussian noise through Bayesian analysis of the true dimensionality of the data, i.e., the number of activation and non-Gaussian noise sources. This enables us to carry out probabilistic modeling and achieves an asymptotically unique decomposition of the data. It reduces problems of interpretation, as each final independent component is now much more likely to be due to only one physical or physiological process. We also describe other improvements to standard ICA, such as temporal prewhitening and variance normalization of timeseries, the latter being particularly useful in the context of dimensionality reduction when weak activation is present. We discuss the use of prior information about the spatiotemporal nature of the source processes, and an alternative-hypothesis testing approach for inference, using Gaussian mixture models. The performance of our approach is illustrated and evaluated on real and artificial FMRI data, and compared to the spatio-temporal accuracy of results obtained from classical ICA and GLM analyses.

Algorithms↗

Discriminative components of data.

A simple probabilistic model is introduced to generalize classical linear discriminant analysis (LDA) in finding components that are informative of or relevant for data classes. The components maximize the predictability of the class distribution which is asymptotically equivalent to 1) maximizing mutual information with the classes, and 2) finding principal components in the so-called learning or Fisher metrics. The Fisher metric measures only distances that are relevant to the classes, that is, distances that cause changes in the class distribution. The components have applications in data exploration, visualization, and dimensionality reduction. In empirical experiments, the method outperformed, in addition to more classical methods, a Renyi entropy-based alternative while having essentially equivalent computational cost.

Algorithms↗

Clustered blockwise PCA for representing visual data.

Principal Component Analysis (PCA) is extensively used in computer vision and image processing. Since it provides the optimal linear subspace in a least-square sense, it has been used for dimensionality reduction and subspace analysis in various domains. However, its scalability is very limited because of its inherent computational complexity. We introduce a new framework for applying PCA to visual data which takes advantage of the spatio-temporal correlation and localized frequency variations that are typically found in such data. Instead of applying PCA to the whole volume of data (complete set of images), we partition the volume into a set of blocks and apply PCA to each block. Then, we group the subspaces corresponding to the blocks and merge them together. As a result, we not only achieve greater efficiency in the resulting representation of the visual data, but also successfully scale PCA to handle large data sets. We present a thorough analysis of the computational complexity and storage benefits of our approach. We apply our algorithm to several types of videos. We show that, in addition to its storage and speed benefits, the algorithm results in a useful representation of the visual data.

Algorithms↗

Metabolic signalling in defence and stress: the central roles of soluble redox couples.

Plant growth and development are driven by electron transfer reactions. Modifications of redox components are both monitored and induced by cells, and are integral to responses to environmental change. Key redox compounds in the soluble phase of the cell are NAD, NADP, glutathione and ascorbate--all of which interact strongly with reactive oxygen. This review takes an integrated view of the NAD(P)-glutathione-ascorbate network. These compounds are considered not as one-dimensional 'reductants' or 'antioxidants' but as redox couples that can act together to condition cellular redox tone or that can act independently to transmit specific information that tunes signalling pathways. Emphasis is placed on recent developments highlighting the complexity of redox-dependent defence reactions, and the importance of interactions between the reduction state of soluble redox couples and their concentration in mediating dynamic signalling in response to stress. Signalling roles are assessed within the context of interactions with reactive oxygen, phytohormones and calcium, and the biochemical reactions through which redox couples could be sensed are discussed.

Adaptation, Physiological↗

Use of IR Biotyper as a feasible methodology to type Klebsiella pneumoniae.

UNLABELLED: Klebsiella pneumoniae is one of the most frequently reported healthcare-associated pathogens. The current gold standard approach to perform the epidemiological typing of these bacteria is Whole Genome Sequencing (WGS), which is an expensive and challenging procedure. IR Biotyper (Bruker Daltonics, GmbH) is a new equipment based on Fourier transform infrared spectroscopy, which allows a rapid, low-cost, and user-friendly method to type bacterial isolates. However, there is a need for studies that evaluate the efficacy of the IR Biotyper. The aim of this study was to evaluate the capability of IR Biotyper to type K. pneumoniae according to sequence type (ST) and capsular type-using K locus (KL)-as well as to develop a classifier using machine learning. Seventy-three isolates of K. pneumoniae previously characterized by WGS were selected for IR Biotyper analysis using principal component analysis for dimensionality reduction, Euclidean, and unweighted pair group method with arithmetic mean (UPGMA) for clustering method, and spectra were analyzed in the 1,300-800 cm⁻¹ wavenumber range. Among these, 54 isolates were used to create a classifier, and 19 were used to validate the classifier. When considering the ST, ST307 was grouped in the same cluster as ST11. When KL was considered for the analysis, the clusters were 100% correctly grouped according to their KL type. Furthermore, the classifier developed was able to classify the isolates according to KL with a high concordance. This study showed that KL correlates well with KL for typing K. pneumoniae isolates using the IR Biotyper. Additionally, IR Biotyper demonstrated to be a cost-effective method and a promising tool to classify isolates within minutes. IMPORTANCE: Klebsiella pneumoniae is a major cause of severe hospital infections, and controlling its spread requires quick identification and comparison of bacterial strains. WGS is accurate but expensive, slow, and technically demanding. In this study, we evaluated the IR Biotyper, a device that uses infrared light to analyze bacteria and group them by capsule type-a key feature linked to their spread. The IR Biotyper matched WGS results with high accuracy, delivering results in minutes instead of days. This fast, affordable method can help hospitals detect outbreaks earlier and respond more effectively. Our findings suggest that the IR Biotyper is a valuable tool for routine use in microbiology laboratories, supporting epidemiological surveillance and outbreak control.

Klebsiella pneumoniae↗

Applications of support vector machines to cancer classification with microarray data.

Microarray gene expression data usually have a large number of dimensions, e.g., over ten thousand genes, and a small number of samples, e.g., a few tens of patients. In this paper, we use the support vector machine (SVM) for cancer classification with microarray data. Dimensionality reduction methods, such as principal components analysis (PCA), class-separability measure, Fisher ratio, and t-test, are used for gene selection. A voting scheme is then employed to do multi-group classification by k(k - 1) binary SVMs. We are able to obtain the same classification accuracy but with much fewer features compared to other published results.

Algorithms↗

Application of multilevel models to morphometric data. Part 2. Correlations.

Multilevel organization of morphometric data (cells are "nested" within patients) requires special methods for studying correlations between karyometric features. The most distinct feature of these methods is that separate correlation (covariance) matrices are produced for every level in the hierarchy. In karyometric research, the cell-level (i.e., within-tumor) correlations seem to be of major interest. Beside their biological importance, these correlation coefficients (CC) are compulsory when dimensionality reduction is required. Using MLwiN, a dedicated program for multilevel modeling, we show how to use multivariate multilevel models (MMM) to obtain and interpret CC in each of the levels. A comparison with two usual, "single-level" statistics shows that MMM represent the only way to obtain correct cell-level correlation coefficients. The summary statistics method (take average values across each patient) produces patient-level CC only, and the "pooling" method (merge all cells together and ignore patients as units of analysis) yields incorrect CC at all. We conclude that multilevel modeling is an indispensable tool for studying correlations between morphometric variables.

Algorithms↗

The interaction of four genes in the inflammation pathway significantly predicts prostate cancer risk.

It is widely hypothesized that the interactions of multiple genes influence individual risk to prostate cancer. However, current efforts at identifying prostate cancer risk genes primarily rely on single-gene approaches. In an attempt to fill this gap, we carried out a study to explore the joint effect of multiple genes in the inflammation pathway on prostate cancer risk. We studied 20 genes in the Toll-like receptor signaling pathway as well as several cytokines. For each of these genes, we selected and genotyped haplotype-tagging single nucleotide polymorphisms (SNP) among 1,383 cases and 780 controls from the CAPS (CAncer Prostate in Sweden) study population. A total of 57 SNPs were included in the final analysis. A data mining method, multifactor dimensionality reduction, was used to explore the interaction effects of SNPs on prostate cancer risk. Interaction effects were assessed for all possible n SNP combinations, where n = 2, 3, or 4. For each n SNP combination, the model providing lowest prediction error among 100 cross-validations was chosen. The statistical significance levels of the best models in each n SNP combination were determined using permutation tests. A four-SNP interaction (one SNP each from IL-10, IL-1RN, TIRAP, and TLR5) had the lowest prediction error (43.28%, P = 0.019). Our ability to analyze a large number of SNPs in a large sample size is one of the first efforts in exploring the effect of high-order gene-gene interactions on prostate cancer risk, and this is an important contribution to this new and quickly evolving field.

Case-Control Studies↗

High-order interactions among genetic variants in DNA base excision repair pathway genes and smoking in bladder cancer susceptibility.

Cancer is a common multifactor human disease resulting from complex interactions between many genetic and environmental factors. In this study, we used a multifaceted analytic approach to explore the relationship between eight single nucleotide polymorphisms in base excision repair (BER) pathway genes, smoking, and bladder cancer susceptibility in a hospital-based case-control study. Overall, we did not find an association between any single BER gene single nucleotide polymorphism and bladder cancer risk. However, in stratified analysis, the OGG1 S326C variant genotypes in ever smokers (odds ratio, 0.74; 95% confidence interval, 0.56-0.99) and ADP-ribosyltransferase (ADPRT) V762A variant genotypes in never smokers (odds ratio, 0.58; 95% confidence interval, 0.37-0.91) conferred a significantly reduced risk. Using logistic regression, we observed that there was a two-way interaction between ADPRT V762A and smoking status. We next used classification and regression tree analysis to explore high-order gene-gene and gene-environment interactions. We found that smoking is the most important influential factor for bladder cancer risk. Consistent with the above findings, we found that the ADPRT V762A was only significantly involved in bladder cancer risk in never smokers and the OGG1 S326C was only significantly involved in ever smokers. We also observed gene-gene interactions among OGG1 S326C, XRCC1 R194W, and MUTYH H335Q in ever smokers. Using multifactor dimensionality reduction approach, the four-factor model, including smoking status, OGG1 S326C (rs1052133), APEX1 D148E (rs3136820), and ADPRT762 (rs1136410), had the best ability to predict bladder cancer risk with the highest cross-validation consistency (100%) and the lowest prediction error (37.02%; P < 0.001). These results support the hypothesis that genetic variants in BER genes contribute to bladder cancer risk through gene-gene and gene-environmental interactions.

Algorithms↗

The ubiquitous nature of epistasis in determining susceptibility to common human diseases.

There is increasing awareness that epistasis or gene-gene interaction plays a role in susceptibility to common human diseases. In this paper, we formulate a working hypothesis that epistasis is a ubiquitous component of the genetic architecture of common human diseases and that complex interactions are more important than the independent main effects of any one susceptibility gene. This working hypothesis is based on several bodies of evidence. First, the idea that epistasis is important is not new. In fact, the recognition that deviations from Mendelian ratios are due to interactions between genes has been around for nearly 100 years. Second, the ubiquity of biomolecular interactions in gene regulation and biochemical and metabolic systems suggest that relationship between DNA sequence variations and clinical endpoints is likely to involve gene-gene interactions. Third, positive results from studies of single polymorphisms typically do not replicate across independent samples. This is true for both linkage and association studies. Fourth, gene-gene interactions are commonly found when properly investigated. We review each of these points and then review an analytical strategy called multifactor dimensionality reduction for detecting epistasis. We end with ideas of how hypotheses about biological epistasis can be generated from statistical evidence using biochemical systems models. If this working hypothesis is true, it suggests that we need a research strategy for identifying common disease susceptibility genes that embraces, rather than ignores, the complexity of the genotype to phenotype relationship.

Disease↗