Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dimensionality Reduction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Efficient and robust feature extraction by maximum margin criterion.

In pattern recognition, feature extraction techniques are widely employed to reduce the dimensionality of data and to enhance the discriminatory information. Principal component analysis (PCA) and linear discriminant analysis (LDA) are the two most popular linear dimensionality reduction methods. However, PCA is not very effective for the extraction of the most discriminant features, and LDA is not stable due to the small sample size problem. In this paper, we propose some new (linear and nonlinear) feature extractors based on maximum margin criterion (MMC). Geometrically, feature extractors based on MMC maximize the (average) margin between classes after dimensionality reduction. It is shown that MMC can represent class separability better than PCA. As a connection to LDA, we may also derive LDA from MMC by incorporating some constraints. By using some other constraints, we establish a new linear feature extractor that does not suffer from the small sample size problem, which is known to cause serious stability problems for LDA. The kernelized (nonlinear) counterpart of this linear feature extractor is also established in the paper. Our extensive experiments demonstrate that the new feature extractors are effective, stable, and efficient.

Algorithms↗

Predictive neural networks for gene expression data analysis.

Gene expression data generated by DNA microarray experiments have provided a vast resource for medical diagnosis and disease understanding. Most prior work in analyzing gene expression data, however, focuses on predictive performance but not so much on deriving human understandable knowledge. This paper presents a systematic approach for learning and extracting rule-based knowledge from gene expression data. A class of predictive self-organizing networks known as Adaptive Resonance Associative Map (ARAM) is used for modelling gene expression data, whose learned knowledge can be transformed into a set of symbolic IF-THEN rules for interpretation. For dimensionality reduction, we illustrate how the system can work with a variety of feature selection methods. Benchmark experiments conducted on two gene expression data sets from acute leukemia and colon tumor patients show that the proposed system consistently produces predictive performance comparable, if not superior, to all previously published results. More importantly, very simple rules can be discovered that have extremely high diagnostic power. The proposed methodology, consisting of dimensionality reduction, predictive modelling, and rule extraction, provides a promising approach to gene expression analysis and disease understanding.

Animals↗

Beta2-adrenergic receptor and UCP3 variants modulate the relationship between age and type 2 diabetes mellitus.

BACKGROUND: It is widely accepted that Type 2 Diabetes Mellitus (T2DM) and other complex diseases are the product of complex interplay between genetic susceptibility and environmental causes. To cope with such a complexity, all the statistical and conceptual strategies available should be used. The working hypothesis of this study was that two well-known T2DM risk factors could have diverse effect in individuals carrying different genotypes. In particular, our effort was to investigate if a well-defined group of genes, involved in peripheral energy expenditure, could modify the impact of two environmental factors like age and obesity on the risk to develop diabetes. To achieve this aim we exploited a multianalytical approach also using dimensionality reduction strategy and conservative significance correction strategies. METHODS: We collected clinical data and characterised five genetic variants and 2 environmental factors of 342 ambulatory T2DM patients and 305 unrelated non-diabetic controls. To take in account the role of one of the major co-morbidity conditions we stratified the whole sample according to the presence of obesity, over and above the 30 Kg/m2 BMI threshold. RESULTS: By monofactorial analyses the ADRB2-27 Glu27 homozygotes had a lower frequency of diabetes when compared with Gln27 carriers (Odds Ratio (OR) 0.56, 95% Confidence Interval (CI) 0.36 - 0.91). This difference was even more marked in the obese subsample. Multifactor Dimensionality Reduction method in the non-obese subsample showed an interaction among age, ADRB2-16 and UCP3 polymorphisms. In individuals that were UCP3 T-carriers and ADRB2-16 Arg-carriers the OR increased from 1 in the youngest to 10.84 (95% CI 4.54-25.85) in the oldest. On the contrary, in the ADRB2-16 GlyGly and UCP3 CC double homozygote subjects, the OR for the disease was 1.10 (95% CI 0.53-2.27) in the youngest and 1.61 (95% CI 0.55-4.71) in the oldest. CONCLUSION: Although our results should be confirmed by further studies, our data suggests that, when properly evaluated, it is possible to identify genetic factors that could influence the effect of common risk factors.

Adult↗

Polytypism in columnar group 14 halide salts: structures of (Et2NH2)3Pb3X9 x nH2O (X = Cl, Br) and (beta-alaninium)2SnI4.

The crystal structures of three hybrid organoammonium metal halide salts composed of edge-sharing MX(6) octahedra have been determined. The genesis of these structures can be traced to the parent hexagonal MX(2) structure via dimensional reduction and recombination arguments. The structures of (Et(2)NH(2))(3)Pb(3)X(9) x nH(2)O (X = Br, I) contain unique columnar (Pb(3)X(9))(n)(3)(n)(-) structures, built up of edge-shared PbX(6) octahedra. The interaction of the Et(2)NH(2)(+) cations with the parent PbX(2) structures leads to a rearrangement of the lattice into the observed columnar structure. Groups of six Et(2)NH(2)(+) cations are hydrogen bonded to these columns, girdling them at their narrowest points. These hydrogen bonds contribute to the formation of the zigzag nature of the columnar inorganic framework. The resultant structures are recombinate analogues (polytypes) of the (Pb(3)X(9))(n)(3)(n)(-) stacks that would be obtained by the dimensional reduction process of the parent layer PbX(2) structure into simple edge-shared ribbons of PbX(6) octahedra. These structures can be described in terms of the stacking of planar bibridged Pb(3)X(8)(2-) units decorated with a single halide ion at a terminal lead ion site. In a similar fashion, (beta-alaH)(2)Sn(2)I(6) contains corrugated (Sn(2)I(6))(n)(2)(n)(-) columns (beta-ala = beta-alanine), with the cations sitting in the clefts of the columns.

Journal Article↗

Mandibular skeletal dysmorphology in micrognathic mice.

The primary manifestations of micrognathia were microglossia, midline fusion of the right and left sides of the mandible, total absence of incisor and molar toothbuds and, in many cases, absence or perhaps premature resorption of Meckel's cartilage. In addition, there was altered osteogenesis as evidenced by disrupted trabecular patterns, as well as an overall dimensional reduction of the mandible both antero-posteriorly and laterally. Strikingly similar results were reported by Johnson (1926), who studied the progeny of x-irradiated mice. How specifically our results correlate with this much earlier work is a matter for further analysis. It seems clear that the critical factor in the development of micrognathia is not so much an abnormal formation of the bony mandible, but a deficiency of tongue development, specifically its intrinsic musculature. Thus, mandibular micrognathia involves not only a dysmorphology of the first branchial arch, but also the mesenchymal cell migration from the occipital somites. Taken together, the picture is one that suggests an underlying cause that may have its inception at a much earlier developmental stage, when ectomesenchymal migration from the region of the neural tube occurs. In any event, we can report confidently that spontaneous micrognathia in prenatal mice is not a simple dimensional reduction of the lower jaw, but a more complex morphological phenomenon.

Animals↗

Physiological studies of information processing in the normal and Parkinsonian basal ganglia: pallidal activity in Go/No-Go task and following MPTP treatment.

Understanding the role of the basal ganglia in day to day behavior is critical for a better understanding of the role of these structures in pathological states--such as Parkinson's disease. To elucidate this connection, we studied pallidal activity in a monkey performing a delayed release Go/No-Go task and in monkeys treated with the dopaminergic neurotoxin--MPTP. We compared the results with the predictions of the action selection and reinforcement driven dimensionality reduction models of the basal ganglia. The fraction of responding pallidal neurons, as well as the ratio of positive to negative responses, were equal in the Go and the No-Go modes. The fraction of pallidal neurons with significant responses following the trigger signal (19/26) was higher than that following the visual cue (11/26); however, the fraction of negative responses was significantly higher following the cue signal (47%) than that following the trigger signal (22%). Most (80%) of the cue responses were sensitive to the laterality of the cue, whereas only 25% of the responses following the trigger signal were sensitive to the cue or movement direction. Finally, pallidal spiking activity was not correlated in the normal behaving monkey, and became highly synchronized following MPTP treatment. We conclude that pallidal activity in the normal monkey is consistent with the model of action selection, assuming that action is selected following the visual cue. However, the reinforcement driven dimensionality reduction model is consistent with both the Go/No-Go responses and the normal/MPTP correlation studies.

1-Methyl-4-phenyl-1,2,3,6-tetrahydropyridine↗

Filter versus wrapper gene selection approaches in DNA microarray domains.

DNA microarray experiments generating thousands of gene expression measurements, are used to collect information from tissue and cell samples regarding gene expression differences that could be useful for diagnosis disease, distinction of the specific tumor type, etc. One important application of gene expression microarray data is the classification of samples into known categories. As DNA microarray technology measures the gene expression en masse, this has resulted in data with the number of features (genes) far exceeding the number of samples. As the predictive accuracy of supervised classifiers that try to discriminate between the classes of the problem decays with the existence of irrelevant and redundant features, the necessity of a dimensionality reduction process is essential. We propose the application of a gene selection process, which also enables the biology researcher to focus on promising gene candidates that actively contribute to classification in these large scale microarrays. Two basic approaches for feature selection appear in machine learning and pattern recognition literature: the filter and wrapper techniques. Filter procedures are used in most of the works in the area of DNA microarrays. In this work, a comparison between a group of different filter metrics and a wrapper sequential search procedure is carried out. The comparison is performed in two well-known DNA microarray datasets by the use of four classic supervised classifiers. The study is carried out over the original-continuous and three-intervals discretized gene expression data. While two well-known filter metrics are proposed for continuous data, four classic filter measures are used over discretized data. The same wrapper approach is used for both continuous and discretized data. The application of filter and wrapper gene selection procedures leads to considerably better accuracy results in comparison to the non-gene selection approach, coupled with interesting and notable dimensionality reductions. Although the wrapper approach mainly shows a more accurate behavior than filter metrics, this improvement is coupled with considerable computer-load necessities. We note that most of the genes selected by proposed filter and wrapper procedures in discrete and continuous microarray data appear in the lists of relevant-informative genes detected by previous studies over these datasets. The aim of this work is to make contributions in the field of the gene selection task in DNA microarray datasets. By an extensive comparison with more popular filter techniques, we would like to make contributions in the expansion and study of the wrapper approach in this type of domains.

Artificial Intelligence↗

Characterization of a partially denatured state of a protein by two-dimensional NMR: reduction of the hydrophobic interactions in ubiquitin.

A stable, partially structured state of ubiquitin, the A-state, is formed at pH 2.0 in 60% methanol/40% water at 298 K. Detailed characterization of the structure of this state has been carried out by 2D NMR spectroscopy. Assignment of slowly exchanging amide resonances protected from the solvent in the native and A-state shows that gross structural reorganization of the protein has not occurred and that the A-state contains a subset of the interactions present in the native state (N-state). Vicinal coupling constants and NOESY data show the presence of the first two strands of the five-strand beta-sheet that is present in the native protein and part of the third beta-strand. The hydrophobic face of the beta-sheet in the A-state is covered by a partially structured alpha-helix, tentatively assigned to residues 24-34, that is considerably more flexible than the alpha-helix in the N-state. There is evidence for some fixed side-chain--side-chain interactions between these two units of structure. The turn-rich area of the protein, which contains seven reverse turns and a short piece of 3(10) helix, does not appear to be structured in the A-state and is approaching random coil.

Amino Acid Sequence↗

Translocation through the nuclear pore complex: selectivity and speed by reduction-of-dimensionality.

Translocation through the nuclear pore complex (NPC), a large transporter spanning the nuclear envelope, is a passive, diffusion-driven process, paradoxically enhanced by binding. To account for this mystery, several models have been suggested. However, recent experiments with modified NPCs make reconsideration necessary. Here, we suggest that nuclear transport receptors (NTRs) such as the karyopherins, in accordance with their peculiar boat-like structure, act as nanoscopic ferries transporting cargos through the NPC by sliding on a surface of phenylalanine glycine (FG) motifs. The dense array of FG motifs that covers the cytoplasmic filaments of the NPC is thought to continue on the wall of the large channel permeating the central framework of the NPC and on parts of the nuclear filaments to yield a coherent FG surface. Nuclear transport receptors are assumed to bind to the FG surface at filaments or at the channel entrance and then to rapidly search the FG surface by a two-dimensional random walk for the channel exit where they are released. The passage of neutral molecules is restricted to a narrow tube in the center of the central channel by a loose network of peptide chains. The model features virtual gating, is compatible with but not dependent on FG affinity gradients and tolerates deletions and transpositions of FG motifs. Implications of the model are discussed and tests are suggested.

Active Transport, Cell Nucleus↗

Evaluation of reaction rate enhancement by reduction in dimensionality.

The paths followed by ligands as they react with or dissociate from cell surface receptors may include weak association with nonreceptor portions of the surface followed by lateral diffusion in the plane of the membrane to a receptor. The change in dimensionality of the diffusion process by utilization of these nonspecific paths has been invoked by a number of investigators as a mechanism for enhancing reaction rate in biological systems. This paper extends our previous work on the calculation of diffusive rate constants for ligand-receptor paths. We find that they have little effect on rate constants unless the number of free receptors per cell has been reduced to less than or equal to 10(2). This number represents better than 90% occupancy for most eukaryotes, suggesting that the dimensional change mechanism is of limited consequence. We show further that when the free receptor number is low enough for rate enhancement, then the primary parameter of consequence is D'K*/D, where D' and D are the two- and three-dimensional diffusion coefficients, respectively, and K* the nonspecific affinity. A 10-fold rate enhancement with 100 free receptors requires that this parameter be of order 10(-3). This value is barely within the lower limit imposed by currently available experimental information, casting doubt on the relevance of nonspecific paths in cellular systems.

Animals↗

Systems with superabsorbing states

We report on some extensive analyses of a recently proposed model [A. Lipowski, Phys. Rev. E 60, 6255 (1999)] with infinitely many absorbing states. By performing extensive Monte Carlo simulations, we have determined critical exponents and shown strong evidence that this model is not in the directed percolation universality class. The conjecture that this two-dimensional model exhibits a dimensional reduction (behaving as one-dimensional directed percolation) is firmly disproven. The reason for the model not exhibiting standard directed percolation scaling behavior is traced back to the existence of what we call superabsorbing sites, i.e., absorbing sites that cannot be directly activated by the presence of neighboring activity in one or more than one direction. Supporting this claim we present two strong evidences: (i) in one dimension, where superabsorbing sites do not appear at the critical point, the system behaves as directed percolation, and (ii) in a modified two-dimensional variation of the model, defined on a honeycomb lattice, for which superabsorbing sites are very rarely observed, directed percolation behavior is recovered. Finally, a parallel updating version of the model exhibiting a nonequilibrium first-order transition is also reported.

Journal Article↗

Non-Fourier-encoded parallel MRI using multiple receiver coils.

This paper describes a general theoretical framework that combines non-Fourier (NF) spatially-encoded MRI with multichannel acquisition parallel MRI. The two spatial-encoding mechanisms are physically and analytically separable, which allows NF encoding to be expressed as complementary to the inherent encoding imposed by RF receiver coil sensitivities. Consequently, the number of NF spatial-encoding steps necessary to fully encode an FOV is reduced. Furthermore, by casting the FOV reduction of parallel imaging techniques as a dimensionality reduction of the k-space that is NF-encoded, one can obtain a speed-up of each digital NF spatial excitation in addition to accelerated imaging. Images acquired at speed-up factors of 2x to 8x with a four-element RF receiver coil array demonstrate the utility of this framework and the efficiency afforded by it.

Calibration↗

A flexible computational framework for detecting, characterizing, and interpreting statistical patterns of epistasis in genetic studies of human disease susceptibility.

Detecting, characterizing, and interpreting gene-gene interactions or epistasis in studies of human disease susceptibility is both a mathematical and a computational challenge. To address this problem, we have previously developed a multifactor dimensionality reduction (MDR) method for collapsing high-dimensional genetic data into a single dimension (i.e. constructive induction) thus permitting interactions to be detected in relatively small sample sizes. In this paper, we describe a comprehensive and flexible framework for detecting and interpreting gene-gene interactions that utilizes advances in information theory for selecting interesting single-nucleotide polymorphisms (SNPs), MDR for constructive induction, machine learning methods for classification, and finally graphical models for interpretation. We illustrate the usefulness of this strategy using artificial datasets simulated from several different two-locus and three-locus epistasis models. We show that the accuracy, sensitivity, specificity, and precision of a naïve Bayes classifier are significantly improved when SNPs are selected based on their information gain (i.e. class entropy removed) and reduced to a single attribute using MDR. We then apply this strategy to detecting, characterizing, and interpreting epistatic models in a genetic study (n = 500) of atrial fibrillation and show that both classification and model interpretation are significantly improved.

Atrial Fibrillation↗

Nonparametric regression applied to quantitative structure-activity relationships

Several nonparametric regressors have been applied to modeling quantitative structure-activity relationship (QSAR) data. The simplest regressor, the Nadaraya-Watson, was assessed in a genuine multivariate setting. Other regressors, the local linear and the shifted Nadaraya-Watson, were implemented within additive models--a computationally more expedient approach, better suited for low-density designs. Performances were benchmarked against the nonlinear method of smoothing splines. A linear reference point was provided by multilinear regression (MLR). Variable selection was explored using systematic combinations of different variables and combinations of principal components. For the data set examined, 47 inhibitors of dopamine beta-hydroxylase, the additive nonparametric regressors have greater predictive accuracy (as measured by the mean absolute error of the predictions or the Pearson correlation in cross-validation trails) than MLR. The use of principal components did not improve the performance of the nonparametric regressors over use of the original descriptors, since the original descriptors are not strongly correlated. It remains to be seen if the nonparametric regressors can be successfully coupled with better variable selection and dimensionality reduction in the context of high-dimensional QSARs.

Journal Article↗

Identifying critical variables of principal components for unsupervised feature selection.

Principal components analysis (PCA) is probably the best-known approach to unsupervised dimensionality reduction. However, axes of the lower-dimensional space, ie., principal components (PCs), are a set of new variables carrying no clear physical meanings. Thus, interpretation of results obtained in the lower-dimensional PCA space and data acquisition for test samples still involve all of the original measurements. To deal with this problem, we develop two algorithms to link the physically meaningless PCs back to a subset of original measurements. The main idea of the algorithms is to evaluate and select feature subsets based on their capacities to reproduce sample projections on principal axes. The strength of the new algorithms is that the computaion complexity involved is significantly reduced, compared with the data structural similarity-based feature evaluation.

Algorithms↗