Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dimensionality Reduction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Data visualisation and manifold mapping using the ViSOM.

The self-organising map (SOM) has been successfully employed as a nonparametric method for dimensionality reduction and data visualisation. However, for visualisation the SOM requires a colouring scheme to imprint the distances between neurons so that the clustering and boundaries can be seen. Even though the distributions of the data and structures of the clusters are not faithfully portrayed on the map. Recently an extended SOM, called the visualisation-induced SOM (ViSOM) has been proposed to directly preserve the distance information on the map, along with the topology. The ViSOM constrains the lateral contraction forces between neurons and hence regularises the interneuron distances so that distances between neurons in the data space are in proportion to those in the map space. This paper shows that it produces a smooth and graded mesh in the data space and captures the nonlinear manifold of the data. The relationships between the ViSOM and the principal curve/surface are analysed. The ViSOM represents a discrete principal curve or surface and is a natural algorithm for obtaining principal curves/surfaces. Guidelines for applying the ViSOM constraint and setting the resolution parameter are also provided, together with experimental results and comparisons with the SOM, Sammon mapping and principal curve methods.

Algorithms↗

The information content of receptive fields.

The nervous system must observe a complex world and produce appropriate, sometimes complex, behavioral responses. In contrast to this complexity, neural responses are often characterized through very simple descriptions such as receptive fields or tuning curves. Do these characterizations adequately reflect the true dimensionality reduction that takes place in the nervous system, or are they merely convenient oversimplifications? Here we address this question for the target-selective descending neurons (TSDNs) of the dragonfly. Using extracellular multielectrode recordings of a population of TSDNs, we quantify the completeness of the receptive field description of these cells and conclude that the information in independent instantaneous position and velocity receptive fields accounts for 70%-90% of the total information in single spikes. Thus, we demonstrate that this simple receptive field model is close to a complete description of the features in the stimulus that evoke TSDN response.

Action Potentials↗

An ischemia detection method based on artificial neural networks.

An automated technique was developed for the detection of ischemic episodes in long duration electrocardiographic (ECG) recordings that employs an artificial neural network. In order to train the network for beat classification, a cardiac beat dataset was constructed based on recordings from the European Society of Cardiology (ESC) ST-T database. The network was trained using a Bayesian regularisation method. The raw ECG signal containing the ST segment and the T wave of each beat were the inputs to the beat classification system and the output was the classification of the beat. The input to the network was produced through a principal component analysis (PCA) to achieve dimensionality reduction. The network performance in beat classification was tested on the cardiac beat database providing 90% sensitivity (Se) and 90% specificity (Sp). The neural beat classifier is integrated in a four-stage procedure for ischemic episode detection. The whole system was evaluated on the ESC ST-T database. When aggregate gross statistics was used the Se was 90% and the positive predictive accuracy (PPA) 89%. When aggregate average statistics was used the Se became 86% and the PPA 87%. These results are better than other reported.

Automation↗

Stepping out of the box: information processing in the neural networks of the basal ganglia.

The Albin-DeLong 'box and arrow' model has long been the accepted standard model for the basal ganglia network. However, advances in physiological and anatomical research have enabled a more detailed neural network approach. Recent computational models hold that the basal ganglia use reinforcement signals and local competitive learning rules to reduce the dimensionality of sparse cortical information. These models predict a steady-state situation with diminished efficacy of lateral inhibition and low synchronization. In this framework, Parkinson's disease can be characterized as a persistent state of negative reinforcement, inefficient dimensionality reduction, and abnormally synchronized basal ganglia activity.

Animals↗

Classification of the myoelectric signal using time-frequency based representations.

An accurate and computationally efficient means of classifying surface myoelectric signal patterns has been the subject of considerable research effort in recent years. Effective feature extraction is crucial to reliable classification and, in the quest to improve the accuracy of transient myoelectric signal pattern classification, an ensemble of time-frequency based representations are proposed. It is shown that feature sets based upon the short-time Fourier transform, the wavelet transform, and the wavelet packet transform provide an effective representation for classification, provided that they are subject to an appropriate form of dimensionality reduction.

Action Potentials↗

Mapping peroxidase in plant tissues by scanning electrochemical microscopy.

Scanning electrochemical microscopy has been firstly used to map the enzymatic activity in natural plant tissues. The peroxidase (POD) was maintained in its original state in the celery (Apium graveolens L.) tissues and electrochemically visualized under its native environment. Ferrocenemethanol (FMA) was selected as a mediator to probe the POD in celery tissues based on the fact that POD catalyzed the oxidation of FMA by H(2)O(2) to increase FMA(+) concentration. Two-dimensional reduction current profiles for FMA(+) produced images indicating the distribution and activity of the POD at the surface of the celery tissues. These images showed that the POD was widely distributed in the celery tissues, and larger amounts were found in some special regions such as the center of celery stem and around some vascular bundles.

Apium↗

Genetic heterogeneity affects the risk of incident depression, comorbidity, and response to environment: A prospective trajectory study.

BACKGROUND: Depression exhibits significant heterogeneity in its genetic underpinnings. The role of genetic components in the development of depression and its comorbidities remains insufficiently explored. METHODS: First, depression risk loci from a large-scale genome-wide meta-analysis were annotated to Gene Ontology (GO) terms by functional enrichment. GO-based polygenic risk scores (GO-PRS) were then calculated for individuals in the UK Biobank. Principal component analysis (PCA) was applied for dimensionality reduction, followed by cluster analysis to identify genetic subtypes of depression. Multistate models were applied to assess the impact of genetic patterns on the trajectory from healthy status to incident depression, and depression to 26 subsequent diseases, as well as the associations between environmental factors and disease trajectories across genetic subtypes. RESULTS: Participants were categorized into three genetic subtypes: immune-dominant, neuro-dominant, and comprehensive-risk. Significant differences in risk of depression and subsequent diseases, and susceptibility to environmental factors were observed across subtypes. Comprehensive-risk subtype showed higher risks of depression compared to immune-dominant (HR: 1.10, 95% CI: 1.05-1.15) and neuro-dominant subtype (HR: 1.12, 95% CI: 1.08-1.16). Comprehensive-risk subtype exhibited higher risks of transition from depression to subsequent diseases, such as anemia compared to immune-dominant subtype, and diseases of the digestive system compared to neuro-dominant subtype. Environmental factors were more strongly associated with the transition from depression to subsequent diseases in immune-dominant and comprehensive-risk subtypes, including cardiovascular, respiratory, and metabolic diseases. CONCLUSIONS: Our findings highlight the genetic heterogeneity of depression and comorbidities, and shed light on how genetic components modulate responses to environmental factors.

Humans↗

The development of features in object concepts.

According to one productive and influential approach to cognition, categorization, object recognition, and higher level cognitive processes operate on a set of fixed features, which are the output of lower level perceptual processes. In many situations, however, it is the higher level cognitive process being executed that influences the lower level features that are created. Rather than viewing the repertoire of features as being fixed by low-level processes, we present a theory in which people create features to subserve the representation and categorization of objects. Two types of category learning should be distinguished. Fixed space category learning occurs when new categorizations are representable with the available feature set. Flexible space category learning occurs when new categorizations cannot be represented with the features available. Whether fixed or flexible, learning depends on the featural contrasts and similarities between the new category to be represented and the individuals existing concepts. Fixed feature approaches face one of two problems with tasks that call for new features: If the fixed features are fairly high level and directly useful for categorization, then they will not be flexible enough to represent all objects that might be relevant for a new task. If the fixed features are small, subsymbolic fragments (such as pixels), then regularities at the level of the functional features required to accomplish categorizations will not be captured by these primitives. We present evidence of flexible perceptual changes arising from category learning and theoretical arguments for the importance of this flexibility. We describe conditions that promote feature creation and argue against interpreting them in terms of fixed features. Finally, we discuss the implications of functional features for object categorization, conceptual development, chunking, constructive induction, and formal models of dimensionality reduction.

Child↗

Single-Cell Proteomics Reveals Proteome Remodeling and Cellular Heterogeneity During NGF-Induced PC12 Neuronal Differentiation.

Single-cell proteomics enables direct measurement of cellular heterogeneity during dynamic biological processes, but its application to fragile and highly adherent neuronal models remains challenging. Here, we developed and applied an optimized single-cell proteomics workflow to characterize proteome remodeling during nerve growth factor (NGF)-induced differentiation of PC12 cells. To enable reliable single-cell analysis, we implemented gentle dissociation, antiaggregation strategies, and thermal inkjet-based cell dispensing, achieving high accuracy in single-cell isolation. Inclusion of n-dodecyl-β-d-maltoside (DDM) improved recovery of membrane-associated and low-solubility proteins. Coupled with LC-ion mobility-mass spectrometry, this workflow enabled quantification of 2,000-3,000 proteins per cell across the differentiation time course. Single-cell proteomic analysis revealed progressive and heterogeneous proteome remodeling during differentiation. While undifferentiated cells formed a relatively homogeneous population, later stages (Days 4-6) exhibited increased variability, including multimodal protein abundance distributions and separation into distinct subpopulations. Dimensionality reduction, clustering, and non-negative matrix factorization identified multiple coexisting proteomic states within the same time points, reflecting asynchronous differentiation trajectories. These subpopulations were characterized by coordinated differences in pathways related to intracellular trafficking, protein translation, cytoskeletal organization, and neuronal maturation. Comparison with bulk proteomics demonstrated that proteins associated with differentiated neuronal states, including those involved in neurite formation and structural remodeling, are underrepresented in population-averaged measurements but are enriched within specific single-cell subpopulations. Temporal and cluster-resolved analyses further revealed distinct protein expression trajectories, including early decreases in cell cycle and metabolic pathways and later increases in neuronal structural and regulatory proteins. Together, this study establishes an optimized workflow for single-cell proteomics of neuronal systems and demonstrates that NGF-induced PC12 differentiation proceeds through heterogeneous and divergent proteomic states that are not resolved by bulk analysis.

Animals↗

Efficient thrombin generation requires molecular phosphatidylserine, not a membrane surface.

Activation of prothrombin to thrombin is catalyzed by a "prothrombinase" complex, traditionally viewed as factor X(a) (FX(a)) in complex with factor V(a) (FV(a)) on a phosphatidylserine (PS)-containing membrane surface, which is widely regarded as required for efficient activation. Activation involves cleavage of two peptide bonds and proceeds via one of two released intermediates or through "channeling" (activation without the release of an intermediate). We ask here whether the PS molecule itself and not the membrane surface is sufficient to produce the fully active human "prothrombinase" complex in solution. Both FX(a) and FV(a) bind soluble dicaproyl-phosphatidylserine (C6PS). In the presence of sufficient C6PS to saturate both FX(a) and FV(a2) (light isoform of FV(a)), these proteins form a tight (Kd = 0.6 +/- 0.09 nM at 37 degrees C) soluble complex. Complex assembly occurs well below the critical micelle concentration of C6PS, as established in the presence of the proteins by quasi-elastic light scattering and pyrene fluorescence. Ferguson analysis of native gels shows that the complex migrates with an apparent molecular mass only slightly larger than that expected for one FX(a) and one FV(a2), further ruling out complex assembly on C6PS micelles. Human prothrombin activation by this complex occurs at nearly the same overall rate (2.2 x 10(8) M(-1) s(-1)) and via the same reaction pathway (50-60% channeling, with the rest via the meizothrombin intermediate) as the activation catalyzed by a complex assembled on PS-containing membranes (4.4 x 10(8) M(-1) s(-1)). These results question the accepted role of PS membranes as providing "dimensionality reduction" and favor a regulatory role for platelet-membrane-exposed PS.

Arginine↗

An adaptive strategy for single- and multi-cluster gene assignment.

Strict assignment of genes to one class, dimensionality reduction, a priori specification of the number of classes, the need for a training set, nonunique solution, and complex learning mechanisms are some of the inadequacies of current clustering algorithms. Existing algorithms cluster genes on the basis of high positive correlations between their expression patterns. However, genes with strong negative correlations can also have similar functions and are most likely to have a role in the same pathways. To address some of these issues, we propose the adaptive centroid algorithm (ACA), which employs an analysis of variance (ANOVA)-based performance criterion. The ACA also uses Euclidian distances, the center-of-mass principle for heterogeneously distributed mass elements, and the given data set to give unique solutions. The proposed approach involves three stages. In the first stage a two-way ANOVA of the gene expression matrix is performed. The two factors in the ANOVA are gene expression and experimental condition. The residual mean squared error (MSE) from the ANOVA is used as a performance criterion in the ACA. Finally, correlated clusters are found based on the Pearson correlation coefficients. To validate the proposed approach, a two-way ANOVA is again performed on the discovered clusters. The results from this last step indicate that MSEs of the clusters are significantly lower compared to that of the fibroblast-serum gene expression matrix. The ACA is employed in this study for single- as well as multi-cluster gene assignments.

Algorithms↗

Nonlinear mapping networks.

Among the many dimensionality reduction techniques that have appeared in the statistical literature, multidimensional scaling and nonlinear mapping are unique for their conceptual simplicity and ability to reproduce the topology and structure of the data space in a faithful and unbiased manner. However, a major shortcoming of these methods is their quadratic dependence on the number of objects scaled, which imposes severe limitations on the size of data sets that can be effectively manipulated. Here we describe a novel approach that combines conventional nonlinear mapping techniques with feed-forward neural networks, and allows the processing of data sets orders of magnitude larger than those accessible with conventional methodologies. Rooted on the principle of probability sampling, the method employs a classical algorithm to project a small random sample, and then "learns" the underlying nonlinear transform using a multilayer neural network trained with the back-propagation algorithm. Once trained, the neural network can be used in a feed-forward manner to project the remaining members of the population as well as new, unseen samples with minimal distortion. Using examples from the fields of image processing and combinatorial chemistry, we demonstrate that this method can generate projections that are virtually indistinguishable from those derived by conventional approaches. The ability to encode the nonlinear transform in the form of a neural network makes nonlinear mapping applicable to a wide variety of data mining applications involving very large data sets that are otherwise computationally intractable.

Journal Article↗

A comparative study on feature selection methods for drug discovery.

Feature selection is frequently used as a preprocessing step to machine learning. The removal of irrelevant and redundant information often improves the performance of learning algorithms. This paper is a comparative study of feature selection in drug discovery. The focus is on aggressive dimensionality reduction. Five methods were evaluated, including information gain, mutual information, a chi2-test, odds ratio, and GSS coefficient. Two well-known classification algorithms, Naïve Bayesian and Support Vector Machine (SVM), were used to classify the chemical compounds. The results showed that Naïve Bayesian benefited significantly from the feature selection, while SVM performed better when all features were used. In this experiment, information gain and chi2-test were most effective feature selection methods. Using information gain with a Naïve Bayesian classifier, removal of up to 96% of the features yielded an improved classification accuracy measured by sensitivity. When information gain was used to select the features, SVM was much less sensitive to the reduction of feature space. The feature set size was reduced by 99%, while losing only a few percent in terms of sensitivity (from 58.7% to 52.5%) and specificity (from 98.4% to 97.2%). In contrast to information gain and chi2-test, mutual information had relatively poor performance due to its bias toward favoring rare features and its sensitivity to probability estimation errors.

Algorithms↗

General melting point prediction based on a diverse compound data set and artificial neural networks.

We report the development of a robust and general model for the prediction of melting points. It is based on a diverse data set of 4173 compounds and employs a large number of 2D and 3D descriptors to capture molecular physicochemical and other graph-based properties. Dimensionality reduction is performed by principal component analysis, while a fully connected feed-forward back-propagation artificial neural network is employed for model generation. The melting point is a fundamental physicochemical property of a molecule that is controlled by both single-molecule properties and intermolecular interactions due to packing in the solid state. Thus, it is difficult to predict, and previously only melting point models for clearly defined and smaller compound sets have been developed. Here we derive the first general model that covers a comparatively large and relevant part of organic chemical space. The final model is based on 2D descriptors, which are found to contain more relevant information than the 3D descriptors calculated. Internal random validation of the model achieves a correlation coefficient of R(2) = 0.661 with an average absolute error of 37.6 degrees C. The model is internally consistent with a correlation coefficient of the test set of Q(2) = 0.658 (average absolute error 38.2 degrees C) and a correlation coefficient of the internal validation set of Q(2) = 0.645 (average absolute error 39.8 degrees C). Additional validation was performed on an external drug data set consisting of 277 compounds. On this external data set a correlation coefficient of Q(2) = 0.662 (average absolute error 32.6 degrees C) was achieved, showing ability of the model to generalize. Compared to an earlier model for the prediction of melting points of druglike compounds our model exhibits slightly improved performance, despite the much larger chemical space covered. The remaining model error is due to molecular properties that are not captured using single-molecule based descriptors, namely both inter- and intramolecular interactions and crystal packing, for which examples of and reasons for outliers are given.

Journal Article↗

Advances in diversity profiling and combinatorial series design.

Rapid advances in synthetic and screening technology have recently enabled the simultaneous synthesis and biological evaluation of large chemical libraries containing hundreds to tens of thousands of compounds, using molecular diversity as a means to design and prioritize experiments. This paper reviews some of the most important computational work in the field of diversity profiling and combinatorial library design, with particular emphasis on methodology and applications. It is divided into four sections that address issues related to molecular representation, dimensionality reduction, compound selection, and visualization.

Chemistry, Pharmaceutical↗

Combinatorial pharmacogenetics.

Combinatorial pharmacogenetics seeks to characterize genetic variations that affect reactions to potentially toxic agents within the complex metabolic networks of the human body. Polymorphic drug-metabolizing enzymes are likely to represent some of the most common inheritable risk factors associated with common 'disease' phenotypes, such as adverse drug reactions. The relatively high concordance between polymorphisms in drug-metabolizing enzymes and clinical phenotypes indicates that research into this class of polymorphisms could benefit patients in the near future. Characterization of other genes affecting drug disposition (absorption, distribution, metabolism and elimination) will further enhance this process. As with most questions concerning biological systems, the complexity arises out of the combinatorial magnitude of all the possible interactions and pathways. The high-dimensionality of the resulting analysis problem will often overwhelm traditional analysis methods. Novel analysis techniques, such as multifactor dimensionality reduction, offer viable options for evaluating such data.

Animals↗

An association study of the N-methyl-D-aspartate receptor NR1 subunit gene (GRIN1) and NR2B subunit gene (GRIN2B) in schizophrenia with universal DNA microarray.

Dysfunction of the N-methyl-D-aspartate (NMDA) receptors has been implicated in the etiology of schizophrenia based on psychotomimetic properties of several antagonists and on observation of genetic animal models. To conduct association analysis of the NMDA receptors in the Chinese population, we examined 16 reported SNPs across the NMDA receptor NR1 subunit gene (GRIN1) and NR2B subunit gene (GRIN2B), five of which were identified in the Chinese population. In this study, we combined universal DNA microarray and ligase detection reaction (LDR) for the purposes of association analysis, an approach we considered to be highly specific as well as offering a potentially high throughput of SNP genotyping. The association study was performed using 253 Chinese patients with schizophrenia and 140 Chinese control subjects. No significant frequency differences were found in the analysis of the alleles but some were found in the haplotypes of the GRIN2B gene. The interactions between the GRIN1 and GRIN2B genes were evaluated using the multifactor-dimensionality reduction (MDR) method, which showed a significant genetic interaction between the G1001C in the GRIN1 gene and the T4197C and T5988C polymorphisms in the GRIN2B gene. These findings suggest that the combined effects of the polymorphisms in the GRIN1 and GRIN2B genes might be involved in the etiology of schizophrenia.European Journal of Human Genetics (2005) 13, 807-814. doi:10.1038/sj.ejhg.5201418 Published online 20 April 2005.

Adult↗

Predicting interpretability of metabolome models based on behavior, putative identity, and biological relevance of explanatory signals.

Powerful algorithms are required to deal with the dimensionality of metabolomics data. Although many achieve high classification accuracy, the models they generate have limited value unless it can be demonstrated that they are reproducible and statistically relevant to the biological problem under investigation. Random forest (RF) generates models, without any requirement for dimensionality reduction or feature selection, in which individual variables are ranked for significance and displayed in an explicit manner. In metabolome fingerprinting by mass spectrometry, each metabolite can be represented by signals at several m/z. Exploiting a prior understanding of expected biochemical differences between sample classes, we aimed to develop meaningful metrics relevant to the significance both of the overall RF model and individual, potentially explanatory, signals. Pair-wise comparison of related plant genotypes with strong phenotypic differences demonstrated that robust models are not only reproducible but also logically structured, highlighting correlated m/z derived from just a small number of explanatory metabolites reflecting the biological differences between sample classes. RF models were also generated by using groupings of samples known to be increasingly phenotypically similar. Although classification accuracy was often reasonable, we demonstrated reproducibly in both Arabidopsis and potato a performance threshold based on margin statistics beyond which such models showed little structure indicative of either generalizability or further biological interpretability. In a multiclass problem using 25 Arabidopsis genotypes, despite the complicating effects of ecotype background and secondary metabolome perturbations common to several mutations, the ranking of metabolome signals by RF provided scope for deeper interpretability.

Arabidopsis↗