Search PubMed⌕ Search

Biomedical subjects

Shun-ichi Amari

Publications and source records attributed to Shun-ichi Amari.

At least 19 recordsLinked to original sources

Discovering biomarkers from gene expression data for predicting cancer subgroups using neural networks and relational fuzzy clustering.

BACKGROUND: The four heterogeneous childhood cancers, neuroblastoma, non-Hodgkin lymphoma, rhabdomyosarcoma, and Ewing sarcoma present a similar histology of small round blue cell tumor (SRBCT) and thus often leads to misdiagnosis. Identification of biomarkers for distinguishing these cancers is a well studied problem. Existing methods typically evaluate each gene separately and do not take into account the nonlinear interaction between genes and the tools that are used to design the diagnostic prediction system. Consequently, more genes are usually identified as necessary for prediction. We propose a general scheme for finding a small set of biomarkers to design a diagnostic system for accurate classification of the cancer subgroups. We use multilayer networks with online gene selection ability and relational fuzzy clustering to identify a small set of biomarkers for accurate classification of the training and blind test cases of a well studied data set. RESULTS: Our method discerned just seven biomarkers that precisely categorized the four subgroups of cancer both in training and blind samples. For the same problem, others suggested 19-94 genes. These seven biomarkers include three novel genes (NAB2, LSP1 and EHD1 - not identified by others) with distinct class-specific signatures and important role in cancer biology, including cellular proliferation, transendothelial migration and trafficking of MHC class antigens. Interestingly, NAB2 is downregulated in other tumors including Non-Hodgkin lymphoma and Neuroblastoma but we observed moderate to high upregulation in a few cases of Ewing sarcoma and Rabhdomyosarcoma, suggesting that NAB2 might be mutated in these tumors. These genes can discover the subgroups correctly with unsupervised learning, can differentiate non-SRBCT samples and they perform equally well with other machine learning tools including support vector machines. These biomarkers lead to four simple human interpretable rules for the diagnostic task. CONCLUSION: Although the proposed method is tested on a SRBCT data set, it is quite general and can be applied to other cancer data sets. Our scheme takes into account the interaction between genes as well as that between genes and the tool and thus is able find a very small set and can discover novel genes. Our findings suggest the possibility of developing specialized microarray chips or use of real-time qPCR assays or antibody based methods such as ELISA and western blot analysis for an easy and low cost diagnosis of the subgroups.

Biomarkers, Tumor↗

Fisher information for spike-based population decoding.

We evaluate the Fisher information of a population of model neurons that receive dynamical input and interact via spikes. With spatially independent threshold noise, the spike-based Fisher information that summarizes the information carried by individual spike timings has a particularly simple analytical form. We calculate the loss of information caused by abandoning spike timing and study the effect of synaptic connections on the Fisher information. For a simple spatiotemporal input, we derive the optimal recurrent connectivity that has a local excitation and global inhibition structure. The optimal synaptic connections depend on the spatial or temporal feature of the input that the system is designed to code.

Algorithms↗

Removal of ballistocardiogram artifacts from simultaneously recorded EEG and fMRI data using independent component analysis.

Simultaneous recording of electroencephalogram (EEG) and functional magnetic resonance imaging (fMRI) has been studied to identify areas related to EEG events. EEG data recorded in the magnetic resonance (MR) scanner with MR imaging is suffered from two specific artifacts, imaging artifact, and ballistocardiogram (BCG). In this paper, we focus on BCG. In preceding studies, average subtraction was often used for this purpose. However, average subtraction requires an assumption that BCG waveforms are precisely periodic, which seems unrealistic because BCG is a biomedical artifact. We propose the application of independent component analysis (ICA) with a postprocessing of high-pass filtering for the removal of BCG. With this approach, it is not necessary to assume that the BCG waveform is periodic. Empirically, we show that our proposed method removes BCG artifacts as well as does the average subtraction method. Power spectral density analysis of the two approaches shows that, with ICA, distortion of recovered EEG data is also as small as that associated with the average subtraction approach. We also propose a hypothesis for how head movement causes BCGs and show why ICA can remove BCG artifacts arising from this source.

Algorithms↗

Global exponential stability of multitime scale competitive neural networks with nonsmooth functions.

In this paper, we study the global exponential stability of a multitime scale competitive neural network model with nonsmooth functions, which models a literally inhibited neural network with unsupervised Hebbian learning. The network has two types of state variables, one corresponds to the fast neural activity and another to the slow unsupervised modification of connection weights. Based on the nonsmooth analysis techniques, we prove the existence and uniqueness of equilibrium for the system and establish some new theoretical conditions ensuring global exponential stability of the unique equilibrium of the neural network. Numerical simulations are conducted to illustrate the effectiveness of the derived conditions in characterizing stability regions of the neural network.

Algorithms↗

A comparison of descriptive models of a single spike train by information-geometric measure.

In examining spike trains, different models are used to describe their structure. The different models often seem quite similar, but because they are cast in different formalisms, it is often difficult to compare their predictions. Here we use the information-geometric measure, an orthogonal coordinate representation of point processes, to express different models of stochastic point processes in a common coordinate system. Within such a framework, it becomes straightforward to visualize higher-order correlations of different models and thereby assess the differences between models. We apply the information-geometric measure to compare two similar but not identical models of neuronal spike trains: the inhomogeneous Markov and the mixture of Poisson models. It is shown that they differ in the second- and higher-order interaction terms. In the mixture of Poisson model, the second- and higher-order interactions are of comparable magnitude within each order, whereas in the inhomogeneous Markov model, they have alternating signs over different orders. This provides guidance about what measurements would effectively separate the two models. As newer models are proposed, they also can be compared to these models using information geometry.

Action Potentials↗

Singularities affect dynamics of learning in neuromanifolds.

The parameter spaces of hierarchical systems such as multilayer perceptrons include singularities due to the symmetry and degeneration of hidden units. A parameter space forms a geometrical manifold, called the neuromanifold in the case of neural networks. Such a model is identified with a statistical model, and a Riemannian metric is given by the Fisher information matrix. However, the matrix degenerates at singularities. Such a singular structure is ubiquitous not only in multilayer perceptrons but also in the gaussian mixture probability densities, ARMA time-series model, and many other cases. The standard statistical paradigm of the Cramér-Rao theorem does not hold, and the singularity gives rise to strange behaviors in parameter estimation, hypothesis testing, Bayesian inference, model selection, and in particular, the dynamics of learning from examples. Prevailing theories so far have not paid much attention to the problem caused by singularity, relying only on ordinary statistical theories developed for regular (nonsingular) models. Only recently have researchers remarked on the effects of singularity, and theories are now being developed. This article gives an overview of the phenomena caused by the singularities of statistical manifolds related to multilayer perceptrons and gaussian mixtures. We demonstrate our recent results on these problems. Simple toy models are also used to show explicit solutions. We explain that the maximum likelihood estimator is no longer subject to the gaussian distribution even asymptotically, because the Fisher information matrix degenerates, that the model selection criteria such as AIC, BIC, and MDL fail to hold in these models, that a smooth Bayesian prior becomes singular in such models, and that the trajectories of dynamics of learning are strongly affected by the singularity, causing plateaus or slow manifolds in the parameter space. The natural gradient method is shown to perform well because it takes the singular geometrical structure into account. The generalization error and the training error are studied in some examples.

Journal Article↗

Correlation and independence in the neural code.

The decoding scheme of a stimulus can be different from the stochastic encoding scheme in the neural population coding. The stochastic fluctuations are not independent in general, but an independent version could be used for the ease of decoding. How much information is lost by using this unfaithful model for decoding? There are discussions concerning loss of information (Nirenberg & Latham, 2003; Schneidman, Bialek, & Berry, 2003). We elucidate the Nirenberg-Latham loss from the point of view of information geometry.

Brain↗

Difficulty of singularity in population coding.

Fisher information has been used to analyze the accuracy of neural population coding. This works well when the Fisher information does not degenerate, but when two stimuli are presented to a population of neurons, a singular structure emerges by their mutual interactions. In this case, the Fisher information matrix degenerates, and the regularity condition ensuring the Cramér-Rao paradigm of statistics is violated. An animal shows pathological behavior in such a situation. We present a novel method of statistical analysis to understand information in population coding in which algebraic singularity plays a major role. The method elucidates the nature of the pathological case by calculating the Fisher information. We then suggest that synchronous firing can resolve singularity and show a method of analyzing the binding problem in terms of the Fisher information. Our method integrates a variety of disciplines in population coding, such as nonregular statistics, Bayesian statistics, singularity in algebraic geometry, and synchronous firing, under the theme of Fisher information.

Action Potentials↗

Computing with continuous attractors: stability and online aspects.

Two issues concerning the application of continuous attractors in neural systems are investigated: the computational robustness of continuous attractors with respect to input noises and the implementation of Bayesian online decoding. In a perfect mathematical model for continuous attractors, decoding results for stimuli are highly sensitive to input noises, and this sensitivity is the inevitable consequence of the system's neutral stability. To overcome this shortcoming, we modify the conventional network model by including extra dynamical interactions between neurons. These interactions vary according to the biologically plausible Hebbian learning rule and have the computational role of memorizing and propagating stimulus information accumulated with time. As a result, the new network model responds to the history of external inputs over a period of time, and hence becomes insensitive to short-term fluctuations. Also, since dynamical interactions provide a mechanism to convey the prior knowledge of stimulus, that is, the information of the stimulus presented previously, the network effectively implements online Bayesian inference. This study also reveals some interesting behavior in neural population coding, such as the trade-off between decoding stability and the speed of tracking time-varying stimuli, and the relationship between neural tuning width and the tracking speed.

Humans↗

Information processing in a neuron ensemble with the multiplicative correlation structure.

The present study investigates the performance of population codes when the fluctuations in neural activity have mutual correlation with strength being proportional to the neuronal firing rate (multiplicative noise). The neural field is used to calculate the Fisher information, which is decomposed in two parts, one due to the tuning function and spatial correlation, and the other due to the multiplicative structure. Their different characteristics are studied. The paper also investigates three types of maximum likelihood method, namely, decoding by using faithful and unfaithful models and the Center of Mass strategy, and compares their performances in terms of decoding accuracy and computational complexity.

Models, Neurological↗

Self-adaptive blind source separation based on activation functions adaptation.

Independent component analysis is to extract independent signals from their linear mixtures without assuming prior knowledge of their mixing coefficients. As we know, a number of factors are likely to affect separation results in practical applications, such as the number of active sources, the distribution of source signals, and noise. The purpose of this paper to develop a general framework of blind separation from a practical point of view with special emphasis on the activation function adaptation. First, we propose the exponential generative model for probability density functions. A method of constructing an exponential generative model from the activation functions is discussed. Then, a learning algorithm is derived to update the parameters in the exponential generative model. The learning algorithm for the activation function adaptation is consistent with the one for training the demixing model. Stability analysis of the learning algorithm for the activation function is also discussed. Both theoretical analysis and simulations show that the proposed approach is universally convergent regardless of the distributions of sources. Finally, computer simulations are given to demonstrate the effectiveness and validity of the approach.

Computer Simulation↗

From blind signal extraction to blind instantaneous signal separation: criteria, algorithms, and stability.

This paper reports a study on the problem of the blind simultaneous extraction of specific groups of independent components from a linear mixture. This paper first presents a general overview and unification of several information theoretic criteria for the extraction of a single independent component. Then, our contribution fills the theoretical gap that exists between extraction and separation by presenting tools that extend these criteria to allow the simultaneous blind extraction of subsets with an arbitrary number of independent components. In addition, we analyze a family of learning algorithms based on Stiefel manifolds and the natural gradient ascent, present the nonlinear optimal activations (score) functions, and provide new or extended local stability conditions. Finally, we illustrate the performance and features of the proposed approach by computer-simulation experiments.

Algorithms↗

Stochastic reasoning, free energy, and information geometry.

Belief propagation (BP) is a universal method of stochastic reasoning. It gives exact inference for stochastic models with tree interactions and works surprisingly well even if the models have loopy interactions. Its performance has been analyzed separately in many fields, such as AI, statistical physics, information theory, and information geometry. This article gives a unified framework for understanding BP and related methods and summarizes the results obtained in many fields. In particular, BP and its variants, including tree reparameterization and concave-convex procedure, are reformulated with information-geometrical terms, and their relations to the free energy function are elucidated from an information-geometrical viewpoint. We then propose a family of new algorithms. The stabilities of the algorithms are analyzed, and methods to accelerate them are investigated.

Algorithms↗

Analysis of sparse representation and blind source separation.

In this letter, we analyze a two-stage cluster-then-l(1)-optimization approach for sparse representation of a data matrix, which is also a promising approach for blind source separation (BSS) in which fewer sensors than sources are present. First, sparse representation (factorization) of a data matrix is discussed. For a given overcomplete basis matrix, the corresponding sparse solution (coefficient matrix) with minimum l(1) norm is unique with probability one, which can be obtained using a standard linear programming algorithm. The equivalence of the l(1)-norm solution and the l(0)-norm solution is also analyzed according to a probabilistic framework. If the obtained l(1)-norm solution is sufficiently sparse, then it is equal to the l(0)-norm solution with a high probability. Furthermore, the l(1)- norm solution is robust to noise, but the l(0)-norm solution is not, showing that the l(1)-norm is a good sparsity measure. These results can be used as a recoverability analysis of BSS, as discussed. The basis matrix in this article is estimated using a clustering algorithm followed by normalization, in which the matrix columns are the cluster centers of normalized data column vectors. Zibulevsky, Pearlmutter, Boll, and Kisilev (2000) used this kind of two-stage approach in underdetermined BSS. Our recoverability analysis shows that this approach can deal with the situation in which the sources are overlapped to some degree in the analyzed domain and with the case in which the source number is unknown. It is also robust to additive noise and estimation error in the mixing matrix. Finally, four simulation examples and an EEG data analysis example are presented to illustrate the algorithm's utility and demonstrate its performance.

Algorithms↗

Gene interaction in DNA microarray data is decomposed by information geometric measure.

MOTIVATION: Given the vast amount of gene expression data, it is essential to develop a simple and reliable method of investigating the fine structure of gene interaction. We show how an information geometric measure achieves this. RESULTS: We introduce an information geometric measure of binary random vectors and show how this measure reveals the fine structure of gene interaction. In particular, we propose an iterative procedure by using this measure (called IPIG). The procedure finds higher-order dependencies which may underlie the interaction between two genes of interest. To demonstrate the method, we investigate the interaction between the two genes of interest in the data from human acute lymphoblastic leukemia cells. The method successfully discovered biologically known findings and also selected other genes as hidden causes that constitute the interaction. AVAILABILITY: Softwares are currently not available but are possibly made available in future at http://www.mns.brain.riken.go.jp/~nakahara/DNA_pub.html where all the related information is also linked.

Algorithms↗

Neuroscience data and tool sharing: a legal and policy framework for neuroinformatics.

The requirements for neuroinformatics to make a significant impact on neuroscience are not simply technical--the hardware, software, and protocols for collaborative research--they also include the legal and policy frameworks within which projects operate. This is not least because the creation of large collaborative scientific databases amplifies the complicated interactions between proprietary, for-profit R&D and public "open science." In this paper, we draw on experiences from the field of genomics to examine some of the likely consequences of these interactions in neuroscience. Facilitating the widespread sharing of data and tools for neuroscientific research will accelerate the development of neuroinformatics. We propose approaches to overcome the cultural and legal barriers that have slowed these developments to date. We also draw on legal strategies employed by the Free Software community, in suggesting frameworks neuroinformatics might adopt to reinforce the role of public-science databases, and propose a mechanism for identifying and allowing "open science" uses for data whilst still permitting flexible licensing for secondary commercial research.

Computational Biology↗

Sequential Bayesian decoding with a population of neurons.

Population coding is a simplified model of distributed information processing in the brain. This study investigates the performance and implementation of a sequential Bayesian decoding (SBD) paradigm in the framework of population coding. In the first step of decoding, when no prior knowledge is available, maximum likelihood inference is used; the result forms the prior knowledge of stimulus for the second step of decoding. Estimates are propagated sequentially to apply maximum a posteriori (MAP) decoding in which prior knowledge for any step is taken from estimates from the previous step. Not only do we analyze the performance of SBD, obtaining the optimal form of prior knowledge that achieves the best estimation result, but we also investigate its possible biological realization, in the sense that all operations are performed by the dynamics of a recurrent network. In order to achieve MAP, a crucial point is to identify a mechanism that propagates prior knowledge. We find that this could be achieved by short-term adaptation of network weights according to the Hebbian learning rule. Simulation results on both constant and time-varying stimulus support the analysis.

Bayes Theorem↗