Search PubMed⌕ Search

Biomedical subjects

De-Shuang Huang

Publications and source records attributed to De-Shuang Huang.

At least 19 recordsLinked to original sources

Independent component analysis-based penalized discriminant method for tumor classification using gene expression data.

MOTIVATION: Microarrays are capable of determining the expression levels of thousands of genes simultaneously. One important application of gene expression data is classification of samples into categories. In combination with classification methods, this technology can be useful to support clinical management decisions for individual patients, e.g. in oncology. Standard statistic methodologies in classification or prediction do not work well when the number of variables p (genes) far too exceeds the number of samples n. So, modification of existing statistical methodologies or development of new methodologies is needed for the analysis of microarray data. RESULTS: This paper proposes a new method for tumor classification using gene expression data. In this method, we first employ independent component analysis to model the gene expression data, then apply optimal scoring algorithm to classify them. Further speaking, this approach can first make full use of the high-order statistical information contained in the gene expression data. Second, this approach also employs regularized regression models to handle the situation of large numbers of correlated predictor variables. Finally, the predictive models are developed for classifying tumors based on the entire gene expression profile. To show the validity of the proposed method, we apply it to classify four DNA microarray datasets involving various human normal and tumor tissue samples. The experimental results show that the method is efficient and feasible. AVAILABILITY: Matlab scripts are available on request.

Algorithms↗

Identifying protein-protein interfacial residues in heterocomplexes using residue conservation scores.

Identifying protein-protein interfaces is crucial for structural biology. Because of the constraints in wet experiments, many computational methods have been proposed. Without knowing any information about the partner chains, a new method of predicting protein-protein interaction interface residues purely based on evolutionary information in heterocomplexes is proposed here. Unlike traditional approaches using multiple sequence alignment profiles to represent the conservation level for each residue, we make predictions based on the concept of residue conservation scores so that the dimension of the feature vector for each residue can be drastically reduced, at least 20 times less than conventional methods. Based on the representation approach, a simple linear discriminant function is used to make predictions, so the computational complexity of the whole prediction procedure can also be greatly decreased. By testing our approach on 69 heterocomplex chains, experimental results demonstrate the performance of our approach is indeed superior to current existing methods.

Computational Biology↗

The forecast of the postoperative survival time of patients suffered from non-small cell lung cancer based on PCA and extreme learning machine.

In this paper, a new effective model is proposed to forecast how long the postoperative patients suffered from non-small cell lung cancer will survive. The new effective model which is based on the extreme learning machine (ELM) and principal component analysis (PCA) can forecast successfully the postoperative patients' survival time. The new model obtains better prediction accuracy and faster convergence rate which the model using backpropagation (BP) algorithm and the Levenberg-Marquardt (LM) algorithm to forecast the postoperative patients' survival time can not achieve. Finally, simulation results are given to verify the efficiency and effectiveness of our proposed new model.

Algorithms↗

Network analysis of the protein chain tertiary structures of heterocomplexes.

In this paper, the tertiary structures of protein chains of heterocomplexes were mapped to 2D networks; based on the mapping approach, statistical properties of these networks were systematically studied. Firstly, our experimental results confirmed that the networks derived from protein structures possess small-world properties. Secondly, an interesting relationship between network average degree and the network size was discovered, which was quantified as an empirical function enabling us to estimate the number of residue contacts of the protein chains accurately. Thirdly, by analyzing the average clustering coefficient for nodes having the same degree in the network, it was found that the architectures of the networks and protein structures analyzed are hierarchically organized. Finally, network motifs were detected in the networks which are believed to determine the family or superfamily the networks belong to. The study of protein structures with the new perspective might shed some light on understanding the underlying laws of evolution, function and structures of proteins, and therefore would be complementary to other currently existing methods.

Models, Molecular↗

Predicting protein interaction sites from residue spatial sequence profile and evolution rate.

This paper proposes a novel method that can predict protein interaction sites in heterocomplexes using residue spatial sequence profile and evolution rate approaches. The former represents the information of multiple sequence alignments while the latter corresponds to a residue's evolutionary conservation score based on a phylogenetic tree. Three predictors using a support vector machines algorithm are constructed to predict whether a surface residue is a part of a protein-protein interface. The efficiency and the effectiveness of our proposed approach is verified by its better prediction performance compared with other models. The study is based on a non-redundant data set of heterodimers consisting of 69 protein chains.

Algorithms↗

Non-linear cancer classification using a modified radial basis function classification algorithm.

This paper proposes a modified radial basis function classification algorithm for non-linear cancer classification. In the algorithm, a modified simulated annealing method is developed and combined with the linear least square and gradient paradigms to optimize the structure of the radial basis function (RBF) classifier. The proposed algorithm can be adopted to perform non-linear cancer classification based on gene expression profiles and applied to two microarray data sets involving various human tumor classes: (1) Normal versus colon tumor; (2) acute myeloid leukemia (AML) versus acute lymphoblastic leukemia (ALL). Finally, accuracy and stability for the proposed algorithm are further demonstrated by comparing with the other cancer classification algorithms.

Acute Disease↗

A novel approach to extracting features from motif content and protein composition for protein sequence classification.

This paper presents a novel approach to extracting features from motif content and protein composition for protein sequence classification. First, we formulate a protein sequence as a fixed-dimensional vector using the motif content and protein composition. Then, we further project the vectors into a low-dimensional space by the Principal Component Analysis (PCA) so that they can be represented by a combination of the eigenvectors of the covariance matrix of these vectors. Subsequently, the Genetic Algorithm (GA) is used to extract a subset of biological and functional sequence features from the eigen-space and to optimize the regularization parameter of the Support Vector Machine (SVM) simultaneously. Finally, we utilize the SVM classifiers to classify protein sequences into corresponding families based on the selected feature subsets. In comparison with the existing PSI-BLAST and SVM-pairwise methods, the experiments show the promising results of our approach.

Amino Acid Motifs↗

Prediction of inter-residue contacts map based on genetic algorithm optimized radial basis function neural network and binary input encoding scheme.

Inter-residue contacts map prediction is one of the most important intermediate steps to the protein folding problem. In this paper, we focus on the problem of protein inter-residue contacts map prediction based on neural network technique. Firstly, we use a genetic algorithm (GA) to optimize the radial basis function widths and hidden centers of a radial basis function neural network (RBFNN), then a novel binary encoding scheme is employed to train the network for the purpose of learning and predicting the inter-residue contacts patterns of protein sequences got from the protein data bank (PDB). The experimental evidence indicates the utility of our proposed encoding strategy and GA optimized RBFNN. Moreover, the simulation results demonstrate that the network got a better performance for these proteins, whose residue length falls into the area of (100, 300), and the predicted accuracy with a contact threshold of 7 Angstroms scores higher than the other 3 values with 5, 6, and 8 Angstroms.

Algorithms↗

Extracting mode components in laser intensity distribution by independent component analysis.

With increasingly sophisticated laser applications in industry and science, a reliable method to characterize the intensity distribution of the laser beam has become a more and more important task. However, traditional optic and electronic methods can offer only a laser beam intensity profile but, cannot separate the main mode components in the laser beam intensity distribution. Recently, independent component analysis has been a surging and developing method in which the goal is to find a linear representation of a non-Gaussian data set. Such a linear representation seems to be able to capture the essential structure of a laser beam profile. After assembling image data of a laser spot, we propose a new analytical approach to extract laser beam mode components based on the independent component analysis technique. For noise reduction and laser spot area location, wavelet thresholding, Canny edge detection, and the Hough transform are also used in this method before extracting mode components. Finally, the experimental results show that our approach can separate the principal mode components in a real laser beam efficiently.

Journal Article↗

Antinoise approximation of the lidar signal with wavelet neural networks.

We propose a new, to our knowledge, denoising method for lidar signals based on a regression model and a wavelet neural network (WNN) that permits the regression model not only to have a good wavelet approximation property but also to make a neural network that has a self-learning and adaptive capability for increasing the quality of lidar signals. Specifically, we investigate the performance of the WNN for antinoise approximation of lidar signals by simultaneously addressing simulated and real lidar signals. To clarify the antinoise approximation capability of the WNN for lidar signals, we calculate the atmosphere temperature profile with the real signal processed by the WNN. To show the contrast, we also demonstrate the results of the Monte Carlo moving average method and the finite impulse response filter. Finally, the experimental results show that our proposed approach is significantly superior to the traditional methods.

Journal Article↗

A gene selection algorithm based on the gene regulation probability using maximal likelihood estimation.

A novel gene selection algorithm based on the gene regulation probability is proposed. In this algorithm, a probabilistic model is established to estimate gene regulation probabilities using the maximum likelihood estimation method and then these probabilities are used to select key genes related by class distinction. The application on the leukemia data-set suggests that the defined gene regulation probability can identify the key genes to the acute lymphoblastic leukemia (ALL)/acute myeloid leukemia (AML) class distinction and the result of our proposed algorithm is competitive to those of the previous algorithms.

Acute Disease↗

Zeroing polynomials using modified constrained neural network approach.

This paper proposes new modified constrained learning neural root finders (NRFs) of polynomial constructed by backpropagation network (BPN). The technique is based on the relationships between the roots and the coefficients of polynomial as well as between the root moments and the coefficients of the polynomial. We investigated different resulting constrained learning algorithms (CLAs) based on the variants of the error cost functions (ECFs) in the constrained BPN and derived a new modified CLA (MCLA), and found that the computational complexities of the CLA and the MCLA based on the root-moment method (RMM) are the order of polynomial, and that the MCLA is simpler than the CLA. Further, we also discussed the effects of the different parameters with the CLA and the MCLA on the NRFs. In particular, considering the coefficients of the polynomials involved in practice to possibly be perturbed by noisy sources, thus, we also evaluated and discussed the effects of noises on the two NRFs. Finally, to demonstrate the advantage of our neural approaches over the nonneural ones, a series of simulating experiments are conducted.

Algorithms↗

Fast modular network implementation for support vector machines.

Support vector machines (SVMs) have been extensively used. However, it is known that SVMs face difficulty in solving large complex problems due to the intensive computation involved in their training algorithms, which are at least quadratic with respect to the number of training examples. This paper proposes a new, simple, and efficient network architecture which consists of several SVMs each trained on a small subregion of the whole data sampling space and the same number of simple neural quantizer modules which inhibit the outputs of all the remote SVMs and only allow a single local SVM to fire (produce actual output) at any time. In principle, this region-computing based modular network method can significantly reduce the learning time of SVM algorithms without sacrificing much generalization performance. The experiments on a few real large complex benchmark problems demonstrate that our method can be significantly faster than single SVMs without losing much generalization performance.

Algorithms↗

A novel hybrid GA/RBFNN technique for protein sequences classification.

A novel hybrid genetic algorithm (GA)/radial basis function neural network (RBFNN) technique, which selects features from the protein sequences and trains the RBF neural network simultaneously, is proposed in this paper. Experimental results show that the proposed hybrid GA/RBFNN system outperforms the BLAST and the HMMer.

Algorithms↗

A novel Markov pairwise protein sequence alignment method for sequence comparison.

The Smith-Waterman (SW) algorithm is a typical technique for local sequence alignment in computational biology. However, the SW algorithm does not consider the local behaviours of the amino acids, which may result in loss of some useful information. Inspired by the success of Markov Edit Distance (MED) method, this paper therefore proposes a novel Markov pairwise protein sequence alignment (MPPSA) method that takes the local context dependencies into consideration. The numerical results have shown its superiority to the SW for pairwise protein sequence comparison.

Amino Acid Sequence↗

Prediction of protein secondary structure using improved two-level neural network architecture.

In this paper we propose constructing an improved two-level neural network to predict protein secondary structure. Firstly, we code the whole protein composition information as the inputs to the first-level network besides the evolutionary information. Secondly, we calculate the reliability score for each residue position based on the output of the first-level network, and the role of the second-level network is to take full advantage of the residues with a higher reliability score to impact the neighboring residues with a lower one for improving the whole prediction accuracy. Thirdly, considering it is indeed a problem that the target protein can be lost in the multiple sequence alignment we propose to code single sequence into the second-level network. The experimental results show that our proposed method can efficiently improve the prediction accuracy.

Algorithms↗

A constructive approach for finding arbitrary roots of polynomials by neural networks.

This paper proposes a constructive approach for finding arbitrary (real or complex) roots of arbitrary (real or complex) polynomials by multilayer perceptron network (MLPN) using constrained learning algorithm (CLA), which encodes the a priori information of constraint relations between root moments and coefficients of a polynomial into the usual BP algorithm (BPA). Moreover, the root moment method (RMM) is also simplified into a recursive version so that the computational complexity can be further decreased, which leads the roots of those higher order polynomials to be readily found. In addition, an adaptive learning parameter with the CLA is also proposed in this paper; an initial weight selection method is also given. Finally, several experimental results show that our proposed neural connectionism approaches, with respect to the nonneural ones, are more efficient and feasible in finding the arbitrary roots of arbitrary polynomials.

Models, Statistical↗

Attractability and location of equilibrium point of cellular neural networks with time-varying delays.

This paper presents new theoretical results on global exponential stability of cellular neural networks with time-varying delays. The stability conditions depend on external inputs, connection weights and delays of cellular neural networks. Using these results, global exponential stability of cellular neural networks can be derived, and the estimate for location of equilibrium point can also be obtained. Finally, the simulating results demonstrate the validity and feasibility of our proposed approach.

Animals↗