Search PubMed⌕ Search

Biomedical subjects

Alan Wee-Chung Liew

Publications and source records attributed to Alan Wee-Chung Liew.

5 recordsLinked to original sources

Microarray missing data imputation based on a set theoretic framework and biological knowledge.

Gene expressions measured using microarrays usually suffer from the missing value problem. However, in many data analysis methods, a complete data matrix is required. Although existing missing value imputation algorithms have shown good performance to deal with missing values, they also have their limitations. For example, some algorithms have good performance only when strong local correlation exists in data while some provide the best estimate when data is dominated by global structure. In addition, these algorithms do not take into account any biological constraint in their imputation. In this paper, we propose a set theoretic framework based on projection onto convex sets (POCS) for missing data imputation. POCS allows us to incorporate different types of a priori knowledge about missing values into the estimation process. The main idea of POCS is to formulate every piece of prior knowledge into a corresponding convex set and then use a convergence-guaranteed iterative procedure to obtain a solution in the intersection of all these sets. In this work, we design several convex sets, taking into consideration the biological characteristic of the data: the first set mainly exploit the local correlation structure among genes in microarray data, while the second set captures the global correlation structure among arrays. The third set (actually a series of sets) exploits the biological phenomenon of synchronization loss in microarray experiments. In cyclic systems, synchronization loss is a common phenomenon and we construct a series of sets based on this phenomenon for our POCS imputation algorithm. Experiments show that our algorithm can achieve a significant reduction of error compared to the KNNimpute, SVDimpute and LSimpute methods.

Algorithms↗

Dominant spectral component analysis for transcriptional regulations using microarray time-series data.

MOTIVATION: Microarray time-series data provides us a possible means for identification of transcriptional regulation relationships among genes. Currently, the most commonly used method in determining whether or not two genes have a potential regulatory relationship is to measure their expressional similarity using Pearson's correlation coefficient. Although this traditional correlation method has been successfully applied to find functionally correlated genes, it does have many limitations. In the hope of overcoming such circumstances and getting more insights into the transcriptional regulatory issue, we propose an autoregressive (AR)-based technique for detection of potential regulated gene pairs from time-series microarray measurements. RESULTS: We use the well-known AR modeling technique to characterize temporal gene expression data from the Spellman's alpha-synchronized yeast cell-cycle experiment. In this method, time-series expression profiles are decomposed into spectral components and correlations between profiles are then computed in a component-wise sense. We show how these component-wise correlations reveal possible regulatory relationships. Our technique is applied on known transcriptional regulations and is able to identify many of those missed by the traditional correlation method.

Algorithms↗

Cluster analysis of gene expression data based on self-splitting and merging competitive learning.

Cluster analysis of gene expression data from a cDNA microarray is useful for identifying biologically relevant groups of genes. However, finding the natural clusters in the data and estimating the correct number of clusters are still two largely unsolved problems. In this paper, we propose a new clustering framework that is able to address both these problems. By using the one-prototype-take-one-cluster (OPTOC) competitive learning paradigm, the proposed algorithm can find natural clusters in the input data, and the clustering solution is not sensitive to initialization. In order to estimate the number of distinct clusters in the data, we propose a cluster splitting and merging strategy. We have applied the new algorithm to simulated gene expression data for which the correct distribution of genes over clusters is known a priori. The results show that the proposed algorithm can find natural clusters and give the correct number of clusters. The algorithm has also been tested on real gene expression changes during yeast cell cycle, for which the fundamental patterns of gene expression and assignment of genes to clusters are well understood from numerous previous studies. Comparative studies with several clustering algorithms illustrate the effectiveness of our method.

Algorithms↗

Classification of short human exons and introns based on statistical features.

The classification of human gene sequences into exons and introns is a difficult problem in DNA sequence analysis. In this paper, we define a set of features, called the simple Z (SZ) features, which is derived from the Z-curve features for the recognition of human exons and introns. The classification results show that SZ features, while fewer in numbers (three in total), can preserve the high recognition rate of the original nine Z-curve features. Since the size of SZ features is one-third of the Z-curve features, the dimensionality of the feature space is much smaller, and better recognition efficiency is achieved. If the stop codon feature is used together with the three SZ features, a recognition rate of up to 92% for short sequences of length <140 bp can be obtained.

Algorithms↗

An adaptive spatial fuzzy clustering algorithm for 3-D MR image segmentation.

An adaptive spatial fuzzy c-means clustering algorithm is presented in this paper for the segmentation of three-dimensional (3-D) magnetic resonance (MR) images. The input images may be corrupted by noise and intensity nonuniformity (INU) artifact. The proposed algorithm takes into account the spatial continuity constraints by using a dissimilarity index that allows spatial interactions between image voxels. The local spatial continuity constraint reduces the noise effect and the classification ambiguity. The INU artifact is formulated as a multiplicative bias field affecting the true MR imaging signal. By modeling the log bias field as a stack of smoothing B-spline surfaces, with continuity enforced across slices, the computation of the 3-D bias field reduces to that of finding the B-spline coefficients, which can be obtained using a computationally efficient two-stage algorithm. The efficacy of the proposed algorithm is demonstrated by extensive segmentation experiments using both simulated and real MR images and by comparison with other published algorithms.

Algorithms↗