Search PubMed⌕ Search

Biomedical subjects

Wing H Wong

Publications and source records attributed to Wing H Wong.

At least 19 recordsLinked to original sources

A comparative analysis of genome-wide chromatin immunoprecipitation data for mammalian transcription factors.

Genome-wide location analysis (ChIP-chip, ChIP-PET) is a powerful technique to study mammalian transcriptional regulation. In order to obtain a basic understanding of the location data generated for mammalian transcription factors and potential issues in their analysis, we conducted a comparative study of eight independent ChIP experiments involving six different transcription factors in human and mouse. Our cross-study comparisons, to the best of our knowledge the first to analyze multiple datasets, revealed the importance of carefully chosen genomic controls in the de novo identification of key transcription factor binding motifs, raised issues about the interpretation of ubiquitously occurring sequence motifs, and demonstrated the clustering tendency of protein-binding regions for certain transcription factors.

Animals↗

Recursive SVM feature selection and sample classification for mass-spectrometry and microarray data.

BACKGROUND: Like microarray-based investigations, high-throughput proteomics techniques require machine learning algorithms to identify biomarkers that are informative for biological classification problems. Feature selection and classification algorithms need to be robust to noise and outliers in the data. RESULTS: We developed a recursive support vector machine (R-SVM) algorithm to select important genes/biomarkers for the classification of noisy data. We compared its performance to a similar, state-of-the-art method (SVM recursive feature elimination or SVM-RFE), paying special attention to the ability of recovering the true informative genes/biomarkers and the robustness to outliers in the data. Simulation experiments show that a 5%- approximately 20% improvement over SVM-RFE can be achieved regard to these properties. The SVM-based methods are also compared with a conventional univariate method and their respective strengths and weaknesses are discussed. R-SVM was applied to two sets of SELDI-TOF-MS proteomics data, one from a human breast cancer study and the other from a study on rat liver cirrhosis. Important biomarkers found by the algorithm were validated by follow-up biological experiments. CONCLUSION: The proposed R-SVM method is suitable for analyzing noisy high-throughput proteomics and microarray data and it outperforms SVM-RFE in the robustness to noise and in the ability to recover informative features. The multivariate SVM-based method outperforms the univariate method in the classification performance, but univariate methods can reveal more of the differentially expressed features especially when there are correlations between the features.

Algorithms↗

An improved distance measure between the expression profiles linking co-expression and co-regulation in mouse.

BACKGROUND: Many statistical algorithms combine microarray expression data and genome sequence data to identify transcription factor binding motifs in the low eukaryotic genomes. Finding cis-regulatory elements in higher eukaryote genomes, however, remains a challenge, as searching in the promoter regions of genes with similar expression patterns often fails. The difficulty is partially attributable to the poor performance of the similarity measures for comparing expression profiles. The widely accepted measures are inadequate for distinguishing genes transcribed from distinct regulatory mechanisms in the complicated genomes of higher eukaryotes. RESULTS: By defining the regulatory similarity between a gene pair as the number of common known transcription factor binding motifs in the promoter regions, we compared the performance of several expression distance measures on seven mouse expression data sets. We propose a new distance measure that accounts for both the linear trends and fold-changes of expression across the samples. CONCLUSION: The study reveals that the proposed distance measure for comparing expression profiles enables us to identify genes with large number of common regulatory elements because it reflects the inherent regulatory information better than widely accepted distance measures such as the Pearson's correlation or cosine correlation with or without log transformation.

Algorithms↗

Is the future biology Shakespearean or Newtonian?

"Cells do not care about mathematics" thus concluded a biologist friend after a discussion on the future of biology. And indeed, why should they care? But if we exchange the word "cell" with "rock", "Moon" or "electrons", do we have to change the sentence also? Starting from this line of thought, we review some recent developments in understanding the stochastic behavior of biological systems. We emphasize the importance of a molecular Signal Generator in the study of genetic networks.

Biology↗

Expression profiling of serous low malignant potential, low-grade, and high-grade tumors of the ovary.

Papillary serous low malignant potential (LMP) tumors are characterized by malignant features and metastatic potential yet display a benign clinical course. The role of LMP tumors in the development of invasive epithelial cancer of the ovary is not clearly defined. The aim of this study is to determine the relationships among LMP tumors and invasive ovarian cancers and identify genes contributing to their phenotypes. Affymetrix U133 Plus 2.0 microarrays (Santa Clara, CA) were used to interrogate 80 microdissected serous LMP tumors and invasive ovarian malignancies along with 10 ovarian surface epithelium (OSE) brushings. Gene expression profiles for each tumor class were used to complete unsupervised hierarchical clustering analyses and identify differentially expressed genes contributing to these associations. Unsupervised hierarchical clustering analysis revealed a distinct separation between clusters containing borderline and high-grade lesions. The majority of low-grade tumors clustered with LMP tumors. Comparing OSE with high-grade and LMP expression profiles revealed enhanced expression of genes linked to cell proliferation, chromosomal instability, and epigenetic silencing in high-grade cancers, whereas LMP tumors displayed activated p53 signaling. The expression profiles of LMP, low-grade, and high-grade papillary serous ovarian carcinomas suggest that LMP tumors are distinct from high-grade cancers; however, they are remarkably similar to low-grade cancers. Prominent expression of p53 pathway members may play an important role in the LMP tumor phenotype.

Carcinoma, Papillary↗

Reliable prediction of transcription factor binding sites by phylogenetic verification.

We present a statistical methodology that largely improves the accuracy in computational predictions of transcription factor (TF) binding sites in eukaryote genomes. This method models the cross-species conservation of binding sites without relying on accurate sequence alignment. It can be coupled with any motif-finding algorithm that searches for overrepresented sequence motifs in individual species and can increase the accuracy of the coupled motif-finding algorithm. Because this method is capable of accurately detecting TF binding sites, it also enhances our ability to predict the cis-regulatory modules. We applied this method on the published chromatin immunoprecipitation (ChIP)-chip data in Saccharomyces cerevisiae and found that its sensitivity and specificity are 9% and 14% higher than those of two recent methods. We also recovered almost all of the previously verified TF binding sites and made predictions on the cis-regulatory elements that govern the tight regulation of ribosomal protein genes in 13 eukaryote species (2 plants, 4 yeasts, 2 worms, 2 insects, and 3 mammals). These results give insights to the transcriptional regulation in eukaryotic organisms.

Base Sequence↗

mSin3A corepressor regulates diverse transcriptional networks governing normal and neoplastic growth and survival.

mSin3A is a core component of a large multiprotein corepressor complex with associated histone deacetylase (HDAC) enzymatic activity. Physical interactions of mSin3A with many sequence-specific transcription factors has linked the mSin3A corepressor complex to the regulation of diverse signaling pathways and associated biological processes. To dissect the complex nature of mSin3A's actions, we monitored the impact of conditional mSin3A deletion on the developmental, cell biological, and transcriptional levels. mSin3A was shown to play an essential role in early embryonic development and in the proliferation and survival of primary, immortalized, and transformed cells. Genetic and biochemical analyses established a role for mSin3A/HDAC in p53 deacetylation and activation, although genetic deletion of p53 was not sufficient to attenuate the mSin3A null cell lethal phenotype. Consistent with mSin3A's broad biological activities beyond regulation of the p53 pathway, time-course gene expression profiling following mSin3A deletion revealed deregulation of genes involved in cell cycle regulation, DNA replication, DNA repair, apoptosis, chromatin modifications, and mitochondrial metabolism. Computational analysis of the mSin3A transcriptome using a knowledge-based database revealed several nodal points through which mSin3A influences gene expression, including the Myc-Mad, E2F, and p53 transcriptional networks. Further validation of these nodes derived from in silico promoter analysis showing enrichment for Myc-Mad, E2F, and p53 cis-regulatory elements in regulatory regions of up-regulated genes following mSin3A depletion. Significantly, in silico promoter analyses also revealed specific cis-regulatory elements binding the transcriptional activator Stat and the ISWI ATP-dependent nucleosome remodeling factor Falz, thereby expanding further the mSin3A network of regulatory factors. Together, these integrated genetic, biochemical, and computational studies demonstrate the involvement of mSin3A in the regulation of diverse pathways governing many aspects of normal and neoplastic growth and survival and provide an experimental framework for the analysis of essential genes with diverse biological functions.

Animals↗

Sampling motifs on phylogenetic trees.

We present a method to find motifs by simultaneously using the overrepresentation property and the evolutionary conservation property of motifs. This method is applicable to divergent species where alignment is unreliable, which overcomes a major limitation of the current methods. The method has been applied to search regulatory motifs in four yeast species based on ChIP-chip data in Saccharomyces cerevisiae and obtained 20% higher accuracy than the best current methods. We also discovered cis-regulatory elements that govern the tight regulation of ribosomal protein genes in two distantly related insects by using this method. These results demonstrate that our method will be useful for the extraction of regulatory signals in multiple genomes.

Base Sequence↗

The use of oscillatory signals in the study of genetic networks.

The structure of a genetic network is uncovered by studying its response to external stimuli (input signals). We present a theory of propagation of an input signal through a linear stochastic genetic network. We found that there are important advantages in using oscillatory signals over step or impulse signals and that the system may enter into a pure fluctuation resonance for a specific input frequency.

Gene Expression Regulation↗

A boosting approach for motif modeling using ChIP-chip data.

MOTIVATION: Building an accurate binding model for a transcription factor (TF) is essential to differentiate its true binding targets from those spurious ones. This is an important step toward understanding gene regulation. RESULTS: This paper describes a boosting approach to modeling TF-DNA binding. Different from the widely used weight matrix model, which predicts TF-DNA binding based on a linear combination of position-specific contributions, our approach builds a TF binding classifier by combining a set of weight matrix based classifiers, thus yielding a non-linear binding decision rule. The proposed approach was applied to the ChIP-chip data of Saccharomyces cerevisiae. When compared with the weight matrix method, our new approach showed significant improvements on the specificity in a majority of cases.

Algorithms↗

GeneNotes--a novel information management software for biologists.

BACKGROUND: Collecting and managing information is a challenging task in a genome-wide profiling research project. Most databases and online computational tools require a direct human involvement. Information and computational results are presented in various multimedia formats (e.g., text, image, PDF, word files, etc.), many of which cannot be automatically processed by computers in biologically meaningful ways. In addition, the quality of computational results is far from perfect and requires nontrivial manual examination. The timely selection, integration and interpretation of heterogeneous biological information still heavily rely on the sensibility of biologists. Biologists often feel overwhelmed by the huge amount of and the great diversity of distributed heterogeneous biological information. DESCRIPTION: We developed an information management application called GeneNotes. GeneNotes is the first application that allows users to collect and manage multimedia biological information about genes/ESTs. GeneNotes provides an integrated environment for users to surf the Internet, collect notes for genes/ESTs, and retrieve notes. GeneNotes is supported by a server that integrates gene annotations from many major databases (e.g., HGNC, MGI, etc.). GeneNotes uses the integrated gene annotations to (a) identify genes given various types of gene IDs (e.g., RefSeq ID, GenBank ID, etc.), and (b) provide quick views of genes. GeneNotes is free for academic usage. The program and the tutorials are available at: http://bayes.fas.harvard.edu/genenotes/. CONCLUSIONS: GeneNotes provides a novel human-computer interface to assist researchers to collect and manage biological information. It also provides a platform for studying how users behave when they manipulate biological information. The results of such study can lead to innovation of more intelligent human-computer interfaces that greatly shorten the cycle of biology research.

Biology↗

Tight clustering: a resampling-based approach for identifying stable and tight patterns in data.

In this article, we propose a method for clustering that produces tight and stable clusters without forcing all points into clusters. The methodology is general but was initially motivated from cluster analysis of microarray experiments. Most current algorithms aim to assign all genes into clusters. For many biological studies, however, we are mainly interested in identifying the most informative, tight, and stable clusters of sizes, say, 20-60 genes for further investigation. We want to avoid the contamination of tightly regulated expression patterns of biologically relevant genes due to other genes whose expressions are only loosely compatible with these patterns. "Tight clustering" has been developed specifically to address this problem. It applies K-means clustering as an intermediate clustering engine. Early truncation of a hierarchical clustering tree is used to overcome the local minimum problem in K-means clustering. The tightest and most stable clusters are identified in a sequential manner through an analysis of the tendency of genes to be grouped together under repeated resampling. We validated this method in a simulated example and applied it to analyze a set of expression profiles in the study of embryonic stem cells.

Algorithms↗

Role of epidermal growth factor receptor signaling in RAS-driven melanoma.

The identification of essential genetic elements in pathways governing the maintenance of fully established tumors is critical to the development of effective antioncologic agents. Previous studies revealed an essential role for H-RAS(V12G) in melanoma maintenance in an inducible transgenic model. Here, we sought to define the molecular basis for RAS-dependent tumor maintenance through determination of the H-RAS(V12G)-directed transcriptional program and subsequent functional validation of potential signaling surrogates. The extinction of H-RAS(V12G) expression in established tumors was associated with alterations in the expression of proliferative, antiapoptotic, and angiogenic genes, a profile consistent with the observed phenotype of tumor cell proliferative arrest and death and endothelial cell apoptosis during tumor regression. In particular, these melanomas displayed a prominent RAS-dependent regulation of the epidermal growth factor (EGF) family, leading to establishment of an EGF receptor signaling loop. Genetic complementation and interference studies demonstrated that this signaling loop is essential to H-RAS(V12G)-directed tumorigenesis. Thus, this inducible tumor model system permits the identification and validation of alternative points of therapeutic intervention without neutralization of the primary genetic lesion.

Animals↗

CisModule: de novo discovery of cis-regulatory modules by hierarchical mixture modeling.

The regulatory information for a eukaryotic gene is encoded in cis-regulatory modules. The binding sites for a set of interacting transcription factors have the tendency to colocalize to the same modules. Current de novo motif discovery methods do not take advantage of this knowledge. We propose a hierarchical mixture approach to model the cis-regulatory module structure. Based on the model, a new de novo motif-module discovery algorithm, CisModule, is developed for the Bayesian inference of module locations and within-module motif sites. Dynamic programming-like recursions are developed to reduce the computational complexity from exponential to linear in sequence length. By using both simulated and real data sets, we demonstrate that CisModule is not only accurate in predicting modules but also more sensitive in detecting motif patterns and binding sites than standard motif discovery methods are.

Algorithms↗

Clustering analysis of SAGE data using a Poisson approach.

Serial analysis of gene expression (SAGE) data have been poorly exploited by clustering analysis owing to the lack of appropriate statistical methods that consider their specific properties. We modeled SAGE data by Poisson statistics and developed two Poisson-based distances. Their application to simulated and experimental mouse retina data show that the Poisson-based distances are more appropriate and reliable for analyzing SAGE data compared to other commonly used distances or similarity measures such as Pearson correlation or Euclidean distance.

Animals↗

Genomic analysis of mouse retinal development.

The vertebrate retina is comprised of seven major cell types that are generated in overlapping but well-defined intervals. To identify genes that might regulate retinal development, gene expression in the developing retina was profiled at multiple time points using serial analysis of gene expression (SAGE). The expression patterns of 1,051 genes that showed developmentally dynamic expression by SAGE were investigated using in situ hybridization. A molecular atlas of gene expression in the developing and mature retina was thereby constructed, along with a taxonomic classification of developmental gene expression patterns. Genes were identified that label both temporal and spatial subsets of mitotic progenitor cells. For each developing and mature major retinal cell type, genes selectively expressed in that cell type were identified. The gene expression profiles of retinal Müller glia and mitotic progenitor cells were found to be highly similar, suggesting that Müller glia might serve to produce multiple retinal cell types under the right conditions. In addition, multiple transcripts that were evolutionarily conserved that did not appear to encode open reading frames of more than 100 amino acids in length ("noncoding RNAs") were found to be dynamically and specifically expressed in developing and mature retinal cell types. Finally, many photoreceptor-enriched genes that mapped to chromosomal intervals containing retinal disease genes were identified. These data serve as a starting point for functional investigations of the roles of these genes in retinal development and physiology.

Animals↗

Molecular diversity of astrocytes with implications for neurological disorders.

The astrocyte represents the most abundant yet least understood cell type of the CNS. Here, we use a stringent experimental strategy to molecularly define the astrocyte lineage by integrating microarray datasets across several in vitro model systems of astrocyte differentiation, primary astrocyte cultures, and various astrocyterich CNS structures. The intersection of astrocyte data sets, coupled with the application of nonastrocytic exclusion filters, yielded many astrocyte-specific genes possessing strikingly varied patterns of regional CNS expression. Annotation of these astrocyte-specific genes provides direct molecular documentation of the diverse physiological roles of the astrocyte lineage. This global perspective in the normal brain also provides a framework for how astrocytes may participate in the pathogenesis of common neurological disorders like Alzheimer's disease, Parkinson's disease, stroke, epilepsy, and primary brain tumors.

Animals↗

Identification of DNA copy number changes in microdissected serous ovarian cancer tissue using a cDNA microarray platform.

We have established a method for using a cDNA array platform in combination with degenerate oligonucleotide primer polymerase chain reaction (DOP-PCR) and taramide signal amplification (TSA) to identify DNA copy number abnormalities (CNA) in cancer cell lines and cancer cells procured with laser-based microdissection. To determine the sensitivity and specificity for detecting single-copy gain and loss, receiver-operator curve analysis was performed on hybridization signal ratios generated from non-DOP and DOP amplified female and male DNA using a 10,816-element cDNA microarray. A cutoff value of 1.12 and 1.07 average signal ratio for X-chromosomal genes versus autosomal genes provided a sensitivity and specificity of 50 and 79%, respectively, for non-DOP amplified DNA and a sensitivity and specificity of 50 and 72%, respectively, for DOP amplified DNA. We used this approach to identify DNA copy number abnormalities in the ovarian cancer cell line OVCA633, which has previously been shown to have 12p amplification. Transcription profiling of OVCA633 was also performed. Two amplified and overexpressed genes located on 12p11, KRAS2 and LRMP, were identified; these were validated with quantitative real-time PCR. Subsequently, the same approach was used to identify CNAs and gene expression alterations in 11 microdissected serous ovarian adenocarcinoma cases. Validated data revealed amplification and overexpression of ERBB3 and FOS and deletion and underexpression of KRT6 and APXL in more than 50% of the tissue samples. These results show the feasibility of using the cDNA array platform to identify changes in DNA and mRNA copy number simultaneously in microdissected tumor tissues.

Adenocarcinoma↗