Search PubMed⌕ Search

Biomedical subjects

Wing Hung Wong

Publications and source records attributed to Wing Hung Wong.

15 recordsLinked to original sources

TileMap: create chromosomal map of tiling array hybridizations.

MOTIVATION: Tiling array is a new type of microarray that can be used to survey genomic transcriptional activities and transcription factor binding sites at high resolution. The goal of this paper is to develop effective statistical tools to identify genomic loci that show transcriptional or protein binding patterns of interest. RESULTS: A two-step approach is proposed and is implemented in TileMap. In the first step, a test-statistic is computed for each probe based on a hierarchical empirical Bayes model. In the second step, the test-statistics of probes within a genomic region are used to infer whether the region is of interest or not. Hierarchical empirical Bayes model shrinks variance estimates and increases sensitivity of the analysis. It allows complex multiple sample comparisons that are essential for the study of temporal and spatial patterns of hybridization across different experimental conditions. Neighboring probes are combined through a moving average method (MA) or a hidden Markov model (HMM). Unbalanced mixture subtraction is proposed to provide approximate estimates of false discovery rate for MA and model parameters for HMM. AVAILABILITY: TileMap is freely available at http://biogibbs.stanford.edu/~jihk/TileMap/index.htm. SUPPLEMENTARY INFORMATION: http://biogibbs.stanford.edu/~jihk/TileMap/index.htm (includes coloured versions of all figures).

Algorithms↗

A small-molecule inhibitor of Mps1 blocks the spindle-checkpoint response to a lack of tension on mitotic chromosomes.

The spindle checkpoint prevents chromosome loss by preventing chromosome segregation in cells with improperly attached chromosomes [1, 2 and 3]. The checkpoint senses defects in the attachment of chromosomes to the mitotic spindle [4] and the tension exerted on chromosomes by spindle forces in mitosis [5, 6 and 7]. Because many cancers have defects in chromosome segregation, this checkpoint may be required for survival of tumor cells and may be a target for chemotherapy. We performed a phenotype-based chemical-genetic screen in budding yeast and identified an inhibitor of the spindle checkpoint, called cincreasin. We used a genome-wide collection of yeast gene-deletion strains and traditional genetic and biochemical analysis to show that the target of cincreasin is Mps1, a protein kinase required for checkpoint function [8]. Despite the requirement for Mps1 for sensing both the lack of microtubule attachment and tension at kinetochores, we find concentrations of cincreasin that selectively inhibit the tension-sensitive branch of the spindle checkpoint. At these concentrations, cincreasin causes lethal chromosome missegregation in mutants that display chromosomal instability. Our results demonstrate that Mps1 can be exploited as a target and that inhibiting the tension-sensitive branch of the spindle checkpoint may be a way of selectively killing cancer cells that display chromosomal instability.

Bromobenzenes↗

Comparative linkage analysis and visualization of high-density oligonucleotide SNP array data.

BACKGROUND: The identification of disease-associated genes using single nucleotide polymorphisms (SNPs) has been increasingly reported. In particular, the Affymetrix Mapping 10 K SNP microarray platform uses one PCR primer to amplify the DNA samples and determine the genotype of more than 10,000 SNPs in the human genome. This provides the opportunity for large scale, rapid and cost-effective genotyping assays for linkage analysis. However, the analysis of such datasets is nontrivial because of the large number of markers, and visualizing the linkage scores in the context of genome maps remains less automated using the current linkage analysis software packages. For example, the haplotyping results are commonly represented in the text format. RESULTS: Here we report the development of a novel software tool called CompareLinkage for automated formatting of the Affymetrix Mapping 10 K genotype data into the "Linkage" format and the subsequent analysis with multi-point linkage software programs such as Merlin and Allegro. The new software has the ability to visualize the results for all these programs in dChip in the context of genome annotations and cytoband information. In addition we implemented a variant of the Lander-Green algorithm in the dChipLinkage module of dChip software (V1.3) to perform parametric linkage analysis and haplotyping of SNP array data. These functions are integrated with the existing modules of dChip to visualize SNP genotype data together with LOD score curves. We have analyzed three families with recessive and dominant diseases using the new software programs and the comparison results are presented and discussed. CONCLUSIONS: The CompareLinkage and dChipLinkage software packages are freely available. They provide the visualization tools for high-density oligonucleotide SNP array data, as well as the automated functions for formatting SNP array data for the linkage analysis programs Merlin and Allegro and calling these programs for linkage analysis. The results can be visualized in dChip in the context of genes and cytobands. In addition, a variant of the Lander-Green algorithm is provided that allows parametric linkage analysis and haplotyping.

Family Health↗

Functional annotation and network reconstruction through cross-platform integration of microarray data.

The rapid accumulation of microarray data translates into a need for methods to effectively integrate data generated with different platforms. Here we introduce an approach, 2(nd)-order expression analysis, that addresses this challenge by first extracting expression patterns as meta-information from each data set (1(st)-order expression analysis) and then analyzing them across multiple data sets. Using yeast as a model system, we demonstrate two distinct advantages of our approach: we can identify genes of the same function yet without coexpression patterns and we can elucidate the cooperativities between transcription factors for regulatory network reconstruction by overcoming a key obstacle, namely the quantification of activities of transcription factors. Experiments reported in the literature and performed in our lab support a significant number of our predictions.

Algorithms↗

HumanUpstream and MouseUpstream: databases of promoter sequences in the human and mouse genomes.

Large-scale genome annotations, based largely on gene prediction programs, may be inaccurate in their predictions of transcription start sites, so that the identification of promoter regions remains unreliable. Here we focus on the identification of reliable gene promoter regions, critical to the understanding of transcriptional regulation. We report the construction of databases of upstream sequences Human Upstream and Mouse Upstream based on information from both the human and mouse genomes and the database of expressed sequence tags (dbEST). Using the ENSEMBL generic genome annotation system, our approach allows more reliable identification of transcript start sites, and therefore extraction of more reliable promoters regions. The Human Upstream and Human Upstream databases are available free of charge.

Animals↗

Expression of heat shock proteins and heat shock protein messenger ribonucleic acid in human prostate carcinoma in vitro and in tumors in vivo.

Heat shock proteins (HSPs) are thought to play a role in the development of cancer and to modulate tumor response to cytotoxic therapy. In this study, we have examined the expression of hsf and HSP genes in normal human prostate epithelial cells and a range of prostate carcinoma cell lines derived from human tumors. We have observed elevated expressions of HSF1, HSP60, and HSP70 in the aggressively malignant cell lines PC-3, DU-145, and CA-HPV-10. Elevated HSP expression in cancer cell lines appeared to be regulated at the post-messenger ribonucleic acid (mRNA) levels, as indicated by gene chip microarray studies, which indicated little difference in heat shock factor (HSF) or HSP mRNA expression between the normal and malignant prostate cell lines. When we compared the expression patterns of constitutive HSP genes between PC-3 prostate carcinoma cells growing as monolayers in vitro and as tumor xenografts growing in nude mice in vivo, we found a marked reduction in expression of a wide spectrum of the HSPs in PC-3 tumors. This decreased HSP expression pattern in tumors may underlie the increased sensitivity to heat shock of PC-3 tumors. However, the induction by heat shock of HSP genes was not markedly altered by growth in the tumor microenvironment, and HSP40, HSP70, and HSP110 were expressed abundantly after stress in each growth condition. Our experiments indicate therefore that HSF and HSP levels are elevated in the more highly malignant prostate carcinoma cells and also show the dominant nature of the heat shock-induced gene expression, leading to abundant HSP induction in vitro or in vivo.

Animals↗

dChipSNP: significance curve and clustering of SNP-array-based loss-of-heterozygosity data.

MOTIVATION: Oligonucleotide microarrays allow genotyping of thousands of single-nucleotide polymorphisms (SNPs) in parallel. Recently, this technology has been applied to loss-of-heterozygosity (LOH) analysis of paired normal and tumor samples. However, methods and software for analyzing such data are not fully developed. RESULT: Here, we report automated methods for pooling SNP array replicates to make LOH calls, visualizing SNP and LOH data along chromosomes in the context of genes and cytobands, making statistical inference to identify shared LOH regions, clustering samples based on LOH profiles and correlating the clustering results to clinical variables. Application of these methods to prostate and breast cancer datasets generates biologically important results. AVAILABILITY: The software module dChipSNP implementing these methods is available at http://biosun1.harvard.edu/complab/dchip/snp/ SUPPLEMENTARY INFORMATION: The breast cancer data are provided by Andrea L. Richardson, Zhigang C. Wang and James D. Iglehart.

Algorithms↗

Estimation of genotype error rate using samples with pedigree information--an application on the GeneChip Mapping 10K array.

Currently, most analytical methods assume all observed genotypes are correct; however, it is clear that errors may reduce statistical power or bias inference in genetic studies. We propose procedures for estimating error rate in genetic analysis and apply them to study the GeneChip Mapping 10K array, which is a technology that has recently become available and allows researchers to survey over 10,000 SNPs in a single assay. We employed a strategy to estimate the genotype error rate in pedigree data. First, the "dose-response" reference curve between error rate and the observable error number were derived by simulation, conditional on given pedigree structures and genotypes. Second, the error rate was estimated by calibrating the number of observed errors in real data to the reference curve. We evaluated the performance of this method by simulation study and applied it to a data set of 30 pedigrees genotyped using the GeneChip Mapping 10K array. This method performed favorably in all scenarios we surveyed. The dose-response reference curve was monotone and almost linear with a large slope. The method was able to estimate accurately the error rate under various pedigree structures and error models and under heterogeneous error rates. Using this method, we found that the average genotyping error rate of the GeneChip Mapping 10K array was about 0.1%. Our method provides a quick and unbiased solution to address the genotype error rate in pedigree data. It behaves well in a wide range of settings and can be easily applied in other genetic projects. The robust estimation of genotyping error rate allows us to estimate power and sample size and conduct unbiased genetic tests. The GeneChip Mapping 10K array has a low overall error rate, which is consistent with the results obtained from alternative genotyping assays.

Computer Simulation↗

Integrated analysis of microarray data and gene function information.

Microarray data should be interpreted in the context of existing biological knowledge. Here we present integrated analysis of microarray data and gene function classification data using homogeneity analysis. Homogeneity analysis is a graphical multivariate statistical method for analyzing categorical data. It converts categorical data into graphical display. By simultaneously quantifying the microarray-derived gene groups and gene function categories, it captures the complex relations between biological information derived from microarray data and the existing knowledge about the gene function. Thus, homogeneity analysis provides a mathematical framework for integrating the analysis of microarray data and the existing biological knowledge.

Algorithms↗

Array comparative genome hybridization for tumor classification and gene discovery in mouse models of malignant melanoma.

Chromosomal numerical aberrations (CNAs), particularly regional amplifications and deletions, are a hallmark of solid tumor genomes. These genomic alterations carry the potential to convey etiologic and clinical significance by virtue of their clonality within a tumor cell population, their distinctive patterns in relation to tumor staging, and their recurrence across different tumor types. In this study, we showed that array-based comparative genomic hybridization (CGH) analysis of genome-wide CNAs can classify tumors on the basis of differing etiologies and provide mechanistic insights to specific biological processes. In a RAS-induced p19(Arf-/-) mouse model that experienced accelerated melanoma formation after UV exposure, array-CGH analysis was effective in distinguishing phenotypically identical melanomas that differed solely by previous UV exposure. Moreover, classification by array-CGH identified key CNAs unique to each class, including amplification of cyclin-dependent kinase 6 in UV-treated cohort, a finding consistent with our recent report that UVB targets components of the p16(INK4a)-cyclin-dependent kinase-RB pathway in melanoma genesis (K. Kannan, et al., Proc. Natl. Acad. Sci. USA, 21: 2003). These results are the first to establish the utility of array-CGH as a means of etiology-based tumor classification in genetically defined cancer-prone models.

Animals↗

ChipInfo: Software for extracting gene annotation and gene ontology information for microarray analysis.

To date, assembling comprehensive annotation information for all probe sets of any Affymetrix microarrays remains a time-consuming, error-prone and challenging task. ChipInfo is designed for retrieving annotation information from online databases such as NetAffx and Gene Ontology and organizing such information into easily interpretable tabular format outputs. As companion software to dChip and GoSurfer, ChipInfo enables users to independently update the information resource files of these software packages. It also has functions for computing related summary statistics of probe sets and Gene Ontology terms. ChipInfo is available at http://biosun1.harvard.edu/complab/chipinfo/.

DNA Probes↗

Novel mechanisms of T-cell and dendritic cell activation revealed by profiling of psoriasis on the 63,100-element oligonucleotide array.

A global picture of gene expression in the common immune-mediated skin disease, psoriasis, was obtained by interrogating the full set of Affymetrix GeneChips with psoriatic and control skin samples. We identified 1,338 genes with potential roles in psoriasis pathogenesis/maintenance and revealed many perturbed biological processes. A novel method for identifying transcription factor binding sites was also developed and applied to this dataset. Many of the identified sites are known to be involved in immune response and proliferation. An in-depth study of immune system genes revealed the presence of many regulating cytokines and chemokines within involved skin, and markers of dendritic cell (DC) activation in uninvolved skin. The combination of many CCR7+ T cells, DCs, and regulating chemokines in psoriatic lesions, together with the detection of DC activation markers in nonlesional skin, strongly suggests that the spatial organization of T cells and DCs could sustain chronic T-cell activation and persistence within focal skin regions.

Cell Separation↗

Transitive functional annotation by shortest-path analysis of gene expression data.

Current methods for the functional analysis of microarray gene expression data make the implicit assumption that genes with similar expression profiles have similar functions in cells. However, among genes involved in the same biological pathway, not all gene pairs show high expression similarity. Here, we propose that transitive expression similarity among genes can be used as an important attribute to link genes of the same biological pathway. Based on large-scale yeast microarray expression data, we use the shortest-path analysis to identify transitive genes between two given genes from the same biological process. We find that not only functionally related genes with correlated expression profiles are identified but also those without. In the latter case, we compare our method to hierarchical clustering, and show that our method can reveal functional relationships among genes in a more precise manner. Finally, we show that our method can be used to reliably predict the function of unknown genes from known genes lying on the same shortest path. We assigned functions for 146 yeast genes that are considered as unknown by the Saccharomyces Genome Database and by the Yeast Proteome Database. These genes constitute around 5% of the unknown yeast ORFome.

Cell Nucleus↗

Recombinatoric exploration of novel folded structures: a heteropolymer-based model of protein evolutionary landscapes.

The role of recombination in evolution is compared with that of point mutations (substitutions) in the context of a simple, polymer physics-based model mapping between sequence (genotype) and conformational (phenotype) spaces. Crossovers and point mutations of lattice chains with a hydrophobic polar code are investigated. Sequences encoding for a single ground-state conformation are considered viable and used as model proteins. Point mutations lead to diffusive walks on the evolutionary landscape, whereas crossovers can "tunnel" through barriers of diminished fitness. The degree to which crossovers allow for more efficient sequence and structural exploration depends on the relative rates of point mutations versus that of crossovers and the dispersion in fitness that characterizes the ruggedness of the evolutionary landscape. The probability that a crossover between a pair of viable sequences results in viable sequences is an order of magnitude higher than random, implying that a sequence's overall propensity to encode uniquely is embodied partially in local signals. Consistent with this observation, certain hydrophobicity patterns are significantly more favored than others among fragments (i.e., subsequences) of sequences that encode uniquely, and examples reminiscent of autonomous folding units in real proteins are found. The number of structures explored by both crossovers and point mutations is always substantially larger than that via point mutations alone, but the corresponding numbers of sequences explored can be comparable when the evolutionary landscape is rugged. Efficient structural exploration requires intermediate nonextreme ratios between point-mutation and crossover rates.

Biophysical Phenomena↗

IN02, a positive regulator of lipid biosynthesis, is essential for the formation of inducible membranes in yeast.

Expression of the 180-kDa canine ribosome receptor in Saccharomyces cerevisiae leads to the accumulation of ER-like membranes. Gene expression patterns in strains expressing various forms of p180, each of which gives rise to unique membrane morphologies, were surveyed by microarray analysis. Several genes whose products regulate phospholipid biosynthesis were determined by Northern blotting to be differentially expressed in all strains that undergo membrane proliferation. Of these, the INO2 gene product was found to be essential for formation of p180-inducible membranes. Expression of p180 in ino2Delta cells failed to give rise to the p180-induced membrane proliferation seen in wild-type cells, whereas p180 expression in ino4Delta cells gave rise to membranes indistinguishable from wild type. Thus, Ino2p is required for the formation of p180-induced membranes and, in this case, appears to be functional in the absence of its putative binding partner, Ino4p.

Basic Helix-Loop-Helix Proteins↗