Search PubMed⌕ Search

Biomedical subjects

Hongkai Ji

Publications and source records attributed to Hongkai Ji.

6 recordsLinked to original sources

Novel Predictive Spatial Biomarker in Non-Small Cell Lung Carcinoma: The Diversity of Niches Unlocking Treatment Sensitivity (DONUTS).

Probabilistic spatial modelling techniques developed on large-scale tumor-immune Atlases (~35M individually mapped cells; 50,000 high power fields) were used to characterize predictive features of treatment-responsive lung cancer. We identified CD8+FoxP3+ cell density as a robust pre-treatment biomarker for outcomes across disease stages and therapy types. In parallel, single-cell RNAseq studies of CD8+FoxP3+ T-cells revealed an activated, early effector phenotype, substantiating an anti-tumor role, and contrasting with CD4+FoxP3+ T-regulatory cells. A spatial biomarker was developed using an empirical probabilistic model to define the immediate cell neighbors or niche surrounding CD8+FoxP3+ cells and proximity to the tumor-stromal boundary. The resultant 'Diversity of Niches Unlocking Treatment Sensitivity (DONUTS)' are more prevalent than the CD8+FoxP3+ cells themselves, mitigating sampling error in small biopsies. Further, the DONUTS only require four markers, are additive to PD-L1, and associate with tertiary lymphoid structure counts. Taken together, the DONUTS represent a next-generation predictive biomarker poised for clinical implementation.

AstroPath↗

A comparative analysis of genome-wide chromatin immunoprecipitation data for mammalian transcription factors.

Genome-wide location analysis (ChIP-chip, ChIP-PET) is a powerful technique to study mammalian transcriptional regulation. In order to obtain a basic understanding of the location data generated for mammalian transcription factors and potential issues in their analysis, we conducted a comparative study of eight independent ChIP experiments involving six different transcription factors in human and mouse. Our cross-study comparisons, to the best of our knowledge the first to analyze multiple datasets, revealed the importance of carefully chosen genomic controls in the de novo identification of key transcription factor binding motifs, raised issues about the interpretation of ubiquitously occurring sequence motifs, and demonstrated the clustering tendency of protein-binding regions for certain transcription factors.

Animals↗

An improved distance measure between the expression profiles linking co-expression and co-regulation in mouse.

BACKGROUND: Many statistical algorithms combine microarray expression data and genome sequence data to identify transcription factor binding motifs in the low eukaryotic genomes. Finding cis-regulatory elements in higher eukaryote genomes, however, remains a challenge, as searching in the promoter regions of genes with similar expression patterns often fails. The difficulty is partially attributable to the poor performance of the similarity measures for comparing expression profiles. The widely accepted measures are inadequate for distinguishing genes transcribed from distinct regulatory mechanisms in the complicated genomes of higher eukaryotes. RESULTS: By defining the regulatory similarity between a gene pair as the number of common known transcription factor binding motifs in the promoter regions, we compared the performance of several expression distance measures on seven mouse expression data sets. We propose a new distance measure that accounts for both the linear trends and fold-changes of expression across the samples. CONCLUSION: The study reveals that the proposed distance measure for comparing expression profiles enables us to identify genes with large number of common regulatory elements because it reflects the inherent regulatory information better than widely accepted distance measures such as the Pearson's correlation or cosine correlation with or without log transformation.

Algorithms↗

Computational biology: toward deciphering gene regulatory information in mammalian genomes.

Computational biology is a rapidly evolving area where methodologies from computer science, mathematics, and statistics are applied to address fundamental problems in biology. The study of gene regulatory information is a central problem in current computational biology. This article reviews recent development of statistical methods related to this field. Starting from microarray gene selection, we examine methods for finding transcription factor binding motifs and cis-regulatory modules in coregulated genes, and methods for utilizing information from cross-species comparisons and ChIP-chip experiments. The ultimate understanding of cis-regulatory logic in mammalian genomes may require the integration of information collected from all these steps.

Animals↗

TileMap: create chromosomal map of tiling array hybridizations.

MOTIVATION: Tiling array is a new type of microarray that can be used to survey genomic transcriptional activities and transcription factor binding sites at high resolution. The goal of this paper is to develop effective statistical tools to identify genomic loci that show transcriptional or protein binding patterns of interest. RESULTS: A two-step approach is proposed and is implemented in TileMap. In the first step, a test-statistic is computed for each probe based on a hierarchical empirical Bayes model. In the second step, the test-statistics of probes within a genomic region are used to infer whether the region is of interest or not. Hierarchical empirical Bayes model shrinks variance estimates and increases sensitivity of the analysis. It allows complex multiple sample comparisons that are essential for the study of temporal and spatial patterns of hybridization across different experimental conditions. Neighboring probes are combined through a moving average method (MA) or a hidden Markov model (HMM). Unbalanced mixture subtraction is proposed to provide approximate estimates of false discovery rate for MA and model parameters for HMM. AVAILABILITY: TileMap is freely available at http://biogibbs.stanford.edu/~jihk/TileMap/index.htm. SUPPLEMENTARY INFORMATION: http://biogibbs.stanford.edu/~jihk/TileMap/index.htm (includes coloured versions of all figures).

Algorithms↗

Why do human diversity levels vary at a megabase scale?

Levels of diversity vary across the human genome. This variation is caused by two forces: differences in mutation rates and the differential impact of natural selection. Pertinent to the question of the relative importance of these two forces is the observation that both diversity within species and interspecies divergence increase with recombination rates. This suggests that mutation and recombination are either directly coupled or linked through some third factor. Here, we test these possibilities using the recently generated sequence of the chimpanzee genome and new estimates of human diversity. We find that measures of GC and CpG content, simple-repeat structures, as well as the distance from the centromeres and the telomeres predict diversity as well as divergence. After controlling for these factors, large-scale recombination rates measured from pedigrees are still significant predictors of human diversity and human-chimpanzee divergence. Furthermore, the correlation between human diversity and recombination remains significant even after controlling for human-chimpanzee divergence. Two plausible and non-mutually exclusive explanations are, first, that natural selection has shaped the patterns of diversity seen in humans and, second, that recombination rates across the genome have changed since humans and chimpanzees shared a common ancestor, so that current recombination rates are a better predictor of diversity than of divergence. Because there are indications that recombination rates may have changed rapidly during human evolution, we favor the latter explanation.

Animals↗