Search PubMedSearch

Biomedical subjects

George C Tseng

Publications and source records attributed to George C Tseng.

3 recordsLinked to original sources

An embedding-based framework enables statistical testing of gene-set function hypotheses inferred by large language models.

Emerging large language models (LLMs) can infer gene functions directly from gene lists, enabling hypothesis generation without predefined gene sets. However, these LLM-derived predictions are qualitative, and principled statistical validation is lacking. Here, we develop an embedding-based statistical framework that transforms gene and function descriptions into vector representations, enabling statistical testing of gene-gene and gene-function relationships and quantitative prioritization of de novo functional hypotheses inferred by LLMs. We benchmark seven state-of-the-art embedding models using curated and retrieval-augmented literature-derived gene descriptions across diverse biological contexts. OpenAI's text-embedding-3-large and Google's gemini-embedding-001 perform best, capturing gene-gene functional relationships in 88.7-92.5% of Gene Ontology biological processes and approximately 98.6% of canonical pathways. In gene-function association analyses, these models achieve high sensitivity (95.2-98.4%) and specificity (72.7-84.3%). Through contamination analysis and evaluation using experimentally informed protein assembly gene sets, our framework distinguishes biologically meaningful LLM-inferred hypotheses from noise, outperforming confidence-based inference and conventional enrichment analysis. We further develop the open-source R package DEGEmbedR and demonstrate its utility for interpreting a drug perturbation-derived differentially expressed gene (DEG) signature lacking significant conventional enrichment results. Together, these results establish LLM-derived embeddings as a quantitative foundation for functional genomics and the statistical validation of LLM-based gene function inference.

Large Language Models

AI-guided analysis of human pancreatic islet sociology reveals distinct cell compositional changes in type 1 diabetes.

Human pancreatic islets exhibit greater anatomic and cellular heterogeneity than previously appreciated, raising fundamental questions about how their composition varies with age, sex, region, and islet size and how type 1 diabetes (T1D) alters these relationships. Yet these questions remained largely unresolved due to the bottleneck of manual tissue inspection. Here, we developed an integrated artificial intelligence (AI)-guided imaging, processing, and statistical pipeline enabling unbiased, high-throughput analysis of more than 2 million candidate islets from 106 non-diabetic (ND) and T1D donors. We identified age-, region-, sex-, and islet size-dependent differences in islet distribution and composition between ND and T1D donors. Profound β-cell loss in T1D was accompanied by reciprocal α-cell expansion, whereas δ-cells and pancreatic polypeptide cells were largely resilient. Cell area and pseudotime analyses uncovered regional and age-dependent trajectories of islet remodeling across T1D progression, along with distinct patterns of cytoarchitectural reorganization of the endocrine pancreas.

Type 1 diabetes

Model-based multifacet clustering with high-dimensional omics applications.

High-dimensional omics data often contain intricate and multifaceted information, resulting in the coexistence of multiple plausible sample partitions based on different subsets of selected features. Conventional clustering methods typically yield only one clustering solution, limiting their capacity to fully capture all facets of cluster structures in high-dimensional data. To address this challenge, we propose a model-based multifacet clustering (MFClust) method based on a mixture of Gaussian mixture models, where the former mixture achieves facet assignment for gene features and the latter mixture determines cluster assignment of samples. We demonstrate superior facet and cluster assignment accuracy of MFClust through simulation studies. The proposed method is applied to three transcriptomic applications from postmortem brain and lung disease studies. The result captures multifacet clustering structures associated with critical clinical variables and provides intriguing biological insights for further hypothesis generation and discovery.

Humans