Search PubMed⌕ Search

Biomedical subjects

Damien Chaussabel

Publications and source records attributed to Damien Chaussabel.

4 recordsLinked to original sources

Automating candidate gene prioritization with large language models: from naive scoring to literature-grounded validation.

MOTIVATION: Identifying promising therapeutic targets from thousands of genes in transcriptomic studies remains a major bottleneck in biomedical research. While large language models (LLMs) show potential for gene prioritization, they suffer from hallucination and lack systematic validation against expert knowledge. RESULTS: The framework identified 609 sepsis-relevant genes with >94% filtering efficiency, demonstrating strong enrichment for inflammatory pathways including TNF-α signaling, complement activation, and interferon responses. Literature validation yielded 30 ultra-high confidence therapeutic candidates, including both established sepsis genes (IL10, TREM1, S100A9, NLRP3) and novel targets warranting investigation. Benchmark validation against expert-curated databases achieved 71.2% recall, with systematic correlation between computational confidence and evidence quality. The final candidate set balanced discovery (11 novel genes) with validation (19 known genes), maintaining biological coherence throughout the filtering process. This framework demonstrates that rigorous methodology can transform unreliable LLM outputs into systematically validated biological insights. By combining computational efficiency with literature grounding, the approach provides a practical tool for prioritizing experimental validation efforts. The modular design enables adaptation to other diseases through knowledge base substitution, offering a systematic approach to literature-guided biomarker discovery. AVAILABILITY AND IMPLEMENTATION: We developed a two-stage computational framework that combines LLM-based screening with literature validation for systematic gene prioritization. Starting with 10 824 genes from the BloodGen3 repertoire, we applied multi-criteria evaluation for sepsis relevance, followed by retrieval-augmented generation using 6346 curated sepsis publications. A novel faithfulness evaluation system verified that LLM predictions aligned with retrieved literature evidence. Source code and implementation details are available at https://github.com/taushifkhan/llm-geneprioritization-framework, vector database at https://doi.org/10.5281/zenodo.15802241, and Interactive demonstration at https://llm-geneprioritization.streamlit.app/.

Humans↗

Relationships among murine CD11c(high) dendritic cell subsets as revealed by baseline gene expression patterns.

The functional relationships and properties of different subtypes of dendritic cells (DC) remain largely undefined. To better characterize these cells, we used global gene analysis to determine gene expression patterns among murine CD11c(high) DC subsets. CD4(+), CD8alpha(+), and CD8alpha(-) CD4(-) (double negative (DN)) DC were purified from spleens of normal C57/BL6 mice and analyzed using Affymetrix microarrays. The CD4(+) and CD8alpha(+) DC subsets showed distinct basal expression profiles differing by >200 individual genes. These included known DC subset markers as well as previously unrecognized, differentially expressed CD Ags such as CD1d, CD5, CD22, and CD72. Flow cytometric analysis confirmed differential expression in nine of nine cases, thereby validating the microarray analysis. Interestingly, the microarray expression profiles for DN cells strongly resembled those of CD4(+) DC, differing from them by <25 genes. This suggests that CD4(+) and DN DC are closely related phylogenetically, whereas CD8alpha(+) DC represent a more distant lineage, supporting the historical distinction between CD8alpha(+) and CD8alpha(-) DC. However, staining patterns revealed that in contrast to CD4(+) DC, the DN subset is heterogeneous and comprises at least two subpopulations. Gene Ontology and literature mining analyses of genes expressed differentially among DC subsets indicated strong associations with immune response parameters as well as cell differentiation and signaling. Such associations offer clues to possible unique functions of the CD11c(high) DC subsets that to date have been difficult to define as rigid distinctions.

Animals↗

Unique gene expression profiles of human macrophages and dendritic cells to phylogenetically distinct parasites.

Monocyte-derived dendritic cells (DCs) and macrophages (Ms) generated in vitro from the same individual blood donors were exposed to 5 different pathogens, and gene expression profiles were assessed by microarray analysis. Responses to Mycobacterium tuberculosis and to phylogenetically distinct protozoan (Leishmania major, Leishmania donovani, Toxoplasma gondii) and helminth (Brugia malayi) parasites were examined, each of which produces chronic infections in humans yet vary considerably in the nature of the immune responses they trigger. In the absence of microbial stimulation, DCs and Ms constitutively expressed approximately 4000 genes, 96% of which were shared between the 2 cell types. In contrast, the genes altered transcriptionally in DCs and Ms following pathogen exposure were largely cell specific. Profiling of the gene expression data led to the identification of sets of tightly coregulated genes across all experimental conditions tested. A newly devised literature-based clustering algorithm enabled the identification of functionally and transcriptionally homogenous groups of genes. A comparison of the responses induced by the individual pathogens by means of this strategy revealed major differences in the functionally related gene profiles associated with each infectious agent. Although the intracellular pathogens induced responses clearly distinct from the extracellular B malayi, they each displayed a unique pattern of gene expression that would not necessarily be predicted on the basis of their phylogenetic relationship. The association of characteristic functional clusters with each infectious agent is consistent with the concept that antigen-presenting cells have prewired signaling patterns for use in the response to different pathogens.

Animals↗

Mining microarray expression data by literature profiling.

BACKGROUND: The rapidly expanding fields of genomics and proteomics have prompted the development of computational methods for managing, analyzing and visualizing expression data derived from microarray screening. Nevertheless, the lack of efficient techniques for assessing the biological implications of gene-expression data remains an important obstacle in exploiting this information. RESULTS: To address this need, we have developed a mining technique based on the analysis of literature profiles generated by extracting the frequencies of certain terms from thousands of abstracts stored in the Medline literature database. Terms are then filtered on the basis of both repetitive occurrence and co-occurrence among multiple gene entries. Finally, clustering analysis is performed on the retained frequency values, shaping a coherent picture of the functional relationship among large and heterogeneous lists of genes. Such data treatment also provides information on the nature and pertinence of the associations that were formed. CONCLUSIONS: The analysis of patterns of term occurrence in abstracts constitutes a means of exploring the biological significance of large and heterogeneous lists of genes. This approach should contribute to optimizing the exploitation of microarray technologies by providing investigators with an interface between complex expression data and large literature resources.

Cluster Analysis↗