PubMed · 42753698
An open benchmark and language models for AI in aging biology.
Abstract
Over the past two decades, human aging has been characterized across DNA methylation, transcriptomic, proteomic, and clinical modalities, yet no benchmark evaluates whether AI systems can interpret these heterogeneous data types in the context of aging biology. We introduce LongevityBench, an open suite of 17 tasks spanning five biodata domains, and use it to assess 18 frontier AI systems from six developer teams. Despite recent advances in AI, no single model dominates all tasks, with omics-based age prediction being the hardest task regardless of scale. To test whether these gaps can be closed without frontier-scale resources, we fine-tuned a family of five multitask Longevity-LLMs on domain-specific aging data. The compact (0.6B-9B parameters) Longevity-LLMs matched or exceeded far larger frontier systems on LongevityBench, showing that general-purpose language models can be adapted to structured-omics tasks. We publicly release the benchmark, models, and Longevity Claw, an agentic research interface for aging researchers.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alex Zhavoronkov, Vladimir Naumov, Denis Sidorenko, Alex Aliper, Vladimir Aladinskiy, Ramin Hasani, Alexander Amini, Katerina Nasto, Mathieu Reymond, Rim Shayakhmetov, Zulfat Miftakhutdinov, Vadim N Gladyshev, Fedor Galkin. 2026-09-17. An open benchmark and language models for AI in aging biology.. https://doi.org/10.1016/j.cell.2026.08.026
Cite the original work for its findings. Save a collection to share your selection of sources.