Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “LLM”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Human leucocyte migration: studies with an improved skin chamber technique.

An improved skin chamber technique has been developed for the study of localized leucocyte mobilization (LLM). Uniform "windows" of denuded dermis were produced by a suction device applied to the forearm skin, eliciting delineated areas of epidermal separation by blister formation. The acellular blister fluid, roof and basement membrane were removed, and the blister base was covered with a rubber chamber containing autologous serum as leucocyte attractant. Duplicate chambers were harvested at prescribed intervals during the first 24 hours. In 15 healthy individuals, virtually no cells were observed after 2 hours, a median of 1.9 X 10(6) after 4 hours, increasing to 3.8 X 10(7) after 24 hours. Subnormal LLM was demonstrated in three of seven patients with severe bacterial infections and in three of seven leukaemia patients. LLM was normal in eight patients with other malignancies. Ninety to 98 per cent of the cells were polymorphonuclear neutrophils and less than 1 per cent were erythrocytes. In the chamber neutrophils, vacuolization of the cytoplasm was prominent, bactericidal capacity reduced and nitroblue tetrazolium reduction increased, thus indicating functional derangement of emigrated cells compared to peripheral blood neutrophils. Simplicity and good reproducibility should make this method a valuable tool in the study of leucocyte migration.

Adolescent↗

An in vitro stimulation of the effects of chewing sugar-free and sugar-containing chewing gums on pH changes in dental plaque.

The objective of these studies was to simulate the effect of chewing sugar-free and sucrose-containing chewing gums on the return of the pH to neutrality after exposure to sucrose of plaque located on the buccal (BLM) and lingual (LLM) surfaces of the lower molar teeth. In study 1, a 0.5-mm-deep artificial plaque containing Streptococcus oralis cells was exposed to 10% sucrose for one min, and a 0.1-mm-thick film of sucrose-free artificial saliva was then flowed over the plaque surface at the unstimulated salivary film velocities previously found at the BLM and LLM sites. At the time of the pH minimum (pH 4-5), one of three conditions was simulated: (a) a no-gum-chewing control, or chewing for 20 min on either (b) a sugar-free gum or (c) a sucrose-containing gum. The recovery of the plaque pH to resting values was rapid during simulation of chewing a sugar-free gum (SFG), much slower with the no-gum control, and even slower with simulation of chewing a sucrose-containing gum (SCG). The pH recovery was slower with the BLM than the LLM plaque. In study 2, the BLM plaque was exposed to a 2% sucrose solution for 20 min under stimulated salivary conditions, to simulate the consumption of a meal, followed by one of conditions (a), (b), or (c) described above. The pH recovery with simulation of chewing a SCG was faster than with the no-gum control, but much slower than with the SFG simulation.(ABSTRACT TRUNCATED AT 250 WORDS)

Bicarbonates↗

Motor impairment as a predictor of functional recovery and guide to rehabilitation treatment after stroke.

OBJECTIVE: This study tests three hypotheses relevant for the efficient use of rehabilitation services after stroke: (a) the severity of initial motor impairment after stroke predicts discharge motor impairment and self-care mobility scores; (b) identification of those unlikely to show improvement in motor impairment can focus rehabilitzation efforts on use of compensatory techniques and assist devices; and (c) improvement in self-care mobility scores without change in motor impairment, balance, or cognition is a quantitative estimate of the value of teaching compensatory techniques and use of assist devices. METHODS: We studied 171 sequential patients previously independent in the community who were admitted for inpatient rehabilitation within 17 +/- 12 SD days of an initial, unilateral, hemispheric, ischemic stroke. Impairment was assessed using the Fugl-Meyer upper limb motor (ULM), lower limb motor (LLM), and upper plus lower limb total motor (TM) subscores. Disability was assessed using the Functional Independence Measure (FIM), FIM self-care (FIMS), FIM mobility (FIMM), and FIM self-care plus FIM mobility (FIMSM) subscores. Spearman correlation coefficients tested strength of association between dependent and independent variables, stepwise linear regression tested the effects of clinically relevant co-variables, and positive and negative predictive values (PPV, NPV) assessed the clinical relevance of outcome-prediction models. RESULTS: The highest correlations observed were between admission TM scores and the following discharge scores: TM (R = 0.92; p < 0.01), ULM (R = 0.91; p < 0.01), LLM (R = 0.82; p < 0.01), FIMSM (R = 0.67; p < 0.01), FIMM (R = 0.67; p < 0.001), FIM (R = 0.58; p < 0.0001). An admission TM score in the lowest quartile had a PPV of 0.74 for a discharge ULM score in the lowest quartile. An admission TM score in the highest quartile had a PPV of 0.86 for a discharge ULM score in the highest quartile. Similar but weaker PPVs were seen for admission TM scores and discharge LLM scores. Patients without significant change in TM scores (< or = 2 points) had a 17 +/- 9 SD improvement in FIMSM scores. CONCLUSIONS: Admission motor impairment scores (a) predict discharge impairment and activities of daily living mobility functional outcome; and (b) guide treatment toward improving motor impairment versus use of compensatory techniques and assistive devices. The use of compensatory techniques and assistive devices, without change in motor impairment, is associated with a 17 +/- 9 SD improvement in FIMSM score.

Activities of Daily Living↗

Evaluation of a cornea-specialized large language model for diagnostic and management accuracy in complex corneal cases.

PURPOSE: To evaluate whether a cornea-specialized large language model (LLM) enhanced with retrieval-augmented generation (RAG) improves clinicians' diagnostic and management accuracy in complex corneal cases compared to a general-purpose GPT-4o model and unaided clinician performance. METHODS: This prospective, randomized, masked evaluation study involved three cornea trainees who each independently reviewed 39 real-world corneal cases under three experimental conditions: unaided, GPT-4o-assisted, and assisted by a cornea-specialized GPT-4o model. The cornea-specialized model was constructed by embedding over 200 publicly available Wikipedia articles into GPT-4o's RAG framework. Participants provided open-ended diagnoses and selected the next-step management options (multiple choice). They were allowed up to three GPT-4o queries per case, and the AI-assisted arms were randomized to minimize bias. Accuracy for both tasks was compared against expert reference standards using McNemar's test. RESULTS: Diagnostic accuracy was 48.7%, 20.5%, and 38.5% unaided, improving to 69.2%, 46.2%, and 59.0% with general GPT-4o (p<0.04). The cornea-specialized GPT-4o further improved accuracy to 71.8%, 48.7%, and 74.4%, with improvements over unaided performance for all clinicians (p<0.01). For next-step decisions, unaided accuracy was 76.9%, 87.2%, and 59.0%. With the specialized model, Ophthalmologist 3 improved to 71.8% (p<0.05), Ophthalmologist 1 remained high at 82.1%, and Ophthalmologist 2 declined to 64.1% (p<0.05). CONCLUSIONS: A cornea-specialized LLM enhanced with RAG improved diagnostic accuracy in complex corneal cases, particularly among clinicians with lower baseline performance. Effects on management accuracy were inconsistent. Future studies should explore the use of open-ended management tasks and examine whether smaller, curated retrieval corpora yield better model performance.

Humans↗

Effects of different malnutrition techniques on the behavior of rats tested in the elevated T-maze.

The influence of different malnutrition techniques on the behavior of adult animals was investigated in the elevated T-maze (ETM). Control litters (C) were composed by eight pups constantly kept with their mother and fed by a 16%-protein diet ad libitum; protein malnutrition litters (PM) were fed by a 6%-protein diet; protein-calorie malnutrition litters (PCM) were fed with 50% of the 16%-protein diet ingested by C litters; malnutrition by increase in the size of the litter (LLM-number of pups was twice the number of pups in C litters), and malnutrition by separation (SM-litters spent half of the day with non-lactating females). After weaning, all groups received lab chow diet until the test day (70th day). During the test were recorded the basal, avoidance 1, avoidance 2 and escape latencies. The data showed that PM, PCM, LLM and SM animals showed lower increases in avoidance latencies, when compared to their control groups. However, malnutrition did not affect escape latencies. The nature of these alterations seems to be nutritional, as the extra-nutritional factors (i.e. maternal care) differ a lot among the malnutrition techniques. These results suggest that malnutrition, irrespective of the technique, altered the neural mechanisms believed to control defensive behaviors in the ETM.

Animals↗

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models↗

ELISA (Embedding-Linked Interactive Single-cell Agent): an interpretable hybrid generative Artificial Intelligence agent for expression-grounded discovery in single-cell genomics.

Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models remain opaque to natural language. Here, we introduce ELISA (Embedding-Linked Interactive Single-cell Agent), an interpretable framework that unifies single-cell generative pretrained transformer expression embeddings with biomedical bidirectional encoder representations from transformers-based semantic retrieval and large-language model (LLM)-mediated interpretation for interactive single-cell discovery. An automatic query classifier routes inputs to gene marker scoring, semantic matching, or reciprocal rank fusion pipelines depending on whether the query is a gene signature, natural language concept, or mixture of both. Integrated analytical modules perform pathway activity scoring across 60+ gene sets, ligand-receptor interaction prediction using 280+ curated pairs, condition-aware comparative analysis, and cell-type proportion estimation, all operating directly on embedded data without access to the original count matrix. Benchmarked across six diverse scRNA-seq datasets spanning inflammatory lung disease, pediatric and adult cancers, organoid models, healthy tissue, and neurodevelopment, ELISA significantly outperforms CellWhisperer, a classical lexical retriever (BM25), and a random baseline in cell type retrieval (combined permutation test, $p < 2\times 10^{-5}$ for each), with particularly large gains on gene-signature queries (Cohen's $d = 5.98$ for mean reciprocal rank). ELISA replicates published biological findings (mean composite score 0.88), and generates candidate hypotheses through grounded LLM reasoning, bridging the gap between transcriptomic data exploration and biological discovery.

Generative Artificial Intelligence↗

AskBeacon-performing genomic data exchange and analytics with natural language.

MOTIVATION: Enabling clinicians and researchers to directly interact with global genomic data resources by removing technological barriers is vital for medical genomics. AskBeacon enables large language models (LLMs) to be applied to securely shared cohorts via the Global Alliance for Genomics and Health Beacon protocol. By simply "asking" Beacon, actionable insights can be gained, analyzed, and made publication-ready. RESULTS: In the Parkinson's Progression Markers Initiative (PPMI), we use natural language to ask whether the sex-differences observed in Parkinson's disease are due to X-linked or autosomal markers. AskBeacon returns a publication-ready visualization showing that for PPMI the autosomal marker occurred 1.4 times more often in males with Parkinson's disease than females, compared to no differences for the X-linked marker. We evaluate commercial and open-weight LLM models, as well as different architectures to identify the best strategy for translating research questions to Beacon queries. AskBeacon implements extensive safety guardrails to ensure that genomic data is not exposed to the LLM directly, and that generated code for data extraction, analysis and visualization process is sanitized and hallucination resistant, so data cannot be leaked or falsified. AVAILABILITY AND IMPLEMENTATION: AskBeacon is available at https://github.com/aehrc/AskBeacon.

Genomics↗

Agentomics: an agentic system that autonomously develops novel state-of-the-art solutions for biomedical machine learning tasks.

MOTIVATION: Extracting knowledge from biomedical data is crucial for advancing our understanding of biological systems and developing novel therapeutics. The quantity, quality, and resolution of biomedical data constantly evolves, requiring the automation of biomedical machine learning (ML). Existing Automated ML tools lack flexibility, while large language models (LLMs) struggle to consistently deliver reproducible machine learning codebases, and existing LLM Agent-powered solutions lag behind human-engineered ML models. RESULTS: Here, we introduce Agentomics, an autonomous LLM-powered agentic system for end-to-end ML experimentation. Given a biomedical dataset, Agentomics implements various ML modeling strategies, and produces a ready-to-use ML model. Agentomics introduces strict validation checkpoints for standard ML development steps, allowing gradual development on top of working code with defined interfaces and validated artifacts. Further, it offers native support for biomedical foundation models that can be leveraged during experimentation. The generic nature of Agentomics allows the user to create ML solutions for a large variety of datasets and use various LLMs. We evaluate Agentomics across 20 datasets from the domains of Protein Engineering, Drug Discovery, and Regulatory Genomics. When benchmarked against other agentic systems, Agentomics outperformed them in all tested domains. When benchmarked against human expert solutions, Agentomics generated novel state-of-the-art models for 11/20 established benchmark datasets. AVAILABILITY AND IMPLEMENTATION: Agentomics is implemented in Python. Source code and documentation are freely available at: https://github.com/BioGeMT/Agentomics-ML.

Machine Learning↗

Sameness of age cohorts in the mathematics of population growth.

"Considering age groups as part of cohorts, implicit in LLM [a Leslie-Lotka finitist model], causes difficulty, manifested by the extensionality paradox. The proposition made here was that cohorts are empirical temporal entities, while age groups are theoretical entities, references only. Cohort is a multitude of persons born at the same time interval, throughout the totality of their lives. Age group, on the other hand, is a concept only, constituted by the notion of the age interval....A model of household and population growth based on the household composition matrix yields results that are inherently different from these of LLM."

Age Factors↗

Redefining ALS: Large-scale proteomic profiling reveals a prolonged pre-diagnostic phase with immune, muscular, metabolic, and brain involvement.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disorder with a largely unknown duration and pathophysiology of the pre-diagnostic phase, especially for the common non-monogenic form. METHODS: We leveraged the European Prospective Investigation into Cancer and Nutrition (EPIC) cohort with up to 30 years of follow-up to identify incident ALS cases across five European countries. Pre-diagnostic plasma samples from initially healthy participants underwent high-throughput proteomic profiling (7,285 protein markers, SomaScan). Cox proportional hazards models based on 4,567 participants (including 172 incident ALS cases) were used to identify protein biomarkers associated with future ALS diagnosis. Top results were indirectly validated in two independent case-control studies of prevalent ALS (n=417 ALS, 852 controls). Functional annotation included cross-disease comparisons, gene set and tissue enrichment testing, organ-specific proteomic clocks, and the application of large-language models (LLM). FINDINGS: Five proteins (SECTM1, CA3, THAP4, KLHL41, SLC26A7) were identified as significant pre-diagnostic ALS biomarkers (FDR=0.05), detectable approximately two decades before diagnosis. Of these, all except SECTM1 were indirectly validated in independent cohorts of prevalent ALS cases, supporting their clinical significance. Additionally, 22 nominally significant (p<0.05) pre-diagnostic biomarkers were FDR-significant in prevalent ALS with consistent effect directions. Cross-disease comparisons with pre-diagnostic Parkinson's and Alzheimer's disease suggested a largely specific pre-diagnostic ALS biomarker signature. Gene ontology and tissue enrichment highlighted early involvement of immune, muscle, metabolic, and digestive processes. Furthermore, analyses of proteomic clocks revealed accelerated aging in brain-cognition, immune, and muscle tissues before clinical diagnosis. Druggability and LLM analyses revealed possible therapeutic targets and novel strategies, emphasizing translational relevance. INTERPRETATION: Our study provides first evidence of ultra-early molecular changes in common ALS up to two decades prior to clinical onset, mainly affecting immune, muscle, metabolic, digestive, and cognitive systems. Our study nominates several compelling candidates for risk stratification studies and novel therapeutic targets for early intervention. FUNDING: Clinical Research in ALS and Related Disorders for Therapeutic Development (CreATe) Consortium, Cure Alzheimer's Fund, Michael J Fox Foundation, Interdisciplinary Centre for Clinical Research, University M&#xfc;nster.

Journal Article↗

Artificial Intelligence Cannot Replace Peer Reviewers but May Help Editors Triage: A Comparative Analysis of a Large Language Model and Human Reviewer Recommendations at the American Journal of Sports Medicine.

BACKGROUND: The peer review system faces increasing strain from rising manuscript volumes, reviewer fatigue, and well-documented interreviewer disagreement. Large language models (LLMs) have shown potential to support the peer review process, but their ability to replicate editorial decisions at high-impact medical journals and their utility as manuscript screening tools remain unknown. PURPOSE: To compare the agreement between an LLM and the final editorial decision on manuscripts submitted to the American Journal of Sports Medicine and to evaluate the potential of LLMs as a manuscript screening tool. STUDY DESIGN: Cross-sectional agreement study. METHODS: Fifty-four manuscripts randomly selected from submissions to the American Journal of Sports Medicine (September 2024-October 2024) were reviewed by a locally deployed LLM (Ministral 3 14B; Mistral AI) using a standardized prompt. The artificial intelligence (AI) produced a categorical recommendation (reject, cascade, revision, or accept) and a numerical score (0-100) for each manuscript. Agreement with the final editorial decision was assessed by Cohen kappa (4-category model) for pooled human reviewers (n = 139 reviews) and the AI (n = 54). Screening performance was evaluated by positive predictive value (PPV), sensitivity, and specificity. RESULTS: Pooled human reviewers demonstrated fair agreement with the final decision (&#x3ba; = 0.181 [P < .001]; 42.4% agreement), while the AI demonstrated slight, nonsignificant agreement (&#x3ba; = 0.126 [P = .099]; 37.0% agreement). The AI recommended revision for 61.1% of manuscripts, of which 72.7% were ultimately rejected or cascaded, demonstrating systematic "revision bias." When the AI recommended rejection, 54.5% of those manuscripts were ultimately rejected and 27.3% were cascaded; when the AI recommended cascade, 50% were rejected and 50% were cascaded. However, when the AI recommended rejection or cascade (n = 21), 90.5% received a final decision of rejection or cascade (PPV, 90.5%; specificity, 81.8%). Manuscripts with an AI score <70 were rejected or cascaded 88.0% of the time (PPV, 88.0%). CONCLUSION: AI cannot replicate the nuanced judgment of human peer reviewers at a high-impact sports medicine journal. When AI recommended rejection or cascade, 90.5% of manuscripts received that final decision (descriptive PPV, 90.5%; 95% CI, 71.1%-97.3%), suggesting potential utility as an exploratory first-pass screening tool warranting further validation in larger cohorts. However, AI could not reliably distinguish manuscripts destined for outright rejection from those that would be cascaded to a sister journal-an important limitation for editorial triage applications.

Sports Medicine↗

AI-driven CRISPR screening: optimizing gene editing through automation and intelligent decision support.

BACKGROUND: CRISPR-based genetic screening has become a central methodology in functional genomics, enabling systematic interrogation of gene function, genetic interactions and context-dependent vulnerabilities at scale. However, the rapid expansion of screening modalities-including multi-condition designs, combinatorial perturbations, in vivo applications and single-cell readouts-has exposed fundamental limitations of heuristic-driven experimental design and post hoc statistical analysis. MAIN BODY: This Review synthesizes how artificial intelligence is reshaping CRISPR screening by introducing predictive, adaptive and system-level intelligence across the experimental lifecycle. We organize recent advances into two tightly coupled modules. First, machine learning and deep learning (ML/DL) methods optimize experimental design by learning context-dependent perturbation behavior, anticipating confounding effects and enabling iterative, information-efficient screening strategies. Second, large language model-agent (LLM-agent) systems complement these advances by externalizing scientific reasoning, integrating biological knowledge at scale and coordinating analysis and decision-making in human-in-the-loop workflows. CONCLUSIONS: Together, ML/DL and LLM-agent approaches reframe CRISPR screening from a static analytical pipeline into an intelligent experimental system, with important implications for robustness, scalability and biological discovery.

Artificial Intelligence↗

Redefining ALS: Large-scale proteomic profiling reveals a prolonged pre-diagnostic phase with immune, muscular, metabolic, and brain involvement.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disorder with a largely unknown duration and pathophysiology of the pre-diagnostic phase, especially for the common non-monogenic form. METHODS: We leveraged the European Prospective Investigation into Cancer and Nutrition (EPIC) cohort with up to 30 years of follow-up to identify incident ALS cases across five European countries. Pre-diagnostic plasma samples from initially healthy participants underwent high-throughput proteomic profiling (7,285 protein markers, SomaScan). Cox proportional hazards models based on 4,567 participants (including 172 incident ALS cases) were used to identify protein biomarkers associated with future ALS diagnosis. Top results were indirectly validated in two independent case-control studies of prevalent ALS (n=417 ALS, 852 controls). Functional annotation included cross-disease comparisons, gene set and tissue enrichment testing, organ-specific proteomic clocks, and the application of large-language models (LLM). FINDINGS: Five proteins (SECTM1, CA3, THAP4, KLHL41, SLC26A7) were identified as significant pre-diagnostic ALS biomarkers (FDR=0.05), detectable approximately two decades before diagnosis. Of these, all except SECTM1 were indirectly validated in independent cohorts of prevalent ALS cases, supporting their clinical significance. Additionally, 22 nominally significant (p<0.05) pre-diagnostic biomarkers were FDR-significant in prevalent ALS with consistent effect directions. Cross-disease comparisons with pre-diagnostic Parkinson's and Alzheimer's disease suggested a largely specific pre-diagnostic ALS biomarker signature. Gene ontology and tissue enrichment highlighted early involvement of immune, muscle, metabolic, and digestive processes. Furthermore, analyses of proteomic clocks revealed accelerated aging in brain-cognition, immune, and muscle tissues before clinical diagnosis. Druggability and LLM analyses revealed possible therapeutic targets and novel strategies, emphasizing translational relevance. INTERPRETATION: Our study provides first evidence of ultra-early molecular changes in common ALS up to two decades prior to clinical onset, mainly affecting immune, muscle, metabolic, digestive, and cognitive systems. Our study nominates several compelling candidates for risk stratification studies and novel therapeutic targets for early intervention. FUNDING: Clinical Research in ALS and Related Disorders for Therapeutic Development (CreATe) Consortium, Cure Alzheimer's Fund, Michael J Fox Foundation, Interdisciplinary Centre for Clinical Research, University M&#xfc;nster.

Journal Article↗

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design↗

In vivo measurement of flatulence and nutrient digestibility in dogs fed poultry by-product meal, conventional soybean meal, and low-oligosaccharide low-phytate soybean meal.

OBJECTIVE: To determine an optimal window for determining peak flatulence and evaluate the effects of oligosaccharides and supplemental beta-mannanase in soybean meal-based diets on nutrient availability and flatulence. ANIMALS: 6 dogs. PROCEDURES: Dogs were used in a 2 x 3 factorial arrangement of treatments in a 6 x 6 Latin square experiment to evaluate the digestibility, flatulence, and fecal odor metabolites of low-oligosaccharide low-phytate soybean meal (LLM), conventional soybean meal (SBM), and poultry by-product (PBP) meal diets with or without supplemental beta-mannanase (5 g/kg). RESULTS: Enzyme supplementation had no effect on total tract dry matter (DM), nitrogen digestibility, or digestible energy; however, differences between protein sources did exist for total tract DM digestibility and digestible energy. The PBP meal had higher DM digestibility and digestible energy (mean, 0.913 and 4,255 cal/g), compared with soy-based diets (mean, 0.870 and 4,049 cal/g). No differences were detected for any treatment regardless of protein source or addition of supplemental enzyme for any flatulence components analyzed. No differences were detected for all fecal odor metabolites regardless of addition of supplemental enzyme; however, differences between protein sources were detected. The PBP meal had lower concentrations of carboxylic acids and esters and higher concentrations of heterocycles, phenols, thio and sulfides, ketones, alcohols, and indoles than LLM and SBM. CONCLUSIONS AND CLINICAL RELEVANCE: Diets containing < 22.4 g of stachyose/kg and < 2 g of raffinose/kg did not alter digestibility or increase flatulence in dogs.

Animal Feed↗

Large Language Model and Knowledge Graph-Driven AJCC Staging of Prostate Cancer Using Pathology Reports.

Background/Objectives: To develop an automated American Joint Committee on Cancer (AJCC) staging system for radical prostatectomy pathology reports using large language model-based information extraction and knowledge graph validation. Methods: Pathology reports from 152 radical prostatectomy patients were used. Five additional parameters (Prostate-specific antigen (PSA) level, metastasis stage (M-stage), extraprostatic extension, seminal vesicle invasion, and perineural invasion) were extracted using GPT-4.1 with zero-shot prompting. A knowledge graph was constructed to model pathological relationships and implement rule-based AJCC staging with consistency validation. Information extraction performance was evaluated using a local open-source large language model (LLM) (Mistral-Small-3.2-24B-Instruct) across 16 parameters. The LLM-extracted information was integrated into the knowledge graph for automated AJCC staging classification and data consistency validation. The developed system was further validated using pathology reports from 88 radical prostatectomy patients in The Cancer Genome Atlas (TCGA) dataset. Results: Information extraction achieved an accuracy of 0.973 and an F1-score of 0.986 on the internal dataset, and 0.938 and 0.968, respectively, on external validation. AJCC staging classification showed macro-averaged F1-scores of 0.930 and 0.833 for the internal and external datasets, respectively. Knowledge graph-based validation detected data inconsistencies in 5 of 150 cases (3.3%). Conclusions: This study demonstrates the feasibility of automated AJCC staging through the integration of large language model information extraction and knowledge graph-based validation. The resulting system enables privacy-protected clinical decision support for cancer staging applications with extensibility to broader oncologic domains.

artificial intelligence↗

AI-generated familiarity estimates are a useful new source of information about word knowledge in Simplified Chinese.

This study evaluated the usefulness of AI-generated estimates of word familiarity for predicting word difficulty in Simplified Chinese, building on previous research in alphabetic languages. We found that familiarity estimates produced using large language models (LLMs) showed moderate-to-strong correlations with human familiarity ratings. These LLM estimates were the most effective predictors of both word naming and lexical decision times, surpassing traditional metrics such as word frequency and human familiarity ratings, while the latter still provided modest, non-overlapping variance. GPT-4o with English instructions produced superior results compared to the Chinese-centered models currently available. The results imply that LLM familiarity estimates are a valuable resource for Chinese psycholinguistics, supporting work across experimental design, modeling, and norming. We release familiarity estimates for 27,624 words for unrestricted research and educational use.

Humans↗