Search PubMedSearch

SEARCH · Search PubMed

Results for “Algorithms”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

10 recordsLinked to original sources

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

Machine learning-based prediction of unplanned readmission and construction of an online calculator for elderly patients with mild ischemic stroke.

OBJECTIVE: To screen for independent risk factors for unplanned readmission in elderly patients with mild ischemic stroke, and to construct and validate an online risk prediction calculator based on an interpretable machine learning model, thereby providing a promising practical tool for accurate clinical assessment of 30&#x2011;day all&#x2011;cause unplanned readmission risk in this population. METHODS: A prospective cohort study was conducted, including 1050 patients aged&#xa0;&#x2265;&#xa0;60&#xa0;years with mild ischemic stroke admitted between August 2023 and September 2024. Participants were randomly divided into a training set (840 cases) and a test set (210 cases) at a ratio of 8:2. Risk factors were screened by univariate analysis and multivariable Logistic regression. Four machine learning models, namely LightGBM, XGBoost, Random Forest, and K&#x2011;Nearest Neighbors (KNN), were developed and their performance was evaluated using AUC, accuracy, sensitivity, and specificity as metrics. The SHAP framework was used for interpretability analysis, and an online calculator was subsequently developed based on the optimal model. RESULTS: Univariate analysis showed significant differences (P&#xa0;<&#xa0;0.05) in 13 factors including age, smoking, AIP, TyG index, HALP score, etc. Multivariable Logistic regression identified age (OR&#xa0;=&#xa0;9.752), smoking (OR&#xa0;=&#xa0;5.171), AIP (OR&#xa0;=&#xa0;6.691), TyG index (OR&#xa0;=&#xa0;4.393), HALP score (OR&#xa0;=&#xa0;2.831), and&#xa0;&#x2265;&#xa0;2 comorbidities (OR&#xa0;=&#xa0;3.664) as independent risk factors. All four machine learning models demonstrated good predictive performance. Based on a comprehensive evaluation of multiple metrics and computational efficiency, the LightGBM model exhibited the best predictive performance (AUC&#xa0;=&#xa0;0.884, accuracy&#xa0;=&#xa0;0.829, sensitivity&#xa0;=&#xa0;0.812, specificity&#xa0;=&#xa0;0.875). SHAP analysis showed that age, AIP, TyG index, smoking, and HALP score were key predictors. An online calculator developed based on this model enables individualized risk predictions. CONCLUSION: Key risk factors associated with 30&#x2011;day unplanned readmission in elderly patients with mild ischemic stroke were identified. The LightGBM model demonstrated high predictive accuracy, and together with the interpretability analysis and online calculator, offers a practical tool to support clinical risk assessment. However, this tool requires future external validation.

Humans

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence

Clustering patterns of behavioral and metabolic risk factors for noncommunicable diseases in Iran: findings from a national STEPS survey.

BACKGROUND: Noncommunicable diseases (NCDs) are the leading cause of mortality in Iran, driven by behavioral and metabolic risk factors that frequently co-occur. OBJECTIVE: To identify patterns of co-occurring behavioral and metabolic NCD risk factors among Iranian adults and characterize their demographic and socioeconomic correlates. METHODS: This cross-sectional study analyzed data from 16,618 adults aged &#x2265;25&#x2009;years who participated in Iran's 2021 nationally representative STEPS survey. Thirteen behavioral and metabolic variables, including physical activity, nutrition score, smoking frequency, alcohol intake, salt intake, body mass index, blood pressure, fasting plasma glucose, and lipid markers, were entered into a K-means clustering analysis. Clusters were characterized by their risk profiles and demographic/socioeconomic attributes. Multinomial logistic regression examined associations between cluster membership and sociodemographic factors. RESULTS: Five distinct behavioral-metabolic clusters emerged. The smokers-drinkers (SD) cluster (3.1%) comprised mostly older, less-educated men with high smoking and alcohol use. The healthy-low-risk (HLR) cluster (40.3%) showed favorable profiles and included younger, more educated individuals. The physically active (PA) cluster (6.6%) was characterized mainly by younger men with markedly high physical activity levels. The dyslipidemic (DLP) cluster (26.0%) exhibited high dyslipidemia and overweight prevalence, while the hypertensive-diabetic (HTD) cluster (24.0%) had the highest obesity, hypertension, and diabetes rates, common among older urban adults. CONCLUSION: Behavioral and metabolic NCD risk factors in Iran formed five distinct co-occurrence patterns. Nearly half of adults belonged to metabolically high-risk clusters, highlighting the need for targeted prevention strategies that combine lifestyle interventions with screening and management of obesity, hypertension, diabetes, and dyslipidemia.

Humans

Integrative machine learning and transcriptomic analysis reveals molecular mechanisms underlying low survival rate in larval Chinese Bahaba (Bahaba taipingensis).

Chinese Bahaba (Bahaba taipingensis) is a Class I protected marine fish endemic to China. Low larvae survival during artificial breeding severely hinder population recovery. To investigate the molecular mechanism of high mortality in larval fish, this study performed RNA-seq on liver from naturally deceased (ND) and mass-dead (MD) individuals, combined with least absolute shrinkage and selection operator (LASSO) regression and random forest (RF) algorithms to screen for core signature genes. A total of 873 differentially expressed genes (DEGs) were identified, including 112 upregulated and 761 downregulated genes. GO and KEGG enrichment analyses revealed significant enrichment in amino acid metabolism disorders, one&#x2011;carbon folate pool impairment, PPAR signaling abnormalities, ECM-receptor interaction, focal adhesion pathway, indicating widespread metabolic suppression accompanied by extracellular matrix remodeling and signaling disturbances in the livers of MD fish. MAD pre-filtering combined with dual machine learning algorithms yielded 18 robust core signature genes, among which SLC38A4, MMP1, FADD, FKBP5, and APOB were consistently identified as high-frequency core genes by both algorithms. SLC38A4 exhibited the highest importance score in the RF model and was significantly downregulated, making it the primary molecule distinguishing ND from MD phenotypes. ROC curve analysis showed that both models achieved an AUC of 1.000 (95% CI lower bound: 0.610), confirming the precise discriminatory ability of the core genes. GSEA further demonstrated significant enrichment of this core gene set in ND samples. This study provides the first systematic elucidation of the molecular mechanisms underlying liver dysfunction in low survival rate B. taipingensis, characterized by amino acid transport impairment, metabolic reprogramming, and structural remodeling, offering theoretical foundations for health assessment, early mortality risk warning, and artificial breeding conservation of this species.

Animals

Machine learning-ready genomic biomarkers: ATF3 polymorphisms predict postoperative analgesic demand through AI-compatible phenotyping.

PURPOSE: To determine whether ATF3 polymorphisms can serve as genetic biomarkers for machine learning-based precision analgesia by establishing a genotype-phenotype association suitable for predictive modeling of postoperative opioid requirements. METHODS: In a prospective cohort of 167 adults undergoing abdominal surgery, ATF3 SNPs rs3122721 and rs3125293 were genotyped. A structured dataset architecture was developed to represent genetic profiles as input features for supervised learning models, enabling translational analysis of genotype&#x2011;dependent opioid consumption over 72&#xa0;h. RESULTS: Patients with homozygous genotypes of the ATF3 SNPs had significantly higher opioid requirements than non&#x2011;carriers, despite reporting similar subjective pain scores. This consistent genotype&#x2011;dependent pattern provided a clinically relevant phenotype suitable for integration into predictive algorithms. CONCLUSION: ATF3 genotyping offers a promising biomarker for computationally informed precision analgesia. By linking genomic variability to clinically meaningful outcomes within a structured clinical and genomic framework, this approach supports the future development of risk-stratified clinical decision-support systems to optimize postoperative pain management.Trial registration ChiCTR1900021991, registered 30 April 2019. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13755-026-00480-9.

ATF3

Genome-wide characterization of heat shock protein genes reveals thermal stress-responsive candidates in Litopenaeus vannamei.

Heat shock proteins (HSPs) are conserved molecular chaperones involved in protein folding, refolding, aggregation prevention, and degradation of damaged proteins. However, the genomic organization and thermal responsiveness of HSP genes in the Pacific white shrimp (Litopenaeus vannamei) remain incompletely understood. Here, we performed a genome-wide analysis of the HSP gene family and examined its phylogenetic relationships, structural features, duplication patterns, sequence variation, interaction networks, and transcriptional responses to acute heat stress. A total of 34 HSP genes were identified and classified into the HSP90, HSP70, HSP40/DNAJ, HSP60, and small HSP families. Phylogenetic, motif, gene structure, synteny, and subcellular localization analyses revealed evolutionary conservation and structural diversification among family members. Three duplicated gene pairs were identified, comprising two segmental duplications and one tandem duplication. All pairs exhibited Ka/Ks ratios below 1, consistent with purifying selection of varying strength. Sequence analysis identified 295 nonsynonymous single-nucleotide polymorphisms, of which 12 were consistently predicted to be deleterious by multiple algorithms. Protein-protein interaction analysis indicated enrichment of protein-folding and cellular stress-response functions. RT-qPCR analysis showed significant induction of HSPA4, HSP90AA1, TRAP1, BiP, and DNAJA1 after 6, 12, and 24&#xa0;h of exposure to 34&#xa0;&#xb0;C, whereas DNAJC3 was significantly induced only at 12&#xa0;h. All six genes reached their highest transcript abundance at 12&#xa0;h. These findings may provide a genomic framework for HSP genes in L. vannamei and identify candidate genes and variants associated with thermal stress responses.

Animals

Conserved host-exclusive oligonucleotide motifs enriched in pathogenic genes of human oncogenic viruses.

Comparative viral genomics can reveal sequence-level constraints influencing virus-host interactions. Relative minimal absent words (rMAWs) are short oligonucleotide motifs present in viral genomes but completely absent from the host, potentially reflecting selective pressures related to host adaptation and immune evasion. Using the EAGLE algorithm and the GRCh38 human reference genome, we systematically screened for prevalent rMAWs (prMAWs) across six major human oncogenic viruses: Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T-cell leukemia virus type 1 (HTLV-1), and human herpesvirus 8/Kaposi's sarcoma-associated herpesvirus (HHV-8/KSHV). highly conserved 11- and 12-bp prMAWs were identified in EBV, HBV, HTLV-1, and HHV-8/KSHV, with sequence prevalences ranging from 91.5% to 97.9%. Conversely, no short prMAWs were detected in HCV or HPV, likely reflecting differences in genome architecture, mutation rates, and long-term host adaptation to the human host. Importantly, the identified host-exclusive motifs exhibited non-random genomic distribution and were preferentially embedded within viral genes central to replication, persistence, immune modulation, and oncogenesis, including EBNA-1 (EBV), HBx (HBV), Tax-associated regions (HTLV-1), and lytic replication genes of HHV-8/KSHV. Notably, all detected prMAWs were enriched in GC nucleotides and exhibited marked CpG over-representation, suggesting sequence constraints associated with epigenetic regulation and viral persistence. Collectively, these highly conserved, host-exclusive signatures offer promising, candidates for sequence-directed approaches in the diagnosis, monitoring, and investigation of virus-associated cancers.

Humans

Artificial intelligence for anticancer drug discovery from natural products of macroalgae and sponges: A systematic review.

Marine natural products (MNPs) from macroalgae and marine sponges have inspired clinically important anticancer agents, including the cytarabine pharmacophore and the eribulin scaffold, while cyanobacterial dolastatin chemistry supplies the auristatin payloads of several marine-inspired antibody-drug conjugates (ADCs) such as brentuximab vedotin. Artificial intelligence (AI) methods, encompassing both classical machine learning (ML) with hand-engineered features and modern deep learning (DL) with many-layered neural networks, are increasingly supporting key decisions in natural-product anticancer drug discovery, including bioactivity prediction, target identification, absorption, distribution, metabolism, excretion and toxicity (ADMET) filtering, generative analogue design, and the selection of preclinical candidates. DL architectures relevant to this field include graph neural networks, transformer-based molecular generators, diffusion models for protein-ligand docking, and convolutional networks for mass spectrometry, while classical ML contributes interpretable fingerprint-based bioactivity models and molecular networking for dereplication. This review follows a systematic literature review methodology to organize the landscape of AI methods now applied to MNP anticancer discovery, distinguishing ML and DL approaches where relevant, situating them within the chemical context of macroalgal and sponge-derived oncology leads, and critically examining published case studies, including validation level (computational, in vitro, in vivo, clinical). The principal bottleneck for medical translation has shifted partly from algorithmic capability toward data infrastructure and experimental validation. Sparse, heterogeneous, and taxonomically biased bioactivity records limit what current models can learn and reduce the reliability of AI-prioritized candidates entering the preclinical pipeline. A roadmap is proposed that prioritizes open MNP-specific benchmarks, symbiont-aware modeling, and active learning loops with synthesizability and ADMET constraints. These AI workflows may accelerate the prioritization of marine-derived anticancer leads and support earlier, more evidence-based translational decisions in oncology drug development.

Biological Products

Antegrade dissection and re-entry vs retrograde strategy in chronic total occlusion percutaneous coronary intervention: Rationale and design of the ADRENALINE randomized study.

RATIONALE: While antegrade wiring (AW) is the most common initial strategy for chronic total occlusion (CTO) percutaneous coronary intervention (PCI), difficult CTO lesions frequently require either antegrade dissection and re-entry (ADR) or a retrograde strategy. Comparative data between ADR and the retrograde approach remain limited. DESIGN: The Antegrade Dissection vs Retrograde re-ENtry And Load of Interventionalist Effort (ADRENALINE) is a prospective, multicenter randomized study with a superiority design. It is planned to enroll 121 patients with difficult coronary CTO (J-CTO score &#x2265;2) referred for CTO-PCI in accordance with the hybrid algorithm. Subjects undergoing successful AW will be included in the observational arm. Patients with failed or unattempted AW will be randomized 1:1 to ADR or retrograde CTO crossing strategy (n = 74). All patients will undergo pre- and postprocedural laboratory testing (including cardiac troponin T and creatine kinase-MB), cardiac magnetic resonance (CMR) for late gadolinium enhancement, and health status assessment by the Seattle Angina Questionnaire and the Rose Dyspnea Scale. The co-primary endpoints are total procedure time and successful guidewire crossing. Additionally, the relationship between different recanalization strategies and stress among interventional cardiologists will be explored. CONCLUSION: ADRENALINE is the first randomized study of ADR vs retrograde strategy for difficult CTO PCI, assessing procedural outcomes, CMR-detected myocardial infarction, and 3-month quality of life. ENROLMENT STATUS: The first patient was enrolled on July 29, 2025. As of June 14, 2026, 45 patients (26 randomized, 19 observational) of the planned 121 patients have been enrolled. TRIALS REGISTRATION: Clinicaltrials.gov: Identifier, NCT06878729.

Humans