Search PubMedSearch

SEARCH · Search PubMed

Results for “Machine learning algorithms”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

193 records · Page 4Linked to original sources

Adverse Experiences in Brief Meditation Practices: Randomized Controlled Trial.

BACKGROUND: Meditation has become increasingly popular in recent decades. However, relatively little remains known about the prevalence of and risk factors for adverse experiences related to a single meditation practice. OBJECTIVE: The objective of our study was to examine adverse experiences associated with 3 brief, digitally delivered meditation practices (mindfulness, self-compassion, and gratitude) relative to using the internet as usual, as well as to investigate whether preintervention characteristics could predict such outcomes. METHODS: In a secondary analysis of a randomized controlled trial using samples that were representative of the US and UK adult populations with regard to ethnicity, sex, and age, we examined adverse experiences associated with 3 brief (ie, 5 or 10 minutes) meditation practices (ie, mindfulness, self-compassion, and gratitude) relative to using the internet as usual. We also investigated the potential of using preintervention characteristics to predict such outcomes. RESULTS: A total of 5049 participants completed all preintervention measures and were randomly assigned to meditation or control conditions. Across the sample, 4.1% (204/4925) of participants reported having a distressing experience during the intervention, and 7.1% (348/4908) of participants experienced an increase in negative affect from before to after the intervention. The results showed that participants who were randomized to a brief meditation intervention were no more likely to report a distressing experience than those who were randomized to use the internet as usual (odds ratio [OR] 1.05, 95% CI 0.76-1.47; P=.76). The results also showed that participants who were randomized to a brief meditation intervention were less likely to report clinically relevant increases in negative affect relative to using the internet as usual (OR 0.63, 95% CI 0.50-0.80; P<.001). Notably, participants in the 10-minute condition had a significantly higher likelihood of reporting a distressing experience than those in the 5-minute condition (OR 1.42, 95% CI 1.07-1.89; P=.02). Preintervention characteristics showed acceptable discrimination ability to predict a distressing experience (area under the curve=0.73) and slightly lower ability to predict increased negative affect (area under the curve=0.67). CONCLUSIONS: Taken together, we found that the brief, digitally delivered meditation practices tested in this study carry risks of adverse experiences that are comparable to or lower than those of typical activities on the internet; 10-minute condition was more likely to result in distressing experiences than 5-minute condition; and adverse responses to a brief meditation practice can, at least to a certain degree, be predicted using preintervention characteristics. TRIAL REGISTRATION: Open Science Framework 94HKS; https://osf.io/94hks/overview.

Humans

Systematic evaluation of one-dimensional-to-two-dimensional near-infrared spectroscopy transformations with deep learning for quantifying coconut sap adulteration.

Near-infrared (NIR) spectroscopy have limitations when combined with deep learning (DL) algorithms because they rely on low-dimensional datasets. Therefore, we investigated the potential of transforming one-dimensional (1D) NIR spectra into two-dimensional (2D) spectrograms using synchronous and asynchronous techniques and the continuous wavelet transform (CWT) and their effectiveness by integrating with DL for detecting adulteration in coconut sap. NIR spectra (12,500-4000&#xa0;cm-1) were collected from binary mixtures (0%-100%;w/w). The performance of all DL (convolutional neural networks-CNN, AlexNet and ResNet) models was compared with that of partial least squares (PLS). The models were ranked in the mentioned order based on their performances: 2D-CWT&#xa0;>&#xa0;2D-asynchronous > 2D-synchronous > 1D/2D-PLS. The important features of the best model can be explained and visualized using gradient-weighted-class-activation-mapping. The findings highlight that the 1D-to-2D NIR data transformation combined with DL is a highly robust approach because it addresses the feature representation gap in NIR data and effectively captures the spatial-spectral correlations.

Spectroscopy, Near-Infrared

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n&#xa0;=&#xa0;549) and a validation set (n&#xa0;=&#xa0;236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60&#xa0;mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60&#xa0;mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans

Non-destructive prediction of lead content in oilseed rape leaves by fluorescence hyperspectral technology based on neural network.

Based on fluorescence hyperspectral imaging (FHSI), this study targeted rapid, non-destructive quantification of lead (Pb) content in oilseed rape leaves treated with varying silicon (Si) concentrations, acquiring fluorescence spectra over the 484.43-1001.61&#xa0;nm wavelength range. To optimize spectral data quality, preprocessing methods (Savitzky-Golay smoothing, first derivative, detrending) were comprehensively compared. Characteristic wavelengths were then selected via interval variable iterative shrinkage, which effectively compressed data dimensionality and reduced computational load. A hybrid SE-CL1DA model, fusing a 1D convolutional neural network, a long short-term memory network and SE attention mechanism was constructed, with Bayesian optimization tuning hyperparameters to boost stability. The BO-SE-CL1DA outperformed both traditional machine learning and insufficiently optimized deep learning model (Rp2=0.9609, RMSE&#xa0;=&#xa0;0.0377&#xa0;mg/kg, RPD&#xa0;=&#xa0;5.1736), thus enabling accurate Pb estimation, supporting Si-regulated heavy metal stress management and facilitating agricultural contamination monitoring.

Plant Leaves

Reliability-aware hierarchical learning for Chagas disease screening from 12-lead ECGs: tackling label uncertainty and class imbalance.

Objective.Chagas disease, a neglected tropical disease (NTD) with significant cardiovascular impact, remains underdiagnosed in resource-limited regions. Electrocardiogram (ECG) screening offers a low-cost tool for detecting cardiac involvement, yet algorithm development is challenged by label noise, data scarcity, and the latent nature of infection. This study proposes a robust ECG-based screening framework that explicitly addresses these constraints.Approach.We introduce aReliability-Aware Hierarchical Learningstrategy that calibrates supervision according to data provenance, prioritizing serology-confirmed labels over noisy self-reports. To mitigate data scarcity, we compare a specialized convolutional neural network (CNN) trained from scratch with a transfer learning approach based on a Spatio-Temporal ECG foundation Model (FM). Performance is evaluated across varying data scales, and the representation structure is analyzed to interpret model behavior.Main results.On the official hidden test set of the George B. Moody PhysioNet/Computing in Cardiology Challenge 2025, our approach achieved a Challenge Score of 0.163. We observe that while the specialized CNN performs competitively in data-rich regimes, the FM exhibits superior robustness in extreme low-resource settings. Furthermore, performance reaches a plateau imposed by underlying disease physiology. Bimodal score distributions suggest that models distinguish established cardiomyopathy from indeterminate infection, which remains electrophysiologically indistinguishable from healthy controls.Significance.These findings clarify both the potential and intrinsic limits of ECG-based AI screening for NTD-associated cardiac involvement. Reliability-aware supervision and data-efficient transfer learning provide a practical framework toward scalable and clinically meaningful ECG screening systems in resource-constrained environments.

Humans

Integrated multi-omics analyses identify an RAS-SLC11A2-associated molecular framework linking iron metabolism with PCOS-related cardiometabolic risk.

INTRODUCTION: PCOS is a common endocrine disorder with elevated cardiometabolic risk, yet the role of the renin-angiotensin system (RAS)-iron metabolism axis in this comorbidity remains unclear. We explored its underlying mechanisms and evaluated the therapeutic potential of gentiopicroside. METHODS: Integrated multi-omics analyses combining transcriptomics, single-cell RNA sequencing, Mendelian randomization, machine learning, molecular docking, and in vitro functional assays were performed to identify shared molecular pathways and therapeutic targets across PCOS, hypertension, NAFLD, and T2DM. RESULTS: SLC11A2 was consistently dysregulated in PCOS transcriptomic datasets, and associated with iron metabolism, inflammatory response and oxidative stress pathways. Genetic analyses validated RAS-related regulation in hypertension susceptibility and revealed shared genetic architecture between PCOS and cardiometabolic traits. Network and single-cell analyses characterized SLC11A2-associated molecular patterns in disease-relevant cell types; machine learning identified disease-classifying molecular signatures. Gentiopicroside alleviated inflammatory and oxidative stress phenotypes, including reduced IL-6 expression and reactive oxygen species accumulation. CONCLUSION: This study defines an RAS-SLC11A2 molecular framework linking iron metabolism dysregulation to PCOS-related cardiometabolic risk, elucidating the mechanisms connecting ovarian dysfunction, inflammation, oxidative stress and hypertension, and supports gentiopicroside as a promising therapeutic candidate.

Humans

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning

Clinical applications of digital twin technology in In Vitro Fertilisation.

BACKGROUND: Digital twin technology, originating from aerospace and manufacturing industries, has emerged as a transformative tool in healthcare. In vitro fertilisation (IVF) faces persistent challenges including suboptimal embryo selection, unpredictable treatment outcomes, and limited personalisation of protocols. Despite advances in assisted reproductive technology, existing literature exhibits fragmentation: artificial intelligence applications in embryo selection, ovarian stimulation, and endometrial assessment have been developed independently without systematic integration into comprehensive treatment frameworks. Digital twin technology offers unprecedented opportunities to create virtual replicas of biological systems, enabling real-time monitoring, predictive modelling, and personalised treatment strategies. AIM: This narrative review aims to critically examine the current applications of digital twin technology in IVF, evaluate its potential benefits and limitations, synthesize existing evidence into an integrative conceptual model, and identify future directions for implementation in reproductive medicine. METHOD: A comprehensive narrative review was conducted using PubMed, Scopus, Web of Science, and IEEE Xplore databases. A narrative review approach was selected over systematic review to accommodate the heterogeneity of evidence types in this emerging field, including theoretical frameworks, simulation studies, and proof-of-concept implementations that would be excluded from systematic reviews. Search terms included "digital twin," "IVF," "in vitro fertilisation," "assisted reproductive technology," "embryo selection," and "predictive modelling." Studies published between 2015 and 2025 were included, focusing on original research articles, systematic reviews, and proof-of-concept studies describing digital twin applications in reproductive medicine. RESULTS: Digital twin technology in IVF demonstrates significant potential across multiple domains including embryo development simulation, ovarian response prediction, endometrial receptivity modelling, and personalised stimulation protocols. Current applications integrate artificial intelligence, machine learning algorithms, time-lapse imaging, and omics data to create comprehensive virtual models. Early evidence suggests improvements in embryo selection accuracy, ovarian response prediction, and treatment protocol optimization, though large-scale randomized controlled trials remain limited. Implementation challenges include data integration complexity, computational requirements, regulatory considerations, and validation requirements. CONCLUSION: Digital twin technology represents a paradigm shift in IVF practice, offering personalised, predictive, and precision medicine approaches. This review synthesizes existing evidence to propose an integrative conceptual model for digital twin implementation across the IVF treatment spectrum, identifies critical knowledge gaps, and establishes research priorities to advance clinical translation. Despite current limitations, continued advancement promises improved success rates and patient outcomes.

Humans

Challenges and future directions in AI-driven biomaterials for microbiome-associated oral infectious diseases: A systematic review.

Oral biofilm-induced antimicrobial resistance is the core pathogenic mechanism of microbiome-associated oral infectious diseases (dental caries, periodontitis, peri-implantitis, and endodontic infection). Traditional therapies and biomaterials are limited by poor biofilm penetration, drug resistance induction, single functionality, and inadequate adaptation to dynamic oral microenvironmental changes (e.g., pH fluctuations, salivary rinsing, masticatory stimulation). Artificial intelligence (AI) has transformed the field by integrating materials science, microbiology, and stomatology data. Via machine learning, deep learning, and multi-physics simulation, AI optimizes biomaterial physicochemical properties, decodes microenvironmental signals, constructs precise sensing-response loops, and supports the full chain of material design, performance prediction, and action simulation, advancing treatment from empirical intervention to precision regulation. This systematic review retrieved literature from PubMed, Embase, and Web of Science (January 2016-January 2026) using keywords across three dimensions: AI, biomaterials, and oral microbiome. Following inclusion/exclusion criteria, 99 articles were included. It elaborates on five core mechanisms of AI-driven oral biomaterials (precise oral microbiome analysis, targeted material design/optimization, performance prediction/simulation, targeted delivery/intervention, effect evaluation/dynamic regulation), analyzes their applications in microbiome-targeted biomaterial research and development (R&D) and clinical practice for the four major oral infectious diseases, addresses technical bottlenecks (insufficient targeting specificity and precision of biomaterials, poor stability and durability in complex oral microenvironments, inadequate biofilm disruption capacity, and clinical translation obstacles), and proposes future directions (multimodal design to enhance targeting specificity, structural and component optimization to improve stability/durability, development of multi-mechanism synergistic biofilm disruption strategies, strengthening translational research for clinical application, and deep integration of AI in the full chain of biomaterial R&D). This work provides comprehensive theoretical and practical support for the R&D, optimization, and clinical translation of AI-driven microbiome-targeted oral biomaterials.

Humans

Alternative genetic codes in bacteria and archaea identified with a fast k-mer-based algorithm.

The genetic code is conserved across all domains of life and is often described as universal. Nevertheless, many exceptions to the "universal" code have now been documented, most of these through manual or semiautomated inspection of highly conserved genes. Modern bioinformatics tools improved our ability to find alternative genetic codes but remain computationally expensive, preventing widespread use on thousands of new species identified by sequencing environmental samples. Here, I report a >100-fold accelerated method for inferring the genetic code directly from assembled genomes and apply it to thousands of previously uncharacterized assemblies from archaea and bacteria. I describe three candidate genetic code variations, one of which, an alternative genetic code used by a family of Asgard archaea, is a unique example of sense codon reassignments for this domain. Identifying genetic code variations is important for understanding evolution of the standard code and improving accuracy of protein databases and open reading frame identification.

Genetic Code

Meniscal preservation in the age of biologics: toward a quantitative decision algorithm for personalized repair.

BACKGROUND: Despite advances in arthroscopic repair and biologic augmentation, surgical indication for meniscal tears remains heterogeneous. No standardized framework currently integrates biomechanical, clinical, and biological determinants to guide repair versus resection. PURPOSE: To develop a quantitative decision model-the Meniscal Preservation Score (MPS)-that unifies biomechanical and biological evidence to stratify reparability potential and standardize treatment selection in meniscal surgery. METHODS: A systematic evidence synthesis conducted in accordance with PRISMA 2020 reporting standards of studies published from 2000 to 2025 in PubMed, Embase, and Scopus identified key determinants of meniscal healing. Five consistent predictors-patient age, vascularity, tear morphology, associated pathology, and activity profile-were weighted through a two-round modified Delphi consensus among ten experienced knee surgeons. The resulting 0-9-point MPS was incorporated into a stepwise decision tree linking lesion morphology, biological context, and surgical strategy. Conceptual validation used 50 simulated cases and a retrospective cohort of 45 patients to test agreement between algorithm recommendations and expert surgical decisions. RESULTS: The MPS achieved 86% concordance with expert judgment in simulation and 84% agreement in clinical validation. In this retrospective exploratory cohort, cases in which surgical management was concordant with MPS recommendations demonstrated higher mean IKDC scores at 24&#xa0;months and lower observed reoperation rates. These findings should be interpreted as associative rather than causal, as treatment allocation was not controlled and discordant cases may have represented inherently more complex pathology. CONCLUSION: The MPS represents an evidence-informed decision-support framework designed to systematize reparability assessment. While exploratory analyses suggest structural coherence with expert reasoning, prospective implementation and external validation are required before clinical adoption as a predictive tool. LEVEL OF EVIDENCE: conceptual model with exploratory validation.

Humans

From prediction to mechanism: Explainable AI uncovers plasma and CSF proteomic signatures of Alzheimer's disease.

Alzheimer's disease (AD) plasma and cerebrospinal fluid (CSF) proteomics can distinguish AD from cognitively normal controls, but the generalizability of machine learning performance and the recurrence of biological signals across datasets require cautious interpretation. We developed an explainable artificial intelligence framework spanning two fluids and four ADNI proteomic datasets, covering 2082 modality specific samples, all analysed internally within ADNI. Phase 1 analysed plasma using a 119 analyte NULISA and targeted UPENN panel (n&#xa0;=&#xa0;727; 216&#xa0;CE, 511 controls). Phase 2 extended the analysis to CSF using SOMAscan7k, TMT-MS and targeted SET2, with Elecsys A&#x3b2;42, A&#x3b2;40, total tau and p-tau181 as anchor biomarkers. Only SOMAscan was subject-independent relative to Phase 1 plasma; TMT-MS and SET2 overlapped with Phase 1 for 96.0% and 97.7% of subjects and therefore are not independent replication cohorts. Under subject-level splits with fold internal preprocessing, we compared Elastic Net, Explainable Boosting Machines and gradient boosted trees with SHAP-based explanations. Among the candidate pipelines, we selected the pipeline with the highest held-out test ROC AUC for each platform; the selected values were 0.927 in plasma and 0.954-0.973 across the three CSF datasets. Because the same held out test performance was used for pipeline selection and headline reporting, these are optimistically selected single-holdout estimates, not unbiased estimates of generalizable or clinical performance. Explanations identified five recurring biological axes within ADNI: cholinergic (ACHE), tau/14-3-3 (YWHAG, YWHAZ, YWHAB, YWHAE), neuro-axonal (NEFL, NEFH), microglial/complement (CHIT1, SMOC1, CHI3L1, C7, CFH) and synaptic (NPTXR, NPTX2, DLG4, SYT5, VSNL1, ELAVL2). CSF analyses showed synaptic vesicle-cycle enrichment (q&#xa0;=&#xa0;2&#xa0;&#xd7;&#xa0;10-6), and CSF YWHAG correlated strongly with total tau (&#x3c1;&#xa0;=&#xa0;0.87). Cross-fluid directional concordance was modest overall (54-57%) but increased to 73-80% among mapped analyte/protein rows reaching q&#xa0;<&#xa0;0.05 in CSF. These findings provide hypothesis-generating, internally supported evidence within ADNI. Independent external cohorts with locked pipelines are required to evaluate generalizable performance and biological reproducibility; the overlapping TMT-MS and SET2 analyses should not be interpreted as independent replication.

Alzheimer Disease

Fundamentals of pacemakers ECG interpretation - part 2.

BACKGROUND: Modern pacemakers incorporate arrhythmia-response algorithms, ventricular pacing minimization protocols, and safety mechanisms that generate ECG patterns indistinguishable from pathological AV block, sensing malfunction, or device-mediated tachycardia. Failure to recognize these algorithm-driven signatures leads to unnecessary interventions, misdiagnosis, and inappropriate device reprogramming. This manuscript is the second in a two-part series on pacemaker ECG interpretation. METHODS: We conducted a narrative review of peer-reviewed literature and device-specific documentation on algorithm-driven ECG behavior, synthesizing evidence across arrhythmia recognition, upper rate physiology, ventricular pacing minimization, mode switching, safety mechanisms, and hysteresis algorithms. RESULTS: Pacemaker-mediated tachycardia produces regular paced wide-complex tachycardia locked at the upper tracking rate, initiated by any event with retrograde VA conduction. Ventricular tachycardia is identified by QRS morphology diverging from the known paced pattern, absent pacing spikes, and AV dissociation. Upper rate Wenckebach behavior mimics Mobitz type I AV block; 2:1 upper rate response mimics second-degree AV block. Ventricular pacing minimization algorithms produce isolated nonconducted P waves and prolonged AV intervals that simulate pathological conduction disease. Mode switching causes abrupt rate drops misidentified as output failure. Ventricular safety pacing generates a conspicuously short, fixed AV interval. Three discrete pacing artifacts reflect AV-sequential cardiac resynchronization therapy (CRT), ventricular safety pacing in CRT, or His-bundle pacing with backup RV output. Rate and AV hysteresis produce pauses and wandering AV intervals mimicking oversensing or Wenckebach periodicity. CONCLUSIONS: Recognizing algorithm-driven ECG patterns requires knowledge of device timing intervals and refractory periods, which lets clinicians distinguish programmed behavior from true malfunction or cardiac arrhythmia.

Humans

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Bioactive peptides for meat quality and preservation: Integrating peptidomics and computational screening.

Bioactive peptides generated from meat proteins, fermented meat products, and slaughter by-products have attracted increasing attention as functional molecules for improving meat quality and preservation. In meat systems, peptides can be produced through endogenous postmortem proteolysis, microbial fermentation, gastrointestinal digestion, or controlled enzymatic hydrolysis of underutilized animal by-products. These peptides are closely associated with key meat science endpoints, including postmortem tenderization, oxidative stability, color retention, flavor development, microbial inhibition, and the valorization of processing by-products. However, although high-resolution peptidomics has greatly expanded the identification of meat-derived peptide sequences, their translation into practical meat applications remains limited by matrix interactions, processing stability, sensory constraints, safety concerns, and insufficient validation in real meat systems. This review synthesizes recent advances in meat-related peptidomics and computational screening, including sequence-based prediction, machine learning, molecular docking, molecular dynamics, stability assessment, and safety-oriented filtering. Particular attention is given to how these approaches can prioritize peptides with antioxidant, antimicrobial, flavor-modulating, and preservation-related functions under meat-specific technological constraints. By integrating peptide generation pathways, mass spectrometry-based identification, in silico prioritization, and meat quality endpoints, this review proposes a stage-gated framework for translating meat-derived bioactive peptides from discovery to application. Future research should strengthen matrix-specific validation, standardized peptidomic reporting, and safety assessment to support the use of bioactive peptides in meat quality improvement, clean-label preservation, and circular utilization of meat industry by-products.

Animals

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

PaNDA: Efficient Optimization of Phylogenetic Diversity in Networks.

Phylogenetic diversity (PD) plays an important role in biodiversity, conservation, and evolutionary studies by measuring the diversity of a set of taxa based on their phylogenetic relationships. In phylogenetic trees, a subset of k taxa with maximum PD can be found by a simple and efficient greedy algorithm. However, this algorithmic tractability is lost when considering phylogenetic networks, which incorporate reticulate evolutionary events such as hybridization and horizontal gene transfer. To address this challenge, we introduce PaNDA (Phylogenetic Network Diversity Algorithms), the first software package and interactive graphical user-interface for exploring, visualizing, and maximizing diversity in phylogenetic networks. PaNDA includes a novel algorithm to find a subset of k taxa with maximum diversity, running in polynomial time for networks of bounded scanwidth, a measure of tree-likeness of a network that grows slower than the well-known level measure. This algorithm considers the variant of PD on networks in which the branch lengths of all paths from the root to the selected taxa contribute towards their diversity. We demonstrate the scalability of this algorithm on simulated networks, successfully analyzing level-15 networks with up to 200 taxa in seconds. We also provide a proof-of-concept analysis using a phylogenetic network on Xiphophorus species, illustrating how the tool can support diversity studies based on real genomic data. The software is easily installable and freely available at https://github.com/nholtgrefe/panda. Additionally, we extend the definition of PD to semi-directed phylogenetic networks, which are mixed graphs increasingly used in phylogenetic analysis to model uncertainty of the root location. We prove that finding a subset of k taxa with maximum diversity remains NP-hard on semi-directed networks, but do present a polynomial-time algorithm for networks with bounded level.

network