Search PubMedSearch

SEARCH · Search PubMed

Results for “traceability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

15 recordsLinked to original sources

Trustworthy Agentic AI in Bioinformatics: From Workflow Automation to Traceable and Validated Biological Inference.

Agentic artificial intelligence is extending bioinformatics beyond conversational assistance by enabling systems to select tools, execute code, revise analytical plans, and interpret biological data. These capabilities may accelerate research, but they also redistribute decisions that determine whether biological conclusions are valid. We conducted a targeted, structured PubMed search in July 2026 and identified 11 peer-reviewed agentic bioinformatics systems for descriptive review based on predefined eligibility criteria for analytical decision-making, tool or code execution, iterative evaluation, or coordinated agent activity. The evidence base covered single-cell transcriptomics, microbial genomics, cancer genomics, and omics applications, together with methodological literature on reproducibility and biological validation. We examined how current systems report delegated authority, provenance, validation, evidence, abstention, and human oversight. Existing platforms implement safeguards such as sandboxed execution, restricted commands, interaction logs, evidence identifiers, automated checks, critic agents, quality scores, and expert assessment. However, published reports rarely provide a connected account linking the original biological question to samples, reference resources, analytical decisions, computational actions, statistical results, supporting evidence, validation outcomes, and final claims. We distinguish inherited bioinformatics errors, errors amplified through autonomous action, and emergent failures arising from memory, retrieval, tool interaction, or agent coordination. We further propose a multidimensional decision-rights profile, consequence-sensitive validation gates, and a claim-to-evidence provenance architecture organized through the Traceable History of Research Evidence, Agent Actions, and Decisions in Bioinformatics (THREAD-Bio) framework. Illustrative cases show that technically successful execution may still support misleading inference. Trustworthy agentic bioinformatics therefore requires claims to remain reconstructible, challengeable, validated, and proportionate to the evidence.

accountable autonomy

PheBee: A Graph-Aware System for Scalable, Traceable, and Semantic Phenotyping.

OBJECTIVES: Phenotype-driven workflows in clinical and translational research require standardized ontology-based representation, ontology-aware cohort discovery, and provenance inspection for each assertion. Existing approaches optimize either for semantic traversal or scalable batch analytics, but not both. We describe PheBee, a hybrid system that links semantic assertions to scalable evidence storage via a deterministic identifier, preserving provenance while supporting ontology-aware discovery at cohort scale. MATERIALS AND METHODS: PheBee represents phenotype assertions in a knowledge graph as ontology-linked nodes with clinical modifier context (e.g., negated, family history), and stores supporting evidence records in a scalable row-oriented evidence table for cohort-scale access. The two layers are connected by a deterministic identifier enabling stable joins across repeated ingestions without duplicating high-volume evidence in the graph. We evaluated PheBee using synthetic datasets designed to exercise end-to-end ingestion and query workflows. RESULTS: Functional evaluation validated hierarchical term expansion, qualifier-aware retrieval, duplicate-free assertion handling under re-ingestion, and privacy-conscious management of subjects shared across multiple research projects. At scale (10,000 subjects producing 12M evidence records) PheBee completed ingestion in ~30 minutes and responded to interactive queries within 6 seconds under concurrent load. DISCUSSION: PheBee exposes a unified API for ontology-aware cohort discovery with hierarchical term expansion, subject-centric retrieval of phenotypes and clinical modifiers, and evidence and provenance queries. Its data model aligns with GA4GH Phenopackets, facilitating interoperability with phenotype exchange standards. CONCLUSION: By combining ontology-aware semantics with scalable, provenance-bearing evidence storage, PheBee provides a practical open-source foundation for phenotype-driven research workflows that demand both semantic precision and cohort-scale traceability.

cohort studies

MetaServe: a lightweight, metadata-aware governance and delivery layer for pre-publication research omics data.

BACKGROUND: Institutional research teams and core facilities routinely manage pre-publication omics datasets that span heterogeneous file types, nested project structures, and multiple downstream uses. Public repositories mainly support post-publication dissemination, while workflow systems and enterprise data platforms do not directly provide a lightweight governance and delivery layer for internal research assets. RESULTS: We present MetaServe, an open-source governance and delivery layer for pre-publication research assets in institutional multi-omics settings. MetaServe registers and delivers heterogeneous assets, including sequencing files, processed matrices, imaging data, analysis-ready objects, tabular files, and documents, without requiring repository-grade standardization. Its metadata-aware design combines file-type recognition, partial automatic extraction for selected formats, manually supplied project and biological annotations, and indexed faceted retrieval. MetaServe supports authenticated web download, viewer-oriented handoff for compatible services such as cellxgene, and path-manifest export for downstream workflows under shared-storage assumptions. The current implementation combines role-based controls, explicit file-level sharing, path-constrained delivery, and operational traceability to support controlled institutional access. MetaServe has been deployed at the Chinese Institutes for Medical Research (CIMR) as part of an institutional multi-omics data-management system. CONCLUSIONS: MetaServe provides a practical layer between institutional storage and downstream analytical platforms for pre-publication research data. Its contribution is the integration of lightweight metadata-aware registration, permission-aware retrieval, and controlled delivery for heterogeneous institutional omics assets. Rather than replacing workflow engines, public repositories, or enterprise-scale research data platforms, MetaServe offers a deployable governance layer for core facilities and collaborative teams that need structured discovery and traceable delivery before public deposition or manuscript release.

Metadata

Environmental Release of Genetically Intervened Microorganisms: Towards a New Narrative.

The deliberate release of genetically engineered microorganisms for environmental applications has remained largely blocked since the early days of recombinant DNA technology, when limited ecological knowledge, lack of success stories and public apprehension shaped a culture of caution and restrictive regulation. Despite profound advances in microbial ecology, synthetic biology and genetic design, current frameworks still rely on outdated assumptions and legacy regulations that equate engineered microbes with inherent danger and demand unrealistic forms of absolute containment. This review examines how laboratory-trained microorganisms exist on a continuum with naturally evolved life, and that their risks are neither categorically different nor greater. Rather than pursuing unachievable containment, governance should shift towards traceability, stewardship and long-term monitoring through genomic barcodes, digital twins and transparent oversight. The vision moves from domination and control to care and partnership recognizing engineered microbes as live amendments capable of restoring degraded ecosystems. Achieving this transformation requires new terminology, phased field-trial frameworks, improved scaling methods, and the integration of epistemological perspectives that emphasize reciprocity and coexistence with nature. Reframing biotechnology in this way could finally unlock the capacity of engineered microorganisms to contribute responsibly and effectively to planetary repair in an era of escalating environmental crises.

Microorganisms, Genetically-Modified

Selective Extraction of Genomic DNA From Animal Tissues Using a Hydrophobic Magnetic Ionic Liquid.

The development of green and efficient methods for genomic DNA extraction from animal tissues is crucial for molecular diagnostics, food traceability, and genetic research. Conventional methods often involve toxic reagents, multiple centrifugation steps, and are time-consuming. In this study, a hydrophobic magnetic ionic liquid (MIL), N-octyl-4-dimethylaminopyridinium hexafluorophosphate MIL ([C8DMAP][PF6]‑Ni MIL), was synthesized and applied for the selective extraction of genomic DNA from various animal tissues. The material exhibited strong paramagnetic behavior, high thermal stability, and excellent hydrophobicity, enabling rapid phase separation under an external magnetic field. A mechanical shaking-assisted extraction method was developed, and key parameters including temperature, time, shaking speed, and [C8DMAP][PF6]-Ni MIL dosage were systematically optimized. The method demonstrated high selectivity for DNA over proteins, RNA, and amino acids, with a maximum recovery rate of 78.06 ± 1.91%. Compared to a commercial DNA extraction kit, the [C8DMAP][PF6]-Ni MIL-based approach provided higher yields from several tissues, including mouse liver, brain, and rabbit lung. Furthermore, the [C8DMAP][PF6]-Ni MIL could be reused for at least six cycles while maintaining extraction efficiency. This work not only provides a high-performance material for DNA extraction, but also demonstrates a sustainable and easily retrievable liquid-phase separation strategy, offering a generalizable platform for complex sample pretreatment.

Animals

Evaluation of Vaccinia Virus Infection in Mice Using Two-Reporter Recombinant Virus.

The family Poxviridae comprises multiple viruses with large double-stranded (ds) DNA genomes that can infect numerous vertebrate and invertebrate hosts, including humans. The development of genetic engineering methods for Vaccinia virus (VACV), the prototypic member in the family, have allowed the manipulation of the genomes of poxviruses for the generation of recombinant (r)VACV expressing easily traceable luciferase and/or fluorescent reporter genes. These recombinant viruses have significantly contributed to progress in the field of poxvirus research and accelerated the development of novel prophylactic vaccines and therapeutic antiviral treatments. Recently, we described two reporter rVACV expressing luciferase (Nluc) and fluorescent (GFP or Scarlet) proteins to easily track viral infections in different systems, overcoming the limitations associated with the use of rVACV expressing a single luciferase or fluorescent reporter gene. Here, we describe the experimental procedures to carry out in vitro, in vivo and ex vivo studies using these novel bireporter-expressing rVACV, which also represent an excellent option to study the biology of VACV, including the use of these reporter viruses for testing new antivirals and vaccines, using cultured cells and/or well-characterized animal models of infection.

Animals

Use of Rift Valley Fever Virus Expressing NanoLuc Luciferase for the Assessment of Neutralizing Antibodies and Antivirals.

Rift Valley fever (RVF) is an arboviral zoonotic disease affecting many African countries with the potential to spread to other geographical areas. In this chapter we describe the use of a replication-competent recombinant (r)RVFV expressing NanoLuc Luciferase (Nluc) for in vitro studies. The determination of parameters such as neutralizing antibodies in serum samples, or the antiviral activity of drugs is usually carried out using standard assays based on the assessment of cytopathic effect on cell cultures. The use of a virus encoding a traceable reporter protein allows to correlate the presence or absence of infection with the detection of the product in the infected cultures, thus tracking the level of RVFV infection in an objective, quantitative manner. In addition to this quantitative measurement of results, our protocol offers two other advantages, such as a shorter time to read, given that 48 h post-infection the production of the reporter protein is enough to give an accurate result, and the use of an attenuated virus, which reduces the risk of exposure.

Rift Valley fever virus

Measurement of low-density lipoprotein cholesterol and other circulating lipids in Brazil: a systematic literature review.

Accurate laboratory assessment of circulating lipids underpins cardiovascular risk stratification, yet clinical interpretation depends not only on the assays but on the formula chosen to estimate low-density lipoprotein cholesterol (LDL-C). This review integrates the 2019-2025 evidence on laboratory methods for triglycerides (TG), total cholesterol (TC), and high-density lipoprotein cholesterol (HDLC), and on the formulas estimating LDL-C, VLDL-C, and non-HDL cholesterol, to determine how these should be measured, reported, and harmonized in Brazil, where lipid thresholds are adapted from international consensus. A PRISMA 2020 systematic search (PROSPERO CRD420251241064) of PubMed/MEDLINE, Scopus, SciELO, LILACS, Web of Science, and Embase retrieved 57,915 records; after removing 38,210 duplicates, 19,705 titles/abstracts were screened, 312 full texts assessed, and 25 sources included. Enzymatic colorimetric assays remain standard for TG, TC, and HDLC. For LDL-C, Martin/Hopkins classifies more accurately than Friedewald (89.6% vs 83.2% correct categorization in 5,051,467 patients), particularly at high TG and low LDL-C, while Sampson/NIH and modified Sampson/NIH extend reliable estimation into hypertriglyceridemia and very low LDL-C; direct measurement is reserved for TG beyond the validated range. Although the review centers on the Friedewald, Martin/Hopkins, and Sampson/NIH families that dominate guideline practice, other published equations exist and are addressed in context. In Brazil, atherogenic-lipid thresholds are risk-based decision limits rather than reference intervals; national surveys describe lipid distributions but were not designed to establish them. Analytical standardization through traceability programs, multicenter validation of formulas, and-where the distribution-based construct applies (HDLC, pediatrics)-nationally derived reference intervals are priorities for equitable cardiovascular risk assessment in Brazil.

Humans

Toward ethical provenance tracking: The GA4GH model data access agreement (DAA).

PURPOSE: Standardizing contractual clauses that govern data access enables research institutions to responsibly steward genomic and related health data while enabling its efficient downstream reuse. METHODS: We describe a document analysis study using both qualitative and comparative law analytical approaches to identify the most common categories of clauses from 29 different data access agreements used by human biomedical research consortia globally. We furthermore characterized the legal positions and standard practices for each common element of the agreement and synthesized across them to develop model clauses. A total of 3 discussion sessions were organized virtually to refine the clauses among members of the Ethical Provenance Subgroup of the Global Alliance for Genomics and Health. RESULTS: We developed 15 unique data access clauses corresponding to the most common legal elements identified in the sampled agreements. CONCLUSION: Model clauses can be used to drive administrative efficiencies and institutional compliance for managing access to human genomic data for research. Additional machine-readable consents and software solutions are needed to support traceable "ethical provenance" of human genomic data and communicate data use conditions throughout the data's life-cycle.

Humans

Profiler: an open web platform for multi-omics analysis.

MOTIVATION: High-throughput multi-omics technologies produce increasingly large and heterogeneous datasets that are difficult to analyze without advanced computational expertise. Existing bioinformatics tools are often fragmented or limited to specific omics types, hindering reproducibility and accessibility. There is a critical need for an integrated, user-friendly, and scalable platform capable of supporting multi-omics analyses across different data modalities. RESULTS: We present Profiler, an open-source, modular platform that unifies data import, quality control, preprocessing, statistical testing, machine and deep learning, biomarker discovery, pathway and drug-target enrichment, and survival modeling within a single reproducible environment. Built in Python with Streamlit, Profiler is available as both a web-based platform deployed on high-performance computing and a desktop version for local execution, enabling flexible usage across computational infrastructures. Profiler supports diverse omics modalities, including proteomics, transcriptomics, lipidomics, and electroencephalogram data. Through applications to glioblastoma proteomic, pancancer, and multi-omics datasets, Profiler reproduced known molecular subtypes, revealed potential therapeutic targets, and generated fully traceable analysis reports within minutes. By integrating advanced analytics behind an intuitive interface, Profiler democratizes multi-omics analysis and provides a robust, scalable foundation for systems biology and precision medicine research. AVAILABILITY AND IMPLEMENTATION: Profiler is open-source and freely available via its web platform (https://prism-profiler.univ-lille.fr) and GitHub (web version: https://github.com/yanisZirem/Profiler_v1_requests_datatests, desktop version: https://github.com/yanisZirem/prism-profiler), and archived on Zenodo (DOI: https://doi.org/10.5281/zenodo.17478158).

Software

FANTASIA suite: a reproducible and configurable framework for embedding-based functional annotation of proteins.

Embedding-based annotation transfer is increasingly used for protein function inference due to protein language models capture sequence, structural, and functional signals that may extend beyond conventional pairwise similarity. However, systematic application of these approaches requires control over model choice, reference composition, lookup parameters, evidence traceability, and output formats. We developed the FANTASIA suite, a configurable framework for embedding-based functional annotation of proteins. The suite combines a database-backed implementation for reproducible and extensible analyses with a portable flat-file implementation for rapid local annotation and pipeline integration. Using non-model and model-organism proteomes, we show that larger neighbourhood sizes remain practical for proteome-scale analyses and that taxonomy and sequence-identity filtering support leakage-aware benchmarking. We also compare the supported models with baseline methods through external CAFA5 evaluation and provide practical guidance based on empirical evidence variables. FANTASIA provides a controlled, scalable, and reproducible framework for extending functional annotation across the rapidly expanding diversity of sequenced organisms.

Software

ClarID: A Human-Readable and Compact Identifier Specification for Biomedical Metadata Integration.

BACKGROUND: In biomedical research, subjects and biospecimens are commonly tracked using simple IDs or UUIDs, which guarantee uniqueness but convey no embedded semantic information. Contextual metadata (such as tissue type, diagnosis, or assay) is often stored separately, making integration, cohort selection, and downstream analysis cumbersome. While structured barcoding systems exist in large consortia (e.g., TCGA, GTEx) or domain-specific contexts (e.g., SPREC, GOLD), no unified, extensible framework currently spans both subjects and biosamples in a human- and machine-readable way. METHODS: We developed ClarID, a domain-agnostic specification that supports two identifier formats: (i) a human-readable form (e.g., 'CNAG_Test-HomSap-00001-LIV-TUM-RNA-C22.0-TRT-P1W' that encodes key metadata such as project, species, subject_id, tissue, assay, disease, timepoint and duration (from that event); and (ii) a compact version named 'stub' (e.g., 'CT01001LTR0N401T1W') optimized for filenames, pipelines, and labeling.ClarID is implemented through an open-source command-line tool, ClarID-Tools, which processes tabular metadata files (CSV/TSV) and uses a YAML-based codebook to generate, decode, and validate identifiers, as well as to create and read QR codes. The tool supports bulk and single-sample processing and allows easy integration with institutional workflows. RESULTS: To demonstrate ClarID's utility, we applied it to datasets from the Genomic Data Commons (GDC), generating interpretable identifiers for more than 113,000 clinical records (subjects) and 4,255 biospecimen records. All materials, including pre-processing scripts, input and encoded data, are publicly available and fully reproducible via the accompanying GitHub repository and Google Colab. CONCLUSIONS: ClarID fills a critical gap between opaque accession numbers and rich metadata schemas by embedding key context directly into structured identifiers. It enhances traceability, facilitates downstream analysis, and remains adaptable to project-specific needs through a configurable codebook. The accompanying ClarID-Tools software is freely available, together with full documentation and reproducible pipelines, at https://github.com/CNAG-Biomedical-Informatics/clarid-tools.

Biosample identifiers

Genomic Tracking of Market-Derived Bull Shark Fins Back to Source Population of Origin.

International trade of shark fins remains difficult to monitor because products are rarely labelled to species and are often highly processed, resulting in severely degraded DNA. For several shark species listed under Appendix II of the Convention on International Trade in Endangered Species of Wild Fauna and Flora (CITES), this limits external verification of source populations supplying global trade hubs. Here, we assess whether nuclear genomic approaches can be applied to market-derived bull shark (Carcharhinus leucas) fins to determine their population of origin. We analysed dried fin trimmings collected from retail vendors in Hong Kong SAR, one of the world's largest dried shark fin trade hubs, using a targeted DArTcap single nucleotide polymorphism (SNP) panel, originally developed for population genomic studies of this species. Despite substantial DNA degradation, genomic libraries were successfully obtained for most samples, yielding sufficient SNP data to perform robust provenance and sex assignment. Using a Bayesian mixed-stock analysis, most fin samples were assigned to the Indo-West Pacific (71.4%), with smaller contributions from the western Atlantic (22.6%) and eastern Pacific (3.0%). Genetic sex assignment revealed twice as many males as females, although results indicated a conservative bias towards male assignment due to the limited number of X-linked markers available in degraded samples. Our results demonstrate that genome-wide targeted approaches can be effectively applied to highly processed shark fin products to infer population sources and sex composition. This study provides proof-of-concept for integrating genomics into shark trade monitoring, highlighting its potential to improve traceability, support CITES implementation and inform conservation and fisheries management, particularly for species with well-resolved population structure.

Animals

Artificial intelligence for translational personalized neoantigen cancer vaccine development.

Personalized neoantigen cancer vaccine is a promising strategy for precision immunotherapy by targeting patient-specific and mutation-derived tumor antigens. Early clinical studies have demonstrated the feasibility, safety, and immunogenicity of these vaccines across multiple solid tumors, with encouraging outcomes particularly when combined with immune checkpoint blockade. However, broader clinical translation remains limited by sequential bottlenecks across the vaccine development pipeline, including false-positive neoantigen selection,  imperfect modeling of antigen processing and HLA presentation, limited prediction of T-cell receptor recognition, and challenges in formulation, delivery, and manufacturing. Artificial intelligence and advanced computational workflows are increasingly integrated into this pipeline to improve candidate prioritization and support more reproducible decision-making. In this review, we summarize clinical progress and key translational barriers in personalized neoantigen vaccination, and discuss how AI-enabled approaches may contribute across four major stages: multi-omics integration for neoantigen discovery, processing-aware HLA presentation prediction, structure-aware and TCR-informed immunogenicity modeling, and data-driven formulation optimization, particularly for lipid nanoparticle-based delivery systems. These approaches are able to help narrow biological and chemical search spaces, improve prioritization, and provide mechanistic insights into antigen presentation and immune recognition rather than replacing experimental validation. This articlefurther addresses future implementation challenges, including dataset diversity, model interpretability, prospective benchmarking, manufacturing traceability, and evolving regulatory frameworks for individualized mRNA cancer immunotherapies. Integrating computational innovation with rigorous immunological validation, scalable manufacturing, and regulatory oversight will be essential for advancing personalized neoantigen vaccines toward broader clinical implementation.

Cancer Vaccines

Digital and computational morphology in hematology: current platforms, clinical evidence, and future requirements.

INTRODUCTION: Morphologic examination of peripheral blood and bone marrow remains central to the diagnosis and classification of hematologic disorders. Conventional optical microscopy, however, is labor-intensive, dependent on operator expertise, and affected by interobserver variability. Digital morphology has developed from automated image acquisition and cell pre-classification into a broader field that includes whole-slide imaging, remote review, quantitative morphometry, and artificial intelligence-based analysis. CONTENT: This review examines current applications of digital morphology in peripheral blood, bone marrow aspirates, malaria detection, and body-fluid analysis. Commercial platforms are evaluated with particular attention to the distinction between raw automated pre-classification, expert digital post-classification, and comparison with independent optical microscopy. Digital systems generally perform well for common mature leukocyte populations but remain less reliable for rare or diagnostically critical cells, including blasts, abnormal lymphoid cells, plasma cells, and intermediate maturation stages. Research systems increasingly extend analysis from individual-cell classification to whole-slide, specimen-level, and patient-level assessment. SUMMARY: Digital morphology can improve standardization, image traceability, remote consultation, education, proficiency testing, quality assurance, and selected aspects of laboratory workflow. Its clinical value depends on appropriate validation, transparent reporting of reference methods, recognition of algorithm-specific failure modes, and clearly defined criteria for expert review and conventional microscopy. Human expertise remains essential not only for validating results but also for adapting cell taxonomies and interpretive rules to evolving classifications of hematologic diseases. OUTLOOK: Future progress will require representative multicenter datasets, harmonized morphologic terminology, external validation, interoperability with laboratory information systems, and continuous monitoring after software or hardware updates. Integration of morphology with quantitative hematology, flow cytometry, cytogenetics, genomics, and clinical data may support more comprehensive computational diagnosis. Digital platforms may also broaden access to specialist expertise, training, and quality programs in resource-limited institutions and regions, provided that infrastructure, governance, and professional competency are adequately supported.

artificial intelligence