Search PubMedSearch

SEARCH · Search PubMed

Results for “data governance”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

MetaServe: a lightweight, metadata-aware governance and delivery layer for pre-publication research omics data.

BACKGROUND: Institutional research teams and core facilities routinely manage pre-publication omics datasets that span heterogeneous file types, nested project structures, and multiple downstream uses. Public repositories mainly support post-publication dissemination, while workflow systems and enterprise data platforms do not directly provide a lightweight governance and delivery layer for internal research assets. RESULTS: We present MetaServe, an open-source governance and delivery layer for pre-publication research assets in institutional multi-omics settings. MetaServe registers and delivers heterogeneous assets, including sequencing files, processed matrices, imaging data, analysis-ready objects, tabular files, and documents, without requiring repository-grade standardization. Its metadata-aware design combines file-type recognition, partial automatic extraction for selected formats, manually supplied project and biological annotations, and indexed faceted retrieval. MetaServe supports authenticated web download, viewer-oriented handoff for compatible services such as cellxgene, and path-manifest export for downstream workflows under shared-storage assumptions. The current implementation combines role-based controls, explicit file-level sharing, path-constrained delivery, and operational traceability to support controlled institutional access. MetaServe has been deployed at the Chinese Institutes for Medical Research (CIMR) as part of an institutional multi-omics data-management system. CONCLUSIONS: MetaServe provides a practical layer between institutional storage and downstream analytical platforms for pre-publication research data. Its contribution is the integration of lightweight metadata-aware registration, permission-aware retrieval, and controlled delivery for heterogeneous institutional omics assets. Rather than replacing workflow engines, public repositories, or enterprise-scale research data platforms, MetaServe offers a deployable governance layer for core facilities and collaborative teams that need structured discovery and traceable delivery before public deposition or manuscript release.

Metadata

Longitudinal Clinical, Physiological, and Molecular Profiling of Female Patients With Metastatic Cancer: Protocol and Feasibility of a Multicenter High-Definition Oncology Study.

PURPOSE: A substantial proportion of patients receiving genomically matched therapies do not achieve clinical benefit, underscoring the influence of nongenetic factors on cancer outcomes. High-Definition Oncology (HDO) proposes integrating longitudinal, multimodal patient data-spanning clinical, molecular, physiological, and behavioral domains-to enable truly individualized cancer care. This manuscript describes the HDO study design, framework, and feasibility results in women with metastatic cancer. METHODS: We initiated a prospective, multicenter observational study (HDO study; ClinicalTrials.gov identifier: NCT06590506) enrolling 300 female patients with newly diagnosed metastatic breast, lung, or colorectal cancer. Here, we report the study design, standardized workflows, prespecified feasibility criteria, and early internal pilot results. Eleven data modalities are collected longitudinally, including tumor and germline genomics, germline epigenomics, gut microbiome, blood and stool metabolomics and proteomics, exposome characterization, wearable-derived physiological monitoring, digital footprint assessment, medical imaging, and patient-reported outcomes. Standardized workflows govern clinical procedures, data acquisition, biospecimen processing, and quality control across all participating sites. RESULTS: Feasibility was evaluated in the first 30 participants (10% of planned accrual). Patients completed 100% of scheduled clinical visits, 97.4% of planned plasma collections, 80.7% of stool samples, and all tumor biopsies. Wearable devices captured activity, heart rate, sleep, and blood oxygen saturation data during 95.0%, 84.2%, 90.6%, and 70.7% of total patient-days, respectively. Biospecimens met predefined quality control metrics across all molecular modalities. Engagement with mobile applications for pain and emotion reporting exceeded 80%. CONCLUSION: The HDO study demonstrates the feasibility of comprehensive, longitudinal, multimodal data collection in women with metastatic cancer. This internal pilot establishes an integrated framework for future analyses aimed at characterizing disease trajectories, defining molecular and physiological determinants of outcomes, and developing patient-specific computational models.

Humans

A comparison of mail, telephone, and home interview strategies for household health surveys.

The method of data collection in household health surveys can be a major determinant of cost and data quality. A survey strategy can comprise mail, telephone, or home interview methods, individually or in combination to follow up non-respondents. The purpose of this study in Montreal was to compare cost and data quality of various strategies. Strategies which began with mail or telephone contact, followed by the two other methods, provided response rates as high as a home interview strategy (all between 80 and 90 per cent), for one-half the cost of home interviews when used as the sole method. The telephone response rate was higher than the mail response rate. Comparing different follow-up approaches to strategies beginning with mail or telephone, it proved less costly, and equally effective, to use home interviewing as a last resort for persistent non-respondents. Validity of response (comparing individual responses with records of a government health insurance data bank) and willingness to answer sensitive questions were greatest in mail strategy.

Canada

Conference report: the third Bacterial Genome Sequencing Pan-European Network conference.

The third Bacterial Genome Sequencing Pan-European Network conference, held in Engelberg, Switzerland (12-15 January 2026), brought together experts from six European countries to discuss the implementation of bacterial genome sequencing in clinical microbiology and public health. Key themes included regulatory frameworks (In Vitro Diagnostic Regulation, General Data Protection Regulation), standardization, quality control, data sharing, economic evaluation, and the integration of artificial intelligence and long-read sequencing into diagnostic workflows. Across presentations, panel discussions, and workshops, participants emphasized that successful implementation of genome sequencing requires more than technical capacity: it depends on robust validation, sustainable funding, interoperable data standards, ethical governance, and interdisciplinary collaboration. The meeting highlighted that sequencing should remain question-driven and clinically meaningful, balancing cost, turnaround time, and public health impact. Overall, the conference reinforced the need for coordinated European efforts to advance responsible, standardized, and sustainable genomic surveillance and diagnostics.

bacterial genome sequencing

Dataset Readiness Assessment With Large Language Model (DRAFT-LLM): A Multi-Axis Audit Guided by LLM.

This article details the Dataset Readiness Assessment for Training (DRAFT), a systematic method for determining whether a high-dimensional biological dataset is suitable for developing reliable, equitable (i.e., the extent to which model performance, error patterns, and potential benefits or harms are evaluated and found to be acceptably distributed across relevant demographic, biological, clinical, and contextual subgroups), and scientifically meaningful machine-learning models, and DRAFT Large Language Model (DRAFT-LLM), its optional human-in-the-loop extension for calibrating study-specific audits through structured, critically reviewed LLM guidance. Standard model validation often fails to detect when apparent performance is driven by spurious correlations, technical artifacts, or hidden stratification, leading to irreproducible and inequitable findings. DRAFT-LLM addresses this gap by shifting the focus from model tuning to structured dataset auditing, organized around Support Protocols 1 to 4 that capture the scientific intent, data structure, and governance constraints of a given study. These Support Protocols: (1) elicit and formalize investigator input into a study intake and dataset card; (2) compute standardized dataset statistics and structural summaries suitable for downstream analysis and LLM context; (3) configure the language model using form-based responses, safety guardrails, and governance rules; and (4) generate personalized instructions, prompts, and code templates for running DRAFT audits. Basic Protocols 1 to 3 are instantiated from this support layer for generalization, equity, and stability: they are reusable execution patterns whose concrete behavior is determined by the cards, statistics, and configurations defined in the Support Protocols. DRAFT-LLM and DRAFT are demonstrated in this article through an end-to-end case study on The Cancer Genome Atlas (TCGA). © 2026 Wiley Periodicals LLC. Support Protocol 1: Study intake and dataset card construction Support Protocol 2: Dataset structure and advanced summary statistics for LLM context Support Protocol 3: LLM configuration using structured form responses Support Protocol 4: Generation of personalized instructions for DRAFT audits Basic Protocol 1: Generalization audit Basic Protocol 2: Equity audit Basic Protocol 3: Stability audit.

Large Language Models

The prevalence and severity of major disabling conditions--a reappraisal of the government social survey on the handicapped and impaired in Great Britain.

This paper re-examines the data gathered for the Government Social Survey on the Handicapped and Impaired in Great Britain. The underlying cause of disablement is considered in conjunction with severity and prevalence. When these are taken together a picture emerges in which stroke, arthritis, and circulatory disorders are the most frequent cause of severe disability in the community. An attempt is also made to examine the way in which the survey might be biased by the non-inclusion of those in institutions.

Adult

Chemogenomic maps reveal a PRDX1-dependent iron-damage axis in the DNA damage response.

The DNA damage response (DDR) is a sophisticated network of cellular pathways whose perturbation leads to genome instability and is a key hallmark of oncogenesis. Here, we present data from 32 genome-scale loss-of-function CRISPR interference chemical-genetic screens with inhibitors targeting core constituents of the DDR machinery (PARP, ATR, ATM, DNAPK and WEE1), as both single agents and in combination with poly(ADP-ribose) polymerase inhibitors. These experiments identify >1,000 genes whose perturbation modifies the DDR and provides a rich resource to the DDR community. In addition, this compendium of functional genomics data reveals key principles governing the DDR and highlights a strong chemical-genetic interaction between loss of activity of the peroxiredoxin PRDX1 and all tested DDR inhibitors through a mechanism involving iron availability mediated by an MRGBP-PAX7-IREB2 axis. Our data position PRDX1 as a key suppressor of DNA damage accumulation and potential druggable target in combination with DDR inhibitors.

Journal Article

The legal setting in prescribing drugs.

By training and experience physicians generally prescribe drugs with only the patient in mind. In today's environment it is essential that the physician stop and think of himself also. We are all aware of malpractice problems, which begin in the physician's office or at the hospital and, if severe, end in court. The ophthalmologist must now consider what a disgruntled patient can find in writing to use in court. Hospital records and physician's charts are now just the beginning point. The Physician's Desk Reference (PDR), the published literature, the package insert, the promotional material of pharmaceutical companies, the internal records of drug companies, the pharmacists' records, and the data now obtainable from government agencies all provide happy hunting grounds for a competent plaintiff's attorney. The ophthalmologist needs to be keenly aware of these potentially adverse resources and how they can be used against him. He also needs to know the alternatives available to him. Once these matters are brought into clear perspective, both the patient and the ophthalmologist will be better off.

Drug Prescriptions

Worldwide Innovative Network Consortium: Building a Common Global Cancer Database.

This review shares the ongoing work of the global Worldwide Innovative Network (WIN) Consortium for Precision Medicine to synthesize emerging cancer treatment data and to define the requirements for a common global cancer database that can truly support precision oncology. We performed a narrative review of emerging cancer treatment data, molecular profiling technologies, and existing clinicogenomic databases, focusing on how tumors are characterized, how subgroups are defined, and how demographic, lifestyle, and environmental factors are captured. The growth in molecular profiling technologies and the development of new targeted therapies are transforming cancer care. Tumors, regardless of tissue origin, are increasingly defined as composites of multiple, often rare, subgroups, each with distinct biology and likely response to specific therapies, based on multidimensional profiling of the tumor and its microenvironment. The solution lies in building vast databases that capture racial and ethnic diversity, reflected in genomic data, as well as diet and lifestyle factors that may have epigenetic impact on gene expression and post-translational modifications. A truly inclusive and informative data set must reflect global diversity, and there are multiple examples of demography-dependent differences in genomic signals. With members caring for and studying patients with cancer across five continents, WIN is actively exploring pathways to create a global cancer database, rich in clinical and molecular detail, granular enough for precise analysis, and large enough to power artificial intelligence-driven insights, provided appropriate data quality, validation, and governance frameworks are in place. This review surveys the current landscape and outlines practical paths forward to achieve this goal.

Humans

A decentralized future for the open-science databases.

The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR's federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.

data accessibility

Uniform basic data sets for health statistical systems.

The United States approach to coordinating health statistics involves introduction of multipurpose basic data sets describing health status and the health care system. Standard reporting procedures have been used for many years for vital statistics. Recently designated data sets cover health manpower, inpatient facilities, short-stay hospital discharges, and use of ambulatory care services. A data set for long-term health care is in the design stage. Advantages of this approach in the United States and internationally are: basic comparisons can be made between health care settings are geographic areas while maintaining the variety and flexibility of existing public and private information systems; shared local, regional, and national data systems can be set up; and better coordination can be achieved between government-sponsored general-purpose and administrative data systems. Problem areas are: avoiding undue proliferation, e.g. of disease-specific data sets; adhering to the principle of minimal requirements; linking data sets and coordinating them with census and other social indicators; promoting widespread use; assuring data quality; establishing mechanisms for review and revision; and extending the concept internationally.

Data Display

Health and safety information for regulatory purposes--an industrial point of view.

This article describes how one company (The Dow Chemical Company) is managing the issue of increasing demands by regulatory agencies for health and environmental information. Described are: an interdisciplinary organization linked by a resource and communication network; methods of evaluating information requests and establishing priorities for response; and problems in communicating with the regulators. The need for a responsive technical dialogue between government and industry is stressed.

Data Collection

Privacy, security, and reliability risks of artificial intelligence in healthcare: a systematic review of empirical evidence.

BACKGROUND: Artificial intelligence (AI) is increasingly integrated into healthcare information systems, supporting clinical decision-making, imaging analysis, and predictive modeling. While these applications offer operational and clinical benefits, they also introduce emerging risks to patient privacy, data security, and system reliability. OBJECTIVE: To systematically review empirical evidence on privacy breaches, security vulnerabilities, and misuse associated with AI applications in healthcare settings. METHODS: PubMed, Embase, Web of Science, Scopus, IEEE Xplore, and ACM Digital Library were searched for empirical studies published between January 2015 and November 2025 that evaluated AI use or misuse in clinical diagnosis, treatment, or decision-making. Two reviewers independently screened studies and extracted data using a standardized form. Findings were synthesized narratively due to heterogeneity in study designs, AI methods, and reported outcomes. RESULTS: Of 7,285 records identified through database searches and 205 through citation screening, 22 empirical studies met the inclusion criteria, spanning multiple clinical domains and data modalities, predominantly medical imaging applications. Five recurring threat categories were identified: patient re-identification, membership inference, unauthorized access and adversarial exploitation, input manipulation, and misuse or overinterpretation of AI outputs. Across studies, AI models were shown to encode latent biometric signals across diverse data types, limiting the effectiveness of traditional anonymization and synthetic data approaches. Adversarial attacks and input manipulation were also shown to compromise diagnostic performance and system integrity. CONCLUSION: This systematic review provides empirical evidence suggesting that contemporary AI systems in healthcare introduce privacy and security risks that may challenge traditional assumptions about data protection. These findings underscore the need for privacy- and security-by-design approaches and governance frameworks that address risks across the AI lifecycle.

Humans

Genetic control of immunoregulatory circuits. Genes linked to the Ig locus govern communication between regulatory T-cell sets.

Antigen-stimulated Ly1:Qa1+ cells induce a nonimmune set of T-acceptor cells (surface phenotype Ly123+Qa1+) to participate in the generation of specific suppressive activity. The experiments reported here were designed to test the possibility that the interaction between T-inducer and T-acceptor cells might be governed by genes linked to the Ig locus. We find that inducer:acceptor interactions occur only if the inducer and acceptor T-cell sets are obtained from donor that are identical at the Ig locus and are independent of the Ig locus expressed on the B cells used for assay of T-helper activity. In addition, experiments using inducer and acceptor T cells from the congenic recombinant BAB. 14 strain show that T-T interactions are not governed by Ig-CH genes, per se. These data indicate that T-inducer: T-acceptor interactions are governed by Ig-linked genes that may control expression of VH-like structures on T cells, or control expression of as yet unidentified cell-surface molecules.

Animals

Institutional data commons: a federated Data Use Certification-aware architecture for secure and scalable data use in biomedical data ecosystems.

BACKGROUND: Modern biomedical data ecosystems increasingly rely on global cloud platforms to coordinate access to large-scale genomic and clinical datasets. However, operational governance remains largely investigator-centric, shifting the responsibility for complex security, compliance, and infrastructure management to individual laboratories. As data volumes and regulatory requirements expand, this approach fails to scale across the research enterprise. This disjointed approach creates a substantial governance burden and can slow down scientific progress. In centralized cloud environments, investigators face siloed identity management and high costs, leading to inefficient data use and increased risk when integrating local and global datasets. MATERIALS AND METHODS: We examine limitations in the current infrastructure and propose reframing institutional data commons as governance-aware intermediaries to ensure secure, efficient and sustainable use of controlled-access biomedical data. RESULTS: This federated architecture decouples storage from authorization, enabling dynamic access linked to active certifications, whether data are analyzed in situ on global platforms or in local governance-aware institutional access environments. DISCUSSION: Shifting governance from investigators to institutional infrastructure ensures that biomedical research remains both secure and economically sustainable.

biomedical data ecosystems

Latent allotypes: a window to a genetic enigma.

Recent data concerning the expression of latent allotypes (allotypes present in low concentration and not anticipated from breeding data) are presented. Using various sensitive assays, latent allotypes of groups a, b and d have been observed in sera, IgG samples, isolated antibody fractions and on lymphoid cell surfaces. Recently, molecules bearing latent group a allotypes have been isolated from IgG samples and from specific antibody fractions by immunoadsorbent chromatography and subsequently typed for VH and CH allotypes. In several preparations latent group d allotypes were observed. IgG clearance rates were measured by intravenous injection of rabbits with differentially radiolabelled (125I and 131I) allotype matched and non-matched IgG samples. In no instance was allotype matched IgG cleared faster than non-matched, although the converse was true in several rabbits, suggesting in vivo recognition of allotypes as a possible regulatory mechanism. Genetic models that can account for latent allotypes require the presence of information in the genome that is not expressed under normal conditions, and, furthermore, these models must include inherited regulation mechanisms to govern synthesis from this information. Available data do not require that all Ig genes are present in all members of a species, but may suggest the presence of genes for limited groups of allotypes.

Animals