Search PubMedSearch

SEARCH · Search PubMed

Results for “Research data management”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Pithos - a scalable and secure data container for FAIR-compliant research data management in life sciences.

Modern research techniques have led to exponential growth in the volume and complexity of scientific data. Consequently, managing these volumes securely and efficiently has become a major challenge. While all research domains face these challenges, life science research is particularly affected because current approaches often rely on a large set of different file formats, with metadata stored in separated databases or spreadsheets. This leads to fragmented datasets, orphaned data, and compromised research reproducibility. Traditional solutions also force researchers to choose between security and accessibility, with encrypted files preventing selective access and indexed formats lacking adequate security for sensitive data. These limitations are particularly problematic in large-scale genomic studies where researchers must decompress multi-gigabyte files to access specific regions, creating computational bottlenecks and inefficient network usage when working with cloud-stored datasets. We introduce Pithos, a next-generation file format specifically designed for scientific data management in distributed cloud environments. The format uses content-defined chunking to enable efficient deduplication across distributed storage systems, thereby reducing storage costs and bandwidth requirements. The append-only structure ensures data immutability and allows for incremental updates without compromising content. Benchmark results show that Pithos outperforms existing solutions in read and write performance, with comparable or improved storage efficiency.

Biological Science Disciplines

MetaServe: a lightweight, metadata-aware governance and delivery layer for pre-publication research omics data.

BACKGROUND: Institutional research teams and core facilities routinely manage pre-publication omics datasets that span heterogeneous file types, nested project structures, and multiple downstream uses. Public repositories mainly support post-publication dissemination, while workflow systems and enterprise data platforms do not directly provide a lightweight governance and delivery layer for internal research assets. RESULTS: We present MetaServe, an open-source governance and delivery layer for pre-publication research assets in institutional multi-omics settings. MetaServe registers and delivers heterogeneous assets, including sequencing files, processed matrices, imaging data, analysis-ready objects, tabular files, and documents, without requiring repository-grade standardization. Its metadata-aware design combines file-type recognition, partial automatic extraction for selected formats, manually supplied project and biological annotations, and indexed faceted retrieval. MetaServe supports authenticated web download, viewer-oriented handoff for compatible services such as cellxgene, and path-manifest export for downstream workflows under shared-storage assumptions. The current implementation combines role-based controls, explicit file-level sharing, path-constrained delivery, and operational traceability to support controlled institutional access. MetaServe has been deployed at the Chinese Institutes for Medical Research (CIMR) as part of an institutional multi-omics data-management system. CONCLUSIONS: MetaServe provides a practical layer between institutional storage and downstream analytical platforms for pre-publication research data. Its contribution is the integration of lightweight metadata-aware registration, permission-aware retrieval, and controlled delivery for heterogeneous institutional omics assets. Rather than replacing workflow engines, public repositories, or enterprise-scale research data platforms, MetaServe offers a deployable governance layer for core facilities and collaborative teams that need structured discovery and traceable delivery before public deposition or manuscript release.

Metadata

Building a Digital Health Research Platform to Enable Recruitment, Enrollment, Data Collection, and Follow-Up for a Highly Diverse Longitudinal US Cohort of 1 Million People in the All of Us Research Program: Design and Implementation Study.

BACKGROUND: Longitudinal cohort studies have traditionally relied on clinic-based recruitment models, which limit cohort diversity and the generalizability of research outcomes. Digital research platforms can be used to increase participant access, improve study engagement, streamline data collection, and increase data quality; however, the efficacy and sustainability of digitally enabled studies rely heavily on the design, implementation, and management of the digital platform being used. OBJECTIVE: We sought to design and build a secure, privacy-preserving, validated, participant-centric digital health research platform (DHRP) to recruit and enroll participants, collect multimodal data, and engage participants from diverse backgrounds in the National Institutes of Health's (NIH) All of Us Research Program (AOU). AOU is an ongoing national, multiyear study aimed to build a research cohort of 1 million participants that reflects the diversity of the United States, including minority, health-disparate, and other populations underrepresented in biomedical research (UBR). METHODS: We collaborated with community members, health care provider organizations (HPOs), and NIH leadership to design, build, and validate a secure, feature-rich digital platform to facilitate multisite, hybrid, and remote study participation and multimodal data collection in AOU. Participants were recruited by in-person, print, and online digital campaigns. Participants securely accessed the DHRP via web and mobile apps, either independently or with research staff support. The participant-facing tool facilitated electronic informed consent (eConsent), multisource data collection (eg, surveys, genomic results, wearables, and electronic health records [EHRs]), and ongoing participant engagement. We also built tools for research staff to conduct remote participant support, study workflow management, participant tracking, data analytics, data harmonization, and data management. RESULTS: We built a secure, participant-centric DHRP with engaging functionality used to recruit, engage, and collect data from 705,719 diverse participants throughout the United States. As of April 2024, 87% (n=613,976) of the participants enrolled via the platform were from UBR groups, including racial and ethnic minorities (n=282,429, 46%), rural dwelling individuals (n=49,118, 8%), those over the age of 65 years (n=190,333, 31%), and individuals with low socioeconomic status (n=122,795, 20%). CONCLUSIONS: We built a participant-centric digital platform with tools to enable engagement with individuals from different racial, ethnic, and socioeconomic backgrounds and other UBR groups. This DHRP demonstrated successful use among diverse participants. These findings could be used as best practices for the effective use of digital platforms to build and sustain cohorts of various study designs and increase engagement with diverse populations in health research.

Humans

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis.

Suckermouth catfishes, with their evolved powerful features, have become notorious invasive species, causing significant damage to aquatic ecosystems. However, the lack of high-quality genomes severely restricts research on this group within the field. In this study, we de novo assembled the chromosome-level genome assembly of Pterygoplichthys pardalis using multiple platforms of sequencing data, including Illumina short reads, Nanopore long reads, and Hi-C sequencing reads, resulting in a 1.51 Gb genome assembly. Multiple evaluations, including read mapping ratio (98.52%), transcript mapping ratio (99.61%), conserved BUSCO gene set (98.8%), and N50 score (49.47 Mb), indicated the high continuity and accuracy of the genome assembly we generated. Genome annotation found that 0.97 Gb of genome sequences are repetitive sequences, accounting for 64.47% of the genome assembly. Further, 23,859 protein-coding genes were successfully predicted, 92.92% of which could be annotated in functional databases. This high-quality genome assembly of P. pardalis provides a valuable resource for understanding the genetic underpinnings of P. pardalis's invasive success and offers critical data for future fisheries research and management.

Animals

Toward ethical provenance tracking: The GA4GH model data access agreement (DAA).

PURPOSE: Standardizing contractual clauses that govern data access enables research institutions to responsibly steward genomic and related health data while enabling its efficient downstream reuse. METHODS: We describe a document analysis study using both qualitative and comparative law analytical approaches to identify the most common categories of clauses from 29 different data access agreements used by human biomedical research consortia globally. We furthermore characterized the legal positions and standard practices for each common element of the agreement and synthesized across them to develop model clauses. A total of 3 discussion sessions were organized virtually to refine the clauses among members of the Ethical Provenance Subgroup of the Global Alliance for Genomics and Health. RESULTS: We developed 15 unique data access clauses corresponding to the most common legal elements identified in the sampled agreements. CONCLUSION: Model clauses can be used to drive administrative efficiencies and institutional compliance for managing access to human genomic data for research. Additional machine-readable consents and software solutions are needed to support traceable "ethical provenance" of human genomic data and communicate data use conditions throughout the data's life-cycle.

Humans

Institutional data commons: a federated Data Use Certification-aware architecture for secure and scalable data use in biomedical data ecosystems.

BACKGROUND: Modern biomedical data ecosystems increasingly rely on global cloud platforms to coordinate access to large-scale genomic and clinical datasets. However, operational governance remains largely investigator-centric, shifting the responsibility for complex security, compliance, and infrastructure management to individual laboratories. As data volumes and regulatory requirements expand, this approach fails to scale across the research enterprise. This disjointed approach creates a substantial governance burden and can slow down scientific progress. In centralized cloud environments, investigators face siloed identity management and high costs, leading to inefficient data use and increased risk when integrating local and global datasets. MATERIALS AND METHODS: We examine limitations in the current infrastructure and propose reframing institutional data commons as governance-aware intermediaries to ensure secure, efficient and sustainable use of controlled-access biomedical data. RESULTS: This federated architecture decouples storage from authorization, enabling dynamic access linked to active certifications, whether data are analyzed in situ on global platforms or in local governance-aware institutional access environments. DISCUSSION: Shifting governance from investigators to institutional infrastructure ensures that biomedical research remains both secure and economically sustainable.

biomedical data ecosystems

Bioinformatics in crop research: using genomic data for crop improvement.

Sustainable crop development aims to maintain or increase yields while reducing environmental impact and managing the challenges imposed by climate change. As the global population grows and arable land becomes scarcer, the integration of molecular breeding with bioinformatics has emerged as an effective strategy for long-term crop improvement. Bioinformatics enables researchers to analyze and interpret the vast quantities of genetic data generated by high-throughput sequencing, making it possible to identify molecular markers, candidate genes, and regulatory networks linked to specific agronomic traits, which breeders then translate into focused, ecologically sustainable breeding programs. This approach has enabled major progress across several fronts: the identification of genes conferring resistance to biotic stressors (pests, pathogens) and abiotic stressors (drought, salinity, heat); the development of nutrient-efficient, low-input crop varieties; the improvement of agronomic performance and nutritional quality through identification of yield- and quality-related genes; and the conservation and deployment of genetic diversity to safeguard long-term breeding sustainability. By combining genomic data with precision breeding techniques, researchers are developing crops that are better adapted to a growing population and a changing climate, positioning the integration of molecular breeding and bioinformatics as a central pillar of future global food security.

bioinformatics

Facilitators and Barriers to Volunteers' Involvement in Palliative Care: A Qualitative Meta-Synthesis.

OBJECTIVE: This study aims to systematically synthesize qualitative evidence on facilitators and barriers to volunteer involvement in palliative care services, providing insights to inform strategies for strengthening volunteer support systems. METHODS: PubMed, Web of Science, Embase, Cochrane Library, Medline, EBSCO, ProQuest, China National Knowledge Infrastructure, Wanfang, VIP, and Sinomed were searched from inception to December 2025 to identify qualitative studies examining factors influencing volunteer participation in palliative care. Methodological quality was assessed using the Joanna Briggs Institute Critical Appraisal Checklist for Qualitative Research. Data were analyzed using Thomas and Harden's thematic synthesis approach and managed using NVivo 12.0 software, following the Enhancing Transparency in Reporting the Synthesis of Qualitative Research (ENTREQ) guidelines. RESULTS: Thirty-one studies involving 1042 participants were included, yielding 68 findings. Facilitators included intrinsic motivation and meaning-making at the individual level; supportive relationships and teamwork at the interpersonal level; structured support and professional recognition at the organizational level; social recognition and resource integration at the community level; and institutional safeguards and governmental incentives at the policy level. Barriers included emotional burden and limited competencies at the individual level; relationship conflicts and insufficient collaboration at the interpersonal level; management deficiencies at the organizational level; community resource imbalances at the community level; and inadequate regulations and incentives at the policy level. CONCLUSION: Volunteer participation in palliative care is influenced by multiple interacting factors. Strengthening training and support systems, enhancing team collaboration, and improving institutional frameworks may help sustain volunteer engagement and improve the quality of palliative care services.

Palliative Care

The Polish Konik Horse: A Multidisciplinary Review of Its Origin, Genetics, Ecology, Health, Behaviour and Reproductive Biology.

The Polish Konik horse (PKH) is one of Europe's best-known native conservation breeds. Traditionally associated with the extinct Eurasian tarpan and conservation grazing, the breed has recently become the subject of multidisciplinary research encompassing genetics, ecology, health, behaviour and reproduction. This narrative review summarises current knowledge on the biological characteristics and contemporary scientific significance of the PKH. Literature published between 2005 and 2026 was identified through searches of PubMed, Scopus, Web of Science and Google Scholar and narratively synthesised. Available evidence suggests that, despite severe historical bottlenecks, the PKH has retained considerable genetic diversity and its characteristic maternal and paternal founder-line structure. Recent molecular studies have revised traditional concepts of the breed's origin, while ecological research supports its important role in conservation grazing and wetland restoration. Behavioural and reproductive studies indicate stable temperament, high reproductive efficiency and adaptation to extensive management systems. However, current knowledge is derived predominantly from observational studies, with relatively few comparative investigations and limited genomic and longitudinal data. The PKH represents a valuable model for research on conservation genetics, environmental adaptation, animal welfare, reproductive biology and ecosystem management. Further interdisciplinary studies are needed to strengthen the evidence base for conservation and breeding strategies.

Polish Konik horse

Rehabilomics Strategies Enabled by Cloud-Based Rehabilitation: Scoping Review.

BACKGROUND: Rehabilomics, or the integration of rehabilitation with genomics, proteomics, metabolomics, and other "-omics" fields, aims to promote personalized approaches to rehabilitation care. Cloud-based rehabilitation offers streamlined patient data management and sharing and could potentially play a significant role in advancing rehabilomics research. This study explored the current status and potential benefits of implementing rehabilomics strategies through cloud-based rehabilitation. OBJECTIVE: This scoping review aimed to investigate the implementation of rehabilomics strategies through cloud-based rehabilitation and summarize the current state of knowledge within the research domain. This analysis aims to understand the impact of cloud platforms on the field of rehabilomics and provide insights into future research directions. METHODS: In this scoping review, we systematically searched major academic databases, including CINAHL, Embase, Google Scholar, PubMed, MEDLINE, ScienceDirect, Scopus, and Web of Science to identify relevant studies and apply predefined inclusion criteria to select appropriate studies. Subsequently, we analyzed 28 selected papers to identify trends and insights regarding cloud-based rehabilitation and rehabilomics within this study's landscape. RESULTS: This study reports the various applications and outcomes of implementing rehabilomics strategies through cloud-based rehabilitation. In particular, a comprehensive analysis was conducted on 28 studies, including 16 (57%) focused on personalized rehabilitation and 12 (43%) on data security and privacy. The distribution of articles among the 28 studies based on specific keywords included 3 (11%) on the cloud, 4 (14%) on platforms, 4 (14%) on hospitals and rehabilitation centers, 5 (18%) on telehealth, 5 (18%) on home and community, and 7 (25%) on disease and disability. Cloud platforms offer new possibilities for data sharing and collaboration in rehabilomics research, underpinning a patient-centered approach and enhancing the development of personalized therapeutic strategies. CONCLUSIONS: This scoping review highlights the potential significance of cloud-based rehabilomics strategies in the field of rehabilitation. The use of cloud platforms is expected to strengthen patient-centered data management and collaboration, contributing to the advancement of innovative strategies and therapeutic developments in rehabilomics.

Cloud Computing

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software

Managing workflow executions with WESkit.

SUMMARY: In biomedical research, managing computational workflows across numerous projects-with varying parameters, tools, and environments-creates major challenges in scalability, reproducibility, and collaboration. Here, we present WESkit, an implementation of the Global Alliance for Genomics and Health (GA4GH) Workflow Execution Service (WES) interface, designed to streamline the execution, monitoring, and documentation of data processing workflows. It addresses the complexities involved in managing numerous executions with varying parameters across diverse research projects. Supporting both Snakemake and Nextflow, the system enables consistent automation and centralized monitoring, which benefits research groups aiming for long-term reproducibility and scalable collaboration. Its suitability for larger teams and service units is further enhanced by seamless integration into cloud environments, contributing to the GA4GH cloud framework. AVAILABILITY AND IMPLEMENTATION: The software WESkit is available under MIT license at the GitLab repository (https://gitlab.com/one-touch-pipeline/weskit). The WESkit main repository is archived at Software Heritage (https://archive.softwareheritage.org/browse/origin/directory/?origin_url=https://gitlab.com/one-touch-pipeline/weskit/api.git) and can be found using "one-touch-pipeline/weskit" term in the search section.

Workflow

Implementing a training resource for large-scale genomic data analysis in the All of Us Researcher Workbench.

A lack of representation in genomic research and limited access to computational training create barriers for many researchers seeking to analyze large-scale genetic datasets. The All of Us Research Program provides an unprecedented opportunity to address these gaps by offering genomic data from a broad range of participants, but its impact depends on equipping researchers with the necessary skills to use it effectively. The All of Us Biomedical Researcher (BR) Scholars Program at Baylor College of Medicine aims to break down these barriers by providing early-career researchers with hands-on training in computational genomics through the All of Us Evenings with Genetics Research Program. The year-long program begins with the faculty summit, an in-person computational boot camp that introduces scholars to foundational skills for using the All of Us dataset via a cloud-based research environment. The genomics tutorials focus on genome-wide association studies (GWASs), utilizing Jupyter Notebooks and the Hail computing framework to provide an accessible and scalable approach to large-scale data analysis. Scholars engage in hands-on exercises covering data preparation, quality control, association testing, and result interpretation. By the end of the summit, participants will have successfully conducted a GWAS, visualized key findings, and gained confidence in computational resource management. This initiative expands access to genomic research by equipping early-career researchers from a variety of backgrounds with the tools and knowledge to analyze All of Us data. By lowering barriers to entry and promoting the study of representative populations, the program fosters innovation in precision medicine and advances equity in genomic research.

Humans

Geriatric medicine: model for computer-oriented research analysis.

It is important to have accurate information derived from basic data about geriatric patients, not only for clinical management but to anticipate changes in the demand for services and the taking of preventive action. A model is presented for computer-oriented analysis. The collection of data is designed to cover the specific physical, physiologic and social needs of the patients both in the hospital and in the community. It must also support effective action about these needs which involve not only usage of existing facilities but the planning of preventive measures and adjustment of the provision for care. In addition, it must support clinical research when necessary. Geriatric records have specific problems, not only those of the patient's identity and clinical data but also the maintenance of psychologic and social information so that the requirements for long-term care can be assessed. This type of record facilitates comparison between groups of patients and allows measurement of inter-hospital performance relating to medical care and the planning of social programs in the community. Such records furnish the background of a more detailed statistical analysis for indicating the direction of national policy. By this means the prevalence of medical, psychologic and social demand can be studied and various preventive actions taken, including the monitoring of geriatric facilities.

Computers

Species identification, discovery, and biomonitoring: Strategic priorities for DNA barcoding in Europe, set in a global context.

The International Barcode of Life (iBOL) initiative is building a globally accessible DNA-based system for species identification and discovery. This paper outlines the mission and strategic priorities for the iBOL community in Europe (iBOL Europe), set in a global context. The mission of iBOL Europe is to produce, curate, and provide access to a complete DNA barcode reference library of European eukaryotic biodiversity, catalyzing species discovery and enabling comprehensive, harmonized species identification and biomonitoring, and supporting the global iBOL program. Immediate objectives include completing reference libraries for priority taxa, democratizing access to sequencing technologies, and strengthening a distributed community of practice. Key actions identified span five thematic areas: community building, sample collection and taxonomic verification, sequencing infrastructure, data management, and mainstreaming DNA-based approaches to meet societal needs. The strategy emphasizes integration with European research infrastructures to ensure long-term sustainability and resilience for biodiversity genomics in Europe.

DNA barcoding

Nursing education in crisis: a computer alternative.

The impact of computer-assisted instruction upon the educational process is having its effect. The computer is being used in nursing education today to manage the educational environment, to instruct, to evaluate, to identify problem areas, to gather data, to manipulate data for research purposes and for continued education. As we struggle in our efforts for quality individualized education, we are minutely scrutinizing every aspect of the teaching learning process. We are not only examining how we teach but what we teach and why we are teaching it! Our deeper understanding of the process of education has helped us to delineate nursing process as content. Higher levels of instructional goals -- cognitive, affective and conative -- are resulting in education of the whole individual. The systematization of the learning process with the greater ease of accountability is leading nursing education to blaze promising new trails into the twenty-first century.

Computer-Assisted Instruction

Perspectives of participating neurologists and study nurses - Mixed-methods process evaluation of a web-based program for relapse management in multiple sclerosis (POWER@M2).

BACKGROUND: Relapsing-remitting multiple sclerosis is a chronic inflammatory disease of the central nervous system and the leading cause of disability in young adults. In Germany, 90% of relapses are treated with high-dose intravenous glucocorticoids, despite limited evidence for long-term benefit and international preference for oral administration. Time constraints often hinder informed decision-making. The multicentre Randomized Controlled Trial (RCT) POWER@MS2 (N = 160, 2020-2023), conducted at 18 German MS-centres, aimed to promote self-determined relapse management through a complex intervention (dialogue-based decision aid, nurse-led webinar, online-chat). OBJECTIVE: While RCTs demonstrate effectiveness, process evaluations are essential to understand implementation, mechanisms of impact and contextual factors. This study explored healthcare professionals' experiences and attitudes toward implementing relapse self-management and self-medication in clinical practice. METHODS: A mixed-methods process evaluation followed the UK Medical Research Council- framework. Quantitative data were collected via validated questionnaires at up to three time points and analysed descriptively. Interview guides were developed based on these results. Qualitative data from neurologist and study nurse interviews were thematically analysed. Results were triangulated using a joint display. RESULTS: Data were collected from 55 neurologists and 17 study nurses (quantitative) and from 7 neurologists and 4 nurses (qualitative) (2020-2024). Most neurologists opposed routine steroid use, reserving it for severe relapses. Some voiced concerns about self-management, but informed patients were generally viewed as capable of safe self-medication. Study nurses gave mixed feedback on the intervention, citing overload and improved guidance. CONCLUSION: Clinicians showed openness toward implementing the intervention. Enhancing accessibility and addressing specific concerns may support broader adoption.

Humans

Drug and non-drug factors influencing adverse reaction to pyrazoles.

Some of the factors influencing the likelihood of the appearance of adverse drug reactions in patients are identified and discussed. Ways of quantifying some of the risks of adverse reactions in different patients are demonstrated. It is pointed out that adverse reaction data can give valuable guidance for both day-to-day patient management and for the initiation or guidance of research projects. In this latter connection, this study highlights the difference in time of onset between aplastic anaemia and agranulocytosis, suggesting different mechanisms of reaction. The majority of adverse reactions occur during the first three weeks of treatment and it is during this time that patients must be most carefully supervised. The old patients should be watched with particular care. It is concluded that age, sex, disease being treated, length of treatment and even the geographical location where the patient lives, can affect the time, type, frequency and outcome of an adverse drug reaction.

Adult