Search PubMedSearch

SEARCH · Search PubMed

Results for “data sharing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Examining gaps in institutional policies for clinical genomic data sharing: A cross-jurisdictional study.

The sharing of data generated by clinical genetic and genomic testing without explicit consent is important for timely diagnosis and treatment. While many jurisdictions permit the sharing of identifiable data for direct clinical care, institutional policies vary in how clearly they specify key elements, including when sharing is permitted, what data are covered, and what safeguards apply. Greater clarity around these elements may support responsible data sharing while balancing timely care with transparency and appropriate protections. We conducted a mixed-methods content analysis of data-sharing and privacy policies from 33 clinical genomic institutions across 17 countries and regions. Using a predefined analytical framework, we assessed how policies document key governance elements relevant to sharing without explicit consent. Two independent reviewers extracted information about clinical contexts, data types, justifications, and protections. Although 70% of institutions described circumstances permitting data sharing without explicit consent, most policies did not clearly define the scope or governance of such sharing. Policies also rarely distinguished clinical from research or secondary use and inconsistently specified privacy and security safeguards. While sharing was commonly justified for clinical care (78.3%) or testing services (43.5%), data recipient roles and onward-sharing expectations were often left undefined. This uneven documentation could make it difficult for clinical teams and institutional decision-makers to identify and justify decisions about what is permitted and under what conditions. A guidance framework specifying core governance elements and corresponding protections could help institutions communicate their governance choices more clearly and support comparable baseline practices for responsible data sharing.

Information Dissemination

"It just feels morally not right to Sell the data": Ethical and social perspectives on human genomic data sharing in Uganda-A phenomenological qualitative study.

While genomic data sharing enhances transparency and research efficiency, it also raises significant ethical and social challenges. This study explored stakeholders' perspectives on these issues, particularly around privacy, confidentiality, and equity in collaborative research. A phenomenological qualitative study was conducted between August and December 2023 at Makerere University College of Health Sciences, other research-intensive institutions, and national regulatory bodies. The study engaged 86 participants: 47 key informants (16 researchers, 14 ethics committee members, nine community advisory board members, and eight research regulators) and four deliberative focus group discussions with 39 participants. Interviews were transcribed verbatim, and thematic analysis was conducted using NVivo 14. Three major themes emerged: (1) stakeholders' experiences in genomic research, including their roles as participants, implementers, or overseers; (2) ethical concerns, such as informed consent, third-party data access, inequities between high-income and low- and middle-income country (LMIC) researchers and participants, and the lack of benefit-sharing frameworks; and (3) social implications, including stigma, discrimination, labeling, community perceptions of fairness, and the need for meaningful engagement. Participants emphasized the importance of protecting participant rights, promoting equity, and ensuring robust data governance and security. The theoretical frameworks of principlism and distributive justice provided a valuable lens for examining these concerns, particularly by highlighting the need to safeguard privacy and fairly distribute responsibilities and benefits in global collaborations. Participants also noted that perceptions of fairness are shaped by trust, local context, and past experiences with research factors that are critical for building equitable and respectful partnerships. This study underscores the urgent need to strengthen protections for research participants and promote fairness in genomic data sharing. Policies should, if adopted, emphasize culturally contextualized consent, active community engagement, restricted third-party data access, and strong data protection mechanisms to address existing inequities and prevent misuse.

LMICs

Beacon Reconstruction Attack: Reconstruction of genomes in genomic data-sharing beacons using summary statistics.

MOTIVATION: Genomic data-sharing beacon protocol, developed by the Global Alliance for Genomics and Health, offers a privacy-preserving mechanism for querying genomic datasets while restricting direct data access. Despite their design, beacons remain vulnerable to privacy attacks. This study introduces a novel privacy vulnerability of the protocol: one can reconstruct large portions of the genomes of all beacon participants by only using the summary statistics reported by the protocol. RESULTS: We introduce a novel optimization-based algorithm that leverages beacon responses and SNP correlations for reconstruction. By optimizing for the SNP correlations and allele frequencies, the proposed approach achieves genome reconstruction with a substantially higher F1-score (70%) compared to baseline methods (45%) on beacons generated using individuals from the HapMap and OpenSNP datasets. We show that reconstructed genomes can be used by downstream applications such as in membership inference attacks against other beacons. Our findings reveal that beacons releasing allele frequencies substantially increase the reconstruction risk, underscoring the need for enhanced privacy-preserving mechanisms to protect genomic data. AVAILABILITY AND IMPLEMENTATION: Our implementation is available at https://github.com/ASAP-Bilkent/Beacon-Reconstruction-Attack.

Genomics

Underrepresented voices in a Colorado Biobank: Perspectives from focus groups on motivations, return of results, and data sharing.

Most participants in large cohorts, such as biobanks, are of European descent. This lack of representation has been an ongoing challenge in genomic research. Understanding the perspectives on genomics research and participation in biobanks of historically underrepresented populations could provide insight into ways to better engage with these groups. We conducted a series of virtual and in-person focus groups with individuals who self-identified as American Indian or Alaska Native (AI/AN), African American/Black (AA/B), or Hispanic/Latino (H/L) and who were enrolled in the Colorado Center for Personalized Medicine (CCPM) biobank. The focus group discussions were centered on participant experiences, including but not limited to their motivations, return of results, and data sharing. There was a total of 23 participants across the six focus groups. The majority of participants identified as AI/AN (60.9%), followed by H/L (39.1%), and AA/B (21.7%); many participants identified with multiple race/ethnicities. The motivations for participating in the biobank included the potential to advance science and health, the potential for return of results, to learn more about one's ancestry, and a few indicated that they were interested in helping the biobank be more representative of all populations. Notably, many expressed positive feedback of the focus groups and felt that their views were valued, illustrating the importance of community-centered work. Our findings can be used to guide recruitment and engagement of biobank participants, especially from diverse backgrounds, contributing to enhanced partnerships advancing knowledge and healthcare.

biobank

NoisyFlow: differentially private optimal transport using neural networks for secure biomedical data sharing across multiple institutions.

MOTIVATION: Biomedical models improve when trained on data pooled across institutions, but sensitive patient records (e.g. genomics, clinical data, and medical images) are difficult to share due to privacy constraints. Moreover, data collected at different sites often have shifted distributions because of covariate differences (including batch effects), so privacy-preserving sharing alone cannot simply resolve cross-site mismatch. Methods that protect individuals while explicitly aligning distributions are needed to enable reliable multi-institutional analyses. RESULTS: We present NoisyFlow, a three-stage differentially private framework for cross-institutional harmonization under distribution shift. In stage I, each site learns a differentially private flow-based generator of its local labeled distribution. In stage II, it learns a neural optimal transport map to a shared reference distribution. In stage III, a central server composes the released models to generate reference-aligned pseudo-data for downstream analysis without accessing raw records. Across four biomedical settings spanning single-cell genomics, histopathology, neurogenomics, and wearable sensing, NoisyFlow reduces distribution shift while preserving downstream utility under formal differential privacy guarantees. AVAILABILITY AND IMPLEMENTATION: The implementation of NoisyFlow is available at https://github.com/gersteinlab/NoisyFlow.

Information Dissemination

Identification Matters: How Data Sharing Affects Pupil Honesty and Engagement in Universal School Well-Being Assessments.

PURPOSE: Universal well-being assessments in schools may support early identification of pupils needing mental health support. However, little is known about how privacy and confidentiality concerns influence pupils' acceptability of assessments and willingness to engage authentically. This study examined how hypothetical identification, where responses are linked to pupils and shared with key stakeholders, affects pupils' anticipated honesty and engagement, and whether known help-seeking barriers predict negative responses. METHODS: Cross-sectional data were collected from 12,377 primary (ages 8-10) and secondary pupils (ages 11-17) across 55 schools in England. Pupils reported whether their responses would change if identifiable and shared with school staff, parents/guardians, or external professionals. Responses indicating reduced honesty or likelihood of disengagement were coded as negative. Predictors were examined using mixed-effects logistic regression models, including demographics, school connectedness, and mental well-being. RESULTS: Identification and data sharing influenced pupils' anticipated engagement, particularly in secondary schools. Identification by school staff elicited the highest proportion of negative responses in both phases, whereas external professionals elicited the fewest. Most primary pupils reported they would respond authentically, while a larger proportion of secondary pupils indicated they would respond less honestly or disengage when responses were identifiable and shared. Across primary and secondary samples, low well-being, low school connectedness, and being female were associated with greater likelihood of negative response. DISCUSSION: Pupils' anticipated engagement with well-being assessments is shaped by who accesses their data, with marked developmental differences. Strengthening trust, privacy, and connectedness, and supporting pupils' autonomy, may improve the acceptability and response accuracy.

Humans

ToxiVerse: chemical bioprofiling, toxicity data sharing and customizable predictive modeling.

MOTIVATION: Chemical toxicity assessment is critical for drug development and environmental safety. Computational models have emerged as a promising alternative to animal testing and now play a significant role in efficiently evaluating new chemicals. To address the urgent need for user-friendly machine learning tools in computational toxicology, we developed ToxiVerse, a public web-based platform. RESULTS: ToxiVerse provides automatic chemical bioprofiling, curated toxicity datasets, and a predictive modeling interface designed for researchers who lack programming expertise. The platform comprises three integrated modules: (i) Bioprofiler, which provides chemical descriptors by combining chemical-bioactivity data from PubChem assays with a machine learning-based data gap-filling procedure; (ii) Database, which hosts ∼50 000 curated chemicals covering diverse toxicity endpoints; and (iii) Cheminformatics, which enables dataset upload, chemical curation, and automatic generation of quantitative structure-activity relationship models for toxicity prediction. AVAILABILITY: The tool is accessible at www.toxiverse.com, and source code is available at https://github.com/zhu-research-group/toxiverse.

Quantitative Structure-Activity Relationship

ONCOLINER: A new solution for monitoring, improving, and harmonizing somatic variant calling across genomic oncology centers.

The characterization of somatic genomic variation associated with the biology of tumors is fundamental for cancer research and personalized medicine, as it guides the reliability and impact of cancer studies and genomic-based decisions in clinical oncology. However, the quality and scope of tumor genome analysis across cancer research centers and hospitals are currently highly heterogeneous, limiting the consistency of tumor diagnoses across hospitals and the possibilities of data sharing and data integration across studies. With the aim of providing users with actionable and personalized recommendations for the overall enhancement and harmonization of somatic variant identification across research and clinical environments, we have developed ONCOLINER. Using specifically designed mosaic and tumorized genomes for the analysis of recall and precision across somatic SNVs, insertions or deletions (indels), and structural variants (SVs), we demonstrate that ONCOLINER is capable of improving and harmonizing genome analysis across three state-of-the-art variant discovery pipelines in genomic oncology.

Humans

The Network of National COVID-19 Data Portals: public health equity through collaboration.

The network of the national COVID-19 Data Portals was developed and linked to the COVID-19 Data Portal (https://www.covid19dataportal.org/)inresponsetothe need for rapid data sharing and analysis during the 2020-2022 SARS-CoV-2 pandemic. Built on open-source code developed by the Swedish COVID-19 Data Portal (now the Swedish Pathogens Portal, www.pathogens.se) the network included 12 national portals addressing demand for local open data sharing and access, across data types and resources. It provides a robust case study of national initiatives for FAIR (Findable, Accessible, Interoperable and Reusable) resources and a foundation for future pandemic preparedness across pathogens globally. In this paper we outline the structure of the origins of the network of National COVID-19 Datal Portals, the technical aspects and code originating from the Swedish Portal and provide an overview of the services and tools offered by each Portal. The paper showcases the process and operation of four Portals: Sweden, Poland, Spain, Norway and The Netherlands. In this study, we observe that pandemic response greatly benefits from an established infrastructure that can be quickly mobilised, developed and extended. Collaborations and preparation built on solid foundations over several years, supported by investment in the form of national and international research grants, is key for sustainability, continuation and readiness to deploy such efforts.

COVID-19

PD GENEration: An International Parkinson's Disease Genetic Research Study.

BACKGROUND: PD GENEration (NCT04057794, NCT04994015), sponsored by the Parkinson's Foundation in partnership with Aligning Science Across Parkinson's (ASAP) through the Global Parkinson's Genetics Program (GP2), is an international, observational, clinical research study that offers genetic testing and counseling to people living with Parkinson's disease (PwP) at no financial cost. PD GENEration has aimed to empower PwP and their clinicians with knowledge of their genetic status, to accelerate recruitment into precision medicine trials, and to advance research through data sharing. Since its launch in 2019, the study has expanded to enroll over 32,000 PwP (as of March 31, 2026), from 10 countries across North, Central, and South America, the Caribbean, and Israel. METHODS: Over the course of 6 years, PD GENEration has evolved to accommodate the growing scientific and research needs of the Parkinson's community while also increasing the ability to return genetic test results to PwP at a greater scale. Participants with a diagnosis of Parkinson's disease (PD) may enroll in-person or virtually where informed consent and blood sample collection can occur. Samples are analyzed at a College of American Pathologists/Clinical Laboratory Improvement Amendments (CAP/CLIA)-certified laboratory using whole genome sequencing, with variants curated for a primary panel of seven PD-associated genes. Results are disclosed during a genetic counseling visit, where further testing is offered for two optional additional gene panels. Those who consent undergo analysis of additional genes, and results are returned during a genetic counseling visit for those that test positive for a variant. In addition to returning genetic results to PwP, a central pillar of the study design has been the open sharing of genomic data to advance discovery in PD research in partnership with ASAP and GP2. DISCUSSION: PD GENEration applies a flexible framework, allowing for country specific considerations and the integration of multiple site models, evolving based on participant needs and the prioritization of equity and accessibility. We summarize PD GENEration's implementation and scaling, highlight key accomplishments and lessons learned, and provide guidance for those interested in implementing large-scale clinical genetic testing studies across other diseases and therapeutic domains.

Parkinson’s disease

Global vaccine readiness: equity-by-design in pandemic preparedness and response.

INTRODUCTION: COVID-19 showed that rapid vaccine development and roll-out, while lifesaving, can still yield large, avoidable harms when equity is not considered from the outset. Disparities in vaccine timing and coverage, especially in low-resource settings, amplified health and economic burdens, highlighting the need for preparedness frameworks that combine speed with fairness. AREAS COVERED: We synthesize evidence from literature and policy reports regarding global vaccine roll-out, focusing on avertable mortality under alternative sharing scenarios, procurement design, pooled mechanisms such as COVAX, and the role of distributed manufacturing and delivery capacity. We also examine how transparent data-sharing, effective public communication, genomic surveillance, adaptive trial designs, and modeling hubs can support more responsive and equitable vaccine deployment. Across six reflection points, we translate these lessons into practical priorities for future pandemic readiness, including strengthening healthcare infrastructure, equitable procurement, data transparency, and safeguarding public health decision-making from political and commercial distortion. EXPERT OPINION: We argue that equity-by-design is essential if vaccine innovation is to deliver equitable public health impact. This requires geographically distributed manufacturing, transparency, equity-conditioned advance purchase agreements, and pre-agreed, epidemiology-triggered allocation of vaccines. We recommend institutionalizing disaggregated reporting, standardized data-sharing, greater pathogen genomic sequencing capacity, and communication strategies that support public health protection while countering misinformation.

Humans

Individual Differences in Cognitive Aging Rodent Datasets (ID-CARD): A collaborative platform for behavioral analysis across the lifespan.

Understanding cognitive aging requires approaches that capture individual variability while enabling integration across studies. In rodent models, behavioral data are central to this effort, yet cross-laboratory differences in experimental design limit comparability and constrain secondary analysis. To address this gap, we developed the Individual Differences in Cognitive Aging Rodent Datasets (ID-CARD), a first-of-its-kind collaborative repository aggregating trial-level Morris water maze data from multiple laboratories. ID-CARD is designed to support large-scale, integrative analyses and to facilitate secondary use of existing behavioral data in alignment with emerging data-sharing and transparency initiatives. Rather than imposing retrospective harmonization of experimental protocols, we implemented a normalization and modeling framework that enables comparison of learning trajectories while preserving meaningful variation across studies. Behavioral data from > 5000 rats spanning common strains, both sexes, and multiple ages were normalized in training and performance domains and fit with a logarithmic function to derive an error accumulation rate coefficient (EARC) as a measure of spatial learning. Age was strongly associated with increased EARC, indicating attenuated learning, even after adjusting for non-spatial cue performance. Analyses of goodness of fit revealed systematic structure in learning dynamics, where age was associated with reduced learning-curve conformity after accounting for overall performance. Inter-individual variability in spatial learning also increased with age, with strain-specific interactions. These findings demonstrate that integrated analysis of heterogeneous behavioral datasets can yield robust, individual-level insights into cognitive aging. ID-CARD provides a scalable resource and analytic framework to advance discovery in behavioral neuroscience by enabling reuse, integration, and comparative analysis of existing data.

Cognitive aging

Conference report: the third Bacterial Genome Sequencing Pan-European Network conference.

The third Bacterial Genome Sequencing Pan-European Network conference, held in Engelberg, Switzerland (12-15 January 2026), brought together experts from six European countries to discuss the implementation of bacterial genome sequencing in clinical microbiology and public health. Key themes included regulatory frameworks (In Vitro Diagnostic Regulation, General Data Protection Regulation), standardization, quality control, data sharing, economic evaluation, and the integration of artificial intelligence and long-read sequencing into diagnostic workflows. Across presentations, panel discussions, and workshops, participants emphasized that successful implementation of genome sequencing requires more than technical capacity: it depends on robust validation, sustainable funding, interoperable data standards, ethical governance, and interdisciplinary collaboration. The meeting highlighted that sequencing should remain question-driven and clinically meaningful, balancing cost, turnaround time, and public health impact. Overall, the conference reinforced the need for coordinated European efforts to advance responsible, standardized, and sustainable genomic surveillance and diagnostics.

bacterial genome sequencing

Privacy-hardened and hallucination-resistant synthetic data generation with logic-solvers.

MOTIVATION: Machine-generated or synthetic data is a valuable resource for training artificial intelligence algorithms, evaluating rare workflows, and sharing data under stricter data legislations. However, current statistical and deep learning methods struggle with large data volumes, are prone to hallucinating scenarios incompatible with reality, and seldom quantify privacy meaningfully. RESULTS: Here, we introduce Genomator, a logic solving approach (SAT solving), which efficiently produces private and realistic representations of the original data. We demonstrate the method on genomic data, which arguably is the most complex and private information. We benchmark Genomator against state-of-the-art methodologies (Markov generation, Wasserstein Generative Adversarial Network and Conditional Restricted Boltzmann Machines), demonstrating a 40%-530% accuracy improvement and 57%-172% higher privacy. Genomator is also 3-100 times more efficient, making it the only tested method that scales to whole genomes. We show the universal trade-off between privacy and accuracy, and use Genomator's tuning capability to cater to all applications along the spectrum, from provable private representations of sensitive cohorts, to datasets with indistinguishable pharmacogenomic profiles. Demonstrating the production-scale generation of tuneable synthetic genomes hold great potential for balancing underrepresented populations in medical research and advancing global data exchange. AVAILABILITY AND IMPLEMENTATION: Genomator is available at https://github.com/csiro/genomator.

Algorithms

P2X7 Receptor in Rare Diseases: Shared Molecular Mechanisms and Therapeutic Implications.

Rare diseases (RDs) are individually uncommon but collectively affect a large global population, and the vast majority still lack effective disease-modifying therapies. With advances in genomics and data-sharing platforms, research has increasingly shifted from a single-disease perspective to the search for convergent molecular pathways that might be shared across clinically distinct entities. In this context, the purinergic P2X7 receptor (P2X7R) has emerged as a putative "shared molecular platform" due to its central role in inflammation amplification, cell death and immune regulation. P2X7R is an ATP-gated ion channel with unique structural and functional features: under high extracellular ATP, it not only forms a non-selective cation channel but can also dilate into a "large pore" permeable to macromolecules, thereby triggering Ca2+overload, NLRP3 inflammasome assembly, reactive oxygen species (ROS) production and apoptotic/necrotic-like cell death. This review briefly outlines the epidemiology of RDs and the structural-functional characteristics of P2X7R, then systematically summarizes current evidence linking P2X7R to multiple rare diseases, including Charcot-Marie-Tooth disease, Guillain-Barré syndrome, amyotrophic lateral sclerosis, Huntington's disease, multiple sclerosis, and selected inflammatory and metabolic RDs (CAPS, familial Mediterranean fever, Systemic sclerosis, Dravet syndrome and Gaucher disease). By comparing P2X7R expression and functional alterations, downstream signaling pathways and pharmacological data from animal models across these conditions, we propose that a P2X7R-dependent network centered on a "Ca2+-NLRP3-inflammation/cell death axis" may constitute a common pathogenic backbone for diverse RDs. At the same time, disease-specific spatiotemporal expression patterns of P2X7R in central vs peripheral nervous systems and in immune vs target organ cells confer marked context dependence and "double-edged sword" properties. Finally, we discuss opportunities and challenges for P2X7R-targeted strategies, including the impact of disease stage and sex differences on therapeutic efficacy, and key bottlenecks in translating preclinical findings into clinical benefit. A deeper understanding of both shared and disease-specific roles of P2X7R may provide a conceptual framework and therapeutic entry point for precision stratification and multi-target interventions in rare diseases.

P2X7 receptor

Integrating Biobanking Into Conservation Practice: The Development and Impact of the EAZA Biobank.

Zoological biobanks are becoming essential tools in conservation, offering a means to preserve genetic material and support in situ population management amid accelerating biodiversity loss. With rapid advances in genomics, cryopreservation, and assisted reproduction technologies, biobanks enable a proactive approach to providing insurance against genetic erosion and facilitating future research, supplementation, and genetic rescue. However, to be effective, zoological biobanks must be purposefully designed, strategically integrated into conservation frameworks such as the Convention on Biological Diversity (CBD) Kunming-Montreal Global Biodiversity Framework (KMGBF), and regularly evaluated for coverage and impact. Using the EAZA Biobank as an example, we outline the structure, development, and collaborative foundations that have enabled its rapid growth, built on community support and conservation impact. Leveraging EAZA's institutional network and data-sharing platforms such as ZIMS, the Biobank employs a decentralized, four-hub model of zoological institutions storing samples. A gap analysis, integrating threat status, breeding programs, genomic data repositories, and phylogenetic diversity, highlights current sampling strengths and deficiencies and guides future collection priorities. The integration of specimen-specific genomic data and the EAZA Biobank Cryonetwork of institutions with expertise in storing and generating gametes and cell lines will expand the Biobank's role in population management and conservation. Zoological biobanks must now evolve alongside advances in biotechnology and genomics. Sample collection strategies should serve conservation needs and anticipate future applications in genomics, cryobiology, and conservation medicine, linking biospecimens with the wealth of data generated from them. This approach should be scalable beyond EAZA, forming the foundation of a global standardized biobanking framework. Ultimately, zoological biobanks are not merely repositories of the past-they are essential infrastructures shaping the future potential of species conservation.

EAZA

Exploring the emerging concept of precision rehabilitation: a qualitative study.

PURPOSE: This descriptive qualitative study explored knowledge users' perspectives on precision rehabilitation concepts, barriers, facilitators, and future directions as part of a convergent mixed methods scoping review. MATERIALS AND METHODS: Sixteen clinicians, administrators, and researchers from three North American tertiary care rehabilitation centers were recruited using convenience and snowball sampling to participate in individual semi-structured interviews. Conventional qualitative content analysis followed a deductive thematic approach based on predetermined categories. RESULTS: Analyses revealed three main themes: (1) Although precision rehabilitation shares foundational concepts with precision medicine, there are certain elements, such as personalization, that are uniquely expressed; (2) Rehabilitation-specific facilitators to precision approaches include the use of unobtrusive technology to collect large amounts of data in real-world contexts, while barriers include rehabilitation's typically small, heterogeneous sample sizes; and (3) The future of precision rehabilitation will require collaborative data-sharing to focus on determining care trajectories that enhance functional outcomes. CONCLUSION: Findings provide the first qualitative synthesis of knowledge users perspectives to complement quantitative evidence and inform the emerging field of precision rehabilitation.

Humans

OmniExtract: an automatic data extraction tool based on large language model and prompt engineering.

Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recent advances in large language models (LLMs) have demonstrated strong capabilities in language understanding, and a number of LLM-based tools have been developed for extraction-oriented tasks. However, it's still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files that can adapt to various data extraction tasks. OmniExtract employs a prompt optimization method to refine task-specific prompts and achieve high extraction performance. It also supports comprehensive data extraction from both documents and tables, making it applicable to a broad range of data sources. Evaluation results show that OmniExtract obtains a high accuracy ~90% for three datasets. Furthermore, two additional data extraction applications of OmniExtract in real-world scenarios have been presented, achieving an accuracy of 92.21% and ~90% precision and recall, respectively. Specifically, OmniExtract can handle tabular files of various sizes and formats, and achieve over 99% precision and recall on table information extraction tasks. The data reliability performance shows that OmniExtract is a valuable tool for database updating. An online testing service is available at https://ngdc.cncb.ac.cn/omniextract/. The service can be deployed locally with the code in https://github.com/wyb39/OmniExtract.

Large Language Models