Search PubMedSearch

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

PreDigs: A Database of Context-specific Cell Type Markers and Precise Cell Subtypes for Digestive Cell Annotation.

Research on cell type markers helps investigators explore the diverse cellular composition of gastrointestinal tumors, thereby enhancing our understanding of tumor heterogeneity and its impact on disease progression and treatment response. However, the integration of large-scale datasets and the standardization of cell type identification remain challenging. Here, we developed PreDigs, a user-friendly database of predicted signatures for the digestive system, which offers 124 curated single-cell RNA sequencing datasets, covering over 3.4 million cells, all available for download. After unsupervised clustering, we unified the identification and nomenclature of cell subtype labels, constructing a cell ontology tree with 142 cell types across 8 hierarchical levels. Meanwhile, we calculated three different context-specific cell type markers, including "Cell Markers", "Subtype Markers", and "TPN Markers", based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. PreDigs is accessible at https://www.biosino.org/predigs/.

Humans

Association between SGLT2 inhibitors and reporting of phimosis/paraphimosis: a comparative pharmacovigilance analysis of the WHO database.

PURPOSE: Recent data have discussed occurrence of phimosis with Sodium-glucose co-transporter-2 (SGLT2) inhibitors. However, the potential risk among the different SGLT2 inhibitors is unknown. METHODS: Using Individual Case Safety Reports (ICSRs) registered in the WHO pharmacovigilance database (01/01/2000-30/06/2025), comparisons between the different SGLT2 inhibitors and versus other drugs used in diabetes (DUD) were performed. Results are shown as Reporting Odds Ratios (ROR). RESULTS: Among 11 342 810 ICSRs, 227 were phimosis/paraphimosis with SGLT2 inhibitors, mainly between 45 and 64 years. The higher ROR value was found with empagliflozin followed by dapagliflozin and canagliflozin. ROR for SGLT2 inhibitors was higher that of all other DUD [34.72 (25.86-46.62)]. The reporting risk of phimosis/paraphimosis with SGLT2 inhibitors was also higher than that of each pharmacological class of DUD. CONCLUSION: The results suggest an association between SGLT2 inhibitors use and phimosis ICSRs. Empagliflozin had the higher reporting risk.

Humans

Impact of homologous recombination repair gene mutations on survival in metastatic prostate cancer: A real-world analysis from an observational database.

The prognostic significance of homologous recombination repair gene (HRRg) mutations across the different metastatic prostate cancer stages remains unclear. This retrospective real-world study analyzed 162 metastatic castration-sensitive (mCSPC) and 126 castration-resistant (mCRPC) patients from the ProGène database, stratified by HRRg mutational status. Mutation prevalence was similar in both groups (16.0% in mCSPC vs. 13.5% in mCRPC). HRR-positive mCSPC patients had significantly shorter median overall survival (OS) (24.0months; 95% confidence interval [CI]: 16.0-41.0) compared to HRR-negative patients (45.0months; 95% CI: 34.0-69.0; P=0.04). Notably, BRCA2-mutated patients exhibited a reduced median OS of 24.0months (95% CI: 9.0-40.0; P=0.036) and a faster progression-free survival compared to HRR-negative patients (median PFS=8.0months; 95% CI: 0.0-14.0 vs. 17.0months; 95% CI: 12.0-20.0; P=0.006). These findings suggest that HRRg mutations - especially BRCA2 - are associated with worse prognosis in mCSPC, supporting the value of early genomic screening to guide personalized treatment strategies.

Humans

Global prevalence of hereditary hemorrhagic telangiectasia-associated variants estimated by analysis of large-scale genomic databases.

BACKGROUND: Hereditary hemorrhagic telangiectasia (HHT) is an autosomal dominant disorder with an overwhelming hemorrhagic phenotype. It is mainly caused by variants in the ENG and ACVRL1 genes. HHT prevalence is currently estimated to be 1 in 5000 individuals, but the disease is likely underdiagnosed due to variable clinical presentation, misdiagnosis, and delayed recognition. OBJECTIVES: To estimate the global genetic prevalence of HHT-associated variants in ENG and ACVRL1. METHODS: We analyzed 3 large population-scale genomic databases: gnomAD, All of Us, and Regeneron Genetics Center-Million Exome. We considered known pathogenic and likely pathogenic variants of ENG and ACVRL1 and extended the analysis to potentially pathogenic variants passing the pathogenic criteria established by the guidelines for HHT of the American College of Medical Genetics and Genomics/Association for Molecular Pathology. RESULTS: The genetic prevalence of HHT ranged from 1.753 to 2.555 in 5000 individuals, when considering only pathogenic and likely pathogenic variants, and from 2.874 to 4.327 in 5000 individuals, when also potentially pathogenic variants were considered. CONCLUSION: This study assesses the prevalence of HHT-associated variants in the general population. Our unbiased approach demonstrates that the genetic prevalence of the disease is substantially higher than currently estimated.

Humans

Microbial genomic database of the Yangtze River, the third-longest river on Earth.

Microbes play an important role in mediating the nutrient cycling in the river ecosystem as a hotspot for biogeochemical processes. Due to scattered sampling efforts, however, there is a lack of a systematic study of the diversity of prokaryotic genomes in the Yangtze River, the third longest river on Earth. Here, we collected 602 metagenomic datasets of water, sediment and riparian soil samples spanning the Upper, Middle, and Lower basins of the Yangtze River over a 6,300 km continuum. We reconstructed 8,110 qualified genomes represented by 927 species-level genomes at the 95% ANI threshold, spanning 31 bacterial and five archaeal phyla. We further showed that more than half of these species (61.3% ~ 82.4%) were novel according to the genomic comparison against the curated databases, greatly expanding the known diversity of river prokaryotes. This dataset depicts an overview of microbial genomic diversity in the Yangtze River and provides a resource for in-depth investigation of metabolic potential, ecology, and evolution of riverine microbiomes.

Rivers

MedImg: An Integrated Database for Public Medical Images.

The advancements in deep learning algorithms for medical image analysis have garnered significant attention in recent years. While several studies have shown promising results, with models achieving or even surpassing human performance, translating these advancements into clinical practice is still accompanied by various challenges. A primary obstacle lies in the availability of large-scale, well-characterized datasets for validating the generalization of approaches. To address this challenge, we curated a diverse collection of medical image datasets from multiple public sources, containing 105 datasets and a total of 1,995,671 images. These images span 14 modalities, including X-ray, computed tomography, magnetic resonance imaging, optical coherence tomography, ultrasound, and endoscopy, and originate from 13 organs, such as the lung, brain, eye, and heart. Subsequently, we constructed an online database, MedImg, which incorporates and systematically organizes these medical images to facilitate data accessibility. MedImg serves as an intuitive and open-access platform for facilitating research in deep learning-based medical image analysis, accessible at https://www.cuilab.cn/medimg/.

Humans

Unveiling the BMI Risk Threshold for Osteoarthritis: Multi-Database Causal and Nonlinear Evidence.

OBJECTIVE: To characterize the nonlinear relationship between BMI and osteoarthritis (OA), and to identify BMI thresholds that inform precise prevention strategies. METHODS: This multi-database study integrated Global burden of disease 2021, National Health and Nutrition Examination Survey 2007-2018, and Genome-Wide Association Studies. A generalized additive model was performed to visualize the BMI-OA relationship, adjusting for multiple confounders. We applied segmented logistic regression models to identify potential threshold effects and used Mendelian randomization to estimate the causal effects of BMI on OA subtypes. RESULTS: From 1990 to 2021, the age-standardized prevalence and years lived with disability rates for OA were highest in regions with high SDI. OA prevalence rose nonlinearly with BMI, with breakpoints at 24.00 and 41.58 kg/m2. Each unit increase in BMI was associated with higher odds of OA between 24.00 and 41.58 kg/m2 (OR = 1.022, 95% CI: 1.003-1.041) and above 41.58 kg/m2 (OR = 1.055, 95% CI: 1.022-1.090). Women and individuals aged ≥ 45 years exhibited a higher susceptibility to knee osteoarthritis. BMI was causally associated with knee osteoarthritis (OR = 1.63, 95% CI 1.50-1.77) and hip osteoarthritis (OR = 1.54, 95% CI 1.40-1.70). CONCLUSIONS: These findings suggest that OA risk awareness and weight-management strategies should begin before BMI reaches the high range, particularly among individuals with BMI exceeding 24.00 kg/m2.

Humans

An epidemiological database system.

An epidemiological database system is presented. It is a system built around the concepts of record structure and location, and of processing of the data item values so stored. It is designed to operate in many situations with a minimum of programming, and to be easily expanded to more complex situations by adding new program modules to those already existing and functioning.

Epidemiology

In silico identification of DNMT1 inhibitors from the PlantCyc database through computational approach to assess the anti-cancer potential of nutraceutical compounds in breast cancer.

Breast cancer accounts for a disproportionate share of global cancer-related deaths, with 670,000 fatalities and 2.3 million new diagnoses recorded in women during 2022 alone. Existing treatment modalities carry considerable toxicity burdens, and resistance to available agents remains an unresolved clinical problem. DNA methyltransferase 1 (DNMT1), the enzyme chiefly responsible for maintaining genome-wide methylation patterns during DNA replication, has been mapped out as a high-value target in breast cancer because its dysregulation silences tumour suppressor genes through promoter hypermethylation. The present work involves hierarchical in silico workflow to screen 4549 plant-derived compounds from the PlantCyc database (v16.0.3) against the human DNMT1 catalytic domain (PDB ID: 4WXX). Ten top-scoring compounds were taken forward for molecular docking via AutoDock Vina; Quercetin and Kaempferol both recorded the highest binding affinities at -9.5 kcal/mol, Wogonin (-9.3 kcal/mol) and Xanthohumol (-8.1 kcal/mol) also emerged as strong binders. Pharmacokinetic evaluation using ADMET-AI confirmed that all 10 compounds met Lipinski's rule of five, with human intestinal absorption values at or above 0.98. Wogonin and Xanthohumol were selected for a 100 ns all-atom molecular dynamics (MD) simulation in GROMACS due to their well-rounded ADMET profiles and limited existing data on their specific interactions with DNMT1 in breast cancer. Across all measured trajectory metrics, backbone RMSD, residue fluctuation, radius of gyration, solvent-accessible surface area, and intermolecular hydrogen bond count, Wogonin formed a more stable, compact complex. These findings suggest that Wogonin and Xanthohumol are non-toxic nutraceutical candidates suitable for DNMT1 targeted epigenetic therapy, with computational foundation strong enough to facilitate future in vitro and in vivo validation work.

Humans

A problem-oriented analysis of database models.

Which is the most convenient database model considering specific applications? The goal of this paper is to try to answer this question by the use of a chemical example. Examples of requests describe the problems of insertion, deletion, and updating; these requests are analyzed for the hierarchical model and are expressed in a relational language defined by the authors and in Socrate for the network model.

Chemical Phenomena

The SARS-CoV-2 Integrated Genomic Epidemiology Database (IGED): Linking viral genomes with patient-level metadata to advance statewide genomic surveillance in California.

In July 2021, the California Code of Regulations Title 17 required all laboratories performing SARS‑CoV‑2 whole genome sequencing (WGS) to report their sequencing results to the California Department of Public Health (CDPH). These viral genomic data and patient metadata were compiled into the Integrated Genomic Epidemiology Database (IGED). Linking anonymized viral sequences with patient‑level information enabled monitoring of infectiousness, pathogenicity, transmission dynamics, evolution, and vaccine evasion among emerging SARS‑CoV‑2 lineages. Laboratories performing SARS-CoV-2 WGS transmitted sequencing results to CDPH through Electronic Laboratory Reporting (ELR) and non-ELR pathways. CDPH applied uniform reporting requirements but allowed flexibility in specific data formats to accommodate diverse data systems. To preserve data quality and interoperability across heterogeneous sources, CDPH implemented standardization, validation, and deduplication protocols. Snowflake, a cloud‑based data storage and analytics platform, and Posit Connect, a cloud deployment and automation platform, supported the management, processing, and integration of data within the IGED. The IGED established links between SARS‑CoV‑2 WGS data and epidemiologic metadata for 801,418 sequences, representing 81.7% of all sequences reported in California. Lineages reported to the IGED showed strong concordance with lineage proportions in GISAID. Sequences reported to the IGED had average turnaround times longer than one month, and the majority of sequencing was performed in Southern California and Los Angeles. The IGED enhanced genomic surveillance through predictive modeling and monitoring concerning evolutionary trends such as recombination and saltations in persistent infections. Development of the IGED highlighted the need for standardized data requirements, sustained funding for sequencing, incentives for data submission, and interdisciplinary collaboration to build an effective genomic surveillance system. This framework for linking genomic and epidemiologic data has not only generated critical insights for SARS‑CoV‑2 but also provided the foundation for CDPH and other public health organizations to develop similar IGED‑like systems for other priority pathogens as genomic surveillance expands.

Journal Article

Analysis of Variants' Dynamic Using the CLIMB Database in COVID-19 Patients Admitted to Hospitals of Barts Health NHS Trust.

The COVID-19 pandemic, caused by SARS-CoV-2, has led to significant global health challenges. This study analyzes the dynamics of SARS-CoV-2 variants among patients admitted to Barts Health National Health Service (NHS) Trust hospitals using data from the CLIMB-COVID decentralized digital infrastructure allowing precise identification of SARS-CoV-2 variants. A total of 423 patients admitted between October 2020 and March 2021 were included in the study and divided into two groups: the alpha lineage group, which comprised the B.1.1.7 variant, and the other lineages group, which included all other variants. Whole-genome sequencing of SARS-CoV-2 genomes was conducted using the COVID-CLIMB pipelines. Clinical outcomes, such as mortality rates and deterioration within 28 days, were analyzed. To ensure robust findings, analyzes were adjusted for confounding factors, including age and comorbidities. Our findings revealed a significant increase in mortality with age for the alpha lineage and other lineages. The study underscores the importance of age adjustment in clinical studies to accurately assess the impact of different variants. Consistent genomic sequencing and data completeness are crucial for obtaining reliable results and guiding public health responses. These insights are vital for improving patient outcomes and providing a truthful picture of the pandemic, informing both current and future healthcare strategies.

Humans

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n = 549) and a validation set (n = 236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60 mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60 mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans

PhyloNaP: a user-friendly database of phylogeny for natural product-producing enzymes.

SUMMARY: Phylogenetic analysis is widely used to predict enzyme function, yet building annotated and reusable trees is labor-intensive and requires extensive knowledge about the specific enzymes. Existing resources rarely cover biosynthetic enzymes and lack the context needed for meaningful analysis. We present PhyloNaP, the first large-scale resource dedicated to phylogenies of biosynthetic enzymes. PhyloNaP provides ∼51 000 annotated and interactive trees enriched with chemical, functional, and taxonomic information. Users can classify their own sequences via phylogenetic placement, enabling functional inference in an evolutionary context. A contribution portal allows the community to submit curated trees. By combining scale, breadth of annotation, and interactive functionality, PhyloNaP fills a major gap in bioinformatics resources for enzyme discovery and annotation, with immediate applications to secondary metabolism and beyond. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at https://phylonap.cs.uni-tuebingen.de.

Phylogeny

A decentralized future for the open-science databases.

The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR's federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.

data accessibility