Search PubMedSearch

SEARCH · Search PubMed

Results for “database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

TRAIT: A Comprehensive Database for T-cell Receptor-antigen Interactions.

Comprehensive and integrated resources on interactions between T-cell receptors (TCRs) and antigens are still lacking for adoptive T-cell-based immunotherapies, highlighting a significant gap that must be addressed to fully understand the mechanisms of antigen recognition by T cells. In this study, we present the T-cell receptor-antigen interaction database (TRAIT), a comprehensive database that profiles the interactions between TCRs and antigens. TRAIT stands out due to its comprehensive description of TCR-antigen interactions by integrating sequences, structures, and affinities. It provides millions of experimentally validated TCR-antigen pairs, resulting in an exhaustive landscape of antigen-specific TCRs. Notably, TRAIT emphasizes single-cell omics as a major reliable data source for TCR-antigen interactions and includes millions of reliable non-interactive TCRs. Additionally, it thoroughly demonstrates the interactions between mutations of TCRs and antigens, thereby benefiting affinity optimization of engineered TCRs as well as vaccine design. TCRs on clinical trials are innovatively provided. With the significant efforts made toward elucidating the complex interactions between TCRs and antigens, TRAIT is expected to ultimately contribute superior algorithms and substantial advancements in the field of T-cell-based immunotherapies. TRAIT is freely accessible at https://pgx.zju.edu.cn/traitdb.

Receptors, Antigen, T-Cell

Cilia.Pro database of ciliary proteins from vertebrates, Chlamydomonas, and Caenorhabditis.

Cilia and flagella are microtubule-based organelles that generate force and sense the extracellular environment. In humans, these structures are essential for development, homeostasis, and reproduction, with defects contributing to a wide array of congenital and degenerative disorders. As cilia were present on the last common ancestor of all eukaryotes, research on cilia across model organisms holds significant relevance for understanding human disease. The green alga Chlamydomonas, which diverged from the human lineage with the animal-plant split, shares striking similarities in ciliary structure and function with humans. Two decades ago, our group published the proteome of the Chlamydomonas cilium, identifying hundreds of new ciliary proteins that were organized in an online database. Since then, advances have brought us a more comprehensive understanding of both Chlamydomonas and mammalian cilia. Our database, www.Cilia.Pro, has been continually updated to integrate proteomic, transcriptomic, and genomic data from Chlamydomonas and Caenorhabditis along with humans, and other vertebrates providing a valuable tool for the ciliary research community.

Cilia

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic

Immune Checkpoint Inhibitor-related Adverse Events in Publicly Accessible United States Malpractice Records: A Systematic Legal Database Review, 2015 to 2026.

OBJECTIVES: By 2023, an estimated 56.7% of US patients with advanced or metastatic cancer were eligible for immune checkpoint inhibitor (ICI) therapy. Grade 3 to 4 immune-related adverse events (irAEs) occur in ∼14% to 21% of patients depending on regimen, yet publicly accessible malpractice records involving ICI administration or irAE management have not been systematically described. METHODS: We searched Lexis+ and Westlaw Advantage for publicly accessible US malpractice records filed from January 1, 2015, through March 31, 2026. Sources included jury verdicts and settlement databases, federal and state dockets, and briefs, pleadings, and motions databases. Cases were included only when ICI administration, indication selection, toxicity counseling, toxicity recognition, toxicity monitoring, or irAE management was causally central to the alleged negligence. RESULTS: Among 38 legal matters identified after cross-platform deduplication, 4 met eligibility criteria. These matters reflected 4 pleaded negligence theories: inappropriate ICI indication, toxicity counseling and informed-consent failure, irAE mismanagement, and treatment-related multiorgan toxicity. Publicly visible irAE-centered malpractice records were rare relative to the clinical burden, but the data do not support a national litigation incidence estimate. CONCLUSIONS: Publicly accessible irAE-centered malpractice records appear rare. As ICI use expands, litigation risk may increasingly focus on indication documentation, individualized toxicity counseling, and structured monitoring.

immune checkpoint inhibitors

Application of machine learning to a renal biopsy database.

This pilot study has applied machine learning (artificial intelligence derived qualitative analysis procedures) to yield non-invasive techniques for the assessment and interpretation of clinical and laboratory data in glomerular disease. To evaluate the appropriateness of these techniques, they were applied to subsets of a small database of 284 case histories and the resulting procedures evaluated against the remaining cases. Over such evaluations, the following average diagnostic accuracies were obtained: microscopic polyarteritis, 95.37%; minimal lesion nephrotic syndrome, 96.50%; immunoglobulin A nephropathy, 81.26%; minor changes, 93.66%; lupus nephritis, 96.27%; focal glomerulosclerosis, 92.06%; mesangial proliferative glomerulonephritis, 92.56%; and membranous nephropathy, 92.56%. Although in general the new diagnostic system is not yet as accurate as the histological evaluation of renal biopsy specimens, it shows promise of adding a further dimension to the diagnostic process. When the machine learning techniques are applied to a larger database, greater diagnostic accuracy should be obtained. It may allow accurate non-invasive diagnosis of some cases of glomerular disease without the need for renal biopsy. This may reduce both the cost and the morbidity of the investigation of glomerular disease and may be of particular value in situations where renal biopsy is considered hazardous or contraindicated.

Artificial Intelligence

Database development in a regulatory agency.

A general discussion of the history and problems associated with the collection of information and the development of a database in a regulatory agency is presented. The proceedings of the U.S. Consumer Product Safety Commission for obtaining chemical formulation information for specified consumer products is discussed. Guidelines for database administrators faced with data collection activities in a regulatory agency are provided.

Government Agencies

A multifaceted investigation into the impact of m6A methylation-related genes on pancreatic cancer, integrating insights from various databases and foundational experimental research.

BACKGROUND: Despite advances in surgical techniques, immunotherapy, the mortality rate associated with pancreatic cancer (PC) has been on the rise in recent years. Understanding the importance of RNA N6-methyladenosine (m6A) in PC is critical for prognosis, tumor microenvironment, and immunotherapy efficacy. The study aims to identify m6A methylation regulators that play an important role in the development and progression of PC by mining databases. The effect of insulin-like growth factor-binding protein 3 (IGFBP3) on pancreatic tumors was explored, and the related mechanisms were explored. METHODS: We analyzed the expression of m6A regulators in PC by digging deeper into the datasets of The Cancer Genome Atlas and Gene Expression Omnibus (GEO) databases, and analyzed its relationship with the prognosis of patients with PC, looking for m6A methylation regulators that play an important role in the development and progression of PC. Reuse the ConsensusClusterPlus package, Cox analysis, and unsupervised clustering to delineate three distinct m6A clusters - designated as m6A cluster A, m6A cluster B, and m6A cluster C single-sample gene set enrichment analysis, gene set variation analysis, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes (KEGG) analyses evaluated the different pathway roles of these clusters in the development and progression of PC. Finally, the cell lines with IGFBP3 overexpression and knockdown were constructed by lentivirus transfection, the transfection effect was identified by WB, and the effects of IGFBP3 overexpression/knockdown on the survival and growth of PC cell lines were verified by cell cloning experiments and cell counting kit-8 experiments, and the possible related pathways were explored by KEGG. RESULTS: Most m6A regulatory factors are highly expressed in PC, and their high expression is negatively correlated with the prognosis of patients with PC. Furthermore, m6A regulatory factors may influence the occurrence and development of PC through metabolic pathways, stroma activation pathways, immune regulatory processes, and the immune microenvironment. Finally, the overexpression of IGFBP3 promoted the growth of PC cells, and vice versa. CONCLUSIONS: Most m6A regulatory factors are differentially expressed in PC and are associated with the prognosis of patients with PC, potentially influencing the occurrence and development of PC through pathways such as the immune microenvironment. The overexpression of IGFBP3 can promote the growth of PC cells and vice versa.

IGFBP3

Reporting and representation of population descriptors in public RNA-seq databases.

Diverse and globally representative datasets are essential to genomic science and medicine. Here, we analyzed population descriptor metadata from RNA sequencing (RNA-seq) studies in two major public repositories: the Sequence Read Archive (SRA) and the Database of Genotypes and Phenotypes. We examined geographic and economic characteristics of institutions depositing the data and compared SRA-deposited descriptors to empirical estimates of genetic ancestry and to those reported in publications, analyzing trends over time. We found that 55% of RNA-seq samples were deposited by United States (US) institutions and 90% by institutions in high-income countries. Only 3% of SRA samples were associated with population descriptors, and among those with US Census terms, 69% were labeled as White. Among samples with continental descriptors, 56% were labeled as European. Our analyses emphasize widespread bias in the composition of public RNA-seq datasets and, more generally, a lack of consistent and careful reporting of population descriptors needing urgent improvement.

Humans

Exploring penetrance of clinically relevant variants in over 800,000 humans from the Genome Aggregation Database.

Incomplete penetrance, or absence of disease phenotype in an individual with a disease-associated variant, is a major challenge in variant interpretation. Studying individuals with apparent incomplete penetrance can shed light on underlying drivers of altered phenotype penetrance. Here, we investigate clinically relevant variants from ClinVar in 807,162 individuals from the Genome Aggregation Database (gnomAD), demonstrating improved representation in gnomAD version 4. We then conduct a comprehensive case-by-case assessment of 734 predicted loss of function variants in 77 genes associated with severe, early-onset, highly penetrant haploinsufficient disease. Here, we identify explanations for the presumed lack of disease manifestation in 701 of 734 variants (95%). Individuals with unexplained lack of disease manifestation in this set of disorders are rare, underscoring the need and power of deep case-by-case assessment presented here to minimize false assignments of disease risk, particularly in unaffected individuals with higher rates of secondary properties that result in rescue.

Humans

PreDigs: A Database of Context-specific Cell Type Markers and Precise Cell Subtypes for Digestive Cell Annotation.

Research on cell type markers helps investigators explore the diverse cellular composition of gastrointestinal tumors, thereby enhancing our understanding of tumor heterogeneity and its impact on disease progression and treatment response. However, the integration of large-scale datasets and the standardization of cell type identification remain challenging. Here, we developed PreDigs, a user-friendly database of predicted signatures for the digestive system, which offers 124 curated single-cell RNA sequencing datasets, covering over 3.4 million cells, all available for download. After unsupervised clustering, we unified the identification and nomenclature of cell subtype labels, constructing a cell ontology tree with 142 cell types across 8 hierarchical levels. Meanwhile, we calculated three different context-specific cell type markers, including "Cell Markers", "Subtype Markers", and "TPN Markers", based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. PreDigs is accessible at https://www.biosino.org/predigs/.

Humans

Human liver protein map: a reference database established by microsequencing and gel comparison.

This publication establishes a reference human liver protein map obtained with immobilized pH gradients. By microsequencing, 57 spots or 42 polypeptide chains were identified. By protein map comparison and matching (liver, red blood cell and plasma sample maps), 8 additional proteins were identified. The new polypeptides and previously known proteins are listed in a table and/or labeled on the protein map, thus providing a human liver two-dimensional gel database. This reference map can be used to identify protein spots on other samples such as rectal cancer biopsies.

Amino Acid Sequence

Association between SGLT2 inhibitors and reporting of phimosis/paraphimosis: a comparative pharmacovigilance analysis of the WHO database.

PURPOSE: Recent data have discussed occurrence of phimosis with Sodium-glucose co-transporter-2 (SGLT2) inhibitors. However, the potential risk among the different SGLT2 inhibitors is unknown. METHODS: Using Individual Case Safety Reports (ICSRs) registered in the WHO pharmacovigilance database (01/01/2000-30/06/2025), comparisons between the different SGLT2 inhibitors and versus other drugs used in diabetes (DUD) were performed. Results are shown as Reporting Odds Ratios (ROR). RESULTS: Among 11 342 810 ICSRs, 227 were phimosis/paraphimosis with SGLT2 inhibitors, mainly between 45 and 64 years. The higher ROR value was found with empagliflozin followed by dapagliflozin and canagliflozin. ROR for SGLT2 inhibitors was higher that of all other DUD [34.72 (25.86-46.62)]. The reporting risk of phimosis/paraphimosis with SGLT2 inhibitors was also higher than that of each pharmacological class of DUD. CONCLUSION: The results suggest an association between SGLT2 inhibitors use and phimosis ICSRs. Empagliflozin had the higher reporting risk.

Humans

A clinical trials database as a research tool in health care.

OBJECTIVE: The rapid, efficient, and accurate communication of clinical research findings to both clinicians and researchers is essential to the process of improving medical care. This information should be conveyed in a form that facilitates the interpretation of the complete body of research on a specific condition. Unfortunately, the prevailing system for dissemination of clinical research fails to meet these criteria. This paper proposes a system for cataloging clinical trials and communicating their results in a comprehensive and comprehensible format. METHOD: The system involves the use of a hierarchical matrix structure that allows for the selection and evaluation of related groups of clinical trials. The format of the data is tailored to the requirements of metaanalysis. RESULTS: An example is presented using the treatment of acute Crohn's disease and including a sample metaanalysis of immunosuppressive therapy for this condition. CONCLUSIONS: The hierarchical matrix structure in conjunction with a comprehensive database of clinical trials holds the potential to facilitate access to and interpretation of clinical research.

Acute Disease

Impact of homologous recombination repair gene mutations on survival in metastatic prostate cancer: A real-world analysis from an observational database.

The prognostic significance of homologous recombination repair gene (HRRg) mutations across the different metastatic prostate cancer stages remains unclear. This retrospective real-world study analyzed 162 metastatic castration-sensitive (mCSPC) and 126 castration-resistant (mCRPC) patients from the ProGène database, stratified by HRRg mutational status. Mutation prevalence was similar in both groups (16.0% in mCSPC vs. 13.5% in mCRPC). HRR-positive mCSPC patients had significantly shorter median overall survival (OS) (24.0months; 95% confidence interval [CI]: 16.0-41.0) compared to HRR-negative patients (45.0months; 95% CI: 34.0-69.0; P=0.04). Notably, BRCA2-mutated patients exhibited a reduced median OS of 24.0months (95% CI: 9.0-40.0; P=0.036) and a faster progression-free survival compared to HRR-negative patients (median PFS=8.0months; 95% CI: 0.0-14.0 vs. 17.0months; 95% CI: 12.0-20.0; P=0.006). These findings suggest that HRRg mutations - especially BRCA2 - are associated with worse prognosis in mCSPC, supporting the value of early genomic screening to guide personalized treatment strategies.

Humans

Global prevalence of hereditary hemorrhagic telangiectasia-associated variants estimated by analysis of large-scale genomic databases.

BACKGROUND: Hereditary hemorrhagic telangiectasia (HHT) is an autosomal dominant disorder with an overwhelming hemorrhagic phenotype. It is mainly caused by variants in the ENG and ACVRL1 genes. HHT prevalence is currently estimated to be 1 in 5000 individuals, but the disease is likely underdiagnosed due to variable clinical presentation, misdiagnosis, and delayed recognition. OBJECTIVES: To estimate the global genetic prevalence of HHT-associated variants in ENG and ACVRL1. METHODS: We analyzed 3 large population-scale genomic databases: gnomAD, All of Us, and Regeneron Genetics Center-Million Exome. We considered known pathogenic and likely pathogenic variants of ENG and ACVRL1 and extended the analysis to potentially pathogenic variants passing the pathogenic criteria established by the guidelines for HHT of the American College of Medical Genetics and Genomics/Association for Molecular Pathology. RESULTS: The genetic prevalence of HHT ranged from 1.753 to 2.555 in 5000 individuals, when considering only pathogenic and likely pathogenic variants, and from 2.874 to 4.327 in 5000 individuals, when also potentially pathogenic variants were considered. CONCLUSION: This study assesses the prevalence of HHT-associated variants in the general population. Our unbiased approach demonstrates that the genetic prevalence of the disease is substantially higher than currently estimated.

Humans

Microbial genomic database of the Yangtze River, the third-longest river on Earth.

Microbes play an important role in mediating the nutrient cycling in the river ecosystem as a hotspot for biogeochemical processes. Due to scattered sampling efforts, however, there is a lack of a systematic study of the diversity of prokaryotic genomes in the Yangtze River, the third longest river on Earth. Here, we collected 602 metagenomic datasets of water, sediment and riparian soil samples spanning the Upper, Middle, and Lower basins of the Yangtze River over a 6,300 km continuum. We reconstructed 8,110 qualified genomes represented by 927 species-level genomes at the 95% ANI threshold, spanning 31 bacterial and five archaeal phyla. We further showed that more than half of these species (61.3% ~ 82.4%) were novel according to the genomic comparison against the curated databases, greatly expanding the known diversity of river prokaryotes. This dataset depicts an overview of microbial genomic diversity in the Yangtze River and provides a resource for in-depth investigation of metabolic potential, ecology, and evolution of riverine microbiomes.

Rivers

MedImg: An Integrated Database for Public Medical Images.

The advancements in deep learning algorithms for medical image analysis have garnered significant attention in recent years. While several studies have shown promising results, with models achieving or even surpassing human performance, translating these advancements into clinical practice is still accompanied by various challenges. A primary obstacle lies in the availability of large-scale, well-characterized datasets for validating the generalization of approaches. To address this challenge, we curated a diverse collection of medical image datasets from multiple public sources, containing 105 datasets and a total of 1,995,671 images. These images span 14 modalities, including X-ray, computed tomography, magnetic resonance imaging, optical coherence tomography, ultrasound, and endoscopy, and originate from 13 organs, such as the lung, brain, eye, and heart. Subsequently, we constructed an online database, MedImg, which incorporates and systematically organizes these medical images to facilitate data accessibility. MedImg serves as an intuitive and open-access platform for facilitating research in deep learning-based medical image analysis, accessible at https://www.cuilab.cn/medimg/.

Humans

Unveiling the BMI Risk Threshold for Osteoarthritis: Multi-Database Causal and Nonlinear Evidence.

OBJECTIVE: To characterize the nonlinear relationship between BMI and osteoarthritis (OA), and to identify BMI thresholds that inform precise prevention strategies. METHODS: This multi-database study integrated Global burden of disease 2021, National Health and Nutrition Examination Survey 2007-2018, and Genome-Wide Association Studies. A generalized additive model was performed to visualize the BMI-OA relationship, adjusting for multiple confounders. We applied segmented logistic regression models to identify potential threshold effects and used Mendelian randomization to estimate the causal effects of BMI on OA subtypes. RESULTS: From 1990 to 2021, the age-standardized prevalence and years lived with disability rates for OA were highest in regions with high SDI. OA prevalence rose nonlinearly with BMI, with breakpoints at 24.00 and 41.58 kg/m2. Each unit increase in BMI was associated with higher odds of OA between 24.00 and 41.58 kg/m2 (OR = 1.022, 95% CI: 1.003-1.041) and above 41.58 kg/m2 (OR = 1.055, 95% CI: 1.022-1.090). Women and individuals aged ≥ 45 years exhibited a higher susceptibility to knee osteoarthritis. BMI was causally associated with knee osteoarthritis (OR = 1.63, 95% CI 1.50-1.77) and hip osteoarthritis (OR = 1.54, 95% CI 1.40-1.70). CONCLUSIONS: These findings suggest that OA risk awareness and weight-management strategies should begin before BMI reaches the high range, particularly among individuals with BMI exceeding 24.00 kg/m2.

Humans