Search PubMedSearch

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Demographics, Overlap, and Latency of Severe Cutaneous Adverse Reactions in an FDA Database.

IMPORTANCE: Severe cutaneous adverse reactions (SCARs), including Stevens-Johnson syndrome/toxic epidermal necrolysis (SJS-TEN), drug reaction with eosinophilia and systemic symptoms (DRESS), acute generalized exanthematous pustulosis (AGEP), and generalized bullous fixed drug eruption (GBFDE), are rare but life-threatening drug hypersensitivity syndromes. Due to their low incidence and diagnostic complexity, large-scale characterization of SCAR is challenging. OBJECTIVE: To characterize the demographics, causative agents, trends, latency, and phenotypic overlap of SCAR using a large-scale, sanitized pharmacovigilance dataset from FAERS (FDA Adverse Event Reporting System). DESIGN: Cross-sectional study of spontaneous adverse event reports. Cases were drawn from the U.S. Food and Drug Administration Adverse Event Reporting System (FDA FAERS) from January 2004 to December 2023 and subjected to sanitization and deduplication. Disproportionality analysis was used to characterize causative agents. Machine learning (random forest classifiers) was used to analyze predictors of drug latency and mortality. SETTING: Global pharmacovigilance reports submitted to FAERS. PARTICIPANTS: A total of 56,683 deduplicated SCAR reports were identified, representing 0.33% of reports during the study period. EXPOSURES: Suspected causative drugs, including both small molecules and biologics. MAIN OUTCOMES AND MEASURES: Main outcomes included the frequency and distribution of SCAR syndromes, reporting trends over time, latency from drug start to reaction onset, drug-specific disproportionality (PRR, ROR, IC), and co-reporting between SCAR types and related conditions. RESULTS: A total of 56,683 unique SCAR reports were identified, including SJS-TEN (28,871), DRESS (22,444), AGEP (6,183), and GBFDE (150). We identified 237 drugs with significant disproportionality for SCAR overall. Co-reporting between SCARs was significantly enriched (p < 1e-200), suggesting overlapping phenotypes. Latency varied by drug and syndrome (median: GBFDE 3 days, AGEP 4 days, SJS-TEN 12 days, DRESS 20 days). CONCLUSIONS AND RELEVANCE: SCAR syndromes display distinct but overlapping phenotypes, with variable latency and diverse causative agents. These findings, based on the largest SCAR dataset to date, highlight the need for improved classification frameworks and molecular validation. Large-scale pharmacovigilance, integrated with genomic and histopathologic data, will be critical to improving diagnosis, mechanistic understanding, and clinical management of SCAR.

Acute Generalized Exanthematous Pustulosis

High-variance phenome database reveals important roles of WD40 proteins in the plant pathogenic fungus Fusarium graminearum.

WD40 is a highly conserved protein domain in eukaryotes that functions as a versatile platform for protein-protein interactions and participates in diverse biological processes. We performed a genome-wide functional analysis of WD40 domain-containing proteins in Fusarium graminearum, a phytopathogenic fungus that causes severe yield losses and mycotoxin contamination in major cereal crops. Comprehensive phenotypic profiling of 119 WD40 gene deletion mutants across 22 phenotypic traits established a systematic WD40 phenome dataset, revealing the broad functional involvement of WD40 proteins and a strong correlation between sexual reproduction and virulence. Protein interaction analyses of selected WD40 proteins revealed diverse WD40-mediated interaction patterns and provided further insights into WD40-mediated protein interactions and their roles in protein complex formation. This study provides a foundation for further characterization of WD40 proteins in filamentous fungi.

Fusarium graminearum

MegaPX: fast and space-efficient peptide assignment method using IBF-based multi-indexing.

MOTIVATION: A central problem for metaproteomic analysis is the often-unknown taxonomic composition of the analyzed microbiomes. Using a database search, the standard approach requires prior knowledge of which proteins and taxa to include in the protein reference database or to use tailored metagenome-derived databases, which are expensive and error-prone in their generation. A possible strategy to circumvent this database search issue is de novo sequencing, where peptide sequences are directly identified from mass spectra. However, these sequences must still be mapped back to potentially extensive databases. Here, alignment-based approaches enable robust and precise results, with the potential drawback of high memory usage and long run times. RESULTS: We present MegaPX, a software for rapidly classifying de novo peptide sequences against large protein databases. MegaPX implemented as a C++-based tool, uses an alignment-free, k-mer approach as a taxonomic classification method with the possibility of generating mutated reference databases for error-tolerant searching. It uses various algorithms, including interleaved Bloom filters, to efficiently compute approximate membership queries, ensuring fast processing times while querying and indexing large databases in a multi-indexing fashion. We demonstrate the potential of MegaPX by analyzing different samples, including metaproteomics, against extensive reference databases, highlighting its use as a fast screening tool.

Software

MetagenomicKG: a knowledge graph for metagenomic applications.

MOTIVATION: The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG), and the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), for various analytical purposes. These databases are crucial for the genetic and functional annotation of microbial communities. Nevertheless, the inconsistent nomenclature or identifiers of these databases present challenges for effective integration, representation, and utilization. Knowledge graphs (KGs) offer an appropriate solution by organizing biological entities from different databases to standardized identifiers, allowing their interrelations to be captured into a cohesive network regardless of the naming conventions used in each source. The graph structure not only facilitates the unveiling of hidden patterns but also enriches our biological understanding with deeper insights. Despite KGs having shown potential in various biomedical fields, their application in metagenomics remains underexplored. RESULTS: We present MetagenomicKG, a novel knowledge graph specifically tailored for metagenomic analysis. MetagenomicKG integrates taxonomic, functional, and pathogenesis-related information on the human microbiome sourced from various databases, and further connects these with existing biomedical KGs to expand the biological network. Through various case studies involving the human microbiome, we demonstrate its utility in enabling hypothesis generation regarding the relationships between microbes and diseases, generating sample-specific graph embeddings, and providing robust pathogen prediction. CODE AVAILABILITY: The source code and technical details for constructing the MetagenomicKG and reproducing all analyses are available on GitHub at https://github.com/KoslickiLab/MetagenomicKG. The data used in this manuscript, including the pre-built files and use case input data, are archived on Zenodo with DOI: 10.5281/zenodo.17546861.

Metagenomics

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (&#xb1;5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors

raxtax: a k-mer-based non-Bayesian taxonomic classifier.

MOTIVATION: Taxonomic classification in biodiversity studies is the process of assigning the anonymous sequences of a marker gene (barcode) or whole genomes (metagenomics) to a specific lineage using a reference database that contains named sequences in a known taxonomy. This classification is important for assessing the diversity of biological systems. Taxonomic classification faces two main challenges: first, accuracy is critical as errors can propagate to downstream analysis results; and second, the classification time requirements can limit study size and study design, in particular when considering the constantly growing reference databases. To address these two challenges, we introduce raxtax, an efficient, novel taxonomic classification tool for barcodes that uses common k-mers between all pairs of query and reference sequences. We also introduce two novel uncertainty scores which take into account the fundamental biases of reference databases. RESULTS: We validate raxtax on three widely-used empirical reference databases and show that it is 2.7-100 times faster than competing state-of-the-art tools on the largest database while being equally accurate. In particular, raxtax exhibits increasing speedups with growing query and reference sequence numbers compared to existing tools (for 100&#x2009;000 and 1&#x2009;000&#x2009;000 query and reference sequences overall, it is 1.3 and 2.9 times faster, respectively), and therefore alleviates the taxonomic classification scalability challenge. AVAILABILITY AND IMPLEMENTATION: raxtax is available at https://github.com/noahares/raxtax under a CC-NC-BY-SA license. The scripts and summary metrics used in our analyses are available at https://github.com/noahares/raxtax_paper_scripts. The source code, sequence data, and summarized results of the analyses are available at https://doi.org/10.5281/zenodo.15057027.

Software

Culturally adapted post-diagnostic dementia support for South Asian people living with dementia and caregivers: a rapid review.

BACKGROUND: The number of minority ethnic people living with dementia (PLWD) in the UK is predicted to rise to 50 000 by 2026 and 172 000 by 2051. As the global population ages, there is a greater need to develop culturally appropriate post-diagnostic support for PLWD from minority ethnic backgrounds. METHODS: A rapid review was conducted of culturally adapted post-diagnostic dementia support for South Asian people with dementia and carers. Eight electronic databases were searched from inception until 16 September 2025. Databases included Cumulative Index to Nursing and Allied Health Literature, Excerpta Medica Database, MEDical Literature Analysis and Retrieval System Online, Psychological Information, Turning Research Into Practice, Allied and Complementary Medicine Database, Social Policy and Practice and the Cochrane Database of Systematic Reviews. Two reviewers independently screened the studies. Consistent with rapid review methods, no formal quality assessment of included studies was undertaken. The rapid review adhered to Preferred Reporting Items for Systematic Review and Meta-Analysis guidelines. RESULTS: Twelve studies were included. These included seven carer support programmes focusing on raising awareness and education on dementia and care. These interventions increased carers' knowledge of dementia and confidence in caregiving. Four studies reported on psychosocial interventions: Cognitive Stimulation Therapy, Cognitive Behaviour Therapy and Meditation Therapy, demonstrating benefits for caregiver burden and mental health. One study reported on service-level innovations through a South Asian link nurse, which improved access to services and facilitated the development of culturally appropriate information materials. CONCLUSION: The findings of this rapid review demonstrate the feasibility and perceived value of culturally sensitive psychoeducation, carer training and psychosocial interventions. However, research remains small-scale, methodologically limited, with little focus given to interventions directly supporting PLWD.

Humans

Characterisation of HIV-1 Gag Cytotoxic T-Lymphocyte Epitopes in the Southern African Region-A Systematic Review.

During early HIV-1 infection, robust Cytotoxic T-lymphocyte (CTL) responses are mostly targeted at immunodominant Gag p24 epitopes to reduce HIV-1 viraemia to a set-point. The aim of this study was to review the current body of knowledge on HIV-1 Gag CTL epitopes in the southern African region where subtype C is prevalent. Peer-reviewed records were obtained from three databases: PubMed Central, Web of Science Core Collection, and Scopus, using the following search terms: HIV subtype C Gag epitopes, and HIV clade C Gag epitopes. The search results were restricted to countries within the southern African region, and only data published in English and between the years 2000-2025 were considered for this review. The search from the three databases produced a total of 2103 peer-reviewed records, and 49 records were included in the review. The majority of studies (58.44%) were conducted in South Africa, followed by Botswana (15.58%), Zambia (10.39%), Malawi (7.79%), Zimbabwe (6.49%) and Angola (1.30%). There were no studies identified from other southern African countries. A total of 60 Gag CTL epitopes were identified, of which 17 (28.33%) were located within the matrix protein (p17), 33 (55.00%) within the capsid protein (p24), and 4 (6.67%) within the Gag polyprotein (p2p7p1p6). The commonly detected immunodominant epitopes were mostly located within the Gag p24 protein; and included TPQDLNTML (TL9, Gag p24 48-56) and TSTLQEQIGW (TW10, Gag p24 108-117) present at 16.00% and 13.3%, respectively. The proportion of HLA-A, B and C allotypes in this systematic review were 18%, 78%, and 4%, respectively. The more common HLA-B allotypes that restrict immunodominant Gag epitopes and facilitate better control of HIV-1 were HLA-B*57, -B*58:01, -B*42:01 and -B*81:01. This systematic review has provided important insights into the description of immunodominant Gag epitopes and HLA-I alleles that contribute to the control of HIV-1 viraemia in the southern African region. It has also exposed that some CTL epitopes identified in the southern African studies are not reported on the Los Alamos HIV database (LANL HIV database). This highlights a need to have this database updated with this information as it is used as a reference for epitopes. This review could provide insights into the design of an epitope-based HIV-1 vaccine that would also be effective in the southern African region.

Humans

Active components and potential mechanisms of Wuzhuyu decoction in the treatment of ethanol-induced acute gastric mucosal injury: a network pharmacology and experimental verification.

OBJECTIVE: To investigate the underlying mechanisms and active components of Wuzhuyu decoction (, WD) in alleviating ethanol-induced acute gastric mucosal injury (GMI) using an integrated approach of network pharmacology and experimental verification. METHODS: Sprague-Dawley rats were randomly divided into six groups: control (Con), model (Mod), bismuth potassium citrate (BPC), WD at low (WD-L), medium (WD-M), and high (WD-H) doses. Following seven days of continuous intragastric administration of the respective treatments, an ethanol-induced gastric mucosal injury model was established in all groups except the control group by oral gavage of anhydrous ethanol. The gastric mucosal injury index was evaluated, and pathological changes were assessed viahematoxylin and eosin (HE) staining. Levels of tumor necrosis factor-alpha (TNF-&#x3b1;), interleukin-1 beta (IL-1&#x3b2;), malondialdehyde (MDA), superoxide dismutase (SOD), and glutathione peroxidase (GSH-Px) were measured by enzyme-linked immunosorbent assay (ELISA). The chemical composition was identified by ultra-performance liquid chromatography-tandem mass spectrometry. Active compounds were screened using the Swiss-absorption, distribution, metabolism, and excretion database, and their potential targets were predicted using the Swiss Target Prediction database and bioinformatics annotation database for molecular mechanism. Simultaneously, disease targets related to GMI were retrieved from the online mendelian inheritance in man and GeneCards databases. A protein-protein interaction (PPI) network was constructed, and functional enrichment analyses of gene ontology (GO) and Kyoto encyclopedia of genes and genomes (KEGG) enrichment analyses were performed using the Metascape database. Key predictions from the network pharmacology analysis were subsequently verified through animal experiments. Protein expression levels of B-cell lymphoma-2 (Bcl-2), Bcl-2-associated X protein (Bax), Cleaved Caspase-3, and Cleaved Caspase-9 were analyzed by Western blot. Finally, molecular docking was performed using AutoDock Vina to investigate the interactions between the active components and core targets. RESULTS: WD treatment significantly reduced the gastric mucosal injury index and the levels of TNF-&#x3b1;, IL-1&#x3b2;, MDA, while it increased the activities of SOD and GSH-Px. Histopathological examination revealed marked improvement in gastric tissue morphology. A total of 145 compounds were identified in WD. Network pharmacology analysis identified 440 overlapping targets between WD and GMI. GO and KEGG enrichment analyses highlighted the apoptosis signaling pathway as a key mechanism for WD's protective effect against ethanol-induced GMI. Experimental validation demonstrated that WD treatment reduced the apoptosis of gastric mucosal epithelial cells, promoted the expression of Bcl-2, and inhibited the expression of Bax, Cleaved Caspase-3 and Cleaved Caspase-9. Molecular docking results indicated that dehydroevodiamine, rutaecarpine, evodiamine, hexahydrocurcumin, and isorhamnetin are potential active components in WD that contribute to the inhibition of apoptosis. CONCLUSIONS: WD alleviates ethanol-induced acute GMI, at least in part, by inhibiting the apoptosis. The primary active components responsible for this effect are dehydroevodiamine, rutaecarpine, evodiamine, hexahydrocurcumin, and isorhamnetin.

Drugs, Chinese Herbal

Elucidating the Mechanism of Xiaoqinglong Decoction in Chronic Urticaria Treatment: An Integrated Approach of Network Pharmacology, Bioinformatics Analysis, Molecular Docking, and Molecular Dynamics Simulations.

INTRODUCTION: Xiaoqinglong Decoction (XQLD) is a traditional Chinese medicinal formula commonly used to treat chronic urticaria (CU). However, its underlying therapeutic mechanisms remain incompletely characterized. This study employed an integrated approach combining network pharmacology, bioinformatics, molecular docking, and molecular dynamics simulations to identify the active components, potential targets, and related signaling pathways involved in XQLD's therapeutic action against CU, thereby providing a mechanistic foundation for its clinical application. METHODS: The active components of XQLD and their corresponding targets were identified using the Traditional Chinese Medicine Systems Pharmacology (TCMSP) database. CU-related targets were retrieved from the OMIM and GeneCards databases. Subsequently, core components and targets were determined via protein-protein interaction (PPI) network analysis and component-target-pathway network construction. Topological analyses were performed using Cytoscape software to prioritize core nodes within these networks. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were conducted via the DAVID database to identify enriched biological processes and signaling pathways. Molecular docking was performed to evaluate binding interactions between key components and core targets, while molecular dynamics (MD) simulations were employed to assess the stability of the component-target complexes with the lowest binding energy. Finally, CU-related targets of XQLD were validated using datasets from the Gene Expression Omnibus (GEO) database. RESULTS: A total of 135 active components and 249 potential targets of XQLD were identified, alongside 1,711 CU-related targets. Core components, such as quercetin, kaempferol, beta-sitosterol, naringenin, stigmasterol, and luteolin, exhibited high degree values in the constructed networks. The core targets identified included AKT1, TNF, IL6, TP53, PTGS2, CASP3, BCL2, ESR1, PPARG, and MAPK3. GO and KEGG pathway enrichment analyses revealed the PI3K-Akt signaling pathway as a central regulatory mechanism. Molecular docking studies demonstrated strong binding affinities between active components and core targets, with the stigmasterol-AKT1 complex exhibiting the lowest binding energy (-11.4 kcal/mol) and high stability in MD simulations. Validation using GEO datasets identified 12 core genes shared between CU-related targets and XQLD-associated targets, including PTGS2 and IL6, which were also prioritized as core targets in the network pharmacology analyses. DISCUSSION: This study comprehensively integrates multidisciplinary approaches to clarify the potential molecular mechanisms of XQLD in treating CU, highlighting its multitarget and multipathway synergistic effects. Molecular docking and dynamics simulations confirm the stable interaction between stigmasterol and the core target AKT1. Additionally, GEO dataset analysis verifies the pathogenic relevance of targets such as PTGS2 and IL6, significantly enhancing the credibility of our findings. These results provide a modern scientific basis for the traditional therapeutic effects of XQLD on CU and have important implications for developing multitarget treatments for this condition. However, this study mainly relies on database mining and computational simulations. Further in vitro and in vivo experimental validations are needed to confirm the predicted component-target-pathway interactions. CONCLUSION: This study identifies the active components, potential targets, and pathways through which XQLD exerts therapeutic effects on CU. These findings provide a theoretical foundation for further mechanistic studies and support their clinical application in the treatment of CU.

Molecular Docking Simulation

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine