Search PubMedSearch

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

CamK-DB: A k-mer MinHash fingerprint database for reference-free genotyping of Camellia accessions.

Tea (Camellia sinensis L.), a major global economic crop in Asia, poses challenges for genetic identification because its highly heterozygous, repetitive genome reduces the efficacy of conventional single-nucleotide polymorphism (SNP) and microsatellite markers, and interspecific hybridization further complicates the situation. To address these issues, CamK-DB was developed as a reference-free Camellia fingerprinting database built on MIKE MinHash sketches. We curated 418 candidate resequencing datasets, and built a database using standardized 5× genome-coverage fingerprints. Each accession is stored as a MIKE. jac fingerprint generated with k = 21 and recommended sketch/pre_cnt = 2000. CamK-DB provides a command-line interface for data management and a custom C++ query engine that computes top-10 matches using Jaccard similarity, complemented by a QT-based graphical interface for interactive analysis. This resource offers a robust and scalable framework for precise and routine germplasm identification, genomic phylogenetic inference, and strategic breeding program design. CamK-DB (database and code) is publicly available at https://github.com/sc-zhang/CamK-DB. CamK-DB binaries are provided for Windows 10/11 and Linux (x86_64, glibc ≥ 2.27).

Databases, Genetic

The Mycobacterium tuberculosis Transposon Sequencing Database (MtbTnDB): A Large-Scale Guide to Genetic Conditional Essentiality.

Characterizing genetic essentiality across various conditions is fundamental for understanding gene function. Transposon sequencing (TnSeq) is a powerful technique to generate genome-wide essentiality profiles in bacteria and has been extensively applied to Mycobacterium tuberculosis (Mtb). Dozens of TnSeq screens have yielded valuable insights into the biology of Mtb in vitro, inside macrophages, and in model host organisms. Despite their value, these Mtb TnSeq profiles have not been standardized or collated into a single, easily searchable database. This results in significant challenges when attempting to query and compare these resources, limiting our ability to obtain a comprehensive and consistent understanding of genetic conditional essentiality in Mtb. We address this problem by building a central repository of publicly available Mtb TnSeq screens, the Mtb transposon sequencing database (MtbTnDB). The MtbTnDB is a living resource that encompasses to date ≈150 standardized TnSeq screens, enabling open access to data, visualizations, and functional predictions through an interactive web app (www.mtbtndb.app). We conduct several statistical analyses on the complete database, such as demonstrating that (i) genes in the same genomic neighborhood have similar TnSeq profiles, and (ii) clusters of genes with similar TnSeq profiles are enriched for genes from similar functional categories. We further analyze the performance of machine learning models trained on TnSeq profiles to predict the functional annotation of orphan genes in Mtb. By facilitating the comparison of TnSeq screens across conditions, the MtbTnDB will accelerate the exploration of conditional genetic essentiality, provide insights into the functional organization of Mtb genes, and help predict gene function in this important human pathogen.

DNA Transposable Elements

PAHG: the database of human multi-gene families.

BACKGROUND: In the early vertebrate history, gene duplications, including single-gene, segmental-gene (SSD), and whole-genome duplication (WGD), formed multigene families. Despite efforts to classify metazoan multigene families hierarchically for evolutionary insight, a gap exists in accessible, curated resources for human/vertebrate multigene families. RESULTS: Addressing this, we present the Phylogenomic Analysis of Human Genome (PAHG) database. It focuses on curated multigene families in the human genome, particularly within four paralogons: HOX-bearing (Hsa:2/7/12/17), FGFR-bearing (Hsa:4/5/8/10), MHC-bearing (Hsa:1/6/9/19), and chromosomes 1/2/8/20. CONCLUSION: The current PAHG version details the phylogenetic history of 221 human multigene families (1247 gene members) with 15,231 protein sequences from diverse metazoans. It provides insights into gene duplication timings, co-duplication events, and their relationships with human genome syntenic organization. The PAHG database addresses the lack of accessible resources, offering valuable information on human/vertebrate multigene family evolution. Access the PAHG database at: https://www.pahgncb.com/ and http://pahg.qau.edu.pk/ . This resource enriches our understanding of vertebrate genetic evolution.

Humans

BAV-LLPS: a database of bacterial, archaea, and virus liquid-liquid phase separation proteins.

MOTIVATION: Liquid-liquid phase separation (LLPS) is a key process underlying the formation of biomolecular condensates, such as membrane-less organelles, that compartmentalize biochemical processes inside the cells. While LLPS has been extensively studied in eukaryotes, its role in bacteria, archaea, and viruses remains far less characterized. Recent studies in bacteria have revealed that LLPS-driven condensates play critical roles in RNA processing, stress response, and pathogenicity. Similarly, many viruses exploit LLPS to facilitate crucial steps in their infection cycles, including viral entry, genome replication, assembly, and host immune evasion. RESULTS: In this work, we introduce a hand-curated database of LLPS proteins from bacteria, archaea, and viruses (BAV-LLPS Database). This resource, extended through sequence similarity searches, comprises over 5000 proteins and integrates diverse data including biological annotations, sequence features, predicted disordered regions, LLPS per site probability, and AlphaFold2-based structural models. Additionally, our web server enables users to explore both the curated and homologous derived datasets, providing a platform to uncover evolutionary relationships and intrinsic and differential properties of LLPS proteins across various taxonomic groups. This work seeks to deepen our understanding of LLPS mechanisms beyond eukaryotic organisms, emphasizing their significance across diverse life forms. It also aims to foster the development of specialized predictive tools that will facilitate the exploration and characterization of LLPS processes in a wide array of living organisms, thereby contributing to advancements in both fundamental biological research and applied biomedical sciences. AVAILABILITY AND IMPLEMENTATION: BAV-LLPS DB is freely accessible at https://bav-llps-db.bioinformatica.org/. The data can be retrieved from the website. The source code of the database can be downloaded from https://bav-llps-db.bioinformatica.org/download.

Databases, Protein

An integrated culturomic and genomic database and analysis platform for methanogenic archaea.

Methanogenic archaea research is challenged by limited strain resources, fragmented genomic data, inconsistent genome quality, substantial uncultured lineages, and difficulties in laboratory culturing, hindering advances in biogas production, climate mitigation, and microbial ecology. These archaea play crucial roles in global carbon cycling and anaerobic environments, yet scattered data and unculturable strains limit systematic studies and applications. To address this, we created MethArDB (Methanogenic Archaeal Genome Database), a specialized database for methanogenic archaea, compiling 3919 genomes, 87 host-associated plasmids, and 42 phages, with standardized quality classifications (complete, scaffold, draft), protein sequences, and metadata on geography, habitats, metabolism, and inheritable elements. Integrated MethArCT (Methanogenic Archaeal Culturomics Toolkit) employs a dual-threshold orthologous/paralogous protein analysis to evaluate metabolic pathway completeness, predicting cultivation parameters and suggesting candidate cultivation strategies, including potential medium formulations and conditions, to support strain isolation. Overall, MethArDB and MethArCT form an integrated platform combining genomics and culturomics to facilitate methanogenic archaea research. Database URL:  http://methardb.cn.

Genome, Archaeal

Next-Generation Sequencing-Based High-Resolution Typing of HLA-A, -B, -C and HPA Genes in Jilin Province: Building a Platelet Donor Database and Identifying Novel Alleles.

To systematically analyse HLA-A, -B and -C and human platelet antigen (HPA) genotypes of platelet donors in Jilin Province using next-generation sequencing (NGS) technology, a comprehensive donor database was established. Additionally, potential novel alleles were identified, providing a scientific basis for enhancing the safety of clinical blood transfusions. DNA fragments from 200 platelet donor samples in Jilin Province were amplified using locus-specific primers. Comprehensive sequencing of HLA and HPA genes was performed via NGS. Bioinformatics analysis was employed to process genotyping results and screen for novel genetic variants. Newly discovered alleles were validated by Sanger sequencing to ensure accuracy and reliability. HLA genotyping achieved three-field allele resolution, revealing the highest-frequency alleles are as follows: HLA-A*11:01:01, HLA-B*13:02:01, HLA-C*01:02:01 and C*03:04:01. A novel allele B*49:91 (mutation: E2 24T>C) was identified. For the HPA systems (HPA-1, -2, -3, -5, -6, -15, -21), high heterozygosity was observed in HPA-3 and HPA-15, while no bb homozygosity was detected in HPA-1, -2, -5, -6 or -21. The application of NGS in constructing a platelet HLA/HPA gene database enables high-resolution genotyping, laying a critical foundation for precise platelet matching. This significantly reduces the risk of platelet transfusion refractoriness (PTR) and facilitates the discovery of novel allelic variants. The database provides essential theoretical and practical guidance for future donor screening and personalised transfusion strategies.

Humans

SUPFAM: a database of sequence superfamilies of protein domains.

BACKGROUND: SUPFAM database is a compilation of superfamily relationships between protein domain families of either known or unknown 3-D structure. In SUPFAM, sequence families from Pfam and structural families from SCOP are associated, using profile matching, to result in sequence superfamilies of known structure. Subsequently all-against-all family profile matches are made to deduce a list of new potential superfamilies of yet unknown structure. DESCRIPTION: The current version of SUPFAM (release 1.4) corresponds to significant enhancements and major developments compared to the earlier and basic version. In the present version we have used RPS-BLAST, which is robust and sensitive, for profile matching. The reliability of connections between protein families is ensured better than before by use of benchmarked criteria involving strict e-value cut-off and a minimal alignment length condition. An e-value based indication of reliability of connections is now presented in the database. Web access to a RPS-BLAST-based tool to associate a query sequence to one of the family profiles in SUPFAM is available with the current release. In terms of the scientific content the present release of SUPFAM is entirely reorganized with the use of 6190 Pfam families and 2317 structural families derived from SCOP. Due to a steep increase in the number of sequence and structural families used in SUPFAM the details of scientific content in the present release are almost entirely complementary to previous basic version. Of the 2286 families, we could relate 245 Pfam families with apparently no structural information to families of known 3-D structures, thus resulting in the identification of new families in the existing superfamilies. Using the profiles of 3904 Pfam families of yet unknown structure, an all-against-all comparison involving sequence-profile match resulted in clustering of 96 Pfam families into 39 new potential superfamilies. CONCLUSION: SUPFAM presents many non-trivial superfamily relationships of sequence families involved in a variety of functions and hence the information content is of interest to a wide scientific community. The grouping of related proteins without a known structure in SUPFAM is useful in identifying priority targets for structural genomics initiatives and in the assignment of putative functions. Database URL: http://pauling.mbu.iisc.ernet.in/~supfam.

Amino Acid Sequence

Database-guided thermodynamic-kinetic regulation for PtNi-based intermetallic electrocatalysts.

Pt-Ni alloys exhibit outstanding oxygen reduction reaction (ORR) activity, yet achieving structural ordering to enhance stability remains a persistent challenge. Here, we propose a database-guided screening strategy to identify promoter elements (X) that facilitate ordering. By screening the critical disorder-to-order transition temperature ([Formula: see text]), solid solubility ([Formula: see text]), and diffusion pre-exponential factor ([Formula: see text]) of candidate elements from material databases, six sets of L10-Pt(NiX) nanoparticles were synthesized. The developed L10-PtNiFe catalysts demonstrated a high mass activity (MA) of 4.38 A mgPt-1, retaining 82.1% after 50,000 cycles of accelerated durability testing (ADT) in half-cells, and retaining 79% of its initial MA of 1.1 A mgPt-1 after 30,000 cycles in membrane electrode assembly (MEA). Theoretical calculations reveal that X incorporation broadens the annealing temperature ([Formula: see text]) window by forming Pt(NiX) with higher [Formula: see text] and lowering the kinetic threshold temperature ([Formula: see text]) via enhanced atom mobility. This work establishes a database-guided framework that enables phase transitions previously difficult to access by creating an effective annealing window through coordinated thermodynamic-kinetic regulation, thereby facilitating the formation of ordered PtNi-based intermetallic structures toward durable electrocatalysts.

Journal Article

KSHVbook: An Information-Sharing Database for Kaposi's Sarcoma-Associated Herpesvirus.

Kaposi's sarcoma-associated herpesvirus (KSHV) is a double-stranded DNA virus belonging to the γ-herpesvirus subfamily. KSHV is the causative agent of Kaposi's sarcoma (KS), primary effusion lymphoma (PEL), multicentric Castleman's disease (MCD), and KSHV inflammatory cytokine syndrome (KICS). Since its discovery, research on KSHV has rapidly progressed, but existing information platforms relatively lack comprehensiveness and do not provide efficient analysis tools tailored for KSHV. To further promote the research on KSHV more effectively, we have developed KSHVbook (http://www.kshvbook.com), a specialized information-sharing database dedicated to KSHV. This platform offers extensive information on genes, coding sequences, proteins, and the gene regulatory region. Besides, the KSHVbook includes about 35 010 transcription factor binding sites (TFBSs), 342 010 pairs of KSHV miRNA-host target gene relationships, protein structures predicted by AlphaFold3, qPCR primers, and so on. We also develop analytical tools for viral genome regions, TFBSs, and KSHV miRNA target genes to discover previously unknown biological functions of KSHV. These analytical tools can effectively identify the potential regulatory relationships between host transcription factors and viral genes. Overall, this platform provides a centralized data resource for KSHV research by integrating multiple databases, offering accessible analysis tools, and simplifying data acquisition. The KSHVbook will continue to be updated, and more features can be found on the website.

Herpesvirus 8, Human

ProteoformDB: A Built-In Application to Generate Proteoform Database.

Proteins play essential functions through their complex regulations on cell-type-specific expression, localization, and molecular complexes. Protein complexity is further enhanced by proteoforms, which are the diverse molecular forms that each gene can produce through genomic alterations, transcriptional variations, translational regulations, and protein modifications. Profiling of proteoforms is a promising method for gaining a deeper understanding of the role of proteins in biological pathways and disease mechanisms. Here, we developed ProteoformDB, an application tool for generating proteoform databases, and we cataloged a total of over one million unique single-site human proteoforms. We showed that ProteoformDB can serve as a valuable resource to document the experimentally identified proteoforms in a database, supporting protein characterization in quantitative proteomics for both total protein abundances and modified protein forms.

Humans

TRAIT: A Comprehensive Database for T-cell Receptor-antigen Interactions.

Comprehensive and integrated resources on interactions between T-cell receptors (TCRs) and antigens are still lacking for adoptive T-cell-based immunotherapies, highlighting a significant gap that must be addressed to fully understand the mechanisms of antigen recognition by T cells. In this study, we present the T-cell receptor-antigen interaction database (TRAIT), a comprehensive database that profiles the interactions between TCRs and antigens. TRAIT stands out due to its comprehensive description of TCR-antigen interactions by integrating sequences, structures, and affinities. It provides millions of experimentally validated TCR-antigen pairs, resulting in an exhaustive landscape of antigen-specific TCRs. Notably, TRAIT emphasizes single-cell omics as a major reliable data source for TCR-antigen interactions and includes millions of reliable non-interactive TCRs. Additionally, it thoroughly demonstrates the interactions between mutations of TCRs and antigens, thereby benefiting affinity optimization of engineered TCRs as well as vaccine design. TCRs on clinical trials are innovatively provided. With the significant efforts made toward elucidating the complex interactions between TCRs and antigens, TRAIT is expected to ultimately contribute superior algorithms and substantial advancements in the field of T-cell-based immunotherapies. TRAIT is freely accessible at https://pgx.zju.edu.cn/traitdb.

Receptors, Antigen, T-Cell

Cilia.Pro database of ciliary proteins from vertebrates, Chlamydomonas, and Caenorhabditis.

Cilia and flagella are microtubule-based organelles that generate force and sense the extracellular environment. In humans, these structures are essential for development, homeostasis, and reproduction, with defects contributing to a wide array of congenital and degenerative disorders. As cilia were present on the last common ancestor of all eukaryotes, research on cilia across model organisms holds significant relevance for understanding human disease. The green alga Chlamydomonas, which diverged from the human lineage with the animal-plant split, shares striking similarities in ciliary structure and function with humans. Two decades ago, our group published the proteome of the Chlamydomonas cilium, identifying hundreds of new ciliary proteins that were organized in an online database. Since then, advances have brought us a more comprehensive understanding of both Chlamydomonas and mammalian cilia. Our database, www.Cilia.Pro, has been continually updated to integrate proteomic, transcriptomic, and genomic data from Chlamydomonas and Caenorhabditis along with humans, and other vertebrates providing a valuable tool for the ciliary research community.

Cilia

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic

Immune Checkpoint Inhibitor-related Adverse Events in Publicly Accessible United States Malpractice Records: A Systematic Legal Database Review, 2015 to 2026.

OBJECTIVES: By 2023, an estimated 56.7% of US patients with advanced or metastatic cancer were eligible for immune checkpoint inhibitor (ICI) therapy. Grade 3 to 4 immune-related adverse events (irAEs) occur in ∼14% to 21% of patients depending on regimen, yet publicly accessible malpractice records involving ICI administration or irAE management have not been systematically described. METHODS: We searched Lexis+ and Westlaw Advantage for publicly accessible US malpractice records filed from January 1, 2015, through March 31, 2026. Sources included jury verdicts and settlement databases, federal and state dockets, and briefs, pleadings, and motions databases. Cases were included only when ICI administration, indication selection, toxicity counseling, toxicity recognition, toxicity monitoring, or irAE management was causally central to the alleged negligence. RESULTS: Among 38 legal matters identified after cross-platform deduplication, 4 met eligibility criteria. These matters reflected 4 pleaded negligence theories: inappropriate ICI indication, toxicity counseling and informed-consent failure, irAE mismanagement, and treatment-related multiorgan toxicity. Publicly visible irAE-centered malpractice records were rare relative to the clinical burden, but the data do not support a national litigation incidence estimate. CONCLUSIONS: Publicly accessible irAE-centered malpractice records appear rare. As ICI use expands, litigation risk may increasingly focus on indication documentation, individualized toxicity counseling, and structured monitoring.

immune checkpoint inhibitors

A multifaceted investigation into the impact of m6A methylation-related genes on pancreatic cancer, integrating insights from various databases and foundational experimental research.

BACKGROUND: Despite advances in surgical techniques, immunotherapy, the mortality rate associated with pancreatic cancer (PC) has been on the rise in recent years. Understanding the importance of RNA N6-methyladenosine (m6A) in PC is critical for prognosis, tumor microenvironment, and immunotherapy efficacy. The study aims to identify m6A methylation regulators that play an important role in the development and progression of PC by mining databases. The effect of insulin-like growth factor-binding protein 3 (IGFBP3) on pancreatic tumors was explored, and the related mechanisms were explored. METHODS: We analyzed the expression of m6A regulators in PC by digging deeper into the datasets of The Cancer Genome Atlas and Gene Expression Omnibus (GEO) databases, and analyzed its relationship with the prognosis of patients with PC, looking for m6A methylation regulators that play an important role in the development and progression of PC. Reuse the ConsensusClusterPlus package, Cox analysis, and unsupervised clustering to delineate three distinct m6A clusters - designated as m6A cluster A, m6A cluster B, and m6A cluster C single-sample gene set enrichment analysis, gene set variation analysis, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes (KEGG) analyses evaluated the different pathway roles of these clusters in the development and progression of PC. Finally, the cell lines with IGFBP3 overexpression and knockdown were constructed by lentivirus transfection, the transfection effect was identified by WB, and the effects of IGFBP3 overexpression/knockdown on the survival and growth of PC cell lines were verified by cell cloning experiments and cell counting kit-8 experiments, and the possible related pathways were explored by KEGG. RESULTS: Most m6A regulatory factors are highly expressed in PC, and their high expression is negatively correlated with the prognosis of patients with PC. Furthermore, m6A regulatory factors may influence the occurrence and development of PC through metabolic pathways, stroma activation pathways, immune regulatory processes, and the immune microenvironment. Finally, the overexpression of IGFBP3 promoted the growth of PC cells, and vice versa. CONCLUSIONS: Most m6A regulatory factors are differentially expressed in PC and are associated with the prognosis of patients with PC, potentially influencing the occurrence and development of PC through pathways such as the immune microenvironment. The overexpression of IGFBP3 can promote the growth of PC cells and vice versa.

IGFBP3

Reporting and representation of population descriptors in public RNA-seq databases.

Diverse and globally representative datasets are essential to genomic science and medicine. Here, we analyzed population descriptor metadata from RNA sequencing (RNA-seq) studies in two major public repositories: the Sequence Read Archive (SRA) and the Database of Genotypes and Phenotypes. We examined geographic and economic characteristics of institutions depositing the data and compared SRA-deposited descriptors to empirical estimates of genetic ancestry and to those reported in publications, analyzing trends over time. We found that 55% of RNA-seq samples were deposited by United States (US) institutions and 90% by institutions in high-income countries. Only 3% of SRA samples were associated with population descriptors, and among those with US Census terms, 69% were labeled as White. Among samples with continental descriptors, 56% were labeled as European. Our analyses emphasize widespread bias in the composition of public RNA-seq datasets and, more generally, a lack of consistent and careful reporting of population descriptors needing urgent improvement.

Humans

Exploring penetrance of clinically relevant variants in over 800,000 humans from the Genome Aggregation Database.

Incomplete penetrance, or absence of disease phenotype in an individual with a disease-associated variant, is a major challenge in variant interpretation. Studying individuals with apparent incomplete penetrance can shed light on underlying drivers of altered phenotype penetrance. Here, we investigate clinically relevant variants from ClinVar in 807,162 individuals from the Genome Aggregation Database (gnomAD), demonstrating improved representation in gnomAD version 4. We then conduct a comprehensive case-by-case assessment of 734 predicted loss of function variants in 77 genes associated with severe, early-onset, highly penetrant haploinsufficient disease. Here, we identify explanations for the presumed lack of disease manifestation in 701 of 734 variants (95%). Individuals with unexplained lack of disease manifestation in this set of disorders are rare, underscoring the need and power of deep case-by-case assessment presented here to minimize false assignments of disease risk, particularly in unaffected individuals with higher rates of secondary properties that result in rescue.

Humans

PreDigs: A Database of Context-specific Cell Type Markers and Precise Cell Subtypes for Digestive Cell Annotation.

Research on cell type markers helps investigators explore the diverse cellular composition of gastrointestinal tumors, thereby enhancing our understanding of tumor heterogeneity and its impact on disease progression and treatment response. However, the integration of large-scale datasets and the standardization of cell type identification remain challenging. Here, we developed PreDigs, a user-friendly database of predicted signatures for the digestive system, which offers 124 curated single-cell RNA sequencing datasets, covering over 3.4 million cells, all available for download. After unsupervised clustering, we unified the identification and nomenclature of cell subtype labels, constructing a cell ontology tree with 142 cell types across 8 hierarchical levels. Meanwhile, we calculated three different context-specific cell type markers, including "Cell Markers", "Subtype Markers", and "TPN Markers", based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. PreDigs is accessible at https://www.biosino.org/predigs/.

Humans