Search PubMedSearch

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Common variants at MS4A4/MS4A6E, CD2AP, CD33 and EPHA1 are associated with late-onset Alzheimer's disease.

The Alzheimer Disease Genetics Consortium (ADGC) performed a genome-wide association study of late-onset Alzheimer disease using a three-stage design consisting of a discovery stage (stage 1) and two replication stages (stages 2 and 3). Both joint analysis and meta-analysis approaches were used. We obtained genome-wide significant results at MS4A4A (rs4938933; stages 1 and 2, meta-analysis P (P(M)) = 1.7 × 10(-9), joint analysis P (P(J)) = 1.7 × 10(-9); stages 1, 2 and 3, P(M) = 8.2 × 10(-12)), CD2AP (rs9349407; stages 1, 2 and 3, P(M) = 8.6 × 10(-9)), EPHA1 (rs11767557; stages 1, 2 and 3, P(M) = 6.0 × 10(-10)) and CD33 (rs3865444; stages 1, 2 and 3, P(M) = 1.6 × 10(-9)). We also replicated previous associations at CR1 (rs6701713; P(M) = 4.6 × 10(-10), P(J) = 5.2 × 10(-11)), CLU (rs1532278; P(M) = 8.3 × 10(-8), P(J) = 1.9 × 10(-8)), BIN1 (rs7561528; P(M) = 4.0 × 10(-14), P(J) = 5.2 × 10(-14)) and PICALM (rs561655; P(M) = 7.0 × 10(-11), P(J) = 1.0 × 10(-10)), but not at EXOC3L2, to late-onset Alzheimer's disease susceptibility.

Adaptor Proteins, Signal Transducing

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis

Tandem splice acceptor sites: Profiling their relevance to human disease.

PURPOSE: Interpretation of variation, particularly the creation or disruption of tandem splice acceptor sites (NAGNnAG variants), challenges genomic medicine practice. METHODS: We analyzed the creation and disruption of dinucleotide AG sites within ±30 bases of natural splice-acceptor sites in the GRCh37 human reference genome. These results were compared with variant data from the ClinVar and gnomAD databases, as well as with data from 779 National Institutes of Health Undiagnosed Diseases Program study participants. Using RNA sequencing, we assessed the splicing at NAGNnAG variants for 107 of the Undiagnosed Diseases Program participants and compared the empirical data with SpliceAI predictions. RESULTS: Creation or disruption of NAGNnAG sites within 30 bases of the natural splice acceptor are enriched in ClinVar compared with gnomAD; however, such variants in the 2 databases are rarely differentiated by SpliceAI scores. Empirical evaluation via RNA sequencing analysis supported novel acceptor site usage from -21 to +30; splice-altering variants did not predominate in a specific region or have SpliceAI scores invariantly, suggesting increased spliceogenicity. CONCLUSION: NAGNnAG variants within 30 bp of the natural splice acceptor have a high probability of clinical relevance and are poorly contextualized for clinical utility. Their interpretation benefits from empirical evaluation via RNA analysis.

Humans

Analysis of differentially expressed genes in schizophrenia based on bioinformatics and corresponding mRNA expression levels.

OBJECTIVE: This study aimed to use bioinformatics analysis to identify differentially expressed genes (DEGs) involved in the pathogenesis of schizophrenia and validate their mRNA expression levels through real-time quantitative PCR (qPCR). MATERIAL/METHODS: Datasets from the publicly available Gene Expression Omnibus (GEO) database were analyzed using R software to identify DEGs. Functional enrichment analyses, including Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways, were conducted. A protein-protein interaction (PPI) network was constructed using Cytoscape software to identify key genes with notable expression changes. The expression levels of these key genes were subsequently validated in schizophrenia patients using qPCR to assess potential susceptibility genes. RESULTS: In total, 813 DEGs were identified, with six key genes highlighted through GO analysis and PPI network screening. Among these, HDAC1, UBA52, and FYN demonstrated statistically significant differences in mRNA expression between schizophrenia patients and healthy controls (P&#xa0;<&#xa0;0.05). CONCLUSIONS: This study identified several DEGs potentially linked to the pathogenesis of schizophrenia, suggesting that HDAC1, UBA52, and FYN could serve as candidate susceptibility genes and diagnostic biomarkers. These findings provide new insights and directions for future schizophrenia research.

Humans

Exploring penetrance of clinically relevant variants in over 800,000 humans from the Genome Aggregation Database.

Incomplete penetrance, or absence of disease phenotype in an individual with a disease-associated variant, is a major challenge in variant interpretation. Studying individuals with apparent incomplete penetrance can shed light on underlying drivers of altered phenotype penetrance. Here, we investigate clinically relevant variants from ClinVar in 807,162 individuals from the Genome Aggregation Database (gnomAD), demonstrating improved representation in gnomAD version 4. We then conduct a comprehensive case-by-case assessment of 734 predicted loss of function variants in 77 genes associated with severe, early-onset, highly penetrant haploinsufficient disease. Here, we identify explanations for the presumed lack of disease manifestation in 701 of 734 variants (95%). Individuals with unexplained lack of disease manifestation in this set of disorders are rare, underscoring the need and power of deep case-by-case assessment presented here to minimize false assignments of disease risk, particularly in unaffected individuals with higher rates of secondary properties that result in rescue.

Humans

The planktonic microbiome of the Great Barrier Reef.

Large genome databases have markedly improved our understanding of marine microorganisms1-5. Although these resources have focused on prokaryotes, genomes from many dominant marine lineages, such as Pelagibacter and Prochlorococcus, are conspicuously underrepresented. Here we present the Great Barrier Reef Microbial Genomes Database (GBR-MGD), comprising 5,283 prokaryotic genomes obtained from Great&#xa0;Barrier&#xa0;Reef seawater samples using Nanopore and Illumina sequencing, including a collection of high-quality genomes of underrepresented groups. We show that standard short-read assemblies miss these populations owing to a combination of strain heterogeneity and low-GC-percentage sequencing bias. The GBR-MGD also comprises 20 chromosome-level picoeukaryote and 808,585 viral genomes, including a newly described clade of marine Crassvirales. We demonstrate the utility of the GBR-MGD to identify indicator taxa that can reliably predict the effects of reef management practices, such as the establishment of marine protected zones.

Bacteria

Microbial genomic database of the Yangtze River, the third-longest river on Earth.

Microbes play an important role in mediating the nutrient cycling in the river ecosystem as a hotspot for biogeochemical processes. Due to scattered sampling efforts, however, there is a lack of a systematic study of the diversity of prokaryotic genomes in the Yangtze River, the third longest river on Earth. Here, we collected 602 metagenomic datasets of water, sediment and riparian soil samples spanning the Upper, Middle, and Lower basins of the Yangtze River over a 6,300&#x2009;km continuum. We reconstructed 8,110 qualified genomes represented by 927 species-level genomes at the 95% ANI threshold, spanning 31 bacterial and five archaeal phyla. We further showed that more than half of these species (61.3% ~ 82.4%) were novel according to the genomic comparison against the curated databases, greatly expanding the known diversity of river prokaryotes. This dataset depicts an overview of microbial genomic diversity in the Yangtze River and provides a resource for in-depth investigation of metabolic potential, ecology, and evolution of riverine microbiomes.

Rivers

AskBeacon-performing genomic data exchange and analytics with natural language.

MOTIVATION: Enabling clinicians and researchers to directly interact with global genomic data resources by removing technological barriers is vital for medical genomics. AskBeacon enables large language models (LLMs) to be applied to securely shared cohorts via the Global Alliance for Genomics and Health Beacon protocol. By simply "asking" Beacon, actionable insights can be gained, analyzed, and made publication-ready. RESULTS: In the Parkinson's Progression Markers Initiative (PPMI), we use natural language to ask whether the sex-differences observed in Parkinson's disease are due to X-linked or autosomal markers. AskBeacon returns a publication-ready visualization showing that for PPMI the autosomal marker occurred 1.4 times more often in males with Parkinson's disease than females, compared to no differences for the X-linked marker. We evaluate commercial and open-weight LLM models, as well as different architectures to identify the best strategy for translating research questions to Beacon queries. AskBeacon implements extensive safety guardrails to ensure that genomic data is not exposed to the LLM directly, and that generated code for data extraction, analysis and visualization process is sanitized and hallucination resistant, so data cannot be leaked or falsified. AVAILABILITY AND IMPLEMENTATION: AskBeacon is available at https://github.com/aehrc/AskBeacon.

Genomics

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic

argNorm: normalization of antibiotic resistance gene annotations to the Antibiotic Resistance Ontology (ARO).

SUMMARY: Currently available and frequently used tools for annotating antibiotic resistance genes (ARGs) in genomes and metagenomes provide results using inconsistent nomenclature. This makes the comparison of different ARG annotation outputs challenging. The comparability of ARG annotation outputs can be improved by mapping gene names and their categories to a common controlled vocabulary such as the Antibiotic Resistance Ontology (ARO). We developed argNorm, a command line tool and Python library, to normalize all detected genes across six ARG annotation tools (eight databases) to the ARO. argNorm also adds information to the outputs using the same ARG categorization so that they are comparable across tools. AVAILABILITY AND IMPLEMENTATION: argNorm is available as an open-source tool at: https://github.com/BigDataBiology/argNorm. It can also be downloaded as a PyPI package and is available on Bioconda and as an nf-core module.

Molecular Sequence Annotation

Integration of multi-source gene interaction networks and omics data with graph attention networks to identify novel disease genes.

MOTIVATION: The pathogenesis of diseases is closely associated with genes, and the discovery of disease genes holds significant importance for understanding disease mechanisms and designing targeted therapeutics. However, biological validation of all genes for diseases is expensive and challenging. RESULTS: In this study, we propose DGP-AMIO, a computational method based on graph attention networks, to rank all unknown genes and identify potential novel disease genes by integrating multi-omics and gene interaction networks from multiple data sources. DGP-AMIO outperforms other methods significantly on 20 disease datasets, with an average AUROC and AUPR exceeding 0.9. The superior performance of DGP-AMIO is attributed to the integration of multiomics and gene interaction networks from multiple databases, as well as triGAT, a proposed GAT-based method that enables precise identification of disease genes in directed gene networks. Enrichment analysis conducted on the top 100 genes predicted by DGP-AMIO and literature research revealed that a majority of enriched GO terms, KEGG pathways and top genes were associated with diseases supported by relevant studies. We believe that our method can serve as an effective tool for identifying disease genes and guiding subsequent experimental validation efforts. AVAILABILITY AND IMPLEMENTATION: DGP-AMIO is publicly available at https://github.com/yangkaiyuan1027/DGP-AMIO.

Gene Regulatory Networks

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256&#x2009;Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software

Beacon Reconstruction Attack: Reconstruction of genomes in genomic data-sharing beacons using summary statistics.

MOTIVATION: Genomic data-sharing beacon protocol, developed by the Global Alliance for Genomics and Health, offers a privacy-preserving mechanism for querying genomic datasets while restricting direct data access. Despite their design, beacons remain vulnerable to privacy attacks. This study introduces a novel privacy vulnerability of the protocol: one can reconstruct large portions of the genomes of all beacon participants by only using the summary statistics reported by the protocol. RESULTS: We introduce a novel optimization-based algorithm that leverages beacon responses and SNP correlations for reconstruction. By optimizing for the SNP correlations and allele frequencies, the proposed approach achieves genome reconstruction with a substantially higher F1-score (70%) compared to baseline methods (45%) on beacons generated using individuals from the HapMap and OpenSNP datasets. We show that reconstructed genomes can be used by downstream applications such as in membership inference attacks against other beacons. Our findings reveal that beacons releasing allele frequencies substantially increase the reconstruction risk, underscoring the need for enhanced privacy-preserving mechanisms to protect genomic data. AVAILABILITY AND IMPLEMENTATION: Our implementation is available at https://github.com/ASAP-Bilkent/Beacon-Reconstruction-Attack.

Genomics

Annotation matters: the effect of structural gene annotation on orthology inference.

MOTIVATION: In silico gene annotation, the process of identifying the genes present in a genome, remains a challenging task. As genome assemblies rapidly increase, the corresponding gene models and repertoires often fall short in quality. Despite advances in annotation methods, a lack of community standards means that most published gene annotations result from ad hoc pipelines. As a result, only a few species have nearly complete and accurate gene models. This annotation quality is thought to affect downstream analyses, including orthology inference, often the first step of comparative genomics studies. RESULTS: We show that different annotation methods yield markedly distinct orthology inferences. We compared orthology assignments of gene models obtained by four prominent protein-coding gene model sources: the NCBI Eukaryotic Genome Annotation Pipeline, the Ensembl Gene Annotation System, the UniProt Reference Proteomes, and Augustus 3.4 (an ab initio pipeline). We observe significant discrepancies between sources, namely in the proportion of orthologous genes per genome, the completeness of Hierarchical Orthologous Groups, and the accuracy and recall of the predicted orthologs on a standard orthology benchmark.

Molecular Sequence Annotation

CAGEcleaner: reducing genomic redundancy in gene cluster mining.

SUMMARY: Mining homologous biosynthetic gene clusters (BGCs) typically involves searching colocalised genes against large genomic databases. However, the high degree of genomic redundancy in these databases often propagates into the resulting hit sets, complicating downstream analyses and visualization. To address this challenge, we present CAGEcleaner, a Python-based pipeline with auxiliary bash scripts designed to reduce redundancy in gene cluster hit sets by dereplicating the genomes that host these hits. CAGEcleaner integrates seamlessly with widely used gene cluster mining tools, such as cblaster and CAGECAT, enabling efficient filtering and streamlining BGC discovery workflows. AVAILABILITY AND IMPLEMENTATION: Source code and documentation is hosted at GitHub (https://github.com/LucoDevro/CAGEcleaner) and Zenodo (https://doi.org/10.5281/zenodo.14726119) under an MIT license. For accessibility, CAGEcleaner is installable from Bioconda (https://anaconda.org/bioconda/cagecleaner) and PyPi (https://pypi.org/project/cagecleaner/), and is also available as a Docker image from DockerHub (https://hub.docker.com/r/lucodevro/cagecleaner).

Software