Search PubMedSearch

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Constructing epigenetic regulatory landscapes of plant lncRNAs-an exploration utilizing the novel specialized platform PERlncDB.

Long non-coding RNAs (lncRNAs), once overlooked as transcriptional byproducts, are now recognized for their crucial roles in plant growth, development, and stress responses, with increasing focus on their epigenetic regulation. However, studies investigating epigenomic signals to explore the functions of lncRNAs in plants remain relatively limited. This study collected a comprehensive dataset of over 160 000 high-quality lncRNAs from 19 representative plant species and integrated 6715 ChIP-seq, BS-seq, and RNA-seq datasets to analyze epigenomic patterns at lncRNA loci. Results showed elevated DNA methylation in lncRNA regions. The highest levels occurred in transposable element-associated lncRNAs. Additionally, activating histone modifications at lncRNA loci showed tissue specificity, with epigenetic preferences differed from those at protein-coding gene (PCG) loci. Differential site analysis in epigenetic mutants further highlighted the selective regulation of lncRNA loci by specific epigenetic factors. To facilitate research, we developed PERlncDB, a platform that provides species-specific lncRNA browsing, epigenetic annotation, cross-species conservation analysis, and visualization of epigenomic landscapes. Case studies on MARS and LINC-AP2 emphasized the platform's utility. Conserved epigenetic mechanisms regulating lncRNAs across species, exemplified by a syntenic conserved MET1-regulated lncRNA pair in Arabidopsis and tomato, suggested the stability of regulatory mechanisms underlying lncRNA functions. This work provides critical insights and resources for understanding plant lncRNA epigenetic regulation.

RNA, Long Noncoding

Up-to-date, and taxonomy-curated mcrA reference databases for methanogen community profiling.

The methyl-coenzyme M reductase subunit alpha gene (mcrA) is an important phylogenetic marker for high throughput ecological profiling of methanogenic archaea, central to industrial biological methane production and greenhouse gas emissions. Yet, dedicated reference databases predate current relevant NCBI sequence accumulation and archaeal taxonomic revision. We present three updated mcrA reference databases: (i) one derived from NCBI-catalogued methanogen genomes (1572 sequences); (ii) a database built by expansion of a previously published reference dataset, leveraging the NCBI nucleotide collection (27,942 sequences); (iii) a curated-taxonomy version of the latter. The updated amplicon databases provide a ∼ 3.5-fold sequence richness expansion, extend genus-level richness from 31 to 83 taxa, more than 4-fold species-level richness, and incorporate novel lineages compared with the previous reference dataset (e.g. Thermoplasmatota-encompassed). All databases were formatted to support analysis with relevant contemporary software pipelines and packages. Overall, the generated databases facilitate a highly improved characterization of methanogen diversity and ecology.

Archaea

AutoPM3: enhancing variant interpretation via LLM-driven PM3 evidence extraction from scientific literature.

MOTIVATION: Rare diseases affect over 300 million people worldwide and are often caused by genetic variants. While variant detection has become cost-effective, interpreting these variants-particularly collecting literature-based evidence like ACMG/AMP PM3-remains complex and time-consuming. RESULTS: We present AutoPM3, a method that automates PM3 evidence extraction from literatures using open-source large language models (LLMs). AutoPM3 combines a Text2SQL-based variant extractor and a retrieval-augmented generation (RAG) module, enhanced by a variant-specific retriever and fine-tuned LLM, to separately process tables and text. We curated PM3-Bench, a dataset of 1027 variant-publication evidence pairs from ClinGen. On openly accessible pairs, AutoPM3 achieved 86.1% accuracy for variant hits and 72.5% recall for in trans variants-outperforming other methods, including those using larger models. We uncovered the effectiveness of AutoPM3's key modules, especially for variant-specific retriever and Text2SQL, through the sequential ablation study. AutoPM3 located evidence in 76 s, demonstrating that open-source LLMs can offer an efficient, cost-effective solution for rare disease diagnosis. AVAILABILITY AND IMPLEMENTATION: AutoPM3 is implemented and freely available under the MIT license at https://github.com/HKU-BAL/AutoPM3.

Genetic Variation

Scaling up orphan crop research: genebank genetics highlight geographic structure in cultivated cowpea from 10 617 global accessions.

Vigna unguiculata (L.) Walp. is a dryland legume crop, providing essential food and nutritional security for millions of people across the semi-arid tropics, in Africa, Asia and Latin America. However, as a typical 'orphan crop', cowpea has long remained underrepresented in global genomic research to support crop improvement. Here, we conducted the largest genetic diversity analysis of cowpea to date, comprising 10 617 accessions sourced from seven international collections. Using genotyping-by-sequencing, we characterised the global patterns of genetic diversity, assessed redundancy within and across collections, and examined the geographic structure of the cowpea global allele pool. Our results revealed nine distinct genetic groups with clear geographic associations and fine-scale population differentiation, reflecting dispersal history, regional adaptation and the influence of modern breeding. Duplication across collections was detected, highlighting the need for improved curation and integration of germplasm resources. Landraces from sub-Saharan Africa do not fully capture the genetic diversity present in several other geographic regions, indicating the existence of abundant and untapped genetic resources worldwide. These findings not only provide insights into the genetic structure and evolutionary history of cowpea but also offer a valuable foundation for harnessing global germplasm diversity to enhance breeding potential and accelerate crop improvement.

Vigna

PreDigs: A Database of Context-specific Cell Type Markers and Precise Cell Subtypes for Digestive Cell Annotation.

Research on cell type markers helps investigators explore the diverse cellular composition of gastrointestinal tumors, thereby enhancing our understanding of tumor heterogeneity and its impact on disease progression and treatment response. However, the integration of large-scale datasets and the standardization of cell type identification remain challenging. Here, we developed PreDigs, a user-friendly database of predicted signatures for the digestive system, which offers 124 curated single-cell RNA sequencing datasets, covering over 3.4 million cells, all available for download. After unsupervised clustering, we unified the identification and nomenclature of cell subtype labels, constructing a cell ontology tree with 142 cell types across 8 hierarchical levels. Meanwhile, we calculated three different context-specific cell type markers, including "Cell Markers", "Subtype Markers", and "TPN Markers", based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. PreDigs is accessible at https://www.biosino.org/predigs/.

Humans

KSHVbook: An Information-Sharing Database for Kaposi's Sarcoma-Associated Herpesvirus.

Kaposi's sarcoma-associated herpesvirus (KSHV) is a double-stranded DNA virus belonging to the γ-herpesvirus subfamily. KSHV is the causative agent of Kaposi's sarcoma (KS), primary effusion lymphoma (PEL), multicentric Castleman's disease (MCD), and KSHV inflammatory cytokine syndrome (KICS). Since its discovery, research on KSHV has rapidly progressed, but existing information platforms relatively lack comprehensiveness and do not provide efficient analysis tools tailored for KSHV. To further promote the research on KSHV more effectively, we have developed KSHVbook (http://www.kshvbook.com), a specialized information-sharing database dedicated to KSHV. This platform offers extensive information on genes, coding sequences, proteins, and the gene regulatory region. Besides, the KSHVbook includes about 35 010 transcription factor binding sites (TFBSs), 342 010 pairs of KSHV miRNA-host target gene relationships, protein structures predicted by AlphaFold3, qPCR primers, and so on. We also develop analytical tools for viral genome regions, TFBSs, and KSHV miRNA target genes to discover previously unknown biological functions of KSHV. These analytical tools can effectively identify the potential regulatory relationships between host transcription factors and viral genes. Overall, this platform provides a centralized data resource for KSHV research by integrating multiple databases, offering accessible analysis tools, and simplifying data acquisition. The KSHVbook will continue to be updated, and more features can be found on the website.

Herpesvirus 8, Human

Single nucleotide polymorphism-based validation of exonic splicing enhancers.

Because deleterious alleles arising from mutation are filtered by natural selection, mutations that create such alleles will be underrepresented in the set of common genetic variation existing in a population at any given time. Here, we describe an approach based on this idea called VERIFY (variant elimination reinforces functionality), which can be used to assess the extent of natural selection acting on an oligonucleotide motif or set of motifs predicted to have biological activity. As an application of this approach, we analyzed a set of 238 hexanucleotides previously predicted to have exonic splicing enhancer (ESE) activity in human exons using the relative enhancer and silencer classification by unanimous enrichment (RESCUE)-ESE method. Aligning the single nucleotide polymorphisms (SNPs) from the public human SNP database to the chimpanzee genome allowed inference of the direction of the mutations that created present-day SNPs. Analyzing the set of SNPs that overlap RESCUE-ESE hexamers, we conclude that nearly one-fifth of the mutations that disrupt predicted ESEs have been eliminated by natural selection (odds ratio = 0.82 +/- 0.05). This selection is strongest for the predicted ESEs that are located near splice sites. Our results demonstrate a novel approach for quantifying the extent of natural selection acting on candidate functional motifs and also suggest certain features of mutations/SNPs, such as proximity to the splice site and disruption or alteration of predicted ESEs, that should be useful in identifying variants that might cause a biological phenotype.

Alleles

Animal breeding and disease.

Single-locus disorders in domesticated animals were among the first Mendelian traits to be documented after the rediscovery of Mendelism, and to be included in early linkage maps. The use of linkage maps and (increasingly) comparative genomics has been central to the identification of the causative gene for single-locus disorders of considerable practical importance. The 'score-card' in domestic animals is now more than 100 disorders for which the molecular lesion has been identified and hence for which a DNA test is available. Because of the limited lifespan of any such test, a cost-effective and hence popular means of protecting the intellectual property inherent in a DNA test is not to publish the discovery. While understandable, this practice creates a disconcerting precedent. For multifactorial disorders that are scored on an all-or-none basis or into many classes, the effectiveness of control schemes could be greatly enhanced by selection on estimated breeding values for liability. Genetic variation for resistance to pathogens and parasites is ubiquitous. Selection for resistance can therefore be successful. Because of the technical and welfare challenges inherent in the requirement to expose animals to pathogens or parasites in order to be able to select for resistance, there is a very active search for DNA markers for resistance. The first practical fruits of this research were seen in 2002, with the launch of a national scrapie control programme in the UK.

Animal Diseases

Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE): a cloud-based platform for curating and classifying germline variants.

Variant interpretation in the era of massively parallel sequencing is challenging. Although many resources and guidelines are available to assist with this task, few integrated end-to-end tools exist. Here, we present the Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE), a web- and cloud-based platform for annotation, identification, and classification of variations in known or putative disease genes. Starting from a set of variants in variant call format (VCF), variants are annotated, ranked by putative pathogenicity, and presented for formal classification using a decision-support interface based on published guidelines from the American College of Medical Genetics and Genomics (ACMG). The system can accept files containing millions of variants and handle single-nucleotide variants (SNVs), simple insertions/deletions (indels), multiple-nucleotide variants (MNVs), and complex substitutions. PeCanPIE has been applied to classify variant pathogenicity in cancer predisposition genes in two large-scale investigations involving >4000 pediatric cancer patients and serves as a repository for the expert-reviewed results. PeCanPIE was originally developed for pediatric cancer but can be easily extended for use for nonpediatric cancers and noncancer genetic diseases. Although PeCanPIE's web-based interface was designed to be accessible to non-bioinformaticians, its back-end pipelines may also be run independently on the cloud, facilitating direct integration and broader adoption. PeCanPIE is publicly available and free for research use.

Child

ClinGen recuration of hearing loss-associated genes demonstrates significant changes in gene-disease validity over time.

PURPOSE: The Clinical Genome Resource (ClinGen) Hearing Loss Gene Curation Expert Panel was assembled in 2016 and has since curated 174 gene-disease relationships (GDRs) using ClinGen's semiquantitative framework. ClinGen mandates the timely recuration of all GDRs classified as Disputed, Limited, Moderate, and Strong every 2 to 3 years. METHODS: Thirty-five GDRs met the criteria for recuration within 2 years of original curation. Previous evidence was reevaluated using the latest curation guidelines, and a comprehensive literature review was performed to obtain new evidence. Recurations were approved by the Gene Curation Expert Panel and published on the ClinGen website (www.clinicalgenome.org). RESULTS: Eight of 35 GDRs (22%) changed their classification. Two Moderate and 5 Strong GDRs were upgraded to Definitive because of new case evidence. One Strong was subsumed under another Definitive GDR after evaluation of the lumping/splitting of disease entities. Twenty-seven of 35 patients remained unchanged, with little to no new evidence reported. CONCLUSION: Genes classified as Moderate and Strong were likely to build evidence and change their classification over time, whereas Limited were unlikely to gain evidence. These findings highlight the critical role of recuration in ensuring that genetic tests and research studies incorporate the most recent evidence into their efforts.

Humans

Oncopacket: integration of cancer research data using GA4GH phenopackets.

SUMMARY: Lack of data integration remains a significant impediment to cancer research, and many analyses still require customized software to transform and prepare cancer data. We describe a software package to harmonize genetic and clinical cancer data into the GA4GH Phenopacket schema, an ISO standard for representing clinical case data. We integrated demographic, mutation, morphology, diagnosis, intervention, and survival data using case data from the National Cancer Institute for 12 cancer types. The Phenopacket standard provides a foundation for downstream use, including sophisticated statistical and AI/ML analyses. We demonstrate fitness for purpose by using the integrated data to recapitulate a known association between mutations in the gene encoding isocitrate dehydrogenase 1 and survival time in brain cancer patients. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at: https://github.com/monarch-initiative/oncopacket (archived at 10.5281/zenodo.15353125).

Humans

Functional mapping and annotation of genetic associations with FUMA.

A main challenge in genome-wide association studies (GWAS) is to pinpoint possible causal variants. Results from GWAS typically do not directly translate into causal variants because the majority of hits are in non-coding or intergenic regions, and the presence of linkage disequilibrium leads to effects being statistically spread out across multiple variants. Post-GWAS annotation facilitates the selection of most likely causal variant(s). Multiple resources are available for post-GWAS annotation, yet these can be time consuming and do not provide integrated visual aids for data interpretation. We, therefore, develop FUMA: an integrative web-based platform using information from multiple biological resources to facilitate functional annotation of GWAS results, gene prioritization and interactive visualization. FUMA accommodates positional, expression quantitative trait loci (eQTL) and chromatin interaction mappings, and provides gene-based, pathway and tissue enrichment results. FUMA results directly aid in generating hypotheses that are testable in functional experiments aimed at proving causal relations.

Chromatin

Large-scale functional annotation establishes a reference framework for human LRRK2 variants.

Pathogenic variants in leucine-rich repeat kinase 2 (LRRK2)1are among the most frequent monogenic causes of Parkinson's disease (PD)2 and act through a gain-of-function mechanism of increased kinase activity. LRRK2-targeted therapies are in clinical development, but interpretation of the rapidly expanding catalogue of rare LRRK2 variants remains a barrier to translation. Here, we present functionally annotated data on >350 LRRK2 coding variants using a standardized cellular assay with Rab10 phosphorylation as a readout of kinase activity and integrated these data with curated genetic and clinical annotations from the Movement Disorders Society Genetic Mutation Database (MDSGene). Variants differed in activation magnitude, ranging from modest increases (e.g., p.G2019S) to strongly activating substitutions such as p.Y1699C or p.L1795F. Activating variants occurred across the full length of LRRK2, although the largest effects clustered within the ROC-COR regulatory hub, where structural analysis identified subdomains forming an allosteric scaffold controlling kinase output. All known/established pathogenic variants showed increased activity, whereas benign and likely benign variants remained within the wild-type range. Functional effect sizes correlated with pathway activation in patient-derived immune cells, altogether providing a framework for ACMG-based variant interpretation in which kinase activation can support PS3 functional evidence for reclassification of variants.

Protein phosphorylation

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans

PSIA: A Comprehensive Knowledgebase of Plant Self-incompatibility.

Self-incompatibility (SI) is an important genetic mechanism in angiosperms that prevents inbreeding and promotes outcrossing, with significant implications for crop breeding, including genetic diversity, hybrid seed production, and yield optimization. In eudicots, SI is typically governed by a single S-locus containing tightly linked pistil and pollen S-determinant genes. Despite major advances in SI research, a centralized, comprehensive resource for SI-related genomic data remains lacking. To address this gap, we developed the Plant Self-Incompatibility Atlas (PSIA), a systematically curated knowledgebase providing an extensive compilation of plant SI, including genomic resources for SI species, S gene annotations, molecular mechanisms, phylogenetic relationships, and comparative genomic analyses. The current release of PSIA includes over 500 genome assemblies from 469 SI species. Using known S genes as queries, we manually identified and rigorously curated 3700 S genes. PSIA provides detailed S-locus information from assembled genomes of SI species and offers an interactive platform for browsing, BLAST searches, S gene analysis, and data retrieval. Additionally, PSIA serves as a unique platform for comparative genomic studies of S-loci, facilitating exploration of the dynamic processes underlying the origin, loss, and regain of SI. As a comprehensive and user-friendly resource, PSIA will greatly advance our understanding of angiosperm SI and serve as a valuable tool for crop breeding and hybrid seed production. PSIA is freely available at http://www.plantsi.cn.

Self-Incompatibility in Flowering Plants

STABIX: summary-statistic-based GWAS indexing and compression.

MOTIVATION: Genome-wide association studies (GWAS) are widely used to investigate the role of genetics in disease traits, but the resulting file sizes from these studies are large, posing barriers to efficient storage, sharing, and querying. This issue is especially important for biobanks like the UK Biobank that publish GWAS for thousands of traits, increasing the volume of data that must be effectively managed. Current compression and query methods reduce file sizes and allow for quick genomic position-based queries but do not provide utility for quickly finding loci based on their summary statistics. For example, finding all SNVs in a particular p-value range would require decompressing and scanning the whole file. We propose a new tool, STABIX, which introduces summary-statistic-based queries and improves upon the standard bgzip compression and Tabix query tool in both compression ratio and decompression speed. RESULTS: When applied to 10 GWAS files from PanUKBB, STABIX created smaller compressed data and indices than Tabix for all files, where bgzip and tbi files were an average of 1.2 times the size of STABIX compressed files and indexes. In the same 10 files, STABIX per gene decompression was, on average 7× faster than Tabix per gene decompression, and achieved faster per gene decompression times for over 99% of nearly 20,000 genes. AVAILABILITY AND IMPLEMENTATION: Software freely available for download at GitHub: https://github.com/kristen-schneider/stabix/.

Genome-Wide Association Study

A corpus of GA4GH phenopackets: Case-level phenotyping for genomic diagnostics and discovery.

The Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema was released in 2022 and approved by ISO as a standard for sharing clinical and genomic information about an individual, including phenotypic descriptions, numerical measurements, genetic information, diagnoses, and treatments. A phenopacket can be used as an input file for software that supports phenotype-driven genomic diagnostics and for algorithms that facilitate patient classification and stratification for identifying new diseases and treatments. There has been a great need for a collection of phenopackets to test software pipelines and algorithms. Here, we present Phenopacket Store. Phenopacket Store v.0.1.19 includes 6,668 phenopackets representing 475 Mendelian and chromosomal diseases associated with 423 genes and 3,834 unique pathogenic alleles curated from 959 different publications. This represents the first large-scale collection of case-level, standardized phenotypic information derived from case reports in the literature with detailed descriptions of the clinical data and will be useful for many purposes, including the development and testing of software for prioritizing genes and diseases in diagnostic genomics, machine learning analysis of clinical phenotype data, patient stratification, and genotype-phenotype correlations. This corpus also provides best-practice examples for curating literature-derived data using the GA4GH Phenopacket Schema.

Humans

Common variants at ABCA7, MS4A6A/MS4A4E, EPHA1, CD33 and CD2AP are associated with Alzheimer's disease.

We sought to identify new susceptibility loci for Alzheimer's disease through a staged association study (GERAD+) and by testing suggestive loci reported by the Alzheimer's Disease Genetic Consortium (ADGC) in a companion paper. We undertook a combined analysis of four genome-wide association datasets (stage 1) and identified ten newly associated variants with P ≤ 1 × 10(-5). We tested these variants for association in an independent sample (stage 2). Three SNPs at two loci replicated and showed evidence for association in a further sample (stage 3). Meta-analyses of all data provided compelling evidence that ABCA7 (rs3764650, meta P = 4.5 × 10(-17); including ADGC data, meta P = 5.0 × 10(-21)) and the MS4A gene cluster (rs610932, meta P = 1.8 × 10(-14); including ADGC data, meta P = 1.2 × 10(-16)) are new Alzheimer's disease susceptibility loci. We also found independent evidence for association for three loci reported by the ADGC, which, when combined, showed genome-wide significance: CD2AP (GERAD+, P = 8.0 × 10(-4); including ADGC data, meta P = 8.6 × 10(-9)), CD33 (GERAD+, P = 2.2 × 10(-4); including ADGC data, meta P = 1.6 × 10(-9)) and EPHA1 (GERAD+, P = 3.4 × 10(-4); including ADGC data, meta P = 6.0 × 10(-10)).

ATP-Binding Cassette Transporters