Search PubMedSearch

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface

Distinct mutational landscapes for germline and somatic cancer variants in forty tumor suppressor genes.

Germline and somatic cancer variants in tumor suppressor genes (TSGs) share loss-of-function mechanisms, but studies of a few genes (DICER1 and CEBPA) have demonstrated differences in variant consequence and location. To systematically assess whether TSGs display distinct mutational patterns, we leveraged large public genetic databases and compared 32,941 high-quality pathogenic/likely pathogenic (P/LP) germline variants in ClinVar, with 12,907 oncogenic/likely oncogenic (O/LO) somatic tumor variants from cBioPortal across 40 TSGs. Only 3,863 (9.2%) variants were shared. Eighteen TSGs showed significantly different distributions of variant occurrences by molecular consequence, replicated with non-overlapping somatic data from the COSMIC database (chi-squared tests, false discovery rate = 5%). DICER1, TP53, and SMAD4 displayed excess somatic missense events, while nine TSGs (e.g., RB1 and APC) contained excess somatic stop-gain events throughout the coding sequence. Analysis by tumor type revealed excess stop-gain events in tissues exposed to environmental mutagens with corresponding mutation signatures. For several TSGs (WT1), germline variants predispose to tumors (Wilms' tumor) distinct from the majority source of somatic data (myeloid leukemia). Germline and somatic events are also distributed unevenly across cDNA locations, with 103 regions of preferential clustering in 39 TSGs (78 somatic and 25 germline). Twenty somatic clusters contained recurring frameshifts in homopolymer runs, many in tumors with microsatellite instability. Germline clusters contain more germline-exclusive variants, some driving non-cancer phenotypes reflecting genetic pleiotropy. Altogether, germline and somatic variants of TSGs represent unique sets with substantially different patterns shaped by selection pressures from gene-specific and somatic mutational mechanisms. Characterizing these distinctions enables more accurate clinical interpretation of TSG variants.

Humans

Genetic Research on Cardiac Channelopathies in African and African-Descent Populations: A Scoping Review.

Cardiac channelopathies are inherited arrhythmias that can lead to sudden cardiac death. Despite Africa's extensive genomic diversity, African and African-descent populations remain underrepresented in genetic research, creating gaps in variant interpretation and clinical care. This scoping review aims to map the extent, range, and nature of genetic research on cardiac channelopathies in these populations and to identify key geographic, thematic, and methodological gaps. Using the Joanna Briggs Institute scoping review methodology and the Population-Concept-Context framework, systematic searches in PubMed, Embase, and Web of Science identified original human studies on cardiac channelopathies with genetic data. Extracted variables included study characteristics, populations, types of channelopathies, and reported genes and variants. Forty-four studies met the inclusion criteria. Most studies originated from the United States and South Africa, while West, Central, and East Africa were largely underrepresented. US Black individuals and South African individuals of continental African or African-descended ancestry (excluding populations of European descent such as Cape Afrikaner people) were the most studied groups, with other continental African groups rarely included. Long QT syndrome was the predominant focus, and SCN5A, KCNQ1, and KCNH2 were the most frequently analyzed genes. Many of the genetic variants discussed remained of uncertain significance due to limited functional validation and the underrepresentation of African genomes in reference databases. Genetic research on cardiac channelopathies in populations of African ancestry is limited, restricting variant interpretation, counseling, and risk prediction. Broader African inclusion, expanded gene screening, and functional studies are essential to improve diagnostics and promote equity in genomic medicine.

Humans

A Case Report of a Pedigree with Distal Hereditary Motor Neuropathy Caused by a Homozygous c.1124G>A Variant an the /*9Vaccinia-Related Kinase 1 Gene.

This study aimed to analyze the clinical phenotypes, neurophysiological characteristics, and pathogenicity of gene variants in a pedigree with distal hereditary motor neuropathy (dHMN) caused by VRK1 variants, and to provide evidence to support clinical diagnosis and genetic counseling for this disease. We report a Chinese consanguineous family with dHMN caused by a homozygous c.1124G>A variant in the VRK1 gene. Clinical and electrophysiological data of the proband were collected. Whole-exome sequencing (WES) and validation by Sanger sequencing were performed to identify the variant site, and pathogenicity interpretation was conducted in accordance with American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP) guidelines. The proband was a 24-year-old male who presented with 2 years of progressive weakness and atrophy of the distal lower limbs, accompanied by slender upper limbs and no sensory disturbance. Electrophysiological examination showed decreased compound muscle action potential (CMAP) amplitude in motor nerves of both upper and lower limbs, indicating peripheral neurogenic damage, while sensory nerve conduction was normal. Genetic testing detected a homozygous c.1124G>A (p.Trp375Ter) variant in the VRK1 gene. His parents and elder sisters were heterozygous carriers, and the pedigree conformed to autosomal recessive inheritance. According to ACMG guidelines, this variant was classified as pathogenic (evidence: PVS1, PM2, PP1). The homozygous VRK1 c.1124G>A variant causes adult-onset dHMN, rather than pontocerebellar hypoplasia type 1A (PCH1A) as annotated in some genetic databases. This pedigree presents distinctive phenotypes, including slender upper limbs and diffusely decreased CMAP amplitudes in both upper and lower limbs, thereby expanding the clinical and genetic spectrum of VRK1-related dHMN in the Chinese population.

Humans

Influence of ADRB2 variants on bronchodilator response and asthma control in a mixed population.

OBJECTIVE: Given that b2 agonists constitute the primary treatment for asthma and that treatment response varies as a result of polymorphisms in the ADRB2 gene, we sought to investigate the associations between ADRB2 gene variants and bronchodilator response (BDR) in asthma patients. METHODS: A genetic database comprising 813 individuals was analyzed for variants in the ADRB2 gene. A longitudinal analysis of severe asthma patients was performed to evaluate changes in BDR over time. RESULTS: The rs1042713, rs1042714, and rs1042717 variants were associated with age-related changes in BDR in patients with severe asthma. The G allele (rs1042714) and the A allele (rs1042717) were associated with uncontrolled asthma, with carriers of the G46/G79/A252 alleles showing a higher risk of difficult-to-control asthma. Notably, no association was found between these variants and ADRB2 expression levels. CONCLUSIONS: Our findings suggest that a genetic panel including ADRB2 variants, as well as age-related differences in BDR, is a useful complementary tool in asthma management.

Humans

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common ∼4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV

TRAIT: A Comprehensive Database for T-cell Receptor-antigen Interactions.

Comprehensive and integrated resources on interactions between T-cell receptors (TCRs) and antigens are still lacking for adoptive T-cell-based immunotherapies, highlighting a significant gap that must be addressed to fully understand the mechanisms of antigen recognition by T cells. In this study, we present the T-cell receptor-antigen interaction database (TRAIT), a comprehensive database that profiles the interactions between TCRs and antigens. TRAIT stands out due to its comprehensive description of TCR-antigen interactions by integrating sequences, structures, and affinities. It provides millions of experimentally validated TCR-antigen pairs, resulting in an exhaustive landscape of antigen-specific TCRs. Notably, TRAIT emphasizes single-cell omics as a major reliable data source for TCR-antigen interactions and includes millions of reliable non-interactive TCRs. Additionally, it thoroughly demonstrates the interactions between mutations of TCRs and antigens, thereby benefiting affinity optimization of engineered TCRs as well as vaccine design. TCRs on clinical trials are innovatively provided. With the significant efforts made toward elucidating the complex interactions between TCRs and antigens, TRAIT is expected to ultimately contribute superior algorithms and substantial advancements in the field of T-cell-based immunotherapies. TRAIT is freely accessible at https://pgx.zju.edu.cn/traitdb.

Receptors, Antigen, T-Cell

circASbase: A Comprehensive Database of Alternative Splicing Events in circRNAs.

Although extensive evidence has underscored the critical role of alternative splicing (AS) in generating mature circular RNA (circRNA) isoforms and augmenting their functional diversity, a significant gap remains in the availability of specialized databases housing circRNA AS events. To bridge this gap, we develop circASbase, a pioneering and comprehensive database that catalogs 452,129 AS events in 884,047 full-length circRNAs from 581 samples across 13 species, and provides rich annotations to facilitate understanding the splicing regulation of circRNA. Our findings reveal substantial differences between circRNAs and linear transcripts regarding the distribution and occurrence of AS events, highlighting the unique regulatory landscape of circRNAs. These special splicing events result in functional differences of circRNAs by affecting internal ribosome entry sites, N6-methyladenosine sites, open reading frames, protein features, microRNA targets, and more. In summary, circASbase not only meets the urgent need of the research community for data repositories, but also represents a significant advancement in our understanding of circRNA biology. With its user-friendly interfaces and web-based visualization tools, circASbase is poised to become an indispensable resource for researchers exploring the regulatory mechanisms and functional roles of AS events in circRNAs. This database will continuously drive new insights and discoveries in the field, setting the stage for further advancements in circRNA research. circASbase is freely available at http://reprod.njmu.edu.cn/cgi-bin/circASbase/.

Alternative Splicing

REvolutionH-tl 2.0: A fast and robust tool for decoding evolutionary gene histories.

REvolutionH-tl is a fast, scalable, and integrated software platform for inferring orthology relationships, gene trees, species trees, and reconciled evolutionary scenarios directly from sequence data. Built upon the formal framework of best match graphs (BMGs), REvolutionH-tl predicts orthogroups and orthologous gene pairs with high accuracy, requiring neither precomputed trees nor multiple external tools. The software reconstructs event-labeled gene and species trees, seamlessly integrating reconciliation to produce fast, accurate, and biologically insightful evolutionary scenarios. Through extensive benchmarking on synthetic datasets with known ground truth, REvolutionH-tl outperforms or matches the accuracy of established tools such as OrthoFinder, Proteinortho, RAxML, GeneRax, and RANGER-DTL, while achieving significantly lower runtimes. A key innovation of REvolutionH-tl is its built-in support for detailed, publication-ready visualizations, which allow users to explore genome evolution dynamics, orthogroup composition, and reconciliation results with clarity and ease. These visual features position REvolutionH-tl as the first platform of its kind to combine analytical precision with intuitive interpretability. The software is open-source, cross-platform, and freely available at https://pypi.org/project/revolutionhtl/, providing a robust solution for large-scale evolutionary analyses in comparative genomics.

Software

Estimation of carrier frequencies of autosomal and X-linked recessive genetic conditions based on gnomAD v4.0 data in different ancestries.

PURPOSE: Monogenic rare diseases contribute significantly to infant deaths and pediatric hospitalizations and cause burden to the patients and their families. The American College of Medical Genetics and Genomics recommended in 2021 that carrier screening of autosomal recessive and X-linked conditions with a carrier frequency of ≥1/200 and a severe or moderate phenotype should be offered when planning or during pregnancy. In November 2023 gnomAD v4.0 was released. It contains in total 807,162 individuals, being nearly 5× larger than previous versions, which have been used to estimate gene carrier frequencies (GCF). METHODS: We utilized gnomAD v4.0 (GRCh38) to calculate the GCFs for available genetic ancestry groups for variants having pathogenic or likely pathogenic classification (>80% of submissions) in ClinVar. We calculated GCF separately for exomes and genomes, combined data, and at-risk couple frequencies (ACF) per genetic ancestry group. RESULTS: In total, 324 genes had a GCF ≥1/200 in at least 1 ancestry subgroup. The number of genes with GCF ≥1/200 varied greatly between subgroups. ACFs were more similar, Ashkenazi Jewish having the highest ACF of 6.11%. CONCLUSION: Improved understanding of carrier risks and updated carrier screening content would allow patients to make more informed reproductive decisions.

Humans

eccDNABase: A Comprehensive and High-Quality Database for Extrachromosomal Circular DNA.

Extrachromosomal circular DNA (eccDNA) refers to small, circular DNA molecules that originate from chromosomal sequences and are prevalent across nearly all eukaryotic organisms. In humans, eccDNAs are widely distributed in normal tissues, cancerous tissues, and body fluids, where they play important roles in tumorigenesis and are often associated with poor clinical outcomes. Given their biological and clinical significance, a well-integrated and high-quality database is essential for advancing eccDNA-related research. To address this need, we developed eccDNABase, a comprehensive and curated resource for browsing, searching, and analyzing eccDNAs across multiple species. The database systematically catalogs eccDNA-disease associations from diverse tissues and organisms. Currently, eccDNABase contains 1,875,452 eccDNA-disease associations, encompassing 8,398 ecDNA entries across nine species, 63 diseases, and healthy individuals. Each entry provides detailed information, including eccDNA ID, type, chromosomal localization, species, tissue or cell line source, disease name and Disease Ontology ID, overlap length and percentage with genes, oncogene overlap, detection method, and links to literature and source databases. Given its extensive and curated datasets, eccDNABase serves as a valuable resource for both basic and translational research, offering deeper insights into the role of eccDNA in health and disease. The database is publicly accessible at http://cgga.org.cn/eccDNABase/.

Humans

Frequency of variants in Mendelian Alzheimer's disease genes within the Alzheimer's Disease Sequencing Project.

BackgroundPrior studies examined variants within presenilin-2 (PSEN2), presenilin-1 (PSEN1), and amyloid precursor protein (APP) genes. However, previously-reported clinically-relevant variants and other predicted damaging missense (DM) variants have not been characterized in a newer release of the Alzheimer's Disease Sequencing Project (ADSP).ObjectiveTo characterize previously-reported clinically-relevant variants and DM variants in PSEN2, PSEN1, APP within the participants from the ADSP.MethodsWe identified rare variants (MAF&#x2009;<&#x2009;1%) in PSEN2, PSEN1, and APP in 14,641 individuals with whole genome sequencing and 16,849 individuals with whole exome sequencing available (Ntotal&#x2009;=&#x2009;31,490). We additionally curated variants from ClinVar, OMIM, and Alzforum and report carriers of variants in clinical databases as well as predicted DM variants in these genes.ResultsWe detected 31 previously-reported clinically-relevant variants with alternate alleles observed within the ADSP: 4 variants in PSEN2, 25 in PSEN1, and 2 in APP. The overall variant carrier rate for the 31 clinically-relevant variants in the ADSP was 0.3%. We observed that 79.5% of the variant carriers were cases compared to 3.9% were controls. In those with AD, the mean age of onset of AD among carriers of these clinically-relevant variants was 19.6&#x2009;&#xb1;&#x2009;1.4 years earlier compared with noncarriers (p&#x2009;=&#x2009;7.8&#x2009;&#xd7;&#x2009;10-57). Additionally, we identified 197 rare variants (MAF&#x2009;<&#x2009;1%) within ADSP participants not reported in known clinical databases.ConclusionsA small proportion of individuals in the ADSP are carriers of a previously-reported clinically-relevant variant allele for AD and these participants have significantly earlier age of AD onset compared to noncarriers.

Humans

Characterizing trends in clinical genetic testing: A single-center analysis of EHR data from 1.8 million patients over two decades.

A lack of structural data in electronic health records (EHRs) makes assessing the impact of genetic testing on clinical practice challenging. We extracted clinical genetic tests from the EHRs of more than 1.8 million patients seen at Vanderbilt University Medical Center from 2002 to 2022. With these data, we quantified the use of clinical genetic testing in healthcare and described how testing patterns and results changed over time. We assessed trends in types of genetic tests, tracked usage across medical specialties, and introduced a new measure, the genetically attributable fraction (GAF), to quantify the proportion of observed phenotypes attributable to a genetic diagnosis over time. We identified 104,392 tests and 19,032 molecularly confirmed diagnoses. The proportion of patients with genetic testing in their EHRs increased from 1.0% in 2002 to 6.1% in 2022, and testing became more comprehensive with the growing use of multi-gene panels. The number of unique diseases diagnosed with genetic testing increased from 51 in 2002 to 509 in 2022, and there was a rise in the number of variants of uncertain significance. The phenome-wide GAF for 6,505,620 diagnoses made in 2022 was 0.46%, and the GAF was greater than 5% for 74 phenotypes, including pancreatic insufficiency (67%), chorea (64%), atrial septal defect (24%), microcephaly (17%), paraganglioma (17%), and ovarian cancer (6.8%). Our study provides a comprehensive quantification of the increasing role of genetic testing at a major academic medical institution and demonstrates its growing utility in explaining the observed medical phenome.

Humans

Comprehensive evaluation of ACMG/AMP-based variant classification tools.

MOTIVATION: The American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP) guidelines represent the gold standard for clinical variant interpretation. Despite the widespread adoption of ACMG/AMP guidelines, a comprehensive comparison of the software tools designed to implement them has been lacking. This represents a significant gap, as clinicians require evidence-based guidance on which tools to use in their practice. RESULTS: We benchmarked four ACMG/AMP-based tools (Franklin, InterVar, TAPES, Genebe) selected from 22 tools, and compared their performance with LIRICAL, a top-performing phenotype-driven tool, using 151 expert-curated datasets from Mendelian disorders. Selection criteria included free availability, VCF compatibility, operational reliability, and not being disease-specific. Our evaluation framework assessed top-N accuracy (N&#x2009;=&#x2009;1, 5, 10, 20, 50), retention rates, precision, recall, F1 scores, and area under the curve (AUC). Statistical validation employed bootstrap confidence intervals (n&#x2009;=&#x2009;1000) and Friedman tests. LIRICAL (68.21%) and Franklin (61.59%) demonstrated superior top-10 variant prioritization accuracy in Mendelian disorders, significantly outperforming other tools (P&#x2009;=&#x2009;.0000). Results demonstrate that tools with advanced phenotypic integration significantly outperform those relying primarily on genomic features. AVAILABILITY AND IMPLEMENTATION: All data and source code required to reproduce the findings of this study are openly available in the Code Ocean repository at https://doi.org/10.24433/CO.6562438.v1.

Software

PlantPan: A comprehensive multi-species plant pan-genome database.

The pan-genome represents the complete genomic diversity of specific species, serving as a valuable resource for studying species evolution, crop domestication, and guiding crop breeding and improvement. While there are several single-species-specific plant pan-genome databases, the availability of multi-species pan-genome databases is limited. Additionally, variations in methods and data types used for plant pan-genome analysis across different databases hinder the comparison and integration of pan-genome information from various projects at multi-species or single-species levels. To tackle this challenge, we introduce PlantPan, a comprehensive database housing the results of pan-genome analysis for 195 genomes from 11 plant species. PlantPan aims to provide extensive information, including gene-centric and sequence-centric pan-genome information, graph-based pan-genome, pan-genome openness profiles, gene functions and its variation characteristics, homologous genes, and gene clusters across different species. Statistically, PlantPan incorporates 9&#x2009;163&#x2009;011 genes, 694&#x2009;191 gene clusters, 526&#x2009;973&#x2009;370 genome variations, and 1&#x2009;616&#x2009;089 non-redundant genome variation groups at the species level, 33&#x2009;455,098 genome synteny, and 177&#x2009;827 non-redundant genome synteny groups at the species level. Regarding functional genes, PlantPan contains 5&#x2009;222&#x2009;720 genes related to transcription factors, 395&#x2009;247 literature-reported resistance genes, 455&#x2009;748 predicted microbial/disease resistance genes, and 1&#x2009;612&#x2009;112 genes related to molecular pathways. In summary, PlantPan is a vital platform for advancing the application of pan-genomes in molecular breeding for crops and evolutionary research for plants.

Genome, Plant

cfMethDB: A Comprehensive cfDNA Methylation Data Resource for Cancer Biomarkers.

Cancer is a major global health threat, and early detection is crucial for improving patient outcomes. DNA methylation in circulating cell-free DNA (cfDNA) has emerged as a promising biomarker for non-invasive cancer diagnosis. However, the integration and utilization of existing cfDNA methylation data have been limited, hindering comprehensive research efforts, particularly in the discovery of cfDNA methylation biomarkers. To address this challenge, we introduced cfMethDB, a comprehensive database dedicated to cfDNA methylation in cancer that encompasses 4828 publicly available datasets. Through standardized analysis, we identified 1,048,770 differentially methylated cytosines (DMCs) as candidate biomarkers across seven cancer types. With cfMethDB, we not only identified known cfDNA methylation biomarkers, but also discovered several genes, such as ZIC4, that could be novel biomarkers. Moreover, cfMethDB offers a suite of user-friendly tools, including biomarker evaluation, pan-cancer search, and end motif analysis. We hope that cfMethDB will serve as a valuable platform for the discovery of novel cancer cfDNA methylation biomarkers and facilitate cancer research and clinical applications. cfMethDB is publicly available at https://cfmethdb.hzau.edu.cn/home.

Humans

Microbiome Datahub: an open-access platform integrating environmental metadata, taxonomy, and functional annotation for comprehensive metagenome-assembled genome datasets.

BACKGROUND: Metagenome-assembled genomes (MAGs) provide crucial insights into the genomic diversity of uncultured microbes. However, MAG datasets deposited in public repositories such as INSDC are often difficult to reuse due to heterogeneous quality, inconsistent taxonomic and functional annotations, and insufficiently curated environmental metadata. While secondary MAG databases such as MGnify, IMG/M, and SPIRE provide standardized resources, they reconstruct MAGs de novo from public metagenomic reads and therefore do not represent the original MAGs reported in publications. RESULTS: To address this gap, we developed Microbiome Datahub, an open-access platform that systematically aggregates and re-annotates original MAGs from INSDC. We collected 214,427 MAGs, predicted genes by DFAST, performed quality assessment with CheckM, standardized taxonomic assignments with GTDB-Tk, inferred 27 phenotypic traits using Bac2Feature, assigned proteins to MBGD ortholog clusters and KEGG Orthology IDs using PZLAST, and annotated environmental metadata with the Metagenome and Microbes Environmental Ontology. Across these MAGs, the average completeness was 80.5% and contamination 1.8%; notably, the most frequent values were&#x2009;>95% completeness and&#x2009;<1% contamination, indicating that the majority of MAGs are of high quality. Comparative analyses showed that Microbiome Datahub provides phylogenetically and environmentally diverse MAGs: while the majority originated from vertebrate gut environments, a substantial number were also recovered from other habitats such as groundwater, including nearly 10,000 MAGs from the Patescibacteria. Inference of 27 phenotypic traits, including optimum growth temperature, further revealed ecological differentiation across phyla. Protein clustering revealed 56 million identity 40% clusters, with the majority unique compared with MGnify and GlobDB, and&#x2009;~19% of proteins unassigned to MBGD ortholog clusters, underscoring their novelty. CONCLUSIONS: Microbiome Datahub integrates MAG genome sequences, gene and protein predictions, quality metrics, environmental and taxonomic annotations, ortholog cluster assignments, and phenotype predictions, all accessible via a web interface, API, and bulk downloads. By combining original MAGs with curated metadata and functional annotations, Microbiome Datahub constitutes a comprehensive and reusable resource that will accelerate microbiome and microbial genomics research. Video Abstract.

Metagenome

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software