Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Unique signatures of highly constrained genes across publicly available genomic databases.

PURPOSE: Publicly available genomic databases are critical in understanding human genetic variation. They also provide unique insights into patterns of genetic constraints and their relationship with human disease. METHODS: We utilized one of the largest publicly available databases, Genome Aggregate Database, to determine genes that are highly constrained for only loss-of-function, only missense, and both loss-of-function/missense variants. We identified their unique signatures and explored their causal relationship with human diseases. Those genes were also evaluated for chromosomal location, tissue-level expression, Gene Ontology analysis, and gene family categorization using multiple publicly available databases. RESULTS: We identified unique patterns of inheritance, protein size, and enrichment in distinct molecular pathways for those constrained genes associated with human disease. In addition, we identified genes that are currently not known to cause human disease, which may be excellent gene discovery candidates. CONCLUSION: We elucidate biological pathways of highly constrained genes that expand our understanding of critical cellular proteins. The findings can also advance research in rare diseases.

Humans

PlantPan: A comprehensive multi-species plant pan-genome database.

The pan-genome represents the complete genomic diversity of specific species, serving as a valuable resource for studying species evolution, crop domestication, and guiding crop breeding and improvement. While there are several single-species-specific plant pan-genome databases, the availability of multi-species pan-genome databases is limited. Additionally, variations in methods and data types used for plant pan-genome analysis across different databases hinder the comparison and integration of pan-genome information from various projects at multi-species or single-species levels. To tackle this challenge, we introduce PlantPan, a comprehensive database housing the results of pan-genome analysis for 195 genomes from 11 plant species. PlantPan aims to provide extensive information, including gene-centric and sequence-centric pan-genome information, graph-based pan-genome, pan-genome openness profiles, gene functions and its variation characteristics, homologous genes, and gene clusters across different species. Statistically, PlantPan incorporates 9 163 011 genes, 694 191 gene clusters, 526 973 370 genome variations, and 1 616 089 non-redundant genome variation groups at the species level, 33 455,098 genome synteny, and 177 827 non-redundant genome synteny groups at the species level. Regarding functional genes, PlantPan contains 5 222 720 genes related to transcription factors, 395 247 literature-reported resistance genes, 455 748 predicted microbial/disease resistance genes, and 1 612 112 genes related to molecular pathways. In summary, PlantPan is a vital platform for advancing the application of pan-genomes in molecular breeding for crops and evolutionary research for plants.

Genome, Plant

An integrated culturomic and genomic database and analysis platform for methanogenic archaea.

Methanogenic archaea research is challenged by limited strain resources, fragmented genomic data, inconsistent genome quality, substantial uncultured lineages, and difficulties in laboratory culturing, hindering advances in biogas production, climate mitigation, and microbial ecology. These archaea play crucial roles in global carbon cycling and anaerobic environments, yet scattered data and unculturable strains limit systematic studies and applications. To address this, we created MethArDB (Methanogenic Archaeal Genome Database), a specialized database for methanogenic archaea, compiling 3919 genomes, 87 host-associated plasmids, and 42 phages, with standardized quality classifications (complete, scaffold, draft), protein sequences, and metadata on geography, habitats, metabolism, and inheritable elements. Integrated MethArCT (Methanogenic Archaeal Culturomics Toolkit) employs a dual-threshold orthologous/paralogous protein analysis to evaluate metabolic pathway completeness, predicting cultivation parameters and suggesting candidate cultivation strategies, including potential medium formulations and conditions, to support strain isolation. Overall, MethArDB and MethArCT form an integrated platform combining genomics and culturomics to facilitate methanogenic archaea research. Database URL:  http://methardb.cn.

Genome, Archaeal

Global prevalence of hereditary hemorrhagic telangiectasia-associated variants estimated by analysis of large-scale genomic databases.

BACKGROUND: Hereditary hemorrhagic telangiectasia (HHT) is an autosomal dominant disorder with an overwhelming hemorrhagic phenotype. It is mainly caused by variants in the ENG and ACVRL1 genes. HHT prevalence is currently estimated to be 1 in 5000 individuals, but the disease is likely underdiagnosed due to variable clinical presentation, misdiagnosis, and delayed recognition. OBJECTIVES: To estimate the global genetic prevalence of HHT-associated variants in ENG and ACVRL1. METHODS: We analyzed 3 large population-scale genomic databases: gnomAD, All of Us, and Regeneron Genetics Center-Million Exome. We considered known pathogenic and likely pathogenic variants of ENG and ACVRL1 and extended the analysis to potentially pathogenic variants passing the pathogenic criteria established by the guidelines for HHT of the American College of Medical Genetics and Genomics/Association for Molecular Pathology. RESULTS: The genetic prevalence of HHT ranged from 1.753 to 2.555 in 5000 individuals, when considering only pathogenic and likely pathogenic variants, and from 2.874 to 4.327 in 5000 individuals, when also potentially pathogenic variants were considered. CONCLUSION: This study assesses the prevalence of HHT-associated variants in the general population. Our unbiased approach demonstrates that the genetic prevalence of the disease is substantially higher than currently estimated.

Humans

Microbial genomic database of the Yangtze River, the third-longest river on Earth.

Microbes play an important role in mediating the nutrient cycling in the river ecosystem as a hotspot for biogeochemical processes. Due to scattered sampling efforts, however, there is a lack of a systematic study of the diversity of prokaryotic genomes in the Yangtze River, the third longest river on Earth. Here, we collected 602 metagenomic datasets of water, sediment and riparian soil samples spanning the Upper, Middle, and Lower basins of the Yangtze River over a 6,300 km continuum. We reconstructed 8,110 qualified genomes represented by 927 species-level genomes at the 95% ANI threshold, spanning 31 bacterial and five archaeal phyla. We further showed that more than half of these species (61.3% ~ 82.4%) were novel according to the genomic comparison against the curated databases, greatly expanding the known diversity of river prokaryotes. This dataset depicts an overview of microbial genomic diversity in the Yangtze River and provides a resource for in-depth investigation of metabolic potential, ecology, and evolution of riverine microbiomes.

Rivers

The SARS-CoV-2 Integrated Genomic Epidemiology Database (IGED): Linking viral genomes with patient-level metadata to advance statewide genomic surveillance in California.

In July 2021, the California Code of Regulations Title 17 required all laboratories performing SARS‑CoV‑2 whole genome sequencing (WGS) to report their sequencing results to the California Department of Public Health (CDPH). These viral genomic data and patient metadata were compiled into the Integrated Genomic Epidemiology Database (IGED). Linking anonymized viral sequences with patient‑level information enabled monitoring of infectiousness, pathogenicity, transmission dynamics, evolution, and vaccine evasion among emerging SARS‑CoV‑2 lineages. Laboratories performing SARS-CoV-2 WGS transmitted sequencing results to CDPH through Electronic Laboratory Reporting (ELR) and non-ELR pathways. CDPH applied uniform reporting requirements but allowed flexibility in specific data formats to accommodate diverse data systems. To preserve data quality and interoperability across heterogeneous sources, CDPH implemented standardization, validation, and deduplication protocols. Snowflake, a cloud‑based data storage and analytics platform, and Posit Connect, a cloud deployment and automation platform, supported the management, processing, and integration of data within the IGED. The IGED established links between SARS‑CoV‑2 WGS data and epidemiologic metadata for 801,418 sequences, representing 81.7% of all sequences reported in California. Lineages reported to the IGED showed strong concordance with lineage proportions in GISAID. Sequences reported to the IGED had average turnaround times longer than one month, and the majority of sequencing was performed in Southern California and Los Angeles. The IGED enhanced genomic surveillance through predictive modeling and monitoring concerning evolutionary trends such as recombination and saltations in persistent infections. Development of the IGED highlighted the need for standardized data requirements, sustained funding for sequencing, incentives for data submission, and interdisciplinary collaboration to build an effective genomic surveillance system. This framework for linking genomic and epidemiologic data has not only generated critical insights for SARS‑CoV‑2 but also provided the foundation for CDPH and other public health organizations to develop similar IGED‑like systems for other priority pathogens as genomic surveillance expands.

Journal Article

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus

The Saccharomyces Genome Database-a history of ideas and accomplishments, 1994-2026.

The Saccharomyces Genome Database (SGD) is one of the longest-running and most consequential biological databases in the world. Founded in the early 1990s at Stanford University under the visionary leadership of David Botstein and developed under the long-term technical direction of J. Michael Cherry, SGD has served for more than three decades not only as the authoritative knowledge center for the budding yeast Saccharomyces cerevisiae, but also as the source for much of the fundamentals of eukaryotic biology. This history traces the arc of a remarkable intellectual and scientific project: beginning with the challenge of building the very first integrated eukaryotic genome database and evolving across 30 years into a global knowledge hub for genetics, functional genomics, and human disease research. The history is organized chronologically, with each section highlighting the central ideas, technical developments, and concrete accomplishments of that period.

Databases, Genetic

Exploring penetrance of clinically relevant variants in over 800,000 humans from the Genome Aggregation Database.

Incomplete penetrance, or absence of disease phenotype in an individual with a disease-associated variant, is a major challenge in variant interpretation. Studying individuals with apparent incomplete penetrance can shed light on underlying drivers of altered phenotype penetrance. Here, we investigate clinically relevant variants from ClinVar in 807,162 individuals from the Genome Aggregation Database (gnomAD), demonstrating improved representation in gnomAD version 4. We then conduct a comprehensive case-by-case assessment of 734 predicted loss of function variants in 77 genes associated with severe, early-onset, highly penetrant haploinsufficient disease. Here, we identify explanations for the presumed lack of disease manifestation in 701 of 734 variants (95%). Individuals with unexplained lack of disease manifestation in this set of disorders are rare, underscoring the need and power of deep case-by-case assessment presented here to minimize false assignments of disease risk, particularly in unaffected individuals with higher rates of secondary properties that result in rescue.

Humans

The planktonic microbiome of the Great Barrier Reef.

Large genome databases have markedly improved our understanding of marine microorganisms1-5. Although these resources have focused on prokaryotes, genomes from many dominant marine lineages, such as Pelagibacter and Prochlorococcus, are conspicuously underrepresented. Here we present the Great Barrier Reef Microbial Genomes Database (GBR-MGD), comprising 5,283 prokaryotic genomes obtained from Great Barrier Reef seawater samples using Nanopore and Illumina sequencing, including a collection of high-quality genomes of underrepresented groups. We show that standard short-read assemblies miss these populations owing to a combination of strain heterogeneity and low-GC-percentage sequencing bias. The GBR-MGD also comprises 20 chromosome-level picoeukaryote and 808,585 viral genomes, including a newly described clade of marine Crassvirales. We demonstrate the utility of the GBR-MGD to identify indicator taxa that can reliably predict the effects of reef management practices, such as the establishment of marine protected zones.

Bacteria

Role of ctDNA Tumor Fraction in Selecting Immunotherapy-Based Regimens in Advanced Non-Small Cell Lung Cancer.

PURPOSE: Immune checkpoint blockers (ICB) have transformed advanced non-small cell lung cancer (aNSCLC) treatment, but identifying patients who benefit from adding chemotherapy remains challenging, especially in PD-L1 &#x2265; 50%. PD-L1 is an imperfect biomarker, highlighting the need for better selection tools. EXPERIMENTAL DESIGN: Liquid biopsy (LBx) assessment was performed using hybrid capture-based next-generation sequencing of plasma cell-free DNA. LBx data, molecular profile, and clinicopathologic data were collected. The predictive and prognostic values of tumor fraction (TF) were assessed using a deidentified nationwide (US-based) NSCLC clinicogenomic database [Clinico-Genomic Database (CGDB)]. An independent cohort with aNSCLC from Gustave Roussy was used to validate the findings and to study the correlation of circulating tumor DNA (ctDNA) TF and total metabolic tumor volume and its molecular correlates. RESULTS: In the CGDB database (n = 965), elevated ctDNA TF was prognostic for worse outcomes on ICBs and, when &#x2265;5%, predictive of benefit from ICB + chemotherapy [HR for real-world progression-free survival 0.58 (0.41-0.82); P = 0.002]. The 5% cutoff for TF was validated in an independent cohort from Gustave Roussy. In 283 patients with paired PET scans, ctDNA TF correlated with metabolic tumor volume (rho = 0.46; P < 0.001) and was influenced by TP53/RB1 mutations. CONCLUSIONS: ctDNA TF integrates disease burden and biology. Patients with high ctDNA TF derive greater benefit from chemoimmunotherapy, supporting its use as a biomarker to guide treatment intensification.

Humans

Mapping the oral microbiome opens links to periodontitis.

Many microbiome analysis techniques can only detect the microbes present in the reference genome database used. In this issue of Cell Host & Microbe, Cha et al. establish an improved genome database of the human oral microbiome, which they use to discover a connection between periodontitis and an enigmatic bacterial phylum.

Humans

CAGEcleaner: reducing genomic redundancy in gene cluster mining.

SUMMARY: Mining homologous biosynthetic gene clusters (BGCs) typically involves searching colocalised genes against large genomic databases. However, the high degree of genomic redundancy in these databases often propagates into the resulting hit sets, complicating downstream analyses and visualization. To address this challenge, we present CAGEcleaner, a Python-based pipeline with auxiliary bash scripts designed to reduce redundancy in gene cluster hit sets by dereplicating the genomes that host these hits. CAGEcleaner integrates seamlessly with widely used gene cluster mining tools, such as cblaster and CAGECAT, enabling efficient filtering and streamlining BGC discovery workflows. AVAILABILITY AND IMPLEMENTATION: Source code and documentation is hosted at GitHub (https://github.com/LucoDevro/CAGEcleaner) and Zenodo (https://doi.org/10.5281/zenodo.14726119) under an MIT license. For accessibility, CAGEcleaner is installable from Bioconda (https://anaconda.org/bioconda/cagecleaner) and PyPi (https://pypi.org/project/cagecleaner/), and is also available as a Docker image from DockerHub (https://hub.docker.com/r/lucodevro/cagecleaner).

Software

Real-World Treatment Patterns and Clinical Outcomes After First-Line Therapy in Patients with KRAS G12C-Mutant Advanced Non-Small-Cell Lung Cancer in the United States.

BACKGROUND: Approximately 13% of NSCLC cases have KRAS G12C mutations. As therapeutic strategies targeting KRAS G12C-mutant NSCLC evolve, it is important to understand clinical presentation and current outcomes for these patients. METHODS: This retrospective study used data from two US nationwide databases, an electronic health records (EHR) database and a clinico-genomic database (CGDB) of EHR data linked to data from comprehensive genomic profiling tests. Eligible patients had advanced NSCLC, initiated first-line therapy from August 2018 to December 2022, and had KRAS test results. Clinicopathologic characteristics, treatments, real-world progression-free survival (rwPFS), and overall survival (OS) were analyzed. RESULTS: There were 1227 patients with KRAS G12C-mutant NSCLC in the EHR database and 447 in the CGDB. First-line regimen was platinum-based chemotherapy plus pembrolizumab for 46% and pembrolizumab monotherapy for 20%. Less than 40% of patients received second-line therapy. Median (95% CI) OS for KRAS G12C-mutant NSCLC patients in the EHR was 17.0 (15.2-18.9) months. Variables significantly associated with shorter OS included PD-L1 <1%, brain metastases, STK11 co-mutation, and poor performance status. Patients treated with platinum-based chemotherapy plus pembrolizumab had median rwPFS of 5.3 (4.5-7.3) months and OS of 12.8 (11.1-17.3) months in the CGDB; median OS was 15.6 (12.5-18.6) months in the EHR. Patients with PD-L1 &#x2265; 50% treated with pembrolizumab monotherapy had median rwPFS of 4.6 (3.0-15.6) months and OS of 20.4 (10.3-38.5) months in the CGDB; median OS was 22.1 (18.7-30.7) in the EHR. CONCLUSIONS: These data provide a real-world benchmark of outcomes for patients with KRAS G12C-mutant NSCLC receiving the current standard of care and indicate an unmet need for more effective first-line therapies.

KRAS G12C

Screening and identification of key genes related to the immune microenvironment of rectal cancer influenced by radiotherapy based on bioinformatics methods.

OBJECTIVE: Radiotherapy (RT) plays a crucial role in the comprehensive treatment of rectal cancer. However, the impact of radiotherapy on the tumor microenvironment (TME), especially its effect on immune cell infiltration and immune-related gene expression, has not been fully studied. This study aims to screen and analyze key genes related to the immune microenvironment of rectal cancer influenced by radiotherapy based on bioinformatics methods for the purpose of identifying potential biomarkers and providing new insights for the personalized therapy of rectal cancer. METHODS: Using data from the Public Gene Expression Database (GEO) and the Cancer Genomics Database (TCGA), the impact of radiotherapy on the immune microenvironment of rectal cancer was explored using bioinformatics tools. Through screening differentially expressed genes (DEGs), correlation analysis, TIMER database analysis, immune infiltration score, and correlation analysis between key genes and prognosis, the effects of radiotherapy on the immune microenvironment of rectal cancer were investigated. RESULTS: Totally 7 upregulated and 4 downregulated differentially expressed genes were identified, among which MASP1, LTK, SLC9A3R2 were negatively correlated with myeloid suppressor cell infiltration (MDSCs), while ZP2 was positively correlated. The expression of MASP1 and SLC9A3R2 was closely related to the level of immune cell infiltration and played significant roles in the immune microenvironment. High expression of MASP1 was significantly correlated with survival benefits from immune checkpoint inhibitor therapy, while SLC9A3R2 was closely related to the efficacy of PD-L1 inhibitors and CTLA4 inhibitors. CONCLUSIONS: MASP1 and SLC9A3R2, as two key genes that may be related to the immune microenvironment of rectal cancer radiotherapy, deserve further exploration of their roles in the mechanism. The combination of radiotherapy and immunotherapy holds promising prospects in the treatment of rectal cancer, and exploration of related mechanisms will provide new strategies and targets for the treatment of various tumors and rectal cancer.

Bioinformatics

Private detection of relatives in forensic genomics using homomorphic encryption.

BACKGROUND: Forensic analysis heavily relies on DNA analysis techniques, notably autosomal Single Nucleotide Polymorphisms (SNPs), to expedite the identification of unknown suspects through genomic database searches. However, the uniqueness of an individual's genome sequence designates it as Personal Identifiable Information (PII), subjecting it to stringent privacy regulations that can impede data access and analysis, as well as restrict the parties allowed to handle the data. Homomorphic Encryption (HE) emerges as a promising solution, enabling the execution of complex functions on encrypted data without the need for decryption. HE not only permits the processing of PII as soon as it is collected and encrypted, such as at a crime scene, but also expands the potential for data processing by multiple entities and artificial intelligence services. METHODS: This study introduces HE-based privacy-preserving methods for SNP DNA analysis, offering a means to compute kinship scores for a set of genome queries while meticulously preserving data privacy. We present three distinct approaches, including one unsupervised and two supervised methods, all of which demonstrated exceptional performance in the iDASH 2023 Track 1 competition. RESULTS: Our HE-based methods can rapidly predict 400 kinship scores from an encrypted database containing 2000 entries within seconds, capitalizing on advanced technologies like Intel AVX vector extensions, Intel HEXL, and Microsoft SEAL HE libraries. Crucially, all three methods achieve remarkable accuracy levels (ranging from 96% to 100%), as evaluated by the auROC score metric, while maintaining robust 128-bit security. These findings underscore the transformative potential of HE in both safeguarding genomic data privacy and streamlining precise DNA analysis. CONCLUSIONS: Results demonstrate that HE-based solutions can be computationally practical to protect genomic privacy during screening of candidate matches for further genealogy analysis in Forensic Genetic Genealogy (FGG).

Humans

Effect of Chang'an decoction on ulcerative colitis by regulating T helper 17 cells and regulatory T cellsRab27 in the p53/high mobility group box 1 pathway.

OBJECTIVE: To explore the effect of Chang'an decoction (, CAD) of ameliorating the immune imbalances in ulcerative colitis (UC) by regulating Rab27 in the P53/high mobility group box 1 pathway. METHODS: The functions and important signaling pathways of the Rab27- and UC-related genes were analyzed viathe use of microarray data from the gene expression omnibus database, gene ontology database, Kyoto encyclopedia of genes and genomes database and gene set enrichment analysis. Dextran sulfate sodium salt-induced colitis mouse model was used to verify the bioinformatics results. Colon length, body weight, and disease activity index were measured. Hematoxylin and eosin staining was applied to validate the histopathology. Tight junction proteins were detected by immunohistochemistry. The proportions of T helper 17 cells (Th17) and regulatory T cells (Treg) in mesenteric lymph nodes were measured viaflow cytometry. Proinflammatory cytokines like interleukin (IL) 17 (IL-17), IL-21 and IL-22 and anti-inflammatory cytokines like transforming growth factor &#x3b2; and IL-10 in the serum and colon of mice were detected by enzyme-linked immunosorbent assay and quantitative real-time polymerase chain reaction, respectively. The expression levels of high mobility group box 1 (HMGB1), P53 and phospho- P53 (P-P53) in colonic tissues were detected by immunofluorescence and Western blotting. RESULTS: Bioinformatics analysis revealed that compared with normal tissues, the expression of Rab27 was significantly increased in UC tissues. Receiver operating characteristic curve showed that Rab27 has the potential to be used as a biomarker for the diagnosis of disease activity. Enrichment analysis showed that UC and Rab27 were mainly associated with small molecule transport, nutrient metabolism, transmembrane transport and the downstream pathway of P53. According to animal experiments, the expression of Rab27 was increased in UC tissues, which aggravated the colonic pathological damage, activated the expression of HMGB1, and also leaded to the imbalance of Th17 and Treg cells. After CAD intervention, Rab27 overexpression, weight loss, colon shortening, and pathological damage were substantial reduced, the expression of tight junction proteins, zona occludens 1 and Occludin were increased. The effect of CAD at high-dose was more obvious. In addition, CAD upgraded the number of Treg cells and the production of TGF-&#x3b2; and IL-10, while decreasing the number of Th17 cells and the expression of inflammatory cytokines (IL-17, IL-21, and IL-22). Moreover, colon inflammation was alleviated by CAD, as indicated by the regulation of HMGB1 and P-P53 expression. CONCLUSION: The expression of Rab27, HMGB1 and P-P53 could be decreased by CAD, and the balance of Th17 and Treg cells as well as their related cytokines could be regulated by CAD.

Animals

Genome-based predictions of metabolic preferences and substrate phenotypes in psychrotrophic bacteria from permafrost environments.

Genomes reveal vast functional potential, but harbor genomic noise that obscures prediction of metabolic and environmental preferences. Genomic databases are skewed towards clinically relevant and easily cultivated bacteria, limiting predictions for diverse and underrepresented environmental taxa. Psychrotrophic bacteria, which can survive and grow in cold, nutrient-limited, dry, and saline environments, are especially underrepresented despite their relevance for understanding microbial responses to changing cold environments and potential biotechnological value given growth at low temperatures. Assembling complete genomes of 48 isolates from Alaskan permafrost, seasonally frozen active layer soils, and terrestrial ice, we used Kyoto Encyclopedia of Genes and Genomes (KEGG) ortholog annotations to evaluate the predictability of metabolic resource-use traits observed using phenotypic tests. Genome-predicted values for glycolytic versus gluconeogenic catabolic preference index, or sugar-acid preference (SAP), explained over 50% of the variance in empirically observed SAP. SAP was inversely correlated to genomic GC content, which follows phylum-level trends, indicating that coarse metabolic preference covaries with phylogeny. Regularized elastic net models offered a more granular view, linking KEGG genes to specific substrate utilization and sensitivity phenotypes and yielding moderate but reproducible accuracy (AUC 0.70-0.79) for 11 substrates, demonstrating that specific substrate responses may be predictable from relatively small subsets of KO genes. These results extend recent advances, such as the SAP metric, and highlight associations among genomic GC content, phylum, and broad metabolic strategy. Linking genomic content to phenotype using isolates is a necessary step toward predictive models of microbial function in environmental communities, and this work can be used for hypothesis generation, with applications towards more expansive data sets.IMPORTANCECold region soils and ice host psychrotrophic bacteria with metabolic traits and adaptations that enable persistence in harsh, resource-limited environments. However, these taxa are underrepresented in genomic reference databases dominated by well-studied, mesophilic organisms. This gap limits inference of ecological strategies and our ability to predict how these microbes may influence the large, thaw-vulnerable carbon reservoirs in permafrost. Here, we show that genomic GC content is associated with the sugar-versus-acid catabolic preference (SAP) of isolates across major phyla, suggesting that broad genomic features may provide a coarse signal of metabolic strategy. We demonstrate that a modified SAP metric, using binary (positive/negative) substrate utilization rather than detailed growth rate measurements, is moderately predictive, thus extending its application to slow-growing or difficult-to-culture taxa. Together, these advances broaden the toolkit for linking genome content to resource-use traits (phenotype) in poorly characterized, cold-adapted bacteria and offer a tractable entry point to broad prediction and hypothesis generation.

Genome, Bacterial