Search PubMedSearch

SEARCH · Search PubMed

Results for “Global genome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

High-level terpene production via a novel Actinomycetota-derived MVA pathway in E. coli.

The heterologous production of terpene in microbial hosts is often limited by inefficient and unstable pathway expression, creating a major bottleneck for industrial-scale synthesis. While E. coli as a chassis offers significant advantages, such as rapid growth, ease of cultivation, and genetic tractability. Its endogenous supply of terpenoid precursors remains a critical constraint, fundamentally restricting high-yield production. To address this challenge, we developed a genomically integrated Mevalonate (MVA) pathway from Actinomycetota in E. coli BL21(DE3) to enhance terpene precursor supply. Our approach began with an in silico multi-layer global genome mining analysis of 25,261 Actinomycetota genomes to identify a series of MVA pathway enzymes with potentially high catalytic efficiency, created a high-efficiency chassis E. coli MVA platform (ecMVA-1 and ecMVA-2) for terpene precursor synthesis. Its functionality was validated by testing eight distinct TSs. Among them, the fermentation of artemisinin precursor amorphadiene using a 5-liter bioreactor yielded 947.80 mg/L. These results indicated that E. coli (MVA) is well-suited for TS studies in the laboratory as well as holding significant promise for industrial applications. In addition, this in silico approach offers a new perspective for metabolic engineering and provides potential reservoir of diverse chassis for the industrial production of terpenoid-derived compounds.

Actinomycetota

Integrated Genome Mining and Bioactivity-Guided Isolation of Antimicrobial Peptides from Bacillus amyloliquefaciens BS4.

Bacterial resistance remains a critical global health challenge, driving the continuous search for novel antimicrobial agents. Bacillus amyloliquefaciens is a recognized repository of bioactive metabolites; however, its full biosynthetic potential requires integrated genomic and experimental validation. This study characterized the antimicrobial profile of B. amyloliquefaciens BS4 through a hybrid pipeline. Genome sequencing and de novo assembly revealed a 3.9 Mb chromosome with a G + C content of 46.14%. Functional annotation identified 3,887 coding sequences, including pathways for siderophore biosynthesis and a complete bacilysin biosynthetic cluster. BGC analysis using antiSMASH v7.1.0 and BAGEL4 identified 18 biosynthetic gene clusters, while similarity network analysis via BiG-SCAPE highlighted unique singleton BGCs, indicating untapped biosynthetic diversity. Although in silico screening via Macrel predicted two putative cationic antimicrobial peptides (AMPs), bioactivity-guided purification utilizing sequential RP-HPLC, and de novo sequencing revealed a distinct set of four active peptides. Notably, three of these sequences were identified as fragments derived from the BclA exosporium protein family, highlighting the structural proteome as a non-canonical source of antimicrobials. The purified fractions exhibited activity against M. luteus and E. coli, while displaying no significant hemolytic activity or cytotoxicity, even above the MIC values. Molecular docking further supported the interaction of these candidates with bacterial targets. Overall, this hybrid strategy effectively uncovers the antimicrobial complexity of BS4, revealing 'cryptic' peptide candidates with therapeutic potential.

Bacillus amyloliquefaciens BS4

Unraveling the diversity, function, and virus-host interactions of archaeal proviruses.

Archaea, the third domain of life, play critical roles in global biogeochemical cycles. However, archaeal proviruses integrated into host genomes remain largely unexplored. To bridge this gap, we conducted a large-scale mining of genomes spanning all presently known 21 archaeal phyla for their proviruses. We identified 770 archaeal proviruses across 12 archaeal phyla and 84 families, which clustered into 655 viral operational taxonomic units (vOTUs). Among these, 86.1% of the vOTUs were novel at the species level, and 69.3% could not be classified at the family level, substantially expanding the known diversity of archaeal viruses. Additionally, phylogenomic analysis supported the proposal of 16 putative novel viral families, further extending the current taxonomy landscape of archaeal viruses. Notably, 21.8% of the identified proviruses were predicted to adopt a lytic lifestyle, suggesting that these proviruses may retain the capacity to enter the lytic cycle under appropriate conditions. Host prediction indicated only 14 out of the 655 vOTUs might have potential across-lineage infection abilities. We detected 63 anti-defense genes encoded by 61 provirus genomes, such as anti-CRISPR and anti-RM, suggesting an ongoing evolutionary arms race between hosts and proviruses. However, only 10 auxiliary metabolic genes (AMGs) were identified, suggesting a limited impact of proviruses in the modulation of host metabolism through AMGs. This study establishes a systematic global genomic atlas of archaeal proviruses, advancing our understanding of their distribution and diversity while providing a foundation for future research into how proviruses regulate archaeal metabolism and ecosystem functioning.

anti-defense system

Multi-Omics and Integrative Analytics in Natural Products Discovery.

Natural products (NPs) have long been an essential source of new bioactive compounds for drug discovery; however, traditional methods for screening and isolating these compounds can be slow and often yield diminishing returns. Fortunately, advanced multi-omics and computational approaches present powerful solutions to these challenges. This review highlights innovative methodologies that integrate metabolomics, genomics, transcriptomics, and proteomics with bioinformatics and analytical chemistry to accelerate NP discovery. For instance, untargeted metabolomics platforms like high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) and Global Natural Products Social (GNPS) molecular networking allow for comprehensive profiling of new compounds, while targeted isotope-labeling strategies enhance this process. Additionally, genome and metagenome mining tools such as antibiotics and secondary metabolite analysis shell (antiSMASH), Deep Biosynthetic Gene Cluster (DeepBGC), and Pipeline for Reconstructing Integrated Syntheses of Metabolites (PRISM) quickly identify biosynthetic gene clusters (BGCs) in both cultured and uncultured organisms, often using heterologous expression to validate products. Transcriptomic analyses, including RNA sequencing (RNA-seq), co-expression networks, and fluxomics, help clarify how pathways are regulated, while quantitative proteomics techniques like tandem mass tags/isobaric tags for relative and absolute quantitation (TMT/iTRAQ) and label-free methods, along with chemoproteomics approaches such as cellular thermal shift assay and thermal proteome profiling (TPP), uncover molecular targets and their mechanisms of action. This review also places significant emphasis on the role of artificial intelligence (AI) and machine learning (ML) in integrating multi-omics data, spanning activities from constructing gene-metabolite correlation networks to leveraging knowledge graphs and graph neural networks for data fusion and functional prediction. Finally, this review concludes by discussing the synergistic benefits of multi-omics for natural-product discovery, addressing current technical challenges, and exploring future directions toward high-throughput, intelligent data integration for next-generation NP research.

Biological Products

From sequence space to ecological function: microbiome-derived antimicrobial peptides as community effectors and therapeutic leads.

Antimicrobial peptide research has long centred on host defence molecules, yet microbiomes themselves encode a diverse and increasingly important repertoire of peptide-based antimicrobials. These microbiome-derived antimicrobial peptides include bacteriocins, ribosomally synthesised and post-translationally modified peptides, cryptic short open reading frame-encoded peptides, embedded antimicrobial regions within larger proteins, and selected peptide antibiotics recovered from human, animal, plant and environmental microbiomes. Recent advances in genome mining, metagenomics, and machine learning have greatly expanded the scale of discovery, moving the field from a handful of landmark exemplars to large candidate catalogues spanning the global microbiome. In the clearest cases, these molecules are not only anti-infective leads but ecological effectors: they mediate microbial competition, enforce colonisation resistance, and influence community structure within densely occupied niches. The present review synthesises the field across discovery classes, microbiome sources, ecological roles, and translational bottlenecks, emphasizing a central limitation of the field: candidate catalogues are expanding at extraordinary scale, while evidence for native expression, producer assignment, ecological function, and in vivo relevance remains limited for the vast majority of predicted molecules. Progress will depend on workflows that connect sequence level prediction to biological context through expression support, producer assignment, community level validation, and perturbation-based approaches that distinguish ecological association from causal function. Microbiome-derived antimicrobial peptides are best understood not only as promising therapeutic leads, but also as molecular mediators of microbial social life whose ecological origins are central to their interpretation and future application.

Microbiota

Genome mining based on transcriptional regulatory networks uncovers a novel locus involved in desferrioxamine biosynthesis.

Bacteria produce a plethora of natural products that are in clinical, agricultural and biotechnological use. Genome mining has uncovered millions of biosynthetic gene clusters (BGCs) that encode their biosynthesis, the vast majority of them lacking a clear product or function. Thus, a major challenge is to predict the bioactivities of the molecules these BGCs specify, and how to elicit their expression. Here, we present an innovative strategy whereby we harness the power of regulatory networks combined with global gene expression patterns to predict BGC functions. Bioinformatic analysis of all genes predicted to be controlled by the iron master regulator DmdR1 combined with co-expression data, led to identification of the novel operon desJGH that plays a key role in the biosynthesis of the iron overload drug desferrioxamine (DFO) B in Streptomyces coelicolor. Deletion of either desG or desH strongly reduces the biosynthesis of DFO B, while that of DFO E is enhanced. DesJGH most likely act by changing the balance between the DFO precursors. Our work shows the power of harnessing regulation-based genome mining to functionally prioritize BGCs, accelerating the discovery of novel bioactive molecules.

Deferoxamine

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis

The 2026 Bundibugyo Ebola Outbreak: A Warning for Global Preparedness for Future Epidemics.

Dear Editor, The 2026 Bundibugyo Ebolavirus (BDBV) outbreak has once again demonstrated that the threat of emerging diseases remains a major global health challenge. The outbreak, first detected in the Democratic Republic of Congo (DRC) and spread to Uganda, is not only a regional crisis but also a test of the world's preparedness for pathogens with epidemic potential. Unlike Zaire Ebolavirus (EBOV), which has benefited from effective vaccines and treatments in recent years, BDBV still lacks a licensed vaccine or specific treatment[1]. As of June 6, a total of 515 laboratory-confirmed cases and 91 deaths have been reported in DRC, while Uganda has reported 19 laboratory-confirmed cases and two deaths. The occurrence of unexplained deaths among both the community and healthcare workers, along with prior reports of an unidentified hemorrhagic fever, suggest that the outbreak has been likely originated in March 2026 or even earlier. Accordingly, the virus is believed to have spread unnoticed for several weeks before being identified through genomic sequencing in mid-May 2026[2]. The resurgence of Ebola in Africa results from a complex interaction of environmental, social, and political factors. Deforestation, the development of mining activities, the expansion of agriculture, and increased human contact with wildlife have elevated the likelihood of spillovers from wildlife reservoirs, particularly fruit bats, which are considered the most likely natural hosts of ebolaviruses. Moreover, weak disease surveillance systems and limited access to health services have delayed the identification of early cases. The similarity of the initial symptoms of Ebola to other endemic diseases in the region, such as malaria, makes early diagnosis difficult and provides ample opportunity for transmission to spread. Insecurity, misinformation, attacks on healthcare facilities, and armed conflict in the region have also posed serious challenges to the implementation of contact tracing programs and rapid response to the epidemic[3,4]. One of the most critical challenges highlighted by this outbreak is the weakness of diagnostic capacities in the affected areas. The initial 2007 outbreak of BDBV proved that delayed lab confirmation paralyzes public health responses[5]. Now, dealing with a much larger outbreak in 2026, the persistence of this challenge highlights a dangerous failure to invest in diagnostic infrastructure over the last 19 years. Many health facilities do not have access to molecular laboratories, rapid sample transport systems, and biosafety infrastructure[6]. These limitations delay the diagnosis and isolation of patients, thus perpetuating disease transmission. Investment in the development of mobile laboratories, rapid point-of-care diagnostic tests, and digital reporting systems can dramatically reduce the time to diagnosis and response to an outbreak. The BDBV outbreak shows that laboratory preparedness must be considered an essential part of global health security. Furthermore, the early detection of emerging pathogens depends not only on diagnostic technologies but also on the expertise of local scientists who are able to recognize unusual epidemiological and laboratory patterns. During the current outbreak, suspected Ebola cases initially tested negative using common diagnostic tests (designed for Zaire Ebola Virus), which delayed the identification of the BDBV. Specifically, field-based diagnostics in Bunia were calibrated exclusively to detect the EBOV responsible for recent Congolese outbreaks. Consequently, patient samples collected throughout late April and early May yielded negative results, requiring cross-country transport to Kinshasa for genomic confirmation[2]. This experience revealed a major vulnerability in outbreak preparedness: diagnostic tools designed for known threats may be ineffective in detecting less common or unexpected pathogens. Therefore, strengthening local scientific capacities, developing genomic surveillance, and expanding access to flexible and adaptable diagnostic platforms should be considered as a top priority for global health security. The lack of a licensed vaccine for BDBV was one of the most significant challenges of this epidemic. While the rVSV-ZEBOV vaccine has played a significant role in controlling Zaire ebolavirus, there is no licensed vaccine for BDBV. In response to this outbreak, efforts to develop mRNA-based vaccines, adenoviral vectors, rVSV-based vaccines, and multipotent vaccines have been accelerated[7]. However, the experience of this epidemic has shown that the development of medical products for rare diseases continues to face financial and investment constraints. This challenge highlights the need for sustained support from governments and international institutions for research and development of pathogens with epidemic potential. The 2026 Bundibugyo outbreak provides several key lessons for the global community. First, early detection and rapid diagnosis are the most important factors in containing the epidemic. The 19-year interval between the 2007 BDBV outbreak and the 2026 outbreak underscores persistent shortcomings in investment toward decentralized, pan-ebolavirus diagnostic infrastructure, with diagnostic delays hindering timely outbreak identification in both instances. Second, the trust and active participation of local communities are as important as medical interventions. Additionally, the rapid cross-border transmission dynamics between the DRC and Uganda demonstrate that blanket travel restrictions and border closures are impractical. As communities in the Great Lakes region routinely cross national borders for trade and healthcare, coordinated regional surveillance and timely information sharing are likely to be more effective than broad border closures in mitigating disease transmission[8]. Third, the protection of health workers must be a priority in preparedness plans. Fourth, a "One Health" approach is essential for simultaneous monitoring of humans, animals, and the environment. Although BDBV is not a new pathogen, the lack of licensed medical interventions and limited investment in research reflect many of the vulnerabilities associated with the concept of "Disease X."[9]. Unlike Zaire Ebola Virus, for which licensed vaccines and monoclonal antibody therapies are available, BDBV forces public health responses to rely almost entirely on non-pharmaceutical interventions such as isolation and infection control[10]. This gap reflects the structural inequity in global health research and development funding, with pathogens affecting resource-limited regions receiving insufficient attention until they spark an international emergency[2]. The BDBV outbreak proves that global epidemic preparedness cannot be pathogen-selective; it requires proactive investment in broad-spectrum countermeasures and resilient frontline health systems[8]. In conclusion, the 2026 BDBV outbreak is a serious wake-up call for the global health system. The epidemic revealed that gaps in surveillance systems, diagnostic capacities, vaccine development, and preparedness for emerging diseases persist. Investing in health infrastructure, developing Pan-Ebolavirus vaccines, strengthening laboratories, expanding the One-Health approach, and supporting research on emerging zoonotic pathogens must be at the top of global health security priorities. Otherwise, the BDBV outbreak may be just a prelude to larger crises to come.

Ebolavirus

Horizontal plasmid transfer promotes antibiotic resistance in selected bacteria in Chinese frog farms.

The emergence and dissemination of antibiotic resistance genes (ARGs) in the ecosystem are global public health concerns. One Health emphasizes the interconnectivity between different habitats and seeks to optimize animal, human, and environmental health. However, information on the dissemination of antibiotic resistance genes (ARGs) within complex microbiomes in natural habitats is scarce. We investigated the prevalence of antibiotic resistant bacteria (ARB) and the spread of ARGs in intensive bullfrog (Rana catesbeiana) farms in the Shantou area of China. Antibiotic susceptibilities of 361 strains, combined with microbiome analyses, revealed Escherichia coli, Edwardsiella tarda, Citrobacter and Klebsiella sp. as prevalent multidrug resistant bacteria on these farms. Whole genome sequencing of 95 ARB identified 250 large plasmids that harbored a wide range of ARGs. Plasmid sequences and sediment metagenomes revealed an abundance of tetA, sul1, and aph(3″)-Ib ARGs. Notably, antibiotic resistance (against 15 antibiotics) highly correlated with plasmid-borne rather than chromosome-borne ARGs. Based on sequence similarities, most plasmids (62%) fell into 32 distinct groups, indicating a potential for horizontal plasmid transfer (HPT) within the frog farm microbiome. HPT was confirmed in inter- and intra-species conjugation experiments. Furthermore, identical mobile ARGs, flanked by mobile genetic elements (MGEs), were found in different locations on the same plasmid, or on different plasmids residing in the same or different hosts. Our results suggest a synergy between MGEs and HPT to facilitate ARGs dissemination in frog farms. Mining public databases retrieved similar plasmids from different bacterial species found in other environmental niches globally. Our findings underscore the importance of HPT in mediating the spread of ARGs in frog farms and other microbiomes of the ecosystem.

Animals

Molecular mechanisms and breeding strategies for heat tolerance in vegetable crops under global warming.

Extreme heat driven by climate change poses a catastrophic threat to global vegetable production, undermining nutritional security because of the heightened physiological sensitivity and succulent tissues of these crops. This review synthesizes the multistage impacts of heat stress across critical developmental phases-from germination to reproduction-emphasizing morphological impairments (such as leaf wilting and floral abortion) and physiological disruptions (including photosynthetic inhibition and oxidative damage). We systematically dissect thermotolerance mechanisms in vegetables, highlighting transcriptional reprogramming by HSFs, WRKY, and NAC transcription factors; chaperone-mediated proteostasis via HSPs; epigenetic remodeling; Ca2+-ROS signaling pathways; and the role of phase separation dynamics. Importantly, we propose six strategic pathways to develop heat-resilient vegetables: harnessing natural variation through pan-genome-driven allele mining; employing biotechnological interventions such as CRISPR-mediated editing and synthetic promoters; engineering multistress tolerance by targeting conserved 'core response' pathways; exploiting epigenetic memory to achieve transgenerational resilience; optimizing source-sink dynamics with ''Climate-Responsive Carbon Optimization; and applying plant growth regulators and nanotechnology to enhance thermotolerance. Together, these strategies chart a clear roadmap for climate-smart vegetable breeding and call for interdisciplinary collaboration to translate molecular discoveries into practical breeding approaches for sustainable food systems under escalating thermal extremes.

Journal Article

Genomic prospecting and biochemical characterization of a novel thermostable 3-quinuclidinone reductase from hot spring metagenomes for efficient biocatalysis.

This study presents the discovery and characterization of a novel thermophilic 3-quinuclidinone reductase (ScQR) identified through metagenomic mining of hot spring environments. ScQR, a member of the short-chain dehydrogenase/reductase (SDR) superfamily, was heterologously expressed in Escherichia coli, and its catalytic properties were systematically characterized. The enzyme demonstrates exceptional thermal stability, retaining 86% of its activity after 48 hours at 70°C. Furthermore, K+ and Mg²+ ions significantly enhanced ScQR's activity at specific concentrations. Structural analysis revealed that ScQR adopts a typical SDR fold with a conserved catalytic triad (S141-Y155-K159), and it is NAD(H) dependent. Enzyme assays indicated that ScQR is highly stereoselective for (R)-3-quinuclidinol, with no activity against its enantiomer, (S)-3-quinuclidinol. The enzyme exhibits optimal activity at pH 9 and 85°C, making it a promising candidate for industrial applications requiring high thermal stability. Molecular dynamics simulations further revealed that ScQR preserves global structural integrity up to 360 K, whereas higher temperatures induce destabilization, predominantly in the C-terminal region and residues 95-100. In addition, structure-guided computational design enabled by LigandMPNN and UniKP yielded three ScQR variants with improved substrate affinity and catalytic efficiency while maintaining the overall fold and function. This work underscores the power of metagenomics with structure-driven protein design in discovering novel enzymes with unique catalytic properties from extreme environments and establishes ScQR as a promising biocatalyst for biotechnological and pharmaceutical applications.IMPORTANCEThis study reports the discovery of ScQR, a novel thermophilic 3-quinuclidinone reductase identified via metagenomic mining. ScQR represents one of the most heat-resistant members of the SDR superfamily discovered to date, maintaining 86% activity after 48 hours at 70°C. These findings establish ScQR as a robust biocatalyst for high-temperature pharmaceutical applications and demonstrate a scalable workflow for optimizing enzymes from extreme environments, offering significant value to the fields of biocatalysis and protein engineering.

computational design

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus

Genomic and biosynthetic landscape of high-temperature Daqu microbiome.

As the core starter for Chinese Baijiu, high-temperature Daqu is produced through open solid-state fermentation with recurrent inoculation by mature Daqu, forming a rich yet largely untapped reservoir of genomes and bioactive compounds. This study constructs the High-temperature Daqu Fermentation Microbiome catalog using 463 metagenomes spanning the full fermentation cycle. The catalog comprises 4,264 metagenome-assembled genomes that are dereplicated into 252 representative genome-based species, 82 % of which are absent from current global food microbiome databases. It further contains 14.3 million non-redundant genes, of which 17.3 % are novel, and 17,031 biosynthetic gene clusters, of which 62.63 % are novel, thereby substantially expanding the known genomic and biosynthetic space of food microbiomes. Genome-resolved analyses revealed a U-shaped ecological trajectory, shifting from early Bacillus velezensis-enriched assemblages to transient dominance of lactic acid bacteria during peak thermogenesis, before returning in late fermentation to thermotolerant, spore-forming Bacillota and Actinomycetota. In parallel, biosynthetic potential was further organized into four recurrent, stage-enriched profiles, from RiPP-rich thermogenic states to mature-state assemblages enriched in PKS-, NRPS-, and terpene-related capacities, with Bacillus, Kroppenstedtia, and Saccharopolyspora constituting the principal biosynthetic reservoir. Together, this work uncovers a largely unexplored genomic and biosynthetic reservoir in high-temperature Daqu fermentation, providing a target resource for mining thermotolerant industrial enzymes, flavor-related genes, and bioactive metabolites with biotechnological potential.

Microbiota

Logan: Planetary-Scale Genome Assembly Surveys Life's Diversity.

The breadth of life's diversity is unfathomable, but public nucleic acid sequencing data offers a window into the dispersion and evolution of genetic diversity across Earth. However the rapid growth and accumulation of sequence data have outpaced efficient analysis capabilities. The largest collection of freely available sequencing data is the Sequence Read Archive (SRA), comprising 27.3 million datasets or 5 × 1016 basepairs. To realize the potential of the SRA, we constructed Logan, a massive sequence assembly transforming short reads into long contigs and compressing the data over 100-fold, enabling highly efficient petabase-scale analysis. We created Logan-Search, a k-mer index of Logan for free planetary-scale sequence search, returning matches in minutes. We used Logan contigs to identify >200 million plastic-degrading enzyme homologs, and validate novel enzymes with catalytic activities exceeding current reference standards. Further, we vastly expand the known diversity of proteins (30-fold over UniRef50), plasmids (22-fold over PLSDB), P4 satellites (4.5-fold), and the recently described Obelisk RNA elements (3.7-fold). Logan also enables ecological and biomedical data mining, such as global tracking of antimicrobial resistance genes and the characterization of viral reactivation across millions of human BioSamples. By transforming the SRA, Logan democratizes access to the world's public genetic data and opens frontiers in biotechnology, molecular ecology, and global health.

Journal Article

Uncovering encrypted antimicrobial peptides in health-associated Lactobacillaceae by large-scale genomics and machine learning.

BACKGROUND: Antimicrobial peptides (AMPs) are well known for their broad-spectrum activity and have shown great promise in addressing the antibiotic-resistant crisis. The Lactobacillaceae family, recognized for its health-promoting effects in humans, represents a valuable source of novel AMPs. However, the global prevalence and distribution of AMPs within Lactobacillaceae remains largely unknown, which limits the efficient discovery and development of novel AMPs. RESULTS: We analyzed all available genomes (10,327 genomes), encompassing 38 genera and 515 species, to investigate the biosynthetic potential (indicated by the number of AMP sequences in the genome) of AMP in the Lactobacillaceae family. We demonstrated Lactobacillaceae species had ubiquitous (69.90%) biosynthetic potential of AMPs. Overall, 9601 AMPs were identified, clustering into 2092 gene cluster families (GCFs), which showed strong interspecies specificity (95.27%), intraspecies heterogeneity (93.31%), and habitat uniqueness (95.83%), that greatly expanded on the AMP sequence landscape. Novelty assessment indicated that 1516 GCFs (72.47%) had no similarity to any known AMPs in existing databases. Machine learning predictions suggested that novel AMPs from Lactobacillaceae possessed strong antimicrobial potential, with 664 GCFs having an additive minimum inhibitory concentration (MIC) below 100&#xa0;&#x3bc;M. We randomly synthesized 16 AMPs (with predicted MIC&#x2009;<&#x2009;100&#xa0;&#x3bc;M) and identified 10 AMPs exhibiting varied-spectrum activity against 11 common pathogens. Finally, we identified one Lactobacillus delbrueckii-originated AMP (delbruin_1) having broad-spectrum (all 11 pathogens) and high antimicrobial activity (average MIC&#x2009;=&#x2009;38.56 &#xb5;M), which proved its potential as a clinically viable antimicrobial agent. CONCLUSIONS: We uncovered the global prevalence of AMPs in Lactobacillaceae and proved that Lactobacillaceae is an untapped and invaluable source of novel AMPs to combat the antibiotic-resistance crisis. Meanwhile, we provided a machine learning-guided framework for AMP discovery, offering a scalable roadmap for identifying novel AMPs not only in Lactobacillaceae but also in other organisms. Video Abstract.

Machine Learning

Genetic Analysis of Genomic and Methylomic Variation and Identification of Multi-Trait Mutants in Rice Carried on Chang'e-5.

Global food security is facing challenges from population growth to diminishing arable land. Space mutation breeding holds promise for overcoming the variation limitations in conventional breeding; however, the mutagenic effects of the deep-space environment on rice and the transgenerational inheritance patterns of induced variations remain unclear. In this study, rice seeds carried by the Chang'e-5 spacecraft were used as materials. Whole-genome sequencing and whole-genome bisulfite sequencing were performed on the first (SP1) and second generations (SP2) of space-mutagenized plants after their return to Earth. The results showed that the number of genomic variants in the SP2 generation increased significantly compared with SP1, and SNPs, homozygous sites, and variants in coding regions were more heritable. The genome-wide methylation level was elevated in the SP2 generation, and among differentially methylated cytosines, those in the CG context exhibited the highest heritability. Furthermore, large-scale screening for nitrogen efficiency, tolerance to PEG-induced stress, and germination-stage cold resistant mutants was conducted in the SP2 generation, and phenotypic validation was performed in the third generation (SP3). By integrating multi-omics analyses of representative mutants to mine candidate genes, a number of heritable elite mutants were obtained, and seven candidate genes for key traits were identified. This study systematically elucidates the transgenerational inheritance patterns of deep-space-induced variation in rice. The multi-trait mutants obtained provide valuable germplasm resources for gene cloning and breeding applications in rice.

DNA methylation

From spillover to systems: evidence gaps in One Health preparedness for emerging infectious diseases in Latin America and the Caribbean.

Latin America and the Caribbean are a global hotspot for emerging and re-emerging infectious diseases, yet regional One Health preparedness remains uneven and incompletely operationalized. This narrative Mini Review synthesizes evidence published mainly between 2015 and 2026 on One Health preparedness for emerging infectious diseases in the region, emphasizing how environmental disruption and climate change shape zoonotic and vector-borne spillover risk. Available regional surveys suggest broad professional familiarity with the One Health concept but limited operational implementation, with environmental health frequently identified as the least-integrated domain. We argue that spillover risk-and the failure to detect and contain spillover once it occurs-should be understood as a system-level outcome shaped by ecological disruption, socioeconomic vulnerability, surveillance capacity, and governance, rather than as an isolated biological event: deforestation, agricultural and extractive expansion-including illegal mining and logging-unplanned urbanization, and climate variability generate new human-animal-vector interfaces, while fragmented governance, uneven and poorly decentralized laboratory capacity, and limited reservoir and environmental surveillance leave these interfaces unmonitored. Environmental and climatic drivers are robustly linked to spillover, although the pathways are disease-specific rather than universal, and socioeconomic vulnerability concentrates the resulting burden in Indigenous, rural, and marginalized populations. We identify priority gaps in integrated surveillance, decentralized diagnostics, genomic capacity, reservoir ecology, governance, financing, and equity, and propose an agenda for anticipatory, climate-informed, and context-sensitive preparedness.

Latin America

Molecular biology and integrated strategies for activating cryptic biosynthetic gene clusters toward next-generation antibiotic discovery.

Antimicrobial resistance (AMR) has been identified as one of the 21st century's severest global public health crises. AMR led to an estimated 4.95 million deaths in 2019 and will claim 10 million lives a year by 2050 in the absence of targeted interventions. During the same period, the number of novel antibiotics discovered has decreased drastically as many researchers are rediscovering known antibiotics, non-model microorganisms are poorly understood or difficult to culture and antibiotic research and development investment has declined drastically. However, high-throughput whole genome sequencing and the subsequent application of bioinformatics in bacterial and fungal genomes have shown that a numerous of cryptic or silent biosynthetic gene clusters (BGCs) remain latent at ambient laboratory conditions since their genes are transcriptionally inactive. Cryptic BGCs represent a vast source of unique secondary metabolites, many of which may yield novel antibacterial, antifungal, anti-cancer and other potentially valuable natural products. This review discusses the biological relevance of cryptic BGCs, the major limiting factors that restricts their activation and novel strategies that have been employed to activate them and exploit their potential to produce novel natural products. The review focuses on biological approaches including CRISPR-Cas mediation for the activation of cryptic BGCs, promoter engineering, pathway refactoring, and heterologous expression; biochemical strategies such as Osman, OsMAC, Precursor Feeding, Chemical Elicitation, Epigenetic Regulation and Co-cultivation and technology-based strategies such as Genome mining, Microfluidic Cultivation systems, High-Throughput Screening, Metabolomics, Molecular Networking and Artificial Intelligence and Machine Learning based prediction of BGCs and their metabolites. The use of multi-omics technologies combined with synthetic biology to achieve better discovery, characterization and large-scale production of novel natural products is also discussed herein. Finally, we will talk about the ecological significance and evolutionary advantage of cryptic BGCs' role in interactions between microorganisms, such as competition, communication, symbiosis and environmental adaptability, so as to provide a useful background for accelerating next-generation antibiotics.

CRISPR-Cas activation