Search PubMedSearch

SEARCH · Search PubMed

Results for “pan-genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

PlantPan: A comprehensive multi-species plant pan-genome database.

The pan-genome represents the complete genomic diversity of specific species, serving as a valuable resource for studying species evolution, crop domestication, and guiding crop breeding and improvement. While there are several single-species-specific plant pan-genome databases, the availability of multi-species pan-genome databases is limited. Additionally, variations in methods and data types used for plant pan-genome analysis across different databases hinder the comparison and integration of pan-genome information from various projects at multi-species or single-species levels. To tackle this challenge, we introduce PlantPan, a comprehensive database housing the results of pan-genome analysis for 195 genomes from 11 plant species. PlantPan aims to provide extensive information, including gene-centric and sequence-centric pan-genome information, graph-based pan-genome, pan-genome openness profiles, gene functions and its variation characteristics, homologous genes, and gene clusters across different species. Statistically, PlantPan incorporates 9 163 011 genes, 694 191 gene clusters, 526 973 370 genome variations, and 1 616 089 non-redundant genome variation groups at the species level, 33 455,098 genome synteny, and 177 827 non-redundant genome synteny groups at the species level. Regarding functional genes, PlantPan contains 5 222 720 genes related to transcription factors, 395 247 literature-reported resistance genes, 455 748 predicted microbial/disease resistance genes, and 1 612 112 genes related to molecular pathways. In summary, PlantPan is a vital platform for advancing the application of pan-genomes in molecular breeding for crops and evolutionary research for plants.

Genome, Plant

Pitfalls of bacterial pan-genome analysis approaches: a case study of Mycobacterium tuberculosis and two less clonal bacterial species.

SUMMARY: Pan-genome analysis is a fundamental tool for studying bacterial genome evolution; however, the variety in methods used to define and measure the pan-genome poses challenges to the interpretation and reliability of results. Using Mycobacterium tuberculosis, a clonally evolving bacterium with a small accessory genome, as a model system, we systematically evaluated sources of variability in pan-genome estimates. Our analysis revealed that differences in assembly type (short-read versus hybrid), annotation pipeline, and pan-genome software, significantly impact predictions of core and accessory genome size. Extending our analysis to two additional bacterial species, Escherichia coli and Staphylococcus aureus, we observed consistent tool-dependent biases but species-specific patterns in pan-genome variability. Our findings highlight the importance of integrating nucleotide- and protein-level analyses to improve the reliability and reproducibility of pan-genome studies across diverse bacterial populations. AVAILABILITY AND IMPLEMENTATION: Panqc is freely available under an MIT license at https://github.com/maxgmarin/panqc.

Genome, Bacterial

Comprehensive genomic and computational insights into Brucella suis: pan-genome analysis, evolutionary perspectives, and in-silico vaccine design.

BACKGROUND: Brucella suis is a zoonotic intracellular pathogen responsible for brucellosis, mainly in swine and humans. Although numerous genome sequences are publicly available, an integrative genomic analysis combining pan-genome architecture, structural organization, evolutionary relationships, and vaccine-associated targets remains limited. RESULTS: In this study, we analyzed 91 publicly available B.suis genomes to characterize their pan-genome composition and genomic structure. The pan-genome exhibited an open configuration, indicating continued genomic diversification. A total of 2,146 core genes were identified, representing conserved functions essential for species maintenance, while the accessory genome reflected strain-level variability. Phylogenetic reconstruction based on single-copy orthologs revealed distinct evolutionary clades among the strains. A complementary phylogenetic analysis of pan-genome gene presence-absence patterns further supported clade differentiation and highlighted variation in accessory gene repertoires. Comparative synteny and genome structural analyses demonstrated largely conserved chromosomal organization with localized rearrangements across strains. Screening of the core proteome identified 64 putative antigenic proteins with predicted surface localization and immunogenic properties. Additionally, resistance-associated determinants related to tetracycline and doxycycline were detected in one genome within the dataset. CONCLUSIONS: This comprehensive genomic analysis defines the pan-genome structure, evolutionary relationships, and genome organization of B.suis. The integration of core and pan-genome-based phylogenies provides complementary insights into strain diversification, while the identified conserved antigenic candidates offer a foundation for future experimental validation and rational vaccine development strategies.

Genome, Bacterial

Pan-genome characterization of the maize 4CL gene family and its dynamic responses to abiotic stress.

1.Pan-genome analysis across 26 maize inbred lines identified 13 Zm4CL genes (nine core and four near-core) classified into three evolutionary clades.2.Structural variations (SVs) are significantly associated with the expression and altered conserved protein domains of key Zm4CL genes.3.Zm4CL genes exhibit distinct tissue-specific expression patterns and dynamic enzymatic and transcriptional responses to stresses, particularly cold and drought.4-Coumarate:CoA ligase (4CL) is a key enzyme in the phenylpropanoid pathway and plays important roles in plant growth, development, and responses to environmental stresses. However, a comprehensive pan-genome analysis of the 4CL gene family in maize is still lacking. In this study, 13 Zm4CL genes were identified from a maize pan-genome comprising 26 diverse inbred lines, including nine core genes and four near-core genes. Phylogenetic analysis classified these genes into three evolutionary clades, while Ka/Ks analysis indicated that most members have been maintained under purifying selection, although several genes exhibited greater evolutionary divergence and relatively relaxed evolutionary constraints. Structural variation (SV) analysis revealed significant associations between SVs and the expression of Zm4CL2 and Zm4CL3, while sequence comparisons suggested that SVs were also associated with alterations in conserved protein domains in some genotypes. Transcriptome analyses revealed distinct tissue-specific expression patterns and diverse transcriptional responses to abiotic and biotic stresses. Enzyme activity assays showed that cold stress significantly increased 4CL activity at 12 h, whereas heat, salt, and alkali stresses caused an initial decrease followed by recovery, while drought had no significant effect. Time-course RT-qPCR further validated dynamic expression changes of representative Zm4CL genes under cold and drought stresses. Overall, this study provides a comprehensive pan-genome framework for understanding the evolutionary conservation, regulatory diversification, and stress-responsive characteristics of the maize Zm4CL gene family, providing valuable resources for future functional studies and the genetic improvement of stress tolerance in maize.

Zea mays

Primulina pan-genome reveals differential gene retention following whole-genome duplications and provides insights into edaphic specialization.

Primulina, a genus of >200 species specialized to extreme soils, provides a model for edaphic adaptation. We assemble seven genomes and construct a pan-genome spanning nine species from karst, Danxia, and acidic soils. Comparative analyses reveal that karst-adapted species have smaller genomes. Two lineage-specific whole-genome duplications (WGDs) exhibit biased duplicate loss in large gene families but preferential retention of transcription factors, indicating combined adaptive and nonadaptive forces. Pan-genome analyses identify ion channel and transporter genes enriched in variant hotspots and under positive selection in karst lineages. Candidate genes for drought and salt stress tolerance include ABC transporters and ion channels. Notably, an ABC transporter shows positive selection in karst species and unique structural variation in non-karst species. Together, our findings show that genome downsizing, biased post-WGD retention, and evolution of ion-transport pathways shape adaptation to extreme soils. The Primulina pan-genome provides a resource for dissecting mechanisms underlying edaphic specialization.

Gene Duplication

Effectidor II: a pan-genomic AI-based algorithm for the prediction of type III secretion system effectors.

MOTIVATION: Type III secretion systems are used by many Gram-negative bacteria to inject type 3 effectors (T3Es) directly into eukaryotic cells, promoting disease or provoking immune response. Because of these opposing evolutionary forces, T3E repertoires often vary within taxonomic groups. Identifying the full effector gene repertoire in genomes of related individuals is crucial for determining core and specialized effectors, understanding the disease dynamics, and developing appropriate management strategies against pathogens. It can also help uncover novel T3Es that have recently emerged in a population. Our previously published Effectidor web server successfully addressed the challenge of identifying T3Es in a single bacterial genome. Here, we enriched the web server with various novel capabilities, including the identification of T3Es from multiple genome sequences simultaneously. RESULTS: We present Effectidor II, a web server that relies on machine learning to predict T3E-encoding genes within bacterial pan-genomes. We demonstrate the benefit of learning based on features extracted from the entire sequences comprising the pan-genome and report a novel T3E discovered by it in Xanthomonas euroxanthea. AVAILABILITY AND IMPLEMENTATION: Effectidor II is available at: https://effectidor.tau.ac.il and the source code is available at: https://github.com/naamawagner/Effectidor. A stand-alone version of Effectidor II is available at: https://github.com/naamawagner/Effectidor/tree/StandAlone. The source code for the standalone version and the data used in this work are also provided in https://doi.org/10.5281/zenodo.15081636.

Type III Secretion Systems

Phenotypic and phylogenomic characterization of Lactococcus garvieae isolates from rainbow trout (Oncorhynchus mykiss) in Türkiye.

Lactococcosis is an important bacterial disease of farmed fish and causes substantial economic losses in rainbow trout (Oncorhynchus mykiss) aquaculture. In this study, Lactococcus garvieae isolates recovered from rainbow trout farms in Türkiye were characterized using phenotypic, molecular, and phylogenomic methods. Among 32 presumptive Lactococcus isolates recovered from 127 dead rainbow trout, four were confirmed as L. garvieae and exhibited identical biochemical characteristics, Pulsed Field Gel Electrophoresis (PFGE) profiles, and broad growth tolerance across different pH, salinity, and temperature conditions. All isolates were presumptively classified as resistant to ciprofloxacin and florfenicol, while remaining susceptible to tetracycline and penicillin. Based on the AMR profiles, strain LG2, which exhibited the most susceptible antimicrobial profile among the isolates, was selected for whole-genome sequencing (WGS). WGS of the representative isolate LG2 generated a single 2,214,687-bp chromosomal contig with 38.5% GC content and 99.0% BUSCO completeness. In silico PCR assigned LG2 to serotype I, and the genome contained an intact capsule-associated cps/kps locus. The chromosomal lsa(D) determinant and an mdt(A)-like efflux-associated gene were detected, whereas no plasmid replicons or acquired quinolone or florfenicol resistance genes were identified, indicating discordance between the phenotypic and genomic AMR results. Taxonomic verification of 236 publicly available Lactococcus assemblies yielded 41 verified public L. garvieae genomes, which, together with LG2, formed a 42-genome within-species dataset. LG2 was most closely related to the Turkish isolate OS-37, sharing 99.96% ANI and differing by three core SNPs; both belonged to ST109, whereas the other Turkish isolates belonged to ST139. cgMLST identified a conserved genomic backbone, while pan-genome analysis identified 5,655 gene clusters and an open pan-genome characterized by a large cloud-gene fraction. These findings demonstrate the importance of species verification in Lactococcus population genomics and reveal substantial accessory-genome diversity within L. garvieae. The genomic features of LG2 provide a basis for future pathogenicity and immunogenicity studies, although experimental validation is required. Overall, these findings highlight the importance of local genomic surveillance for understanding L. garvieae population structure and provide a genomic framework for future region-specific vaccine research.

Animals

Polyphasic taxonomic characterization of Brachybacterium netajii sp. nov., a metabolically versatile bacterium isolated from the river Ganges, India.

A comprehensive polyphasic taxonomic strategy was applied to the systematic characterization of strain DNPG3T, which was isolated from the river Ganges, Hooghly, West Bengal, India. The Gram-positive, halotolerant, heavy-metal-tolerant strain exhibited the ability to degrade p-nitrophenol (PNP). Cellular fatty acid analysis revealed that the predominant components were anteiso-C15:0 (24.61%), C11:0 (21.06%), iso-C16:0 (11.89%), C16:0 (11.58%), and anteiso-C17:0 (11.24%). Notably, the presence of C11:0, C10:0 2-OH as major fatty acids differentiate strain DNPG3T from its closely related members of the genus Brachybacterium. The predominant respiratory quinone was identified as menaquinone-7 (MK-7). Analysis of 16S rRNA gene sequence indicated that B. zhongshanense strain JBT was the closest relative of DNPG3T, sharing 97.08% sequence similarity. Genome-based ANI value calculated using the EzBioCloud server revealed that B. zhongshanense JCM 15471T was the closest genomic relative (85.49%). These values were further substantiated by digital DNA-DNA hybridization (dDDH) estimates calculated using the GGDC server. Taxonomic assignment using the GTDB database further indicated that strain DNPG3T constitutes a previously unrecognized species within the genus Brachybacterium. Genome analysis of strain DNPG3T identified eleven genomic islands, along with a rich repertoire of 194 carbohydrate-active enzyme (CAZyme) families, comprising 95 glycoside hydrolases and 53 glycosyltransferases. In addition, five biosynthetic gene clusters were detected. Collectively, these genomic features indicate the involvement of horizontal gene transfer events and highlighted the pronounced metabolic versatility of the strain, underscoring its potential for industrial enzyme production and secondary metabolite biosynthesis. Pan-genome analysis further indicates that the Brachybacterium pan-genome is open, reflecting substantial genetic diversity and ongoing gene acquisition within the genus. Comprehensive biochemical, physiological, chemotaxonomic, and phylogenetic analyses supported the assignment of strain DNPG3T to the genus Brachybacterium while clearly distinguishing it from all currently described species within the genus. Accordingly, strain DNPG3T was proposed to represent a novel species, for which the name Brachybacterium netajii sp. nov. is suggested. The type strain was DNPG3T (= MTCC13125T).

India

Columba: fast approximate pattern matching with optimized search schemes.

MOTIVATION: Aligning sequencing reads to reference genomes is a fundamental task in bioinformatics. Aligners can be classified as lossy or lossless: lossy aligners prioritize speed by reporting only one or a few high-scoring alignments, whereas lossless aligners output all optimal alignments, ensuring completeness and sensitivity. RESULTS: This paper introduces Columba, a high-performance lossless aligner tailored for Illumina sequencing data. Columba processes single or paired-end reads in FASTQ format and outputs alignments in SAM format. By utilizing advanced search schemes and bit-parallel alignment techniques, Columba achieves exceptional speed. Columba is available in two variants. The first, based on the bidirectional FM-index, prioritizes speed. The second, Columba RLC, uses run-length compression using a bidirectional move structure, significantly reducing memory usage for large, repetitive datasets like pan-genomes. Benchmarks on the human genome, as well as bacterial and human pan-genome datasets, demonstrate that Columba is much faster than existing lossless aligners and even competitive with lossy tools. We integrated Columba into the OptiType HLA genotyping pipeline, where it substantially reduced computational time while maintaining accuracy. These results position Columba as a versatile, state-of-the-art tool for high-sensitivity genomic analyses. AVAILABILITY AND IMPLEMENTATION: The source code of Columba is available at https://github.com/biointec/columba under AGPL license. Scripts to reproduce the benchmarks and analyses are available at https://doi.org/10.5281/zenodo.15849246.

Software

Genome-wide cyclin gene evolution in Arabidopsis and Brassica reveals polyploidization-driven duplication and flowering-time associations.

Cyclin genes are plant cell cycle regulators that play essential roles in growth, development, and reproduction. However, the evolutionary dynamics and genomic organization of cyclin genes across the Brassicaceae family remain poorly understood, particularly in the context of allotetraploid genome evolution. Here, we investigated the diversity, expansion mechanisms, and potential functional diversification of cyclin genes across ten Brassicaceae genomes, including four Arabidopsis and six Brassica species. A total of 1087 cyclin genes representing 23 cyclin types were identified. Comparative genomic analyses revealed that cyclin gene expansion was strongly influenced by polyploidization in Brassica species, with 1845 duplication events involving 1063 genes. Whole-genome duplication was the predominant mechanism driving expansion, while both inter- and intra-genomic duplications contributed to gene retention in tetraploid Brassica species, with the highest duplication frequency observed in Brassica juncea. Across genomes, 120 physical gene clusters were identified, including homogeneous and heterogeneous types. Ortholog analysis between progenitor and allotetraploid species identified 852 orthologous pairs involving 366 genes, indicating extensive conservation following allotetraploid formation. Phylogenetic analysis resolved cyclins into three major clades, while expression-based clustering in Brassica napus grouped genes into four major clusters, suggesting functional diversification. Integration of pan-genomic and flowering-time QTL analyses further identified two cyclin genes, Bna21cycA2 and Bna113cycD4, which contain amino acid polymorphisms and represent putative candidate variations potentially associated with flowering-time variation across multiple genomes. These findings provide new insights into the evolutionary expansion, retention, and potential functional divergence of cyclin genes in Brassicaceae and highlight candidate loci for future functional studies and crop improvement.

Evolution, Molecular

Multi-omics revealed the effects of rumen to blood path on early lactation performance in transition dairy cows.

BACKGROUND: The transition period is vitally important to the life cycle of dairy cows. However, the function of the microbiota during both pre- and post-partum and their relationship with ruminal, plasma, and milk metabolites still require systematic investigation. To address this, the 7 highest- and 7 lowest-performing animals among a cohort of 100 dairy cows were selected based on their postpartum energy-corrected milk yield. Rumen fluid and plasma samples were collected during both pre- and post-partum periods, whereas milk samples were obtained postpartum. Shotgun metagenomics of rumen contents in addition to metabolomics of rumen, plasma, and milk samples were performed to evaluate the associations between ruminal microbes and early lactation performance in transition dairy cows. RESULTS: Compared with prepartum cows, postpartum high-yield cows had greater concentrations of ruminal volatile fatty acids and plasma total bile acid. Moreover, plasma urea nitrogen and most amino acids, peptides, and their derivatives in plasma and milk were increased in postpartum high-yield cows, relative to postpartum low-yield cows. Metagenomic analysis revealed that the relative abundances of several species within the Prevotella, Succinimonas, Succinatimonas, and Methanosphaera increased, while other bacteria belong to Alistipes and Bacteroides, and archaeal Methanobrevibacter species decreased in postpartum cows, particularly in postpartum high-yield cows. Co-occurrence network and correlation analysis suggested that Prevotella and Succinatimonas were negatively correlated to Alistipes, Bacteroides, and Methanobrevibacter, potentially contributing to the nutritionally efficient phenotype of postpartum high-yield cows. A metabolic pathway analysis of our metagenomic data revealed that postpartum high-yield cows possessed more microbial genes involved in starch utilization and amino acid synthesis, while a wide range of microbial genes involved in cellulose utilization, acetogenesis, and amino acid degradation were found in prepartum cows with low-yield in postpartum. A structural equation model analysis showed that the increased relative abundances of Prevotella tf.2-5 and Succinatimonas CAG_777 were related to greater concentrations of plasma chenodeoxycholic acid glycine conjugate, milk 5-Methoxytryptophan, and energy-corrected milk yield. Finally, pan-genomic analysis confirmed that Alistipes, Bacteroides, and Methanobrevibacter possess genetic conservation of both hydrogenases and dehydrogenases, which may contribute to energy loss in the rumen via hydrogen dissipation. CONCLUSION: In summary, our findings provide a fundamental understanding of how microbiome-dependent mechanisms contribute to early lactation performance in dairy cows during the transition period. The increased abundance of Prevotella, Succinimonas, and Succinatimonas in postpartum cows suggest that they are important microbes during the transition period and may help in coping with metabolic challenges, while improving nutrient utilization efficiency during this period. Our study underscores the importance of the ruminal microbiome during the transition period and highlights the need for rumen-based nutritional intervention strategies to improve production efficiency in ruminants. Video Abstract.

Animals

A complete hlyCABD-like RTX operon marks a virulence-associated subset of trh-positive Vibrio parahaemolyticus from Hangzhou Bay, China.

Vibrio parahaemolyticus remains a major cause of seafood-associated gastroenteritis, yet routine surveillance still relies largely on the canonical hemolysin markers thermostable direct hemolysin (tdh) and tdh-related hemolysin (trh). To determine whether this framework overlooks accessory virulence determinants in trh-positive lineages, we analyzed 193 V. parahaemolyticus isolates collected between 2022 and 2025 from clinical, environmental, and seafood-associated sources in the Hangzhou Bay region of China. Serotyping identified 45 serotypes, with O10:K4 predominating among clinical isolates. Both clinical and non-clinical populations showed open pan-genomes, although the non-clinical group carried a larger accessory gene pool. We identified a complete hlyCABD-like RTX operon in 10 trh-positive isolates with T3SS2-associated virulence backgrounds. These RTX-positive isolates were distributed across seven sequence types and three of five phylogenetic groups. This distribution was lineage-restricted but non-clonal. In the representative hybrid-assembled genome, the operon occurred within a mosaic genomic region containing additional virulence- and mobility-associated genes, indicating a composite pathogenicity island-like element. In the tested subset, RTX-positive isolates showed significantly greater hemolytic activity than RTX-negative trh-positive isolates. This significant difference was consistently observed in both plate-based and liquid assays, and within the RTX-positive subset, hlyA expression correlated with hemolytic activity, whereas the trh gene and the tlh (thermolabile hemolysin) gene did not. A complete hlyCABD-like RTX operon therefore identifies a hemolysis-associated subset of trh-positive V. parahaemolyticus and supports its further evaluation as an additional target for food safety surveillance.

Vibrio parahaemolyticus

Evolution and Expression Divergence of Legume PAL Genes Suggest Associations with Drought Response and Root Nodule Development.

Comparative genomic analyses provide insight into the mechanisms underlying gene-family evolution and crop adaptation. Here, we used the legume phenylalanine ammonia-lyase (PAL) gene family as a model and integrated pan-genomic, phylogenetic, molecular evolutionary, duplication-mode, and transcriptomic analyses, while developing GFtool for gene family identification. Across 45 genomes, we identified 302 PAL genes and classified them into five Groups. Groups 1-3 represented ancient lineages shared with outgroups, whereas Groups 4 and 5 were legume-specific. Molecular-clock analyses placed the divergence of Group 2 near the Paleocene-Eocene transition, while Groups 4 and 5 diversified from the middle Eocene to the early Oligocene. WGD/segmental duplication broadly contributed to PAL copy-number expansion, whereas tandem duplication was enriched in Group 5 of Papilionoideae. Group 2 genes showed drought-induced expression, whereas Group 5 genes were associated with early root nodule development. GFtool provides a scalable framework for gene-family studies.

Fabaceae

Seqwin: ultrafast identification of signature sequences in microbial genomes.

MOTIVATION: Polymerase chain reaction (PCR) enables rapid, cost-effective diagnostics but requires prior identification of genomic regions that allow sensitive and specific detection of target microbial groups, herein referred to as microbial signature sequences. We introduce Seqwin, an open-source framework designed to automate microbial genome signature discovery. Tens of thousands of microbial genomes are now available for a single species, limiting the application of existing manual and automated approaches for identifying signatures. Modern approaches that are capable of leveraging all available microbial genomes will ensure sensitive and accurate DNA signature identification and enable robust pathogen detection for clinical, environmental, and public health applications. RESULTS: Seqwin builds weighted pan-genome minimizer graphs and uses a traversal algorithm to identify signature sequences that occur frequently in target genomes but remain rare in non-targets. Unlike earlier tools that depend on strict presence or absence of sequences, Seqwin accommodates natural sequence variation and scales to very large genome collections. When applied to genomes from C. difficile, M. tuberculosis, and S. enterica, Seqwin recovered more high-quality signatures than alternative methods with lower computational burden. Seqwin's analysis of nearly 15 000 S. enterica genomes yielded over 200 candidate signatures in three minutes. Seqwin provides an open-source solution for the long-standing need for scalable microbial signature discovery and diagnostic assay design. AVAILABILITY AND IMPLEMENTATION: Seqwin is available on GitHub (https://github.com/treangenlab/Seqwin) and can be installed via Bioconda (https://bioconda.github.io/recipes/seqwin/README.html). Benchmarking datasets, outputs, and scripts are available on Zenodo (https://doi.org/10.5281/zenodo.19874011).

Software

Leaf Rust in Rye: From Pathogen Biology to Host Defense and Resistance Breeding.

Leaf rust (LR), caused by Puccinia recondita f. sp. secalis (Prs), is considered one of the most dangerous rye (Secale cereale L.) diseases, causing yield losses exceeding 35%. This review summarizes all currently available data about this disease: pathogen characteristics (including its life cycle, natural variation, and disease symptoms), resistance resources, and the background of the plant immune response at the genome, transcriptome, and metabolome levels. The research conducted so far has allowed for the identification of dozens of genes that play a significant role in the rye immune response to Prs infection. Among them, genes encoding NBS-LRR proteins (including SECCE1Rv1G0014220, the most likely Pr3 candidate), glycosyltransferase, β-1,3-glucanase, 1-deoxy-D-xylulose 5-phosphate synthase, β-1,3-glucanase, UDP-glycosyltransferase, pathogenesis-related protein 1, ammonium transporter, and cytochrome P450 enzymes are candidates for seedling and all-stage resistance, whereas ScLr_ABC25 currently represents the most promising candidate associated with adult-plant resistance. Among the metabolites differentially accumulated in response to Prs, those related to phenylpropanoids, diterpenoids, and thiamine branches seem to play the most important role in the immune response. Finally, we suggest how the knowledge acquired so far about the rye-Prs interaction can be used in modern breeding programs aimed at obtaining cultivars with enhanced resistance to LR, such as through the use of functional gene markers and/or metabolic biomarker-assisted selection and, in the more distant future, by developing and applying new genomic techniques for precise editing of resistance and susceptibility genes, engineering synthetic immune receptors and decoys, and pan-genomic exploration for identification of rare or lineage-specific resistance alleles. [Formula: see text] Copyright © 2026 The Author(s). This is an open access article distributed under the CC BY-NC-ND 4.0 International license.

Plant Diseases

Pan-analysis of intra- and inter-species diversity reveals a group of highly variable immune receptor genes in rice.

Plant immune receptors and their natural variations play a central role in combating disease-causing pathogens. These immune receptors include intracellular nucleotide-binding leucine-rich repeat (LRR) receptors (NLRs) and cell-surface pattern recognition receptors (PRRs) that can be further classified as receptor-like proteins (RLPs) and receptor-like kinases (RLKs). Although the NLRome has been characterized, the repertoire and extent of diversity of PRRome remain undetermined in rice. In this study, we examined the diversity of immune receptor genes using high-quality genomes of 309 rice accessions from 8 species within the genus Oryza. A total of 376 310 immune receptor genes were identified, including 149 592 NLR-coding genes and 226 718 PRR coding genes. Shannon entropy analysis revealed a set of immune receptors that display significant intra-species and inter-species diversity in rice. In general, RLPs are more variable than RLKs, while NLRs and LRR-RLPs are more variable than LRR-RLKs. Additionally, NLR and PRR genes exhibit contrasting shoot/root expression patterns, with NLRs generally skewed towards root expression. Furthermore, we found that the size of the LRR-RLK gene families correlates with local annual precipitation, suggesting a stronger selection pressure on LRR-RLK genes in rice accessions grown under wet conditions than dry conditions. In sum, this pan-genomic analysis not only reveals the extensive diversity of the immune receptor repertoires in rice but also provides potential target genes for improving disease resistance in rice.

Oryza

New Insights into Genomic Variations and Mutational Events Associated with Plant-Pathogen Interactions.

Plant diseases threaten global food security, causing up to 40% crop yield losses and more than $220 billion in annual economic damage. This review synthesizes recent advances in understanding the genomic variations and mutational events underlying plant-pathogen interactions and durable plant disease resistance. Key insights into evolutionary dynamics, genetic variability, and coadaptive strategies reveal the complexity of host-pathogen relationships and the implications for developing durable disease resistance. Integrative approaches combining genome-wide association studies and functional genomics have uncovered the polygenic and epistatic architecture of quantitative resistance. Advances in pan-genomics and high-throughput sequencing have revealed extensive genetic variability in cultivated/elite germplasm and wild relatives. Emerging technologies, including gene editing, multi-omics, and machine learning, enable predictive modeling of resistance traits and support evolution that informs plant breeding strategies. Collectively, these advances provide a robust framework for developing durable resistance and sustainable crop protection in the face of global agricultural challenges.

Host-Pathogen Interactions

A metabolic atlas of the Klebsiella pneumoniae species complex reveals lineage-specific metabolism and capacity for intra-species co-operation.

The Klebsiella pneumoniae species complex inhabits a wide variety of hosts and environments, and is a major cause of antimicrobial resistant infections. Genomics has revealed the population comprises multiple species/sub-species and hundreds of distinct co-circulating sub-lineage (SLs) that are associated with distinct gene complements. A substantial fraction of the pan-genome is predicted to be involved in metabolic functions and hence these data are consistent with metabolic differentiation at the SL level. However, this has so far remained unsubstantiated because in the past it was not possible to explore metabolic variation at scale. Here, we used a combination of comparative genomics and high-throughput genome-scale metabolic modeling to systematically explore metabolic diversity across the K. pneumoniae species complex (n = 7,835 genomes). We simulated growth outcomes for each isolate using carbon, nitrogen, phosphorus, and sulfur sources under aerobic and anaerobic conditions (n = 1,278 conditions per isolate). We showed that the distributions of metabolic genes and growth capabilities are structured in the population, and confirmed that SLs exhibit unique metabolic profiles. In vitro co-culture experiments demonstrated reciprocal commensalistic cross-feeding between SLs, effectively extending the range of conditions supporting individual growth. We propose that these substrate specializations may promote the existence and persistence of co-circulating SLs by reducing nutrient competition and facilitating commensal interactions. Our findings have implications for understanding the eco-evolutionary dynamics of K. pneumoniae and for the design of novel strategies to prevent opportunistic infections caused by this World Health Organization priority antimicrobial resistant pathogen.

Klebsiella pneumoniae