Search PubMedSearch

SEARCH · Search PubMed

Results for “comparative genome analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Comparative genomic analysis reveals distinct population structure in Legionella anisa.

Legionella anisa has been frequently isolated from engineered water systems; however, its population structure remains understudied compared to Legionella pneumophila. Here, we generated complete genome sequences for four L. anisa isolates recovered from a healthcare facility in Rimouski, Canada. Further the population structure of this species was investigated by performing comparative genomic analyses of the genomes generated in this study together with publicly available L. anisa genomes. Genome-wide phylogenetic analysis revealed the presence of three distinct clades separated by substantial genetic divergence (∼500 SNP), with the Rimouski isolates forming a tightly clustered group, suggesting a clonal lineage. Comparative pangenome analysis indicated moderate core genome conservation accompanied by a highly variable accessory genome (∼50%). The isolates characterized in this study harbored multiple plasmids encoding genes associated with conjugation, heavy metal resistance, and other stress-related functions, suggesting potential roles in environmental persistence. Previous studies have shown that L. anisa can proliferate within protozoan host cells, although outcomes vary depending on the host species. Our isolates showed efficient proliferation within Acanthamoeba castellanii, but not within Vermamoeba vermiformis, under the conditions tested. Together, these findings underscore the genomic diversity of this understudied Legionella species and provide a framework for future investigations regarding environmental persistence and potential pathogenicity.

Legionella anisa, Whole genome sequencing

Comparative genomic analysis of Acer tsinglingense and A. davidii provides insights into nervonic acid biosynthesis, population evolution and genome vulnerability of endangered A. tsinglingense.

Global biodiversity is facing threats from climate change, habitat fragmentation, and anthropogenic activities-pressures that particularly endanger endemic and narrowly distributed species. In this study, the high-quality chromosome-level genomes of two ecologically divergent maples were assembled: the endangered and range-restricted Acer tsinglingense (791.40 Mb) and its widespread congener Acer davidii (1291.99 Mb). Phylogenomic analysis indicates that the two species diverged ~16.3 million years ago, with A. tsinglingense showing notable gene family expansions in secondary metabolite pathways. Notably, the 3-ketoacyl-CoA synthase gene family, which is involved in nervonic acid biosynthesis, underwent significant expansion and tandem duplication in A. tsinglingense, exhibiting high expression in buds. Population genomic analysis revealed that, compared with the widely distributed A. davidii, A. tsinglingense possesses lower genetic diversity, higher harmful mutation load, and signatures of a severe population bottleneck during the Late Pleistocene. Genome-environment association analysis further identified climate-adaptive genomic variations linked to five key environmental factors and projected potential genomic offsets under future climate scenarios. The southern lineage of A. tsinglingense exhibited greater climate sensitivity and genomic vulnerability under strong selective pressures, underscoring its importance as a conservation priority. Our research reveals that metabolic specializations in A. tsinglingense (such as the synthesis of nervonic acid) may confer competitive advantages in specific habitats. However, factors including its restricted distribution, historical population bottlenecks, and accumulated genetic load severely constrain its evolutionary potential to cope with rapid climate change. These findings emphasize the importance of elucidating the genomic basis and mechanisms of endangerment in metabolically specialized and threatened plant species to inform effective conservation strategies.

Genome, Plant

Comparative genomic analysis and functional investigations for MCs catabolism mechanisms and evolutionary dynamics of MCs-degrading bacteria in ecology.

Microcystins (MCs) significantly threaten the ecosystem and public health. Biodegradation has emerged as a promising technology for removing MCs. Many MCs-degrading bacteria have been identified, including an indigenous bacterium Sphingopyxis sp. YF1 that could degrade MC-LR and Adda completely. Herein, we gained insight into the MCs biodegradation mechanisms and evolutionary dynamics of MCs-degrading bacteria, and revealed the toxic risks of the MCs degradation products. The biochemical characteristics and genetic repertoires of strain YF1 were explored. A comparative genomic analysis was performed on strain YF1 and six other MCs-degrading bacteria to investigate their functions. The degradation products were investigated, and the toxicity of the intermediates was analyzed through rigorous theoretical calculation. Strain YF1 might be a novel species that exhibited versatile substrate utilization capabilities. Many common genes and metabolic pathways were identified, shedding light on shared functions and catabolism in the MCs-degrading bacteria. The crucial genes involved in MCs catabolism mechanisms, including mlr and paa gene clusters, were identified successfully. These functional genes might experience horizontal gene transfer events, suggesting the evolutionary dynamics of these MCs-degrading bacteria in ecology. Moreover, the degradation products for MCs and Adda were summarized, and we found most of the intermediates exhibited lower toxicity to different organisms than the parent compound. These findings systematically revealed the MCs catabolism mechanisms and evolutionary dynamics of MCs-degrading bacteria. Consequently, this research contributed to the advancement of green biodegradation technology in aquatic ecology, which might protect human health from MCs.

Humans

Comparative Genomic Analysis of Multidrug-Resistant Escherichia coli Across Poultry-Human-Environmental Interfaces.

The emergence of multidrug-resistant (MDR) Escherichia coli in poultry represents a critical One Health concern, particularly in developing countries. This study employed a comparative genomic approach to investigate the genomic characteristics, antimicrobial resistance (AMR) profiles, virulence determinants, of poultry-derived MDR E. coli isolates from Bangladesh. Whole-genome sequencing of three representative MDR isolates, identified with 83 globally diverse poultry, human, and environmental E. coli genomes. Pangenome analysis identified the characteristic open pangenome of E. coli, with core genes comprising only 4.6% of the combined dataset. Resistome analysis shown diverse AMR determinants, including blaCTX-M, blaTEM, sul, tet, and qnrS1, associated with antibiotic inactivation and efflux mechanisms. Virulence profiling revealed diverse genes involved in adhesion (fim, csg), iron acquisition (ent, fep, chu), motility, and secretion systems, with core virulence genes exhibiting > 90% sequence identity, whereas accessory virulence genes were more variable. Plasmid analysis demonstrated heterogeneous replicon types, predominantly IncF and Col plasmids, indicating their role in horizontal gene transfer. Jaccard similarity indices revealed moderate to high genetic overlap with global strains (~0.63 for virulence genes and ~0.55 for AMR profiles), suggesting shared evolutionary backgrounds. Phylogenomic and MLST identified all Bangladeshi isolates as ST457, clustering within a globally distributed clonal complex linked to ST10 and ST131 lineages. These findings suggest that the three Bangladeshi poultry-derived E. coli isolates are genetically related to globally circulating strains while harboring extensive resistance and virulence determinants, emphasizing poultry as an important reservoir of MDR pathogens and reinforcing the need for strengthened antimicrobial stewardship and genomic surveillance.

Animals

Comparative genomic analysis of Streptococcus parasuis and Streptococcus suis reveals mobile element-associated enrichment of antimicrobial resistance and lack of detectable same-MGE colocalization with virulence-associated genes within stable species boundaries.

Streptococcus suis is a major porcine pathogen and a zoonotic agent that causes meningitis and septicemia in humans. Streptococcus parasuis, a recently recognized close relative, remains poorly characterized with regard to its clinical significance and genomic features. In this study, we generated a single-contig closed genome assembly with genome-wide DNA methylation profiles for S. parasuis strain A1, isolated from a diseased pig in Xinjiang, China, and complemented in silico genomic predictions with isolate-level experimental validation of antimicrobial resistance (AMR) genotypes, virulence genotypes, and phenotypic susceptibility for this reference strain. Using this high-quality genome as a reference anchor, we performed comparative genomic analyses across 195 streptococcal genomes, comprising 15 S. parasuis and 180 S. suis strains, to distinguish genome-level co-occurrence of resistance and virulence determinants from their physical colocalization on the same mobile genetic element (MGE).Species boundaries remained clearly delineated at the genomic level, with a median interspecies average nucleotide identity (ANI) of approximately 86.0%, compared with intraspecies ANI medians of 97.5% for S. parasuis and 96.2% for S. suis. Pangenome analysis identified 12,693 gene clusters, of which 1086 were core clusters, and functional annotation revealed significant differences in accessory gene repertoires between the two species. Within this stable genomic framework, S. parasuis genomes carried a higher AMR gene burden; strain A1 harbored 10 AMR genes, multiple virulence-associated genes, three genomic islands, and eight prophage regions. For strain A1, PCR validation confirmed six AMR genes and six virulence genes, and disk diffusion testing demonstrated a multidrug-resistant phenotype consistent with the genotypic profile.Among 235 predicted mobile elements, 19 harbored AMR genes and seven carried Virulence Factor Database (VFDB) homologs, but none carried both categories simultaneously. This finding reflects a lack of detectable same-MGE colocalization under the applied annotation and assembly framework; it should not be interpreted as evidence of biological physical decoupling. Under a random-placement model, the expected number of co-carrying regions was only 0.57, and the probability of observing zero co-carrying regions was P = 0.55. This negative result should be interpreted with caution, given the limited number of cargo-bearing regions and the predominantly draft status of most genomes. Furthermore, the A1 genome contained multiple restriction-modification systems, showed depletion of several methylation motif families in mobile regions, and had limited CRISPR spacer matching evidence, suggesting prior exposure to the relevant sequence space. None of the genomes met our predefined criteria for whole-genome convergence.Collectively, our results support a model in which S. parasuis accumulates AMR-related genes in a modular fashion via mobile elements within stable species boundaries, with no detectable same-MGE colocalization of AMR and virulence determinants under our analytical pipeline. These findings imply that AMR surveillance strategies for this species should prioritize tracking mobile genetic elements rather than inferring wholesale genomic convergence toward S. suis.

Streptococcus suis

A comparative genomic analysis of left- and right-sided colon cancer using real-world data from the AACR project GENIE BPC dataset.

Left- and Right-sided colon cancers (LCC and RCC) are increasingly recognized as distinct clinicopathological and molecular subtypes with divergent prognoses and therapeutic responses. Leveraging a large, multi-institutional cohort from the AACR Project Genomics Evidence Neoplasia Information Exchange (GENIE) Biopharma Collaborative (BPC) (n = 750; LCC: 363 vs. RCC: 387), we conducted a comprehensive analysis of mutational profiles, tumor mutation burden (TMB), and survival outcomes. Our findings revealed a markedly higher TMB in RCC compared to LCC (6.65 &#xb1; 11.3 vs. 3.17 &#xb1; 4.35; adjusted P = 3.12&#xd7;10-32), suggesting greater genomic instability in RCC. After applying functional annotation filters (PolyPhen > 0.85, SIFT < 0.05), RCC tumors were significantly enriched for mutations in BRAF (23.1% vs. 6.7%), KMT2D (8.6% vs. 3.2%), and SMAD4 (13.1% vs. 7.3%), while TP53 mutations predominated in LCC (40.6% vs. 31.8%). Multivariate Cox regression analysis identified RCC as an independent predictor of poorer overall survival (OS) relative to LCC (HR: 1.30, 95% CI: 1.02-1.66, P = 0.033). Notably, KRAS mutations were associated with significantly worse OS in LCC (HR: 1.68, 95% CI: 1.06-2.70, P = 0.027), while BRAF mutations predicted adverse outcomes in RCC (HR: 1.58, 95% CI: 1.05-2.37, P = 0.028). These results underscore the prognostic value of tumor sidedness and specific genetic alterations in colon adenocarcinoma. Our study highlights the need for sidedness-specific molecular profiling to inform precision oncology strategies in colon cancer management.

BRAF

Chloroplast genome comparative analysis and phylogenetic relationships of 15 Syringa species (Oleaceae).

Syringa is a crucial shrub genus in the family Oleaceae, which has significant ornamental, economic, and medicinal value. However, research on the chloroplast genome (CPG) phylogeny and lineage diversification of this genus remains limited. In this study, all 15 Syringa CPGs exhibited a characteristic quadripartite structure, with genome lengths ranging from 154,019-158,020 bp. These CPGs were highly conserved and moderately differentiated, each containing 130-132 genes. Analysis of inverted repeat (IR) boundaries indicated structural conservation, with six genes: rps19, rpl2, ycf1, trnN, ndhF, and trnH present at the IR/single-copy (SC) junctions. The small single copy (SSC) region displayed greater sequence variability than the IR regions. ycf1, ndhH, trnL-rpl32, ndhF-ycf1, and rbcL-accD were identified as potential molecular markers and rps11, ycf2, and ycf4 may have contributed to the adaptive evolution of Syringa. Phylogenetic reconstruction based on whole CPG data supported the monophyly of the 15 species, which were divided into three distinct subclades. Molecular dating estimated that Syringa diverged from its sister genus approximately 58 million years ago, with most Syringa species diversifying further approximately 47.49 million years ago during the Eocene. Our findings will hopefully stimulate further studies on this genus that may enhance biodiversity knowledge.

Journal Article

Comparative genomic analysis of a novel heat-tolerant and euryhaline strain of unicellular marine cyanobacterium Cyanobacterium sp. DS4 from a high-temperature lagoon.

BACKGROUND: Cyanobacteria have diversified through their long evolutionary history and occupy a wide range of environments on Earth. To advance our understanding of their adaptation mechanisms in extreme environments, we performed stress tolerance characterizations, whole genome sequencing, and comparative genomic analyses of a novel heat-tolerant and euryhaline strain of the unicellular cyanobacterium Cyanobacterium sp. Dongsha4 (DS4). This strain was isolated from a lagoon on Dongsha Island in the South China Sea, a habitat with fluctuations in temperature, salinity, light intensity, and nutrient supply. RESULTS: DS4 cells can tolerate long-term high-temperature up to 50 &#x2103; and salinity from 0 to 6.6%, which is similar to the results previously obtained for Cyanobacterium aponinum. In contrast, most mesophilic cyanobacteria cannot survive under these extreme conditions. Based on the 16S rRNA gene phylogeny, DS4 is most closely related to Cyanobacterium sp. NBRC102756 isolated from Iwojima Island, Japan, and Cyanobacterium sp. MCCB114 isolated from Vypeen Island, India. For comparison with strains that have genomic information available, DS4 is most similar to Cyanobacterium aponinum strain PCC10605 (PCC10605), sharing 81.7% of the genomic segments and 92.9% average nucleotide identity (ANI). Gene content comparisons identified multiple distinct features of DS4. Unlike related strains, DS4 possesses the genes necessary for nitrogen fixation. Other notable genes include those involved in photosynthesis, central metabolisms, cyanobacterial starch metabolisms, stress tolerances, and biosynthesis of novel secondary metabolites. CONCLUSIONS: These findings promote our understanding of the physiology, ecology, evolution, and stress tolerance mechanisms of cyanobacteria. The information is valuable for future functional studies and biotechnology applications of heat-tolerant and euryhaline marine cyanobacteria.

Cyanobacteria

Comparative genomic analysis of key oncogenic pathways in hepatocellular carcinoma among diverse populations.

BACKGROUND/OBJECTIVES: Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality, with significant racial and ethnic disparities in incidence, tumor biology, and clinical outcomes. Hispanic/Latino (H/L) patients tend to be diagnosed at younger ages and more advanced stages than Non-Hispanic White (NHW) patients, yet the molecular mechanisms underlying these disparities remain poorly understood. Key oncogenic pathways, including RTK/RAS, TGF-Beta, WNT, PI3K, and TP53, play pivotal roles in tumor progression, treatment resistance, and response to targeted therapies. However, ethnicity-specific alterations within these pathways remain largely unexplored. This study aims to compare pathway-specific mutations in HCC between H/L and NHW patients, assess tumor mutation burden, and identify ethnicity-associated oncogenic drivers using publicly available datasets. Findings from this analysis may inform precision medicine strategies for improving early detection and targeted therapies in underrepresented populations. METHODS: We conducted a bioinformatics analysis using publicly available HCC datasets to assess mutation frequencies in RTK/RAS, TGF-Beta, WNT, PI3K, and TP53 pathway genes. The study included 547 patients, consisting of 69 H/L patients and 478 NHW patients. Patients were stratified by ethnicity (H/L vs. NHW) to evaluate differences in mutation prevalence. Chi-squared tests were used to compare mutation frequencies, while Kaplan-Meier survival analysis assessed overall survival differences associated with pathway-specific alterations in both populations. RESULTS: Significant differences were observed in the RTK/RAS pathway related genes, particularly in FGFR4 mutations, which were more prevalent in H/L patients compared to NHW patients (4.3% vs. 0.6%, p = 0.02). Additionally, IGF1R mutations exhibited borderline significance (7.2% vs. 2.9%, p = 0.07). In the PI3K pathway, INPP4B alterations were more frequent in H/L patients than in NHW patients (4.3% vs. 1%, p = 0.06), while in the TGF-Beta pathway, TGFBR2 mutations were more common in H/L patients (2.9% vs. 0.4%, p = 0.07), suggesting potential ethnicity-specific variations. Survival analysis revealed no significant differences in overall survival between H/L and NHW patients, indicating that molecular alterations alone may not fully explain survival disparities and suggesting a role for additional factors such as immune response, environmental exposures, or access to targeted therapies. CONCLUSIONS: This study provides one of the first ethnicity-focused analyses of key oncogenic pathway alterations in HCC, revealing distinct molecular differences between H/L and NHW patients. The findings suggest that RTK/RAS (FGFR4, IGF1R), PI3K (INPP4B), and TGF-Beta (TGFBR2) pathway alterations may play a distinct role in HCC among H/L patients, while their prognostic significance in NHW patients remains unclear. These insights emphasize the importance of incorporating ethnicity-specific molecular profiling into precision medicine approaches to improve early detection, targeted therapies, and clinical outcomes in HCC, particularly for underrepresented populations.

PI3K pathway

A python based automated computational framework to classify and comparative genomics analysis of the global diversity of chili leaf curl virus (ChiLCV) strains to understand virus host interactions.

Chili leaf curl virus (ChiLCV) is a Begomovirus chillicapsici that is one of the most devastating viruses impacted on the production of chili in the world, especially in South Asia. In the present study, we combined high-throughput computational genomics with experimental analysis of global diversity. A workflow was created using automated Python scripts to download, curate and process ChiLCV genomes from public database. About 410 complete ChiLCV genomes download from public databases. Using a phylogenetic approach, these isolates were subdivided into 34 strains, belonging to 10 major clades, showing significant genetic diversity. Geographic analysis revealed that Pakistan (207 isolates) and India (148 isolates) were the main sources of ChiLCV diversity and the remainder of the isolates were from Oman, Bangladesh, Iran, Saudi Arabia and Sri Lanka. Recombination was observed as a major evolutionary force as more than twenty recombination events were detected. Analysis of cis-regulatory elements showed a complex structure of the viral promoter, including multiple binding sites for transcription factors, hormone-response elements, light-responsive elements, and stress-responsive elements, indicating a high number of interactions between viral regulatory elements and host signaling pathways. Pangenome analysis showed the presence of a highly dynamic open pangenome made up of strain-specific orthologous groups (species-specific orthogroups). Experimental inoculation of chili plants was also carried out to assess the biological effects of infection, along with phytochemical, FTIR, HPLC, and qPCR analyses.

Begomovirus

Comparative genomic analysis of Artemisia argyi reveals asymmetric expansion of terpene synthases and conservation of artemisinin biosynthesis.

Artemisia argyi, a perennial herb of the Asteraceae family, possesses significant therapeutic and economic value. We present a 7.88&#x2009;Gb chromosome-level haplotype-resolved genome assembly, revealing its unique evolutionary trajectory. The karyotype (2n&#x2009;=&#x2009;34) of A. argyi is that of an autotetraploid, which underwent gametic chromosome fusion prior to species-specific whole-genome duplication (WGD-3). The genome exhibits pronounced multivalent chromosome pairing and frequent recombination among homologous groups. Asymmetrical evolution following WGD-3 is a hallmark feature, evidenced by imbalanced allelic gene loss and widespread neofunctionalization. The terpene synthase (TPS) gene family exemplifies this pattern, having expanded through four duplication events in A. argyi. Recent tandem duplications and allelic functional differentiation have generated substantial gene functional diversity. Notably, we identified a tandem-duplicated six-copy ADS homolog (AarADS)-a key TPS gene in the artemisinin biosynthetic pathway of Artemisia annua (AanADS)-localized exclusively to a single chromosome in A. argyi. Unlike AanADS, which converts farnesyl pyrophosphate (FPP) to amorpha-4,11-diene, AarADS catalyzes FPP to &#x3b1;-bisabolol. Evolutionary analysis suggested that AanADS acquired its specialized function via a derived mutation in the A. annua lineage. This study elucidates the genomic evolution underpinning A. argyi's distinctive medicinal properties.

Alkyl and Aryl Transferases

Genomic Insights Into Multidrug-Resistant Foodborne Serratia liquefaciens Strains Carrying mcr-9 and Comparative Genomic Analysis of Novel Biosynthetic Gene Clusters.

Serratia liquefaciens is an opportunistic nosocomial pathogen with a wide range of antibiotic resistance patterns. This study reports the characterization of the first mcr-9-positive S. liquefaciens strains, 35E-19E1 and CST-066, isolated from meat products in Japan. The strains were screened for the presence of &#x3b2;-lactamases, plasmid-mediated mobile colistin resistance (mcr) genes, and carbapenemase-encoding genes using PCR. Antimicrobial susceptibility was tested using the broth microdilution method. The strains exhibited multidrug resistance (MDR) phenotypes to third-generation cephalosporins, cephamycin, fosfomycin, and other clinically important antimicrobials. Genomic DNA sequencing showed that the genome sizes of CST-066 and 35E-19E1 are 5,529,704 and 5,261,506&#x2009;bps, respectively. mcr-9 was identified on a chromosome within a genetic environment that included the two-component system qseBC, which plays a key role in the signaling network that triggers colistin resistance in Enterobacterales. Downstream genome analysis revealed a 1695-bp eptB-like kdo2-lipid phosphoethanolamine transferase, which is involved in intrinsic polymyxin resistance mechanisms in Serratia spp. The strain 35E-19E1 carries five CRISPR-Cas enzymes that are essential for adaptive immunity in bacteria, allowing defense against invading elements. Functional analysis using subsystem technology revealed that both strains possess subsystem features responsible for invasion and adhesion within the host biomes. Genome mining using antiSMASH and BAGL4 revealed various biosynthetic gene clusters, responsible for secondary metabolite synthesis. Notably, we identified novel gene clusters, mainly nonribosomal peptide synthetases, in both the strains, indicating their potential to produce bioactive compounds. Although the presence of mcr-9 in Serratia may not be of clinical significance because of natural resistance of the strain to polymyxins, we shed light on the genomic characteristics of this MDR pathogen and the potential spread of mcr-9 among other bacterial species. The emergence of mcr-9 in drug-resistant S. liquefaciens provides significant insights, underscoring the need for increased surveillance of this pathogen.

biosynthetic gene cluster

Oxford Nanopore Sequencing of Clinical DNA for Identification and Comparative Genomic Analysis of Erysipelothrix piscisicarius.

The genus Erysipelothrix comprises facultative anaerobic, nonspore-forming, gram-positive bacteria that can cause skin infections and severe diseases such as septicemia and endocarditis in humans. Although E. rhusiopathiae is the primary pathogen, other species may also be involved, necessitating accurate identification. However, 16S rDNA sequencing lacks sufficient resolution to differentiate among Erysipelothrix species. In this study, we used Oxford Nanopore Technology (ONT) to directly sequence low-quality DNA extracted from heart valve tissue of a 66-year-old female patient with a fatal case of septicemia and aortic endocarditis. In contrast to 16S rDNA Illumina sequencing and matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS), which incorrectly identified the pathogen as E. rhusiopathiae, direct sequencing via ONT precisely identified E. piscisicarius as the cause of infection. About 1.47&#x2009;Mb genome was retrieved from nanopore direct sequencing. Within the E. piscisicarius genome, we detected genes associated with virulence. Phylogenetic analysis showed that our strain clustered with a human-derived E. piscisicarius strain from China and swine-derived strains from Brazil. In conclusion, this study demonstrated that ONT can be used to sequence low-quality DNA extracted directly from patient specimens, obtain a draft bacterial genome, and reliably distinguish between pathogenic species.

Aged

Comparative Genomic Analysis of Six Mycoplasma Gallisepticum Strains: Insights into Genetic Diversity and Antibiotic Resistance.

Mycoplasma gallisepticum (MG) is a significant pathogen that causes respiratory diseases, which have had a substantial economic impact on the poultry industry. Despite the resistance of MG to antibiotics, it is imperative to identify genetic diversity in order to develop countermeasures. In this study, the genomes of six MG strains were examined to gain deeper insights into the mutations. The data pertaining to Variant Annotation and Mutation Analysis using SnpEff, along with the calculation of mutation rates as the ratio of total mutations to the length of the genomic regions analyzed, were thoroughly examined. The comprehensive evaluation yielded a total of 25,942 variants across the six strains, underscoring substantial genetic diversity. Notably, strain S6 exhibited a preponderance of frameshift mutations. A notable finding was the presence of a mutation in the MsbA gene shared by all six strains. Furthermore, five of the six strains, with the exception of strain F99 Lab, exhibited a mutation at position 5158, which impacts a multidrug transport system. Notably, strain ATCC exhibits a distinctive mutation at position 942, while strain S6 displays a unique mutation at position 6855, which is linked to efflux ABC transporter components. Furthermore, a substantial degree of genetic variation was observed among the CrmA, GapA, and vlhA genes among the various strains. High-impact changes, such as insertions and deletions, exhibited a higher frequency in CrmA, particularly in strain S6. Conversely, nonsynonymous variations demonstrated a heightened prevalence in GapA, particularly in strain F99 Lab. The vlhA gene exhibited a spectrum of effects, ranging from synonymous mutations to high-impact mutations such as stop-gains and frameshifts, particularly in strains k5111a and k4602. The functional variations observed among the strains can be attributed to these mutations, which have the potential to alter gene expression or protein function. Furthermore, substantial mutations in the dxr and rpoC genes were associated with antibiotic resistance. These mutations underscore the ongoing evolutionary adaptations of M. gallisepticum. Consequently, there is an imperative for the revision of treatment protocols and the formulation of targeted vaccines to regulate resistance within the poultry industry.

Mycoplasma gallisepticum

CAKR: commutative algebra k-mer representations for genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer representations as a nonlinear algebraic framework for analyzing genomic sequences. This representation bridges commutative algebra, algebraic topology, combinatorics, and machine learning to establish a mathematical framework for comparative genomic analysis. We evaluate its effectiveness on three tasks including genetic variant classification, phylogenetic tree reconstruction, and viral classification, typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. In this work, we show that commutative algebra k-mer representations outperform five state-of-the-art sequence analysis methods across twelve primary datasets, with two additional supplementary fragment-placement benchmarks, especially in viral classification, and maintain relatively stable predictive accuracy as dataset size increases, underscoring scalability and robustness.

Genomics

CAKL: Commutative algebra k-mer learning of genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer learning (CAKL) as the first-ever nonlinear algebraic framework for analyzing genomic sequences. CAKL bridges between commutative algebra, algebraic topology, combinatorics, and machine learning to establish a new mathematical paradigm for comparative genomic analysis. We evaluate its effectiveness on three tasks-genetic variant identification, phylogenetic tree analysis, and viral genome classification-typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. Across eleven datasets, CAKL outperforms five state-of-the-art sequence analysis methods, particularly in viral classification, and maintains stable predictive accuracy as dataset size increases, underscoring its scalability and robustness. This work ushers in a new era in commutative algebraic data analysis and learning.

Journal Article

Whole-genome safety assessment of Loigolactobacillus coryniformis WBB05 and identification of a candidate gene for aerobic reuterin production.

This study reports on the safety profile of Loigolactobacillus coryniformis WBB05 for food industry applications and identifies glycerol-3-phosphate oxidase (GlpO) as a candidate gene associated with aerobic reuterin production. The safety of L. coryniformis WBB05 was evaluated through whole-genome sequencing, phenotypic analysis of haemolytic activity and determination of minimum inhibitory concentrations (MICs) of antibiotics. Comparative genomic analysis was performed to identify candidate genetic determinants for aerobic reuterin production. The draft genome (2.83 Mb, 179 contigs) harboured no known virulence factors, acquired antimicrobial resistance (AMR) genes or biogenic amine biosynthetic genes. Prophage analysis identified only one incomplete prophage region, and four CRISPR-Cas systems (212 spacers) were consistent with phage defence capacity. Secondary metabolite analysis revealed biosynthetic gene clusters encoding a coagulin-like bacteriocin. No &#x3b2;-haemolytic activity was observed. The MICs of all antibiotics tested were below the European Food Safety Authority cut-off values except for kanamycin (128&#xa0;mg/L), although no acquired AMR genes were detected. Comparative genomic analysis revealed that L. coryniformis WBB05 possesses two putative copies of GlpO, a gene not detected in publicly available genomes of Limosilactobacillus reuteri, which produces reuterin only under anaerobic conditions. These findings support the use of L. coryniformis WBB05 as a safe adjunct culture for dairy applications and highlight GlpO as a candidate determinant of aerobic reuterin production. Further studies comparing GlpO-positive and GlpO-negative strains under aerobic and anaerobic conditions are warranted to confirm the role of GlpO.

Loigolactobacillus coryniformis

Comparative genomic and proteomic analysis reveals orthogroup structured evolution of tick protease inhibitors.

Protease inhibitors (PIs) play central roles in regulating endogenous proteolysis and host-parasite interactions in ticks. However, the evolutionary architecture underlying their diversification across tick lineages remains insufficiently resolved. Here, we performed a genome-wide comparative analysis of predicted proteomes from 14 tick species to systematically characterize PI repertoires. In total, 4931 putative PIs were identified and grouped into 20 families using the MEROPS classification system. Further, PI families such as Antistasin, WAP-type, and Pacifastin, which have not previously been systematically reported in tick genomes, were classified. Orthogroup inference demonstrated that PI expansion is structured at the level of evolutionary lineages rather than uniformly across families. By stratifying orthogroups according to duplication burden and taxonomic conservation, we identified a broadly conserved single-copy core under strong purifying selection. Motif level analysis of serpin reactive center loops further revealed conservation of inhibitory specificity within single copy orthogroups and diversification of key functional residues in duplication-associated lineages. Integration of secretion prediction and tissue-resolved proteomics from Hyalomma anatolicum and Rhipicephalus microplus demonstrated that evolutionary stratification is reflected at the protein level. Together, these findings provide an orthogroup-resolved evolutionary framework linking duplication dynamics, molecular evolution, and tissue-level protein deployment. This integrative approach offers a systematic basis for prioritizing conserved and diversified PI lineages for future functional and anti-tick intervention studies.

Animals