Search PubMedSearch

SEARCH · Search PubMed

Results for “Introduced Species”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Stability in the face of global decline: a 20-year study of arthropods in an oceanic archipelago.

Insect declines are of global concern, yet no long-term ecological studies (LTER) have confirmed this trend on islands. This study utilises the first available LTER data on island arthropods, targeting epigeal and canopy species from the Azores Archipelago (Portugal), and covering over 20 years in three distinct sampling events from 30 standard sites. We investigate changes in abundance, biomass, and species richness within native forest arthropod communities, focusing on the proportions of endemic and introduced species, and temporal patterns among single-island endemics and forest-dependent endemics. Results reveal significant temporal variability, but overall abundance, biomass, and species richness remain stable across endemic and native non-endemic taxa. Among the species studied, 28% declined, 17% increased, and 55% showed no significant differences. Exotic invasions and related extinctions appear minimal. Forest-dependent endemic species declined below anticipated levels, suggesting that the extinction debt for these species may be less severe than initially expected. Nonetheless, some forest specialists have declined significantly, and seven species, not seen over 20 years, are considered to be extinct. The three-decade-long conservation of Azorean native forests may have contributed to the stability of some populations, thus these findings underscore the need for continued and enhanced conservation efforts of insular forest-associated diversity.

Animals

LYCEUM: learning to call copy number variants on low-coverage ancient genomes.

MOTIVATION: Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduces additional noise into sequencing data; and finally, (iii) the typically low-coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high-coverage read-depth signals, underperform under such conditions. RESULTS: To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high-confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection. AVAILABILITY AND IMPLEMENTATION: LYCEUM is available at https://github.com/ciceklab/LYCEUM.

DNA Copy Number Variations

Adaptation in a keystone grazer under novel predation pressure.

Understanding how species adapt to environmental change is necessary to protect biodiversity and ecosystem services. Growing evidence suggests species can adapt rapidly to novel selection pressures like predation from invasive species, but the repeatability and predictability of selection remain poorly understood in wild populations. We tested how a keystone aquatic herbivore, Daphnia pulicaria, evolved in response to predation pressure by the introduced zooplanktivore Bythotrephes longimanus. Using high-resolution 210Pb-dated sediment cores from 12 lakes in Ontario (Canada), which primarily differed in invasion status by Bythotrephes, we compared Daphnia population genetic structure over time using whole-genome sequencing of individual resting embryos. We found strong genetic differentiation between populations approximately 70 years before versus 30 years after reported Bythotrephes invasion, with no difference over this period in uninvaded lakes. Compared with uninvaded lakes, we identified, on average, 64 times more loci were putatively under selection in the invaded lakes. Differentiated loci were mainly associated with known reproductive and stress responses, and mean body size consistently increased by 14.1% over time in invaded lakes. These results suggest Daphnia populations were repeatedly acquiring heritable genetic adaptations to escape gape-limited predation. More generally, our results suggest some aspects of environmental change predictably shape genome evolution.

Animals

Comparative Responses of Invasive and Native Plant Species to Combined Cd and Microplastic Pollution.

The co-occurrence of heavy metal contamination and biodegradable microplastic (polylactic acid, PLA) pollution poses increasing risks to terrestrial plant communities and soil functioning, yet species-specific responses to combined stress remain poorly understood. Cd and microplastics frequently co-occur in agricultural soils, where microplastics can alter cadmium mobility, bioavailability, and transport pathways, potentially modifying metal toxicity and plant stress responses compared with single-pollutant exposure. We investigated the responses of the invasive Bidens pilosa and the native Solanum nigrum grown in monoculture and mixed culture under combined cadmium (Cd) and biodegradable microplastic (PLA) stress by integrating plant growth, photosynthetic performance, oxidative physiology, and rhizosphere biochemical processes. Combined Cd-MP exposure markedly reduced plant growth, chlorophyll content (SPAD), photosystem II efficiency (Fv/Fm), nitrogen accumulation, biomass production, and rhizosphere enzyme activities associated with carbon, nitrogen, and phosphorus cycling. However, B. pilosa maintained greater physiological stability under stress, characterized by higher antioxidant enzyme activities (SOD, CAT, POD), lower reactive oxygen species (H2O2, O2˙-) accumulation, and reduced lipid peroxidation (MDA), whereas S. nigrum exhibited stronger oxidative damage and functional impairment. Multivariate analyses further revealed that root antioxidant capacity was closely associated with rhizosphere microbial enzyme activity, suggesting a root-centered regulatory mechanism linking plant stress tolerance to soil functioning. Overall, the invasive species showed greater tolerance to combined contamination and maintained relatively higher rhizosphere functional activity than the native species, indicating that multi-pollutant stress may alter competitive interactions between invasive and native plants in contaminated environments.

Cadmium

The impact of non-native trees on galling and herbivory in New York City across space and time.

Cities and suburbs frequently plant native and non-native trees as foundation species, with non-natives cultivated in these areas for centuries while remaining non-invasive. Although previous research has found that native trees often host more arthropods, studies have not simultaneously looked across space and time to determine the consistency of tree origin on urban arthropods. We combined varied methods across spatial and temporal scales in New York City to test if native tree leaves consistently have more insect and mite interactions than long-established non-native trees, predicting stronger effect sizes for specialists (galling arthropods) than generalists (herbivory). We examined (1) congeneric species pairs, controlled for growing conditions and stoichiometry in an arboretum, (2) diverse oaks at a botanical garden, (3) community science records across Brooklyn, and (4) herbarium specimens from 1883 through present across the city. Across spatiotemporal scales, we found consistent results. Specialist interactions were striking: contemporary native trees supported numerous galling species, while only one congeneric non-native species hosted any galls. For generalists, contemporary native trees had equivalent to slightly greater herbivory. Over the last century, herbarium records showed that herbivory increased on non-native trees to nearly the level of natives, whereas native trees increased in gall abundance while non-native trees remained rarely galled. Our results demonstrate the impact of tree origin on tree-arthropod interactions in a real-world urban setting, with far fewer galls even when non-native tree species have been cultivated locally for centuries. Our findings will help city planners and property owners confidently choose native trees to promote arthropod biodiversity.

Trees

Viral tags as keys to advancing invasion genomics.

Invasion genetics and genomics have greatly advanced the study of biological invasions, yet they often fail to resolve population dynamics at the fine spatiotemporal scales characteristic of most invasions. We propose shifting the focus away from the higher-order target species towards their viral symbionts, harnessing these as high-resolution 'genetic tags' to overcome many of these limitations. Owing to their comparably smaller genomes, shorter generation times, and higher mutation rates, most viruses evolve on timescales comparable to the invasion dynamics of their higher-order hosts, potentially better proxying and revealing recent dispersal patterns. We present a conceptual framework outlining how virus evolution may shed light on the contemporary spread of their non-native hosts, opening new avenues for invasion genetics, genomics, and management.

Genomics

Mikania micrantha invasion restructures rhizosphere nitrogen cycling through enzyme activation, microbial recruitment, and allelopathic regulation.

BACKGROUND: Plant invasions profoundly influence terrestrial ecosystems by reshaping nutrient cycling processes. However, the mechanisms through which invasive plants such as Mikania micrantha modulate soil nitrogen (N) cycling and microbial communities remain insufficiently explored. Moreover, comparative studies with indigenous congener are scarce, limiting insights into whether such effects reflect species-specific strategies or genus-wide traits. This study investigates how M. micrantha modulates nitrogen metabolic pathways and rhizosphere microecology using combined metagenomic and metabolomic analyses. RESULTS: Integrated analyses revealed that M. micrantha established a distinctive "high total nitrogen-low mineral nitrogen" profile in the rhizosphere soil. Metagenomic profiling showed consistent enrichment of key ammonium assimilation enzymes, including glutamine synthetase and glutamate dehydrogenase, promoting enhanced incorporation of NH₄⁺ into organic nitrogen pools. In contrast, genes encoding nitrate reductase and nitrate transporters were significantly lower in relative abundance, limiting nitrate assimilation. Mikania micrantha also selectively enriched nitrogen-fixing microbes (notably rhizobia genera) and plant growth-promoting rhizobacteria (PGPR), thereby enhancing biological nitrogen fixation capacity. Metabolomic analysis further identified several allelopathic compounds in invaded soils at higher relative abundance, particularly epicatechin, which exhibited inhibitory effects on nitrifying bacteria. Compared with the congener Mikania cordata, which exerted weaker impacts on soil nitrogen cycling and microbial assembly, M. micrantha deployed a more comprehensive strategy integrating biochemical, microbial, and metabolic regulation. CONCLUSIONS: These findings demonstrate that under greenhouse-controlled conditions, M. micrantha reconfigures rhizosphere nitrogen cycling through a multi-dimensional strategy that couples biochemical regulation, microbial recruitment, and metabolite-mediated interference, thereby suggesting a potential mechanism that may contribute to its ecological advantage in natural settings. Video Abstract.

Rhizosphere

Rapid and repeated evolution of increased competitive ability in a global invader.

Rapid adaptive evolution can increase the competitive ability of invasive species in their non-native ranges. However, whether this increase is a general response and what drives it remain uncertain because the evidence is largely based on studies with limited sampling, inadequate consideration of population co-ancestry, and oversimplified estimates of competitive ability. We conduct a large-scale glasshouse experiment testing the effects of competition and drought on 100 native and 165 non-native populations of Erigeron canadensis, all genotyped to account for co-ancestry. Plants from non-native populations are significantly more competitive against other species than the conspecifics from native populations under both mesic and dry conditions. Genetic clustering indicates that the rapid evolution of competitive ability occurs independently in two out of four clusters in the non-native range. This advantage is present only during interspecific interactions and is absent during intraspecific competition. Repeated evolution of increased competitive ability suggests that adaptation following introduction can reshape species interactions and promote invasion success, even under future drought conditions, highlighting the importance of rapid evolution in determining the ecological impacts of invasive plants.

Biological Evolution

DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings.

We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e., DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 23 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. Model, codes, and data is publicly available at https://github.com/MAGlCS-LAB/DNABERT_S.

Journal Article

Genetic linkage disequilibrium of deleterious mutations in threatened mammals.

The impact of negative selection against deleterious mutations in endangered species remains underexplored. Recent studies have measured mutation load by comparing the accumulation of deleterious mutations, however, this method is most effective when comparing within and between populations of phylogenetically closely related species. Here, we introduced new statistics, LDcor, and its standardized form nLDcor, which allows us to detect and compare global linkage disequilibrium of deleterious mutations across species using unphased genotypes. These statistics measure averaged pairwise standardized covariance and standardize mutation differences based on the standard deviation of alleles to reflect selection intensity. We then examined selection strength in the genomes of seven mammals. Tigers exhibited an over-dispersion of deleterious mutations, while gorillas, giant pandas, and golden snub-nosed monkeys displayed negative linkage disequilibrium. Furthermore, the distribution of deleterious mutations in threatened mammals did not reveal consistent trends. Our results indicate that these newly developed statistics could help us understand the genetic burden of threatened species.

Animals

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings.

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Sequence Analysis, DNA

Mitochondrial Impostors: Prevalence and Impacts of NUMTs on Genetic and Evolutionary Studies in Carnivora.

Nuclear mitochondrial pseudogenes are mitochondria-derived DNA sequences integrated into the nuclear genome, which can introduce errors in species identification, phylogenetic inference, and population genetics. Although nuclear mitochondrial pseudogene contamination has been reported in some Carnivora species, a systematic investigation into the prevalence and impacts of nuclear mitochondrial pseudogenes across an order is still lacking. In this study, 22,102 mitochondrial DNA sequences of 80 Carnivora species from 14 families and 54 genera were retrieved from the public National Center for Biotechnology Information database and further analyzed. Using alignment-based methods, 158 problematic sequences/sequence groups were identified and categorized into four types: nuclear mitochondrial pseudogenes, species misidentification or mislabeling, sequence errors, and anomalous sites. Among families, Felidae exhibited the highest rate of nuclear mitochondrial pseudogene contamination, particularly in species of the genus Panthera. In contrast, no nuclear mitochondrial pseudogene contamination was detected in members of Ursidae and Ailuridae. Phylogenetic analysis revealed multiple independent origins of nuclear mitochondrial pseudogene, with some tracing back to the common ancestor of Carnivora. To mitigate nuclear mitochondrial pseudogene-related errors, rigorous sequence verification strategies, such as sequence alignment and phylogenetic validation, should be implemented. In conclusion, our findings highlight the necessity of nuclear mitochondrial pseudogene awareness in genetic and evolutionary studies of Carnivora and other taxa.

Animals

The gastropod Lottia peitaihoensis as a model to study the body patterning of trochophore larvae.

The body patterning of trochophore larvae is important for understanding spiralian evolution and the origin of the bilateral body plan. However, considerable variations are observed among spiralian lineages, which have adopted varied strategies to develop trochophore larvae or even omit a trochophore stage. Some spiralians, such as patellogastropod mollusks, are suggested to exhibit ancestral traits by producing equal-cleaving fertilized eggs and possessing "typical" trochophore larvae. In recent years, we developed a potential model system using the patellogastropod Lottia peitaihoensis (= Lottia goshimai). Here, we introduce how the species were selected and establish sources and techniques, including gene knockdown, ectopic gene expression, and genome editing. Investigations on this species reveal essential aspects of trochophore body patterning, including organizer signaling, molecular and cellular processes connecting the various developmental functions of the organizer, the specification and behaviors of the endomesoderm and ectomesoderm, and the characteristic dorsoventral decoupling of Hox expression. These findings enrich the knowledge of trochophore body patterning and have important implications regarding the evolution of spiralians as well as bilateral body plans.

Animals

Genetic Tools in the Nakaseomyces clade for Evolutionary Comparisons of Signal Transduction Pathways.

The genus Nakaseomyces provides four species that are closely related but have different characteristics. For example, N. glabratus (formerly known as Candida glabrata) is a common human pathogen, whereas N. bracarensis and N. nivariensis have been isolated in clinical settings but are not common human pathogens. N. delphensis was isolated from fruit and there is no evidence it is pathogenic. Given the differences, we developed the clade as a molecular genetic system where we could introduce plasmids and assess transcriptional output from cloned promoters. We engineered a CRISPR/Cas9 plasmid that allows for rapid Gibson cloning of gRNAs, generated auxotrophic strains for amino acids and nucleotides, and introduced plasmids into each species. We used promoter-YFP plasmids to determine that while there are differences between the species, each species likely has intact thiamine and phosphate (THI and PHO) signal transduction pathways, and that gene expression in N. glabratus and N. bracarensis is more similar to one another than to the other two species. Finally, we determine that N. glabratus, N. bracarensis, and N. nivariensis persist in a murine macrophage for 24 h, whereas N. delphensis does not. This work describes new molecular tools for genetic manipulation in the Nakaseomyces clade and allows for evolutionary questions to be explored.

Signal Transduction

Chromosome scale genomes of two invasive Adelges species enable virtual screening for selective adelgicides.

Two invasive hemipteran adelgids are associated with widespread damage to several North American conifer species. Adelges tsugae, hemlock woolly adelgid, was introduced from Japan and reproduces parthenogenetically in North America, where it has rapidly decimated Tsuga canadensis and Tsuga caroliniana (the eastern and Carolina hemlocks, respectively). Adelges abietis, eastern spruce gall adelgid, introduced from Europe, forms distinctive pineapple-shaped galls on several native spruce species. While not considered a major forest pest, it weakens trees and increases susceptibility to additional stressors. Broad-spectrum insecticides that are often used to control adelgid populations can have off-target impacts on beneficial insects. Whole genome sequencing was performed on both species to aid in development of targeted solutions that may minimize ecological impact. Adelges abietis was sequenced using Illumina Linked-Read technology from 30 pooled individuals, with Hi-C scaffolding performed using data from a single individual collected from the same host plant. Adelges tsugae used Oxford Nanopore long-read sequencing from pooled nymphs. The assembled A. tsugae and A. abietis genomes, pooled from several parthenogenetic females, are 220.75 Mbp and 253.16 Mbp, respectively. Each consists of eight autosomal chromosomes, as well as two sex chromosomes (X1/X2), supporting the XX-XO sex determination system. The genomes are over 96% complete based on BUSCO assessment. Genome annotation identified 11,424 and 12,060 protein-coding genes in A. tsugae and A. abietis, respectively. Comparative analysis of proteins across 29 hemipteran species and 14 arthropod outgroups identified 31,666 putative gene families. Gene family evolution analysis with CAFE revealed lineage-specific expansions in immune-related aminopeptidases (ERAP1) and juvenile hormone binding proteins (JHBP), contractions in juvenile hormone acid methyltransferases (JHAMT), and conservation of nicotinic acetylcholine receptors (nAChR). These genes were explored as candidate families towards a long-term objective of developing adelgid-selective insecticides. Structural comparisons of proteins across seven focal species (Adelges tsugae, Adelges abietis, Adelges cooleyi, Rhopalosiphum maidis, Apis mellifera, Danaus plexippus, and Drosophila melanogaster) revealed high conservation of nAChR and ERAP1, while JHAMT exhibited species-specific structural divergence. The potential of JHAMT as a lineage-specific target for pest control was explored through virtual drug and pesticide screening.

adelgids

PlantPan: A comprehensive multi-species plant pan-genome database.

The pan-genome represents the complete genomic diversity of specific species, serving as a valuable resource for studying species evolution, crop domestication, and guiding crop breeding and improvement. While there are several single-species-specific plant pan-genome databases, the availability of multi-species pan-genome databases is limited. Additionally, variations in methods and data types used for plant pan-genome analysis across different databases hinder the comparison and integration of pan-genome information from various projects at multi-species or single-species levels. To tackle this challenge, we introduce PlantPan, a comprehensive database housing the results of pan-genome analysis for 195 genomes from 11 plant species. PlantPan aims to provide extensive information, including gene-centric and sequence-centric pan-genome information, graph-based pan-genome, pan-genome openness profiles, gene functions and its variation characteristics, homologous genes, and gene clusters across different species. Statistically, PlantPan incorporates 9 163 011 genes, 694 191 gene clusters, 526 973 370 genome variations, and 1 616 089 non-redundant genome variation groups at the species level, 33 455,098 genome synteny, and 177 827 non-redundant genome synteny groups at the species level. Regarding functional genes, PlantPan contains 5 222 720 genes related to transcription factors, 395 247 literature-reported resistance genes, 455 748 predicted microbial/disease resistance genes, and 1 612 112 genes related to molecular pathways. In summary, PlantPan is a vital platform for advancing the application of pan-genomes in molecular breeding for crops and evolutionary research for plants.

Genome, Plant

Predicting genome-wide functional constraints with GPN-Star.

Genomic language models have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences1. However, standard genomic language models adapted from natural language processing often require large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks2-4. Here we introduce a genomic pretrained network with species tree and alignment representations (GPN-Star), which is a biologically grounded genomic language model featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammal and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales show task-dependent advantages of modelling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms previous methods in prioritizing pathogenic and fine-mapped genome-wide association study variants, yields strong enrichments of complex trait heritability and improves power in rare variant association testing5. Extending beyond humans, we train GPN-Star for five model organisms-Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana-demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful and flexible tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95 % or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics