Search PubMedSearch

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Integrated genomic analysis of NF1-associated peripheral nerve sheath tumors: an updated biorepository dataset.

Neurofibromatosis type 1 (NF1) is an inherited neurocutaneous condition that predisposes to the development of peripheral nerve sheath tumors (PNST) including cutaneous neurofibromas (CNF), plexiform neurofibromas (PNF), atypical neurofibromatous neoplasms of uncertain biologic potential (ANNUBP), and malignant peripheral nerve sheath tumors (MPNST). The Johns Hopkins NF1 biospecimen repository promotes the successful advancement of therapeutic developments for NF1-associated PNST through acquisition and genomic analysis of human tumor specimens. RNA sequencing (RNAseq) and whole exome sequencing (WES) data were generated from 73 and 114 primary human tumor samples, respectively. These pre-processed data, standardized for immediate computational analysis, are accessible through the NF Data Portal, allowing immediate interrogation. This dataset combines new and previously released samples, offering a comprehensive view of the entire cohort sequenced. As a dedicated effort to systematically bank tumor samples from people with NF1, in collaboration with molecular geneticists and computational biologists, the Johns Hopkins NF1 biospecimen repository offers access to tissue samples and genomic data to promote the advancement of NF1-related tumor biologic insights and therapies.

Humans

16S rRNA and Metagenomic Datasets of Gastrointestinal Microbiota in Fetal and 7-Day-Old Goat Kids.

The perinatal period (from late gestation to the neonatal stage) in ruminants is a critical phase for fetal organ maturation, where ecological succession of gastrointestinal microbial communities significantly impacts livestock production efficiency. However, research remains insufficient regarding the distribution patterns and functional annotation of microbial communities across different gastrointestinal compartments during this period. This study characterized early microbiota dynamics in Hutianshi Goats using 16S rRNA sequencing (4 fetal goats at 90 ± 10 gestational days) and metagenomics (3 7-day-old goat kids). The fetal goat group generated 852,694 valid reads, yielding 688,277 high-quality reads after chimera removal for downstream analysis. The 7-day-old goat kids group produced 1,081,588,182 final valid reads, after data processing and assembly, 8,561,345 contigs were generated. Gene prediction identified 6,095,352 genes. Multi-database annotations (NR, KEGG, CAZy, etc.) revealed functional potential and antimicrobial resistance traits. The public release of this dataset facilitates academic understanding of microbial community dynamics and host-microbe interactions during this developmental stage, providing both theoretical foundations and data resources for ruminant developmental biology and precision breeding regulation.

Animals

Metagenomic and Transcriptomic Datasets of Plateau Brown Frogs (Rana kukunoris) from the Helan Mountains.

Global climate change has become a primary driving factor behind the biodiversity crisis in amphibians, making it crucial to understand how climate change affects species and their potential responses. The plateau brown frog (Rana kukunoris) is often regarded as an ideal ecological indicator species, yet research on its environmental adaptation mechanisms based on transcriptomic and microbiomic studies remains limited. Therefore, this study investigates the adaptation strategies of the plateau brown frog to environmental changes, providing extensive transcriptomic and the first comprehensive metagenomic dataset from two distinctly different environmental regions (eastern and western slopes of the Helan Mountains). We gathered transcriptomic data from three tissues (blood, liver, and muscle), resulting in 294,962 unigenes and 570,192 transcripts. Metagenomic sequencing identified major bacterial groups, including Firmicutes, Proteobacteria, Bacteroidetes, Spirochetes, and Actinobacteria. In summary, the results of this study can be used to further explore the associations among microbiota, host, and environment, which are crucial for comprehending the mechanisms of environmental adaptation in this species and contributing to the conservation of amphibian biodiversity.

Animals

Hierarchical modeling of tumor subtypes in cell lines using large-scale genomic datasets.

Cancer cell lines (CLs) are widely used to study tumor biology and drug response, yet their translational relevance is often limited by inaccurate subtype annotations. Existing CL-tumor matching approaches are frequently constrained by flat classification schemes, weak subtype definitions, and the exclusion of normal tissue references, leading to potential confounding of tumor-specific and tissue-of-origin signals. To address these limitations, a hierarchical classification (HC) framework is presented in which CLs are aligned with patient tumors across biological resolutions, from organ to molecular subtype. Gene expression profiles from 802 CLs, 5,612 tumors from The Cancer Genome Atlas (TCGA) , and 8,939 non-cancerous tissues were integrated to separate oncogenic signals from tissue-specific signals. Node-specific features were selected using maximum relevance minimum redundancy, and balanced accuracies of 89% in cross-validation and 75%, and 80% on external datasets were achieved. Through the framework, 43 CLs were reassigned, and clinically relevant underrepresented subtypes were identified.

cancer cell lines

Detecting and quantifying circular RNAs in terabyte-scale RNA-seq datasets with CIRI3.

To address recent challenges in circular RNA (circRNA) analysis, we present CIRI3, a tool for circRNA detection and quantification in terabyte-scale RNA-sequencing datasets. Using dynamic multithreaded task partitioning and a blocking search strategy for junction reads, CIRI3 is an order of magnitude faster than existing tools, while providing increased accuracy. We identified differentially spliced circRNAs across 2,535 cancer-related samples, and constructed a pretraining model and a biomarker network provided as the CIRIonco database.

RNA, Circular

BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.

MOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100× more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline).

Genome-Wide Association Study

polars-bio-fast, scalable, and out-of-core operations on large genomic interval datasets.

MOTIVATION: Genomic studies very often rely on computationally intensive analyses of relationships between features, which are typically represented as intervals along a 1D coordinate system (such as positions on a chromosome). In this context, the Python programming language is extensively used for manipulating and analyzing data stored in a tabular form of rows and columns, called a DataFrame. Pandas is the most widely used Python DataFrame package and has been criticized for inefficiencies and scalability issues, which its modern alternative-Polars-aims to address with a native backend written in the Rust programming language. RESULTS: polars-bio is a Python library that enables fast, parallel and out-of-core operations on large genomic interval datasets. Its main components are implemented in Rust, using the Apache DataFusion query engine and Apache Arrow for efficient data representation. It is compatible with Polars and Pandas DataFrame formats. In a real-world comparison (107 versus 1.2×106 intervals), our library runs overlap queries 6.5×, nearest queries 15.5×, count_overlaps queries 38×, and coverage queries 15× faster than Bioframe. On equally sized synthetic sets (107 versus 107), the corresponding speedups are 1.6×, 5.5×, 6×, and 6×. In streaming mode, on real and synthetic interval pairs, our implementation uses 90× and 15× less memory for overlap, 4.5× and 6.5× less for nearest, 60× and 12× less for count_overlaps, and 34× and 7× less for coverage than Bioframe. Multi-threaded benchmarks show good scalability characteristics. To the best of our knowledge, polars-bio is the most efficient single-node library for genomic interval DataFrames in Python. AVAILABILITY AND IMPLEMENTATION: polars-bio is an open-source Python package distributed under the Apache License available for major platforms, including Linux, macOS, and Windows in the PyPI registry. The online documentation is https://biodatageeks.org/polars-bio/ and the source code is available on GitHub: https://github.com/biodatageeks/polars-bio and Zenodo: https://doi.org/10.5281/zenodo.16374290. are available at Bioinformatics online.

Software

SpatialRNA: a Python package for easy application of Graph Neural Network models on single-molecule spatial transcriptomics dataset.

SUMMARY: Image-based spatial transcriptomics (iST) deliver gene expression measurements of RNA transcripts in tissue slices with single-molecule resolution and spatial context preserved. Modern Graph Neural Network (GNN) models are promising methods for capturing the complex molecular and cellular phenotypes in tissues at single-transcript and single-cell levels. A key application of GNNs is the detection of spatial domains or niches, that is, groups of molecules and/or cells that collaboratively work together to produce complex phenotypes. Due to the vast number of detected transcripts in (iST) dataset, applying GNNs on RNA molecule graphs is not trivial. We present a Python package, SpatialRNA, for easy (sub)graph generation from tissue samples and provide comprehensive tutorials for convenient and efficient application of Graph Neural Network models under the PyG framework. This highly scalable tool comprehensively segments tissue into spatial domains, aiding in biological interpretation of iST data and its underlying molecular microenvironments. AVAILABILITY AND IMPLEMENTATION: The SpatialRNA package is freely accessible from online repository https://github.com/ruqianl/spatialrna and can be installed via pip. Comprehensive tutorials, guidance on parameter selection, and complete workflows of case studies are available from the documentation website https://ruqianl.github.io/spatialrna_docs/, and uploaded on Zenodo with a DOI 10.5281/zenodo.17339575.

Neural Networks, Computer

Evaluating Language Models for Biomedical Fact-Checking: A Benchmark Dataset for Cancer Variant Interpretation Verification.

Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. To address this, we developed CIViC-Fact, a benchmark dataset and pipeline for testing automated systems that verify the accuracy of cancer variant claims. CIViC-Fact links structured claims to sentence-level supporting or refuting evidence from full-text articles, and includes expert annotations and explanations. We evaluated multiple language models. Proprietary models performed well without training, but a smaller open-source model, fine-tuned on CIViC-Fact, achieved the highest accuracy (89%). Applying our fact-checking pipeline to real CIViC entries showed that reviewing less than 20% of content, focusing on flagged entries, would be sufficient to catch over half of all errors. This AI-assisted triage greatly accelerates the review process without replacing or reducing expert insight, ensuring that existing careful oversight remains in place while curators can work more efficiently. CIViC-Fact provides a realistic, high-consequence framework for biomedical fact-checking and a path toward more rigorous and efficient knowledgebase curation.

Journal Article

Microbiome Datahub: an open-access platform integrating environmental metadata, taxonomy, and functional annotation for comprehensive metagenome-assembled genome datasets.

BACKGROUND: Metagenome-assembled genomes (MAGs) provide crucial insights into the genomic diversity of uncultured microbes. However, MAG datasets deposited in public repositories such as INSDC are often difficult to reuse due to heterogeneous quality, inconsistent taxonomic and functional annotations, and insufficiently curated environmental metadata. While secondary MAG databases such as MGnify, IMG/M, and SPIRE provide standardized resources, they reconstruct MAGs de novo from public metagenomic reads and therefore do not represent the original MAGs reported in publications. RESULTS: To address this gap, we developed Microbiome Datahub, an open-access platform that systematically aggregates and re-annotates original MAGs from INSDC. We collected 214,427 MAGs, predicted genes by DFAST, performed quality assessment with CheckM, standardized taxonomic assignments with GTDB-Tk, inferred 27 phenotypic traits using Bac2Feature, assigned proteins to MBGD ortholog clusters and KEGG Orthology IDs using PZLAST, and annotated environmental metadata with the Metagenome and Microbes Environmental Ontology. Across these MAGs, the average completeness was 80.5% and contamination 1.8%; notably, the most frequent values were&#x2009;>95% completeness and&#x2009;<1% contamination, indicating that the majority of MAGs are of high quality. Comparative analyses showed that Microbiome Datahub provides phylogenetically and environmentally diverse MAGs: while the majority originated from vertebrate gut environments, a substantial number were also recovered from other habitats such as groundwater, including nearly 10,000 MAGs from the Patescibacteria. Inference of 27 phenotypic traits, including optimum growth temperature, further revealed ecological differentiation across phyla. Protein clustering revealed 56 million identity 40% clusters, with the majority unique compared with MGnify and GlobDB, and&#x2009;~19% of proteins unassigned to MBGD ortholog clusters, underscoring their novelty. CONCLUSIONS: Microbiome Datahub integrates MAG genome sequences, gene and protein predictions, quality metrics, environmental and taxonomic annotations, ortholog cluster assignments, and phenotype predictions, all accessible via a web interface, API, and bulk downloads. By combining original MAGs with curated metadata and functional annotations, Microbiome Datahub constitutes a comprehensive and reusable resource that will accelerate microbiome and microbial genomics research. Video Abstract.

Metagenome

Alterations in ether lipid metabolism in obesity revealed by systems genomics of multi-omics datasets.

Ratios between two metabolites are sensitive indicators of metabolic changes. Lipidomic profiling studies have revealed that plasma ether lipids, a class of glycero- and glycerophospho-lipids with reported health benefits, are negatively associated with obesity. Here, we utilized lipid ratios as surrogate markers of lipid metabolism to explore the processes underlying the inverse relationship between ether lipid metabolism and obesity. Plasma lipidomics data from two independent human cohorts (n&#x2009;=&#x2009;10,339 and n&#x2009;=&#x2009;4,492) were integrated to assess the associations between 82 lipid ratios and obesity-related markers in males and females. Results were externally validated using mouse transcriptomics data from the Hybrid Mouse Diversity Panel (n&#x2009;=&#x2009;152-227 across 74 strains). Genome-wide association studies using imputed genotypes from a population cohort (n&#x2009;=&#x2009;4,492) were performed to examine the genetic architecture of the ratios. Findings showed that waist circumference (WC), body mass index, and waist-hip ratio were inversely associated with total plasmalogens relative to total phospholipids in both sexes. Ratios comprising product-substrate pairs positioned either side of enzymes involved in plasmalogen synthesis and degradation showed positive and negative associations with WC, respectively. Branched-chain fatty acids negatively correlated with WC, while omega-6 polyunsaturated fatty acids exhibited differing associations depending on their position within the pathway. Mouse transcriptomics corroborated these results. Genomics data showed strong associations between ratios containing choline-plasmalogens and single-nucleotide polymorphisms in the transmembrane protein 229B (TMEM229B) gene region. This work demonstrates the utility of lipid ratios in understanding lipid metabolism. By applying the ratios to multi-omic datasets, we identified alterations in enzymatic activity and genetic variants likely affecting ether lipid synthesis in obesity that could not have been obtained from lipidomics data alone. Additionally, we characterized a potential role for TMEM229B, offering new perspectives on ether lipid metabolism and regulation.

Humans

PotatoRTD and TomatoRTD: Comprehensive Reference Transcript Datasets for Accurate Transcriptome Analysis and Isoform Discovery.

Transcriptome annotations provide essential information on transcript locations, sequences and structures, including transcription start, end sites and splice junctions. They underpin key biological analyses such as gene and transcript quantification, and the study of transcriptional and post-transcriptional regulation, including alternative transcription initiation, polyadenylation and splicing. Accurate characterisation of transcript isoforms is critical for understanding how gene expression relates to functional protein products. However, for many species-including Solanaceae crops such as potato and tomato-current annotations suffer from limited isoform coverage, with tens or hundreds of thousands of splice junctions and transcript isoforms missing. This undermines the completeness and accuracy of transcript-level analyses. Here, by generating Iso-seq and RNA-seq on a range of tissues and samples, we have produced transcriptome annotations for both potato and tomato with improved coverage, diversity, accurate splice junctions, and transcript start and end sites. We have also made these high-quality resources accessible through genome browsers. These enhanced annotations will enable more accurate transcriptome analyses, supporting higher-resolution and novel biological discoveries.

Solanum tuberosum

Genetic Ancestry and Colorectal Cancer in the All of Us Dataset.

IMPORTANCE: Genetic ancestry may complement biological, behavioral, and clinical factors in understanding colorectal cancer (CRC) disparities; yet, ancestry-informed analyses in CRC remain limited. OBJECTIVE: To characterize associations of genetic ancestry with CRC burden, age at diagnosis, and age-specific risk, and to develop a multiethnic CRC risk-prediction model. DESIGN, SETTING, AND PARTICIPANTS: This retrospective cohort study used All of Us data from July 1986 to October 2023, with follow-up through last visit or death (median [IQR], 133.1 [57.1-186.5] months); analyses were conducted from February to June 2026. All of Us is a US research cohort with linked electronic health record (EHR) and short-read whole-genome sequencing (srWGS) data. All of Us Research Program participants with srWGS and linked EHR data were included, except those with hereditary polyposis or Lynch syndrome. EXPOSURES: Genetically inferred ancestry categories and principal components. MAIN OUTCOMES AND MEASURES: Any CRC was the primary outcome. Associations were evaluated using Fisher exact tests, cumulative incidence functions with Gray tests, cause-specific and Fine-Gray subdistribution hazard models, and pooled multivariable logistic regression. Prediction models used penalized least absolute shrinkage and selection operator and extreme gradient boosting (XGBoost). RESULTS: Among 316&#x202f;624 participants (median [IQR] age, 56.3 [40.2-68.2] years; 172&#x202f;327 [54.4%] of European ancestry; 191&#x202f;705 female [61.2%]; 121&#x202f;585 male [38.8%]), 2914 (0.9%) developed CRC. European ancestry was associated with higher odds of CRC vs all other ancestries combined (odds ratio, 1.50; 95% CI, 1.39-1.62). The median age at CRC diagnosis was older in European (63.4 [53.9-71.2] years) than in American admixed-Latino, African, East Asian, and Other ancestry groups. In cause-specific hazard models on the attained-age scale, American admixed-Latino (hazard ratio, 1.30; 95% CI, 1.14-1.47) and East Asian (hazard ratio, 1.43; 95% CI, 1.06-1.94) ancestry had higher age-specific CRC hazard than European ancestry, with consistent findings on the subdistribution scale accounting for competing death. The multiethnic XGBoost model performed best (receiver operating characteristic area under the curve, 0.898; 95% CI, 0.882-0.912; precision-recall area under the curve, 0.338; 95% CI, 0.296-0.379) and was well calibrated. CONCLUSIONS AND RELEVANCE: In this cohort study, genetic ancestry was associated with meaningful differences in CRC burden and age-specific risk. These findings suggest that a multiethnic XGBoost model may complement CRC screening as a risk-enrichment tool.

Aged

tidk: a toolkit to rapidly identify telomeric repeats from genomic datasets.

SUMMARY: "tidk" (short for telomere identification toolkit) uses a simple, fast algorithm to scan long DNA reads for the presence of short tandemly repeated DNA in runs, and to aggregate them based on canonical DNA string representation. These are telomeric repeat candidates. Our algorithm is shown to be accurate in genomes for which the telomeric repeat unit is known and is tested across a wide variety of newly assembled genomes to uncover new telomeric repeat units. Tools are provided to identify telomeric repeats de novo, scan genomes for known telomeric repeats, and to visualize telomeric repeats on the assembly. "tidk" is implemented in Rust and is available as a command line tool which can be compiled using the Rust toolchain or downloaded as a binary from bioconda. AVAILABILITY AND IMPLEMENTATION: The "tidk" Rust crate is freely available under the MIT license (https://crates.io/crates/tidk), and the source code is available at https://github.com/tolkit/telomeric-identifier.

Telomere

Meta-analysis models with group structure for pleiotropy detection at gene and variant level using summary statistics from multiple datasets.

Genome-wide association studies (GWASs) have highlighted the importance of pleiotropy in human diseases, where one gene can impact 2 or more unrelated traits. Examining shared genetic risk factors across multiple diseases can enhance our understanding of these conditions by pinpointing new genes and biological pathways involved. Furthermore, with an increasing wealth of GWAS summary statistics available to the scientific community, leveraging these findings across multiple phenotypes could unveil novel pleiotropic associations. Existing selection methods examine pleiotropic associations one by one at a scale of either the genetic variant or the gene, and thus cannot consider all the genetic information at the same time. To address this limitation, we propose a new approach called MPSG (Meta-analysis model adapted for Pleiotropy Selection with Group structure). This method performs a penalized multivariate meta-analysis method adapted for pleiotropy and takes into account the group structure information nested in the data to select relevant variants and genes (or pathways) from all the genetic information. To do so, we implemented an alternating direction method of multipliers algorithm. We compared the performance of the method with other benchmark meta-analysis approaches such as GCPBayes, PLACO, and ASSET by considering as inputs different kinds of summary statistics. We provide an application of our method to the identification of potential pleiotropic genes between breast and thyroid cancers.

Humans

Gene-environment interaction analysis in atopic eczema: evidence from large population datasets and modelling in vitro.

BACKGROUND: Environmental factors play a role in the pathogenesis of complex traits including atopic eczema (AE) and a greater understanding of gene-environment interactions (G*E) is needed to define pathomechanisms for disease prevention. We analysed data from 16 European studies to test for interaction between the 24 most significant AE-associated loci identified from genome-wide association studies and 18 early-life environmental factors. We tested for replication using a further 10 studies and in vitro modelling to independently assess findings. RESULTS: The discovery analysis showed suggestive evidence for interaction (p<0.05) between 7 environmental factors (antibiotic use, cat ownership, dog ownership, breastfeeding, elder sibling, smoking and washing practices) and at least one established variant for AE, 14 interactions in total (maxN=25,339). In replication analysis (maxN=252,040) dog exposure*rs10214237 (on chromosome 5p13.2 near IL7R) was nominally significant (ORinteraction=0.91 [0.83-0.99] P=0.025), with a risk effect of the T allele observed only in those not exposed to dogs. A similar interaction with rs10214237 was observed for siblings in the discovery analysis (ORinteraction=0.84[0.75-0.94] P=0.003), but replication analysis was under-powered ORinteraction=1.09[0.82-1.46]). Rs10214237 homozygous risk genotype is associated with lower IL-7R expression in human keratinocytes, and dog exposure modelled in vitro showed a differential response according to rs10214237 genotype. CONCLUSIONS: Interaction analysis and functional assessment provide evidence that early-life dog exposure may modify the genetic effect of rs10214237 on AE via IL7R, supporting observational epidemiology showing a protective effect for dog ownership. The lack of evidence for other G*E studied here implies that only weak effects are likely to occur.

Atopic eczema

Spatiotemporal genomic analysis and risk assessment of the plasmids carrying&#xa0;blaOXA-48-like genes based on a large-scale international dataset.

BACKGROUND: The spread of OXA-48-like carbapenemases represents a major public health challenge. Although previous studies have investigated OXA-48-like carbapenemases risk factors, nosocomial dissemination, and plasmid dynamics, an integrated plasmid-centered framework combining complete plasmid mining, transmission-unit analysis, phylogenetic reconstruction, and machine learning-based risk assessment remains limited. METHODS: We systematically collected 747 complete plasmid sequences carrying&#xa0;blaOXA-48-like genes from the NCBI database, establishing the largest collections of complete plasmid sequences to date. Using an integrative framework of population genomics, phylogenetic dating, and machine learning, this study aimed to characterize the dissemination patterns, plasmid replicon diversity, transmission units, mobile genetic elements, co-resistance profiles, and risk classification of these plasmid. RESULTS: Plasmids carrying&#xa0;blaOXA-48-like genes&#xa0;were detected across 50 countries on six continents, with blaOXA-48 predominating in Europe, blaOXA-181 in South Asia, and blaOXA-232 largely in Asia. IncL and ColKP3/IncX3 replicons, together with Tn1999.2 and other MGEs, were central drivers of plasmid maintenance and spread. Sixteen transmission units were defined, with AA068_Cluster3 estimated to have originated in the Netherlands around 2005 before expanding to Europe, the Middle East, Asia, and North America. Co-resistance analyses revealed frequent modules involving aminoglycoside and quinolone resistance, with qnrS1 and aph(3'')-Ib most prevalent. Notably, high-risk transposon structures were often identified in non-clinical environments, underscoring their cross-ecological transmission potential. Machine learning-based classification models showed good internal performance for predefined composite-risk categories, with plasmid mobility, clinical/non-clinical source composition, and host background contributing to the classification results. CONCLUSIONS: This study provides a large-scale plasmid-centered genomic analysis of publicly available complete plasmid sequences carrying&#xa0;blaOXA-48-like genes, integrating transmission-unit inference, phylogeographic reconstruction, mobile genetic element and co-resistance profiling, and composite genomic risk stratification. This gene-centered framework may support future One Health-oriented antimicrobial resistance surveillance and prioritization of plasmids with higher dissemination and resistance potential.

Plasmids