Search PubMedSearch

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Community-driven advances in computational mass spectrometry: The perspective of EuBIC-MS members.

Advances in data acquisition, artificial intelligence, and integrative bioinformatics are driving the rapid evolution of computational mass spectrometry, and in turn, transforming modern proteomics, metabolomics, and lipidomics. These developments have greatly increased the scale and complexity of mass spectrometry data, underscoring the importance of evolving accurate, transparent, efficient and reproducible data processing workflows. Addressing these challenges requires collaborative innovation that brings together expertise in software engineering, statistics, and biology. The European Bioinformatics Community for Mass Spectrometry (EuBIC-MS), an initiative of the European Proteomics Association (EuPA), fosters a culture of open, community-driven development through its biennial Developers Meetings and Winter Schools. This commentary summarizes the scientific background and outcomes of the EuBIC-MS Developers Meeting 2025, which took place in Novacella, Italy. Three keynote presentations highlighted major frontiers in the field: deep proteome and phosphoproteome profiling, text mining for protein-protein interaction extraction, and scalable proteomics for AI-driven drug discovery. Seven community-selected hackathons addressed emerging challenges such as single-cell proteomics data analysis, FAIR metadata extraction, deep learning frameworks, R-Python interoperability, and DIA validation. Together, these efforts demonstrate the potential for scientific and technical innovation to arise from open collaboration, and highlight how community-driven initiatives can accelerate progress in computational mass spectrometry. SIGNIFICANCE: Modern proteomics increasingly depends on computational advances to translate complex, high-dimensional data into biological knowledge. The EuBIC-MS Developers Meeting 2025 exemplifies how community-driven collaboration can directly accelerate this process by bringing together experts from bioinformatics, statistics, and experimental proteomics to co-develop open, interoperable, and reproducible analytical tools. By fostering shared software frameworks, transparent benchmarking, and collaborative problem solving, the EuBIC-MS community helps ensure that technological innovation translates into reliable biological insights. This collaborative model strengthens the foundation for quantitative, system-level understanding of proteomes and establishes a sustainable path for integrating artificial intelligence and next-generation data acquisition into routine biological discovery. This commentary shows some current highlights in the field of computational mass spectrometry and community-based approaches undertaken during the most recent Developers Meeting to solve these challenges. The approaches discussed and initiated during the meeting - ranging from deep proteome profiling and phosphosite mapping to text mining, single-cell data analysis, and FAIR metadata extraction - address key bottlenecks that currently limit the biological interpretability and comparability of proteomics data.

Mass Spectrometry

Effect of repeated mass drug administration on the transmission of yaws: a retrospective genomic epidemiology study.

BACKGROUND: Yaws, a neglected tropical disease caused by Treponema pallidum subspecies pertenue (T p pertenue), has evaded eradication, in part due to a high proportion of asymptomatic cases. Repeated mass drug administration (MDA), whereby an entire population is repeatedly treated irrespective of disease, could provide a solution. Here, we aimed to investigate the effect of MDA on the genomic epidemiology of T p pertenue. METHODS: We conducted a retrospective genomic epidemiology study on samples collected during a cluster-randomised trial of mass administration of azithromycin for yaws eradication in the Namatanai District of Papua New Guinea. Participants were in 38 wards (administrative units encompassing several villages) in three local-level government areas (LLGs). The experimental group received an initial round of MDA followed by two further rounds 6 months and 12 months after the first round. The control group received one round of MDA followed by two rounds of treatment targeting clinical cases and contacts only, on the same schedule as the MDA in the experimental group. A follow-up survey on both groups was done 18 months after the first MDA round. Swab samples were collected at each round from ulcerative and nodular skin lesions, and blood was collected by finger-prick for serological testing at 18 months. Metadata on ulcer size (cm) and duration (days) were recorded at each round, and treponemal and non-treponemal antibodies were recorded at 18 months. Samples from swabs positive for T p pertenue underwent library preparation and whole-genome sequencing. We examined the phylogenetic relationships between genomes, linking them with geospatial and patient metadata to understand the impact of MDA on T p pertenue diversity and transmission. FINDINGS: Swabs collected from 297 individuals with active yaws from April 30, 2018, to Nov 2, 2019, yielded 222 good-quality Tp pertenue genomes. We identified 20 sublineages of T p pertenue in the control group and 21 in the experimental group at the beginning of the study. At the end of the study, there were 13 sublineages in the control group and three in the experimental group, of which two persisted in both groups. Three sublineages not detected at baseline were observed in the control group after commencing MDA. The two sublineages that persisted in both groups had non-synonymous mutations in penicillin-binding proteins. One of these sublineages evolved macrolide resistance in three individuals and was associated with lowered treponemal antibody (p=0&#xb7;0036) and longer ulcer duration (p=0&#xb7;015). Despite the study taking place within a small island, sublineages were geographically clustered, with pairs of samples from the same ward (odds ratio 7&#xb7;1, 95% CI 5&#xb7;7-8&#xb7;8; p<0&#xb7;0001) or neighbouring wards (4&#xb7;3, 3&#xb7;3-5&#xb7;4; p<0&#xb7;0001) more likely to share the same sublineages compared with pairs from different LLGs. Additionally, older individuals were more likely to share sublineages than were younger individuals (1&#xb7;5, 1&#xb7;2-1&#xb7;9; p<0&#xb7;0001). INTERPRETATION: Repeated MDA was successful in reducing and maintaining the genetic diversity of T p pertenue at a low level but was associated with the development of macrolide resistance. Yaws re-emergence after MDA was attributed to multiple sublineages, of which the majority were detected in the population before MDA. Participants within the same ward were more likely to share sublineages than those that were more widely geographically separated, suggesting that re-emergence was driven by local transmission. These findings could inform future yaws elimination strategies. FUNDING: European Research Council, EU, Provincial Deputation of Barcelona, Barber&#xe0; Solid&#xe0;ria Foundation, Wellcome, and Fundaci&#xf3; "la Caixa".

Adolescent

Transition of Staphylococcus aureus tetracycline resistance plasmid pT181 from independent multicopy replicon to predominantly integrated chromosomal element over 65 years.

Mobile genetic elements (MGEs), including plasmids, phages and genome islands, are major sources of bacterial genetic diversity. The small plasmid pT181 confers tetracycline resistance in bacterial pathogen Staphylococcus aureus via an efflux pump, TetK. pT181 was one of the earliest sequenced S. aureus plasmids, and has been isolated in both clinical and livestock-associated strains for decades, both as an independent replicon and integrated in the chromosome as part of staphylococcal cassette chromosome mec (SCCmec). Bacterial genome analysis tools and high-quality sequences with metadata are publicly available, but these resources remain underleveraged for examining historical data, especially when studying the spread of MGEs across a species and over time. Using publicly available reads and metadata, we explored the evolution of pT181 over almost seven decades of samples to identify temporal trends in sequence evolution, copy number changes, and spread across S. aureus and beyond. pT181 was prevalent across S. aureus (found in 9.5% of 83,366 genomes tested), with a conserved sequence outside of three hypervariable regions. The history of pT181 since 1954 is characterized by spread across strains, significant variation in plasmid copy number of the independent replicon, and increasing frequency of integration of the plasmid into the S. aureus chromosome. We have identified multiple chromosomal integration locations of the plasmid, including outside of the previously characterized SCCmec. We find that pT181 has been transferred across staphylococcaceae and into a Gram-negative species. The repeated integration of pT181 into the chromosome may indicate co-evolution of the plasmid and the host, potentially to facilitate increased antibiotic resistance.

Journal Article

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models

Performance Profiles of Short DNA Barcode Segments for Family Level Detection of Asteraceae Within Asterales.

Short DNA barcodes may facilitate sequence recovery from degraded material, but their ability to retain target-family identity while excluding related taxa varies among genomic regions. We computationally evaluated 16 nuclear, plastid, and mitochondrial marker regions from 11 Asterales families using 279,956 NCBI locus-record matches and an accession-disjoint discovery/test design. Thirty-one candidate segments of 50-200 bp (mean, 98.55 bp) were screened in discovery data and evaluated for within-Asteraceae sequence recall, differentiation from non-Asteraceae Asterales, in silico primer behavior, phylogenetic placement, and exploratory matching across 808 metadata-defined metagenomic samples. Conserved regions such as matR and rbcL showed high within-Asteraceae identity, whereas ITS1, ITS, and trnH-psbA showed larger differences from related-family backgrounds; ITS2 and ycf1 showed intermediate profiles. Candidate segments were placed within or immediately adjacent to Asteraceae reference branches in segment-specific maximum-likelihood analyses, although support and topology varied among regions. Metadata-defined target-containing groups had higher mean query coverage and identity than background groups; because target presence was not independently verified and no classifier was fitted, these comparisons were descriptive and did not estimate diagnostic accuracy. Definitionally linked sequence statistics were interpreted as structural associations rather than evidence of causal evolutionary mechanisms. These results provide a family-level computational comparison of candidate short segments for Asteraceae detection within Asterales. Species identification, operational marker combinations, threshold robustness, and laboratory performance require validation using taxonomically dense, voucher-linked, and experimentally characterized datasets.

Asteraceae

SQLGEN: a framework for rapid client-server database application development.

SQLGEN is a framework for rapid client-server relational database application development. It relies on an active data dictionary on the client machine that stores metadata on one or more database servers to which the client may be connected. The dictionary generates dynamic Structured Query Language (SQL) to perform common database operations; it also stores information about the access rights of the user at log-in time, which is used to partially self-configure the behavior of the client to disable inappropriate user actions. SQLGEN uses a microcomputer database as the client to store metadata in relational form, to transiently capture server data in tables, and to allow rapid application prototyping followed by porting to client-server mode with modest effort. SQLGEN is currently used in several production biomedical databases.

Computer Communication Networks

Reporting and representation of population descriptors in public RNA-seq databases.

Diverse and globally representative datasets are essential to genomic science and medicine. Here, we analyzed population descriptor metadata from RNA sequencing (RNA-seq) studies in two major public repositories: the Sequence Read Archive (SRA) and the Database of Genotypes and Phenotypes. We examined geographic and economic characteristics of institutions depositing the data and compared SRA-deposited descriptors to empirical estimates of genetic ancestry and to those reported in publications, analyzing trends over time. We found that 55% of RNA-seq samples were deposited by United States (US) institutions and 90% by institutions in high-income countries. Only 3% of SRA samples were associated with population descriptors, and among those with US Census terms, 69% were labeled as White. Among samples with continental descriptors, 56% were labeled as European. Our analyses emphasize widespread bias in the composition of public RNA-seq datasets and, more generally, a lack of consistent and careful reporting of population descriptors needing urgent improvement.

Humans

Inclusion of Multi-Omic Biomarkers Improves Prediction Accuracy of Response, Relapse, and Overall Survival in Acute Myeloid Leukemia Patients Receiving High-Intensity Induction Chemotherapy.

BACKGROUND: Despite advancements in genetic markers for acute myeloid leukemia (AML) risk stratification, outcome prediction remains challenging due to disease heterogeneity and dynamic genetic changes, highlighting the need for reliable biomarkers to improve AML treatment strategies and patient outcomes. To refine outcome predictions, we investigated the use of microbial-derived biomarkers to predict composite complete remission (CRc), relapse, and survival for patients on high- and low-intensity regimens, and to integrate those variables into the widely clinically utilized European Leukemia Network (ELN-2022) genetic risk classification model for high-intensity-treated patients. METHODS: We first developed machine learning models that integrate baseline fecal metabolomics, 16S rRNA-based stool microbiome features, and clinical metadata (sex, antibiotic administration, AML somatic mutations, and cytogenetics) from two cohorts of AML patients (n&#x2009;=&#x2009;83) undergoing remission induction chemotherapy. Univariate tests and sparse canonical correlation analysis were employed for variable selection and to explore fecal metabolite-microbe relationships. A robust machine learning approach using XGBoost was employed, with 100 stratified data splits (80% training, 20% testing) and coarse-to-fine hyperparameter optimization. Variable importance was aggregated across all models to select key predictors. RESULTS: For high-intensity-treated patients, XGBoost models achieved aggregated AUROC scores of 0.719, 0.729, and 0.65 for CRc, relapse, and overall survival, respectively. For low-intensity-treated patients, these models achieved aggregate AUROC scores of 0.945, 0.724, and 0.768 for these same outcomes, respectively. Integrating the biomarkers identified in the high-intensity machine-learning models with the current ELN-2022 AML risk stratification system effectively stratified patients into risk categories, which obtained higher concordance indices and likelihood ratios, demonstrating improved prognostic accuracy for each outcome compared to ELN-2022 alone. CONCLUSIONS: The inclusion of microbial-derived biomarkers serves as a robust prognostic tool to improve outcome prediction in AML patients, highlighting the potential of its integration into AML risk assessment and paving the way for personalized treatment strategies and improved patient outcomes.

Humans

Radiology considerations for the PREMIUM study: a multicenter randomized controlled trial of abbreviated MRI versus ultrasound for liver cancer screening in cirrhosis.

This paper describes the rationale and radiology considerations in the implementation of the Preventing Liver Cancer Mortality through Imaging with Ultrasound versus MRI (PREMIUM) study. PREMIUM is a multicenter, randomized controlled trial sponsored by the Department of Veterans Affairs comparing dynamic contrast-enhanced (DCE) abbreviated MRI (aMRI) plus serum AFP versus ultrasound (US) plus serum AFP for hepatocellular carcinoma (HCC) screening in patients with cirrhosis. PREMIUM aims to randomize 4,700 participants across over 47 Veterans Affairs Medical Centers to semiannual surveillance for up to eight years, with HCC-related mortality as the primary endpoint. To date, 35 sites have been activated with 1,085 patients randomized. To ensure uniform implementation and reporting of per-protocol screening, the PREMIUM Radiology Workgroup developed standardized imaging protocols, structured LI-RADS-based reporting templates, and a centralized training program for radiologists and technologists. They also perform ongoing quality control on both scans and reports. The aMRI protocols utilize multiphasic post-contrast imaging to allow LI-RADS scoring. A non-contrast-enhanced aMRI protocol is available for participants who develop renal impairment or contrast allergy during the study. US protocols conform to US LI-RADS standards. Structured reporting promotes consistency in documentation of findings, visualization scores, and follow-up recommendations. A centralized Image Repository was established, incorporating advanced de-identification methods to remove metadata and pixel-embedded protected health information from imaging files. More than 20,000 curated liver MRI and US exams are anticipated, supporting both trial outcomes and future radiomics and artificial intelligence research. PREMIUM aims to determine whether screening for HCC with a DCE aMRI protocol reduces HCC-related mortality and also facilitates ancillary studies utilizing the Image Repository.

Abbreviated MRI

Integrated &#xb9;H-NMR Metabolomics and Growth Kinetics Uncover Three Distinct Metabolic Scenarios in Lactiplantibacillus pentosus P7 Fermentation of Plant-Derived Prebiotics.

Lactic acid bacteria (LAB) drive a broad range of food and biotechnological fermentations, the outcomes of which depend not only on the bacterial genotype but also on the chemical composition of the fermentation substrate. To resolve how a single strain reorganises chemically distinct plant matrices, we profiled fermentations of Lactiplantibacillus pentosus P7 (GenBank JBLMKZ000000000) on garlic, onion, and kiwifruit extracts prepared in water and 70% ethanol, using growth kinetics combined with solvent-suppressed 500-MHz proton nuclear magnetic resonance metabolomics over 48&#xa0;h, and integrated the data with whole-genome pathway annotations. Three substrate-specific metabolic scenarios emerged. On garlic, P7 grew vigorously, with the water extract exceeding the de Man-Rogosa-Sharpe reference medium at every time point (peak &#x394;OD&#x2086;&#x2080;&#x2080; of 9.38 versus 8.52 at 24&#xa0;h) and accumulating sorbose, rhamnose, and the aromatic amino acids phenylalanine and tryptophan (3.04- to 3.70-fold increases), providing first metabolic evidence consistent with the strain's four-copy aroE shikimate-dehydrogenase expansion. On onion, the lowest cell density coincided with the highest lactate output of the dataset (5.21-fold rise at 48&#xa0;h), transient 5-hydroxymethylfurfural reduction, and accumulation of acetoin and 1,3-propanediol, mapping onto a redundant set of pyridine-nucleotide-dependent oxidoreductases and a pdu-independent diol pathway. On kiwifruit, citrate accumulated 8.9-fold at 16&#xa0;h and then declined, consistent with an intact citCDEFG citrate-lyase operon paired with absence of canonical oxidative tricarboxylic acid enzymes. The optimal extraction solvent was substrate-dependent, water for garlic and ethanol for onion and kiwifruit. Overall, these results show that substrate chemistry, rather than strain identity, dictates which genome-encoded pathways P7 engages, establishing P7 as a versatile, substrate-tunable platform for the functional fermentation and biorefining of furanic-rich substrate streams. Raw NMR data and ISA-Tab metadata are available via MetaboLights with identifier MTBLS14463.

Lactiplantibacillus pentosus

Genetic structure correlates with ethnolinguistic diversity in eastern and southern Africa.

African populations are the most diverse in the world yet are sorely underrepresented in medical genetics research. Here, we examine the structure of African populations using genetic and comprehensive multi-generational ethnolinguistic data from the Neuropsychiatric Genetics of African Populations-Psychosis study (NeuroGAP-Psychosis) consisting of 900 individuals from Ethiopia, Kenya, South Africa, and Uganda. We find that self-reported language classifications meaningfully tag underlying genetic variation that would be missed with consideration of geography alone, highlighting the importance of culture in shaping genetic diversity. Leveraging our uniquely rich multi-generational ethnolinguistic metadata, we track language transmission through the pedigree, observing the disappearance of several languages in our cohort as well as notable shifts in frequency over three generations. We find suggestive evidence for the rate of language transmission in matrilineal groups having been higher than that for patrilineal ones. We highlight both the diversity of variation within Africa as well as how within-Africa variation can be informative for broader variant interpretation; many variants that are rare elsewhere are common in parts of Africa. The work presented here improves the understanding of the spectrum of genetic variation in African populations and highlights the enormous and complex genetic and ethnolinguistic diversity across Africa.

Africa, Southern

Clinical sequelae of gut microbiome development and disruption in hospitalized preterm infants.

Aberrant preterm infant gut microbiota assembly predisposes to early-life disorders and persistent health problems. Here, we characterize gut microbiome dynamics over the first 3&#xa0;months of life in 236 preterm infants hospitalized in three neonatal intensive care units using shotgun metagenomics of 2,512 stools and metatranscriptomics of 1,381 stools. Strain tracking, taxonomic and functional profiling, and comprehensive clinical metadata identify Enterobacteriaceae, enterococci, and staphylococci as primarily exploiting available niches to populate the gut microbiome. Clostridioides difficile lineages persist between individuals in single centers, and Staphylococcus epidermidis lineages persist within and, unexpectedly, between centers. Collectively, antibiotic and non-antibiotic medications influence gut microbiome composition to greater extents than maternal or baseline variables. Finally, we identify a persistent low-diversity gut microbiome in neonates who develop necrotizing enterocolitis after day of life 40. Overall, we comprehensively describe gut microbiome dynamics in response to medical interventions in preterm, hospitalized neonates.

Humans

Whole genome sequence data set of methicillin-resistant Staphylococcus aureus isolated from a milkman associated with cows with subclinical mastitis in Kiruhura district, Uganda.

The whole-genome sequence data set for methicillin-resistant Staphylococcus aureus, which was isolated from a milkman associated with cows with subclinical mastitis in the Kiruhura district of Uganda, is presented here. The assembled genome size was 2822,509 bp, with a 33% GC, 2 Contigs, a Contig N50 of 2818,424, and 1 Contig L50. You can access the genome sequence and related metadata at https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_056782255.1/. This dataset can be used again for resistance gene mapping, comparing genomic analysis, and comprehending genetic diversity among MRSA isolates from Ugandan milkmen.

Antimicrobial-resistant genes Staphylococcus aureu

Global reach and sustained engagement of a structured digital education program in medical mycology: an observational analysis of the 2025 ESCMID-EFISG webinar series.

OBJECTIVES: Evaluate the 2025 European Society of Clinical Microbiology and Infectious Diseases-European Fungal Infection Study Group webinar series to assess digital education as a scalable, equitable model for global professional development in medical mycology. METHODS: This observational study analyzed Zoom metadata across 17 webinars (January-December 2025). Metrics included registration, unique viewers, peak concurrent views, attendance rate, and duration. RESULTS: The series recorded 4631 registrations and 1372 unique participants. Median live attendance was 199 (interquartile range [IQR] 138-269), with a 39.3% attendance rate (IQR 33.7-47.5%) and peak concurrent viewership of 165 (IQR 106-229). Median session duration was 108 minutes. Webinars engaged a median of 60 countries (range 26-89) simultaneously, spanning 128 countries globally. Faculty comprised 67 unique experts from 23 countries with balanced gender representation (52.8% men, 47.2% women), of whom 16.4% (n = 11/67) were affiliated with institutions in low- and middle-income countries. CONCLUSION: Structured digital programs achieve wide global reach and sustained engagement. Strong participation in long-form sessions supports implementing Continuing Medical Education accreditation and unrestricted on-demand access to enhance global health equity.

Antimicrobial resistance

Spatiotemporal patterns of Rift Valley fever virus in Africa: a retrospective genomic epidemiology and phylodynamic modelling study.

BACKGROUND: Rift Valley fever virus (RVFV) is a mosquito-borne zoonotic pathogen causing outbreaks in humans and ruminants across Africa and the Arabian Peninsula. Originally restricted to the Great Rift Valley, RVFV has expanded geographically, prompting its classification by WHO as a pathogen of pandemic potential. We investigated the evolutionary and spatial dynamics of RVFV across Africa. METHODS: We used genomic data generated at the International Livestock Research Institute Nairobi genomic laboratory (BioProject PRJNA1106221) and combined with publicly available datasets retrieved from the National Center for Biotechnology (NCBI) GenBank nucleotide database. In retrieving RVFV genome sequences from the NCBI GenBank, we applied the search terms "Rift Valley fever virus segment L AND 6404[SLEN]", "Rift Valley fever virus segment M AND 3885[SLEN]", and "Rift Valley fever virus segment S AND 1520:1690[SLEN]" for L (Large), M (Medium), and S (Small) segments, respectively. For sequences without additional spatiotemporal information, we searched PubMed to extract the associated sequence metadata. We performed molecular clock analysis, phylogenetic inference, phylodynamic modelling (continuous phylogeographic reconstruction), and landscape phylogeography on the three RVFV genome segments (L, M, and S). We aimed to assess evolutionary rates, dispersal patterns, and environmental drivers. Focus was placed on lineage C, the most widely distributed variant. FINDINGS: The global dataset used in this study consisted of large (n=236), medium (n=237), and small (n=247), which were further filtered to exclude potential reassortants and vaccine strains. Genome sequences retrieved from NCBI GenBank database comprised large (n=180), medium (n=184), and small (n=202). The genome sequences from retrospective human and livestock isolates comprised large (n=56), medium (n=53), and small (n=45) collected in Burundi (2018), Kenya (2007, 2018, 2019, 2021, and 2022), and Rwanda (2018 and 2022). Our dataset revealed that RVFV exhibited low overall genetic diversity. Lineage C, however, showed evidence of active evolution, with substitution rates ranging from 3&#xb7;58&#x2009;&#xd7;&#x2009;10-4 to 9&#xb7;76&#x2009;&#xd7;&#x2009;10-4 substitutions per site per year. This lineage probably originated in Zimbabwe in the mid-1970s and has since expanded across eastern and southern Africa. Phylogeographic reconstructions revealed rapid spread, with diffusion coefficients exceeding 50&#x2009;000 km2 per year. INTERPRETATION: Lineage C appears capable of establishing endemic transmission in new regions, with ongoing diversification observed during interepidemic periods. These observations reinforce the value of continuous genomic surveillance, particularly during cryptic transmission phases when adaptive mutations might emerge. Although further evidence is needed, observed trends in climate variability and land-use change point to the potential benefit of targeted surveillance in settings that could be at increased risk, including urban centres and wetlands. FUNDING: This work was supported by the German Federal Ministry for Economic Cooperation and Development, the Rockefeller Foundation, and the Africa Centres for Disease Control and Prevention.

Rift Valley fever virus

Multidimensional prophage profiling of carbapenem-resistant Enterobacteriaceae in Thailand: a nationwide, multicentre, genomic study.

BACKGROUND: Prophages influence bacterial fitness, resistance, and evolution, yet their epidemiology remains poorly understood in carbapenem-resistant Enterobacteriaceae (CRE). In this nationwide study in Thailand, we aimed to describe prophage repertoires in clinical CRE isolates and to explore their potential relevance for molecular epidemiology. METHODS: We performed a nationwide, retrospective, genomic analysis of all CRE clinical isolates collected through our previous national surveillance study involving 11 hospitals in 11 provinces in Thailand between March 25, 2012, and Jul 21, 2017. Whole-genome sequencing data from 747 CRE isolates were analysed. Intact prophages were identified using PHAge Search Tool Enhanced Release (PHASTER) and clustered by nucleotide sequence similarity. Prophage profiles were compared across multilocus sequence types, carbapenemase genotypes, specimens, geography, and patient demographics (age and sex). FINDINGS: Of the included 747 CRE isolates, 170 (23%) were Escherichia coli and 577 (77%) were Klebsiella pneumoniae. 220 (29%) of 747 strains had been isolated from female patients and 264 (35%) from male patients; metadata on patient sex were missing for 263 (35%) isolates. The median patient age was 63 years (IQR 50-72). 71 (10%) of isolates were from blood, 283 (38%) from sputum, 284 (38%) from urine, and 109 (15%) from other specimens. 374 distinct prophage clusters were identified, with significantly more prophages per genome in K pneumoniae (mean 3&#xb7;01 [SD 1&#xb7;55]) than in E coli (1&#xb7;64 [1&#xb7;46]; p<0&#xb7;0001). Prophage repertoires largely mirrored bacterial multilocus sequence types. However, even within the highly clonal K pneumoniae sequence type 16 lineage, discrete prophage variation was identified, with common profiles observed in geographically dispersed patients. Respiratory K pneumoniae frequently carried a mosaic prophage with environmental signatures and a type VI secretion system, whereas blood-derived E coli harboured a prophage with immune-modulating genes. Distinct prophage clusters were observed across clinical specimens, age groups, carbapenemase genotype, and geographical region. Strains coharbouring blaNDM-1 plus blaOXA-232 (114 [15%] of 747) had the highest prophage loads. INTERPRETATION: The prophage content was shaped by the bacterial lineage, ecological niche, and temporal dynamics, providing an additional layer of epidemiological resolution beyond conventional genome typing. Integrating prophage profiling into molecular surveillance frameworks could help to identify transmission events, improve infectious source attribution, and enhance infection control strategies. FUNDING: Japan Agency for Medical Research and Development.

Female

Comparison of phylogenetic metrics of transmission between symptomatic and asymptomatic tuberculosis in individuals who were incarcerated in Brazil in 2008-24: a retrospective genomic epidemiology study.

BACKGROUND: Tuberculosis control efforts have traditionally targeted symptomatic individuals; however, the role of asymptomatic cases in sustaining transmission is increasingly recognised. We aimed to quantify the contribution of asymptomatic tuberculosis to recent transmission using genomic and epidemiological data from a high-transmission setting. METHODS: We conducted a retrospective genomic epidemiology study of Mycobacterium tuberculosis isolates collected in Mato Grosso do Sul, Brazil, between Aug 25, 2008, and March 19, 2024. Available isolates underwent whole-genome sequencing. Demographic, clinical, incarceration history, and laboratory metadata were obtained from surveillance records. From Jan 1, 2017, to March 19, 2024, active case finding was conducted in the state's three largest prisons (all male-only facilities), during which sputum samples were collected from individuals irrespective of symptoms and tested using GeneXpert and culture. Comparisons of transmission between individuals with and without symptoms were restricted to individuals who were incarcerated and were identified through active case finding and for whom high-quality, M tuberculosis lineage 4 genomes were available. Metrics of recent transmission included phylogenetic clustering, time-scaled haplotype density (THD), local branching index (LBI), and transmission probabilities inferred using Bayesian Reconstruction and Evolutionary Analysis of Transmission Histories. FINDINGS: 4448 tuberculosis cases were notified in Mato Grosso do Sul in 2008-24. After excluding cases for which M tuberculosis isolates were not available or had low sequencing quality, who had contaminated cultures or mixed infection, or who were infected with non-lineage 4 M tuberculosis, we included 2362 lineage 4 M tuberculosis isolates with high-quality genome sequences. 1849 (78&#xb7;3%) of 2362 isolates were part of a genomic cluster. Among 2362 individuals with tuberculosis, 1137 (48&#xb7;1%) were incarcerated at diagnosis. Of these individuals, 505 were identified through active case finding in three male-only prisons. The median age was 30 years (IQR 25-37); 304 (60&#xb7;2%) had mixed ethnicity, 90 (17&#xb7;8%) were White, 56 (11&#xb7;1%) were Black, 13 (2&#xb7;6%) were Indigenous, and six (1&#xb7;2%) were Asian. 277 (54&#xb7;9%) had symptomatic disease and 228 (45&#xb7;1%) had asymptomatic tuberculosis. There were no significant differences between symptomatic and asymptomatic individuals in phylogenetic clustering (213 [76&#xb7;9%] of 277 vs 195 [85&#xb7;5%] of 228; p=0&#xb7;37), THD (median 0&#xb7;39 [IQR 0&#xb7;06-0&#xb7;62] vs 0&#xb7;50 [0&#xb7;09-0&#xb7;65]; p=0&#xb7;12), or LBI (0&#xb7;00863 [0&#xb7;00810-0&#xb7;00988] vs 0&#xb7;00871 [0&#xb7;00829-0&#xb7;01020]; p=0&#xb7;088). Bayesian transmission trees showed no significant difference in the number of secondary infections inferred from symptomatic compared with asymptomatic individuals (p=0&#xb7;56). These findings were consistent across genomic clusters and robust to model assumptions. INTERPRETATION: We identified no differences in transmission between individuals who were symptomatic and those who were asymptomatic using multiple genomic measures. In this high-transmission setting, where systematic screening is implemented, our findings indicate that asymptomatic tuberculosis substantially contributes to tuberculosis transmission at the population level. These results suggest that symptom-based case detection alone is likely to be insufficient to interrupt transmission and highlight the importance of expanded screening strategies in high-risk populations. FUNDING: US National Institutes of Health and the Brazilian National Research Council (CNPq).

Humans

Meningioma transcriptomic landscape demonstrates novel subtypes with regional associated biology and patient outcome.

Meningiomas, although mostly benign, can be recurrent and fatal. World Health Organization (WHO) grading of the tumor does not always identify high-risk meningioma, and better characterizations of their aggressive biology are needed. To approach this problem, we combined 13 bulk RNA sequencing (RNA-seq) datasets to create a dimension-reduced reference landscape of 1,298 meningiomas. The clinical and genomic metadata effectively correlated with landscape regions, which led to the identification of meningioma subtypes with specific biological signatures. The time to recurrence also correlated with the map location. Further, we developed an algorithm that maps new patients onto this landscape, where the nearest neighbors predict outcome. This study highlights the utility of combining bulk transcriptomic datasets to visualize the complexity of tumor populations. Further, we provide an interactive tool for understanding the disease and predicting patient outcomes. This resource is accessible via the online tool Oncoscape, where the scientific community can explore the meningioma landscape.

Meningioma