Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Neuronal database integration: the Senselab EAV data model.

We discuss an approach towards integrating heterogeneous nervous system data using an augmented Entity-Attribute-Value (EAV) schema design. This approach, widely used in implementing electronic patient record systems (EPRSs), allows the physical schema of the database to be relatively immune to changes in domain knowledge. This is because new kinds of facts are added as data (or as metadata) rather than hard-coded as the names of newly created tables or columns. Because the domain knowledge is stored as metadata, a framework developed in one scientific domain can be ported to another with only modest revision. We describe our progress in creating a code framework that handles browsing and hyperlinking of the different kinds of data.

Databases, Factual↗

Representation by standard terminologies of health status concepts contained in two health status assessment instruments used in rheumatic disease management.

Health and functional status data have been shown to have clinical utility in predicting outcome. Various metadata registries in the form of patient self-administered health assessment questionnaires have been incorporated into routine clinical care and clinical research of patients with rheumatic disease. Examples of such health assessment instruments are the Clinical Health Assessment Questionnaire (CLINHAQ) and the Modified Health Assessment Questionnaire (MHAQ). These instruments contain concepts that are an integral part of the health and functional status domain. Using an automated indexing tool we examined the clinical content coverage by SNOMED RT and the Unified Medical Language System (UMLS) Metathesaurus for health and functional status concepts identified in the MHAQ and CLINHAQ. Significant differences existed between the overall representational ability of SNOMED and UMLS for concepts identified in the MHAQ (49%, vs. 77% respectively, p < .005) and for concepts identified in the CLINHAQ (30% vs. 64% respectively p < .005). Representational capability by SNOMED-RT and UMLS for concepts in a given health assessment instrument was carried across four semantic classes of "attitudes", "symptoms", "activities", and "social attributes". The conceptual content coverage of health status assessment concepts contained in the MHAQ and CLINHAQ by SNOMED-RT and UMLS was incomplete but better for UMLS with its panoply of vocabulary sources. This observed overall improved representation by UMLS appeared to be due to better representation of concepts in "activities" and "social attributes" semantic classes. Representation of health or functional status concepts in a computerized medical record should be founded on a universally agreed concept model of that domain. Established functional and health status metadata registries can serve as important sources for concepts and candidate classes within that domain.

Health Status↗

IML: An image markup language.

Image Markup Language is an extensible markup language (XML) schema used to describe both image metadata and annotations. It describes both data pertaining to an entire image, and data that are tied to specific regions or features of the image. Developed for a specific domain in Medical Education, this pa-per describes extensions to take advantage of the Dublin Core metadata standard, and of an XML schema for vector graphics representation. We have developed a prototype system of open source tools implementing an authoring system, a client system, and an image annotation database which can be queried though the Web.

Diagnostic Imaging↗

MeSHmap: a text mining tool for MEDLINE.

Our research goal is to explore text mining from the metadata included in MEDLINE documents. We present MeSHmap our prototype text mining system that exploits the MeSH indexing accompanying MEDLINE records. MeSHmap supports searches via PubMed followed by user driven exploration of the MeSH terms and subheadings in the retrieved set. The potential of the system goes beyond text retrieval. It may also be used to compare entities of the same type such as pairs of drugs or pairs of procedures etc. In addition there is the potential to generate maps of entities (drugs or diseases etc.) such that the strength of the link between two entities in the map represents their similarity as expressed in the MeSH metadata of the MEDLINE documents. Higher level operators have been proposed to support these comparison and mapping functions. This paper motivates and describes MeSHmap. Future work will include user evaluations of the system.

Abstracting and Indexing↗

SQLGEN: a framework for rapid client-server database application development.

SQLGEN is a framework for rapid client-server relational database application development. It relies on an active data dictionary on the client machine that stores metadata on one or more database servers to which the client may be connected. The dictionary generates dynamic Structured Query Language (SQL) to perform common database operations; it also stores information about the access rights of the user at log-in time, which is used to partially self-configure the behavior of the client to disable inappropriate user actions. SQLGEN uses a microcomputer database as the client to store metadata in relational form, to transiently capture server data in tables, and to allow rapid application prototyping followed by porting to client-server mode with modest effort. SQLGEN is currently used in several production biomedical databases.

Computer Communication Networks↗

Automatic query mapping among genomic databases: a pilot exploration.

As databases in the human genome project proliferate, it is important for users of one genomic database to identify similar or inconsistent data in other autonomously developed genomic databases. To do so, the user needs to issue the same query across multiple databases. We describe an approach that allows a query issued against one database to be automatically mapped to an equivalent query against another structurally different database. Our approach features two components: 1) a database designed to capture knowledge (metadata) that describes the correspondences among individual database components and 2) a module that utilizes the metadata to perform query mappings. As a demonstration, we apply our query mapping approach to two chromosome map databases (DB/12 and GDB).

Algorithms↗

Reporting and representation of population descriptors in public RNA-seq databases.

Diverse and globally representative datasets are essential to genomic science and medicine. Here, we analyzed population descriptor metadata from RNA sequencing (RNA-seq) studies in two major public repositories: the Sequence Read Archive (SRA) and the Database of Genotypes and Phenotypes. We examined geographic and economic characteristics of institutions depositing the data and compared SRA-deposited descriptors to empirical estimates of genetic ancestry and to those reported in publications, analyzing trends over time. We found that 55% of RNA-seq samples were deposited by United States (US) institutions and 90% by institutions in high-income countries. Only 3% of SRA samples were associated with population descriptors, and among those with US Census terms, 69% were labeled as White. Among samples with continental descriptors, 56% were labeled as European. Our analyses emphasize widespread bias in the composition of public RNA-seq datasets and, more generally, a lack of consistent and careful reporting of population descriptors needing urgent improvement.

Humans↗

Inclusion of Multi-Omic Biomarkers Improves Prediction Accuracy of Response, Relapse, and Overall Survival in Acute Myeloid Leukemia Patients Receiving High-Intensity Induction Chemotherapy.

BACKGROUND: Despite advancements in genetic markers for acute myeloid leukemia (AML) risk stratification, outcome prediction remains challenging due to disease heterogeneity and dynamic genetic changes, highlighting the need for reliable biomarkers to improve AML treatment strategies and patient outcomes. To refine outcome predictions, we investigated the use of microbial-derived biomarkers to predict composite complete remission (CRc), relapse, and survival for patients on high- and low-intensity regimens, and to integrate those variables into the widely clinically utilized European Leukemia Network (ELN-2022) genetic risk classification model for high-intensity-treated patients. METHODS: We first developed machine learning models that integrate baseline fecal metabolomics, 16S rRNA-based stool microbiome features, and clinical metadata (sex, antibiotic administration, AML somatic mutations, and cytogenetics) from two cohorts of AML patients (n&#x2009;=&#x2009;83) undergoing remission induction chemotherapy. Univariate tests and sparse canonical correlation analysis were employed for variable selection and to explore fecal metabolite-microbe relationships. A robust machine learning approach using XGBoost was employed, with 100 stratified data splits (80% training, 20% testing) and coarse-to-fine hyperparameter optimization. Variable importance was aggregated across all models to select key predictors. RESULTS: For high-intensity-treated patients, XGBoost models achieved aggregated AUROC scores of 0.719, 0.729, and 0.65 for CRc, relapse, and overall survival, respectively. For low-intensity-treated patients, these models achieved aggregate AUROC scores of 0.945, 0.724, and 0.768 for these same outcomes, respectively. Integrating the biomarkers identified in the high-intensity machine-learning models with the current ELN-2022 AML risk stratification system effectively stratified patients into risk categories, which obtained higher concordance indices and likelihood ratios, demonstrating improved prognostic accuracy for each outcome compared to ELN-2022 alone. CONCLUSIONS: The inclusion of microbial-derived biomarkers serves as a robust prognostic tool to improve outcome prediction in AML patients, highlighting the potential of its integration into AML risk assessment and paving the way for personalized treatment strategies and improved patient outcomes.

Humans↗

The BioImage Database Project: organizing multidimensional biological images in an object-relational database.

The BioImage Database Project collects and structures multidimensional data sets recorded by various microscopic techniques relevant to modern life sciences. It provides, as precisely as possible, the circumstances in which the sample was prepared and the data were recorded. It grants access to the actual data and maintains links between related data sets. In order to promote the interdisciplinary approach of modern science, it offers a large set of key words, which covers essentially all aspects of microscopy. Nonspecialists can, therefore, access and retrieve significant information recorded and submitted by specialists in other areas. A key issue of the undertaking is to exploit the available technology and to provide a well-defined yet flexible structure for dealing with data. Its pivotal element is, therefore, a modern object relational database that structures the metadata and ameliorates the provision of a complete service. The BioImage database can be accessed through the Internet.

Copyright↗

Radiology considerations for the PREMIUM study: a multicenter randomized controlled trial of abbreviated MRI versus ultrasound for liver cancer screening in cirrhosis.

This paper describes the rationale and radiology considerations in the implementation of the Preventing Liver Cancer Mortality through Imaging with Ultrasound versus MRI (PREMIUM) study. PREMIUM is a multicenter, randomized controlled trial sponsored by the Department of Veterans Affairs comparing dynamic contrast-enhanced (DCE) abbreviated MRI (aMRI) plus serum AFP versus ultrasound (US) plus serum AFP for hepatocellular carcinoma (HCC) screening in patients with cirrhosis. PREMIUM aims to randomize 4,700 participants across over 47 Veterans Affairs Medical Centers to semiannual surveillance for up to eight years, with HCC-related mortality as the primary endpoint. To date, 35 sites have been activated with 1,085 patients randomized. To ensure uniform implementation and reporting of per-protocol screening, the PREMIUM Radiology Workgroup developed standardized imaging protocols, structured LI-RADS-based reporting templates, and a centralized training program for radiologists and technologists. They also perform ongoing quality control on both scans and reports. The aMRI protocols utilize multiphasic post-contrast imaging to allow LI-RADS scoring. A non-contrast-enhanced aMRI protocol is available for participants who develop renal impairment or contrast allergy during the study. US protocols conform to US LI-RADS standards. Structured reporting promotes consistency in documentation of findings, visualization scores, and follow-up recommendations. A centralized Image Repository was established, incorporating advanced de-identification methods to remove metadata and pixel-embedded protected health information from imaging files. More than 20,000 curated liver MRI and US exams are anticipated, supporting both trial outcomes and future radiomics and artificial intelligence research. PREMIUM aims to determine whether screening for HCC with a DCE aMRI protocol reduces HCC-related mortality and also facilitates ancillary studies utilizing the Image Repository.

Abbreviated MRI↗

Integrated &#xb9;H-NMR Metabolomics and Growth Kinetics Uncover Three Distinct Metabolic Scenarios in Lactiplantibacillus pentosus P7 Fermentation of Plant-Derived Prebiotics.

Lactic acid bacteria (LAB) drive a broad range of food and biotechnological fermentations, the outcomes of which depend not only on the bacterial genotype but also on the chemical composition of the fermentation substrate. To resolve how a single strain reorganises chemically distinct plant matrices, we profiled fermentations of Lactiplantibacillus pentosus P7 (GenBank JBLMKZ000000000) on garlic, onion, and kiwifruit extracts prepared in water and 70% ethanol, using growth kinetics combined with solvent-suppressed 500-MHz proton nuclear magnetic resonance metabolomics over 48&#xa0;h, and integrated the data with whole-genome pathway annotations. Three substrate-specific metabolic scenarios emerged. On garlic, P7 grew vigorously, with the water extract exceeding the de Man-Rogosa-Sharpe reference medium at every time point (peak &#x394;OD&#x2086;&#x2080;&#x2080; of 9.38 versus 8.52 at 24&#xa0;h) and accumulating sorbose, rhamnose, and the aromatic amino acids phenylalanine and tryptophan (3.04- to 3.70-fold increases), providing first metabolic evidence consistent with the strain's four-copy aroE shikimate-dehydrogenase expansion. On onion, the lowest cell density coincided with the highest lactate output of the dataset (5.21-fold rise at 48&#xa0;h), transient 5-hydroxymethylfurfural reduction, and accumulation of acetoin and 1,3-propanediol, mapping onto a redundant set of pyridine-nucleotide-dependent oxidoreductases and a pdu-independent diol pathway. On kiwifruit, citrate accumulated 8.9-fold at 16&#xa0;h and then declined, consistent with an intact citCDEFG citrate-lyase operon paired with absence of canonical oxidative tricarboxylic acid enzymes. The optimal extraction solvent was substrate-dependent, water for garlic and ethanol for onion and kiwifruit. Overall, these results show that substrate chemistry, rather than strain identity, dictates which genome-encoded pathways P7 engages, establishing P7 as a versatile, substrate-tunable platform for the functional fermentation and biorefining of furanic-rich substrate streams. Raw NMR data and ISA-Tab metadata are available via MetaboLights with identifier MTBLS14463.

Lactiplantibacillus pentosus↗

Genetic structure correlates with ethnolinguistic diversity in eastern and southern Africa.

African populations are the most diverse in the world yet are sorely underrepresented in medical genetics research. Here, we examine the structure of African populations using genetic and comprehensive multi-generational ethnolinguistic data from the Neuropsychiatric Genetics of African Populations-Psychosis study (NeuroGAP-Psychosis) consisting of 900 individuals from Ethiopia, Kenya, South Africa, and Uganda. We find that self-reported language classifications meaningfully tag underlying genetic variation that would be missed with consideration of geography alone, highlighting the importance of culture in shaping genetic diversity. Leveraging our uniquely rich multi-generational ethnolinguistic metadata, we track language transmission through the pedigree, observing the disappearance of several languages in our cohort as well as notable shifts in frequency over three generations. We find suggestive evidence for the rate of language transmission in matrilineal groups having been higher than that for patrilineal ones. We highlight both the diversity of variation within Africa as well as how within-Africa variation can be informative for broader variant interpretation; many variants that are rare elsewhere are common in parts of Africa. The work presented here improves the understanding of the spectrum of genetic variation in African populations and highlights the enormous and complex genetic and ethnolinguistic diversity across Africa.

Africa, Southern↗

Clinical sequelae of gut microbiome development and disruption in hospitalized preterm infants.

Aberrant preterm infant gut microbiota assembly predisposes to early-life disorders and persistent health problems. Here, we characterize gut microbiome dynamics over the first 3&#xa0;months of life in 236 preterm infants hospitalized in three neonatal intensive care units using shotgun metagenomics of 2,512 stools and metatranscriptomics of 1,381 stools. Strain tracking, taxonomic and functional profiling, and comprehensive clinical metadata identify Enterobacteriaceae, enterococci, and staphylococci as primarily exploiting available niches to populate the gut microbiome. Clostridioides difficile lineages persist between individuals in single centers, and Staphylococcus epidermidis lineages persist within and, unexpectedly, between centers. Collectively, antibiotic and non-antibiotic medications influence gut microbiome composition to greater extents than maternal or baseline variables. Finally, we identify a persistent low-diversity gut microbiome in neonates who develop necrotizing enterocolitis after day of life 40. Overall, we comprehensively describe gut microbiome dynamics in response to medical interventions in preterm, hospitalized neonates.

Humans↗

Whole genome sequence data set of methicillin-resistant Staphylococcus aureus isolated from a milkman associated with cows with subclinical mastitis in Kiruhura district, Uganda.

The whole-genome sequence data set for methicillin-resistant Staphylococcus aureus, which was isolated from a milkman associated with cows with subclinical mastitis in the Kiruhura district of Uganda, is presented here. The assembled genome size was 2822,509 bp, with a 33% GC, 2 Contigs, a Contig N50 of 2818,424, and 1 Contig L50. You can access the genome sequence and related metadata at https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_056782255.1/. This dataset can be used again for resistance gene mapping, comparing genomic analysis, and comprehending genetic diversity among MRSA isolates from Ugandan milkmen.

Antimicrobial-resistant genes Staphylococcus aureu↗

Global reach and sustained engagement of a structured digital education program in medical mycology: an observational analysis of the 2025 ESCMID-EFISG webinar series.

OBJECTIVES: Evaluate the 2025 European Society of Clinical Microbiology and Infectious Diseases-European Fungal Infection Study Group webinar series to assess digital education as a scalable, equitable model for global professional development in medical mycology. METHODS: This observational study analyzed Zoom metadata across 17 webinars (January-December 2025). Metrics included registration, unique viewers, peak concurrent views, attendance rate, and duration. RESULTS: The series recorded 4631 registrations and 1372 unique participants. Median live attendance was 199 (interquartile range [IQR] 138-269), with a 39.3% attendance rate (IQR 33.7-47.5%) and peak concurrent viewership of 165 (IQR 106-229). Median session duration was 108 minutes. Webinars engaged a median of 60 countries (range 26-89) simultaneously, spanning 128 countries globally. Faculty comprised 67 unique experts from 23 countries with balanced gender representation (52.8% men, 47.2% women), of whom 16.4% (n = 11/67) were affiliated with institutions in low- and middle-income countries. CONCLUSION: Structured digital programs achieve wide global reach and sustained engagement. Strong participation in long-form sessions supports implementing Continuing Medical Education accreditation and unrestricted on-demand access to enhance global health equity.

Antimicrobial resistance↗

Spatiotemporal patterns of Rift Valley fever virus in Africa: a retrospective genomic epidemiology and phylodynamic modelling study.

BACKGROUND: Rift Valley fever virus (RVFV) is a mosquito-borne zoonotic pathogen causing outbreaks in humans and ruminants across Africa and the Arabian Peninsula. Originally restricted to the Great Rift Valley, RVFV has expanded geographically, prompting its classification by WHO as a pathogen of pandemic potential. We investigated the evolutionary and spatial dynamics of RVFV across Africa. METHODS: We used genomic data generated at the International Livestock Research Institute Nairobi genomic laboratory (BioProject PRJNA1106221) and combined with publicly available datasets retrieved from the National Center for Biotechnology (NCBI) GenBank nucleotide database. In retrieving RVFV genome sequences from the NCBI GenBank, we applied the search terms "Rift Valley fever virus segment L AND 6404[SLEN]", "Rift Valley fever virus segment M AND 3885[SLEN]", and "Rift Valley fever virus segment S AND 1520:1690[SLEN]" for L (Large), M (Medium), and S (Small) segments, respectively. For sequences without additional spatiotemporal information, we searched PubMed to extract the associated sequence metadata. We performed molecular clock analysis, phylogenetic inference, phylodynamic modelling (continuous phylogeographic reconstruction), and landscape phylogeography on the three RVFV genome segments (L, M, and S). We aimed to assess evolutionary rates, dispersal patterns, and environmental drivers. Focus was placed on lineage C, the most widely distributed variant. FINDINGS: The global dataset used in this study consisted of large (n=236), medium (n=237), and small (n=247), which were further filtered to exclude potential reassortants and vaccine strains. Genome sequences retrieved from NCBI GenBank database comprised large (n=180), medium (n=184), and small (n=202). The genome sequences from retrospective human and livestock isolates comprised large (n=56), medium (n=53), and small (n=45) collected in Burundi (2018), Kenya (2007, 2018, 2019, 2021, and 2022), and Rwanda (2018 and 2022). Our dataset revealed that RVFV exhibited low overall genetic diversity. Lineage C, however, showed evidence of active evolution, with substitution rates ranging from 3&#xb7;58&#x2009;&#xd7;&#x2009;10-4 to 9&#xb7;76&#x2009;&#xd7;&#x2009;10-4 substitutions per site per year. This lineage probably originated in Zimbabwe in the mid-1970s and has since expanded across eastern and southern Africa. Phylogeographic reconstructions revealed rapid spread, with diffusion coefficients exceeding 50&#x2009;000 km2 per year. INTERPRETATION: Lineage C appears capable of establishing endemic transmission in new regions, with ongoing diversification observed during interepidemic periods. These observations reinforce the value of continuous genomic surveillance, particularly during cryptic transmission phases when adaptive mutations might emerge. Although further evidence is needed, observed trends in climate variability and land-use change point to the potential benefit of targeted surveillance in settings that could be at increased risk, including urban centres and wetlands. FUNDING: This work was supported by the German Federal Ministry for Economic Cooperation and Development, the Rockefeller Foundation, and the Africa Centres for Disease Control and Prevention.

Rift Valley fever virus↗

Multidimensional prophage profiling of carbapenem-resistant Enterobacteriaceae in Thailand: a nationwide, multicentre, genomic study.

BACKGROUND: Prophages influence bacterial fitness, resistance, and evolution, yet their epidemiology remains poorly understood in carbapenem-resistant Enterobacteriaceae (CRE). In this nationwide study in Thailand, we aimed to describe prophage repertoires in clinical CRE isolates and to explore their potential relevance for molecular epidemiology. METHODS: We performed a nationwide, retrospective, genomic analysis of all CRE clinical isolates collected through our previous national surveillance study involving 11 hospitals in 11 provinces in Thailand between March 25, 2012, and Jul 21, 2017. Whole-genome sequencing data from 747 CRE isolates were analysed. Intact prophages were identified using PHAge Search Tool Enhanced Release (PHASTER) and clustered by nucleotide sequence similarity. Prophage profiles were compared across multilocus sequence types, carbapenemase genotypes, specimens, geography, and patient demographics (age and sex). FINDINGS: Of the included 747 CRE isolates, 170 (23%) were Escherichia coli and 577 (77%) were Klebsiella pneumoniae. 220 (29%) of 747 strains had been isolated from female patients and 264 (35%) from male patients; metadata on patient sex were missing for 263 (35%) isolates. The median patient age was 63 years (IQR 50-72). 71 (10%) of isolates were from blood, 283 (38%) from sputum, 284 (38%) from urine, and 109 (15%) from other specimens. 374 distinct prophage clusters were identified, with significantly more prophages per genome in K pneumoniae (mean 3&#xb7;01 [SD 1&#xb7;55]) than in E coli (1&#xb7;64 [1&#xb7;46]; p<0&#xb7;0001). Prophage repertoires largely mirrored bacterial multilocus sequence types. However, even within the highly clonal K pneumoniae sequence type 16 lineage, discrete prophage variation was identified, with common profiles observed in geographically dispersed patients. Respiratory K pneumoniae frequently carried a mosaic prophage with environmental signatures and a type VI secretion system, whereas blood-derived E coli harboured a prophage with immune-modulating genes. Distinct prophage clusters were observed across clinical specimens, age groups, carbapenemase genotype, and geographical region. Strains coharbouring blaNDM-1 plus blaOXA-232 (114 [15%] of 747) had the highest prophage loads. INTERPRETATION: The prophage content was shaped by the bacterial lineage, ecological niche, and temporal dynamics, providing an additional layer of epidemiological resolution beyond conventional genome typing. Integrating prophage profiling into molecular surveillance frameworks could help to identify transmission events, improve infectious source attribution, and enhance infection control strategies. FUNDING: Japan Agency for Medical Research and Development.

Female↗

Comparison of phylogenetic metrics of transmission between symptomatic and asymptomatic tuberculosis in individuals who were incarcerated in Brazil in 2008-24: a retrospective genomic epidemiology study.

BACKGROUND: Tuberculosis control efforts have traditionally targeted symptomatic individuals; however, the role of asymptomatic cases in sustaining transmission is increasingly recognised. We aimed to quantify the contribution of asymptomatic tuberculosis to recent transmission using genomic and epidemiological data from a high-transmission setting. METHODS: We conducted a retrospective genomic epidemiology study of Mycobacterium tuberculosis isolates collected in Mato Grosso do Sul, Brazil, between Aug 25, 2008, and March 19, 2024. Available isolates underwent whole-genome sequencing. Demographic, clinical, incarceration history, and laboratory metadata were obtained from surveillance records. From Jan 1, 2017, to March 19, 2024, active case finding was conducted in the state's three largest prisons (all male-only facilities), during which sputum samples were collected from individuals irrespective of symptoms and tested using GeneXpert and culture. Comparisons of transmission between individuals with and without symptoms were restricted to individuals who were incarcerated and were identified through active case finding and for whom high-quality, M tuberculosis lineage 4 genomes were available. Metrics of recent transmission included phylogenetic clustering, time-scaled haplotype density (THD), local branching index (LBI), and transmission probabilities inferred using Bayesian Reconstruction and Evolutionary Analysis of Transmission Histories. FINDINGS: 4448 tuberculosis cases were notified in Mato Grosso do Sul in 2008-24. After excluding cases for which M tuberculosis isolates were not available or had low sequencing quality, who had contaminated cultures or mixed infection, or who were infected with non-lineage 4 M tuberculosis, we included 2362 lineage 4 M tuberculosis isolates with high-quality genome sequences. 1849 (78&#xb7;3%) of 2362 isolates were part of a genomic cluster. Among 2362 individuals with tuberculosis, 1137 (48&#xb7;1%) were incarcerated at diagnosis. Of these individuals, 505 were identified through active case finding in three male-only prisons. The median age was 30 years (IQR 25-37); 304 (60&#xb7;2%) had mixed ethnicity, 90 (17&#xb7;8%) were White, 56 (11&#xb7;1%) were Black, 13 (2&#xb7;6%) were Indigenous, and six (1&#xb7;2%) were Asian. 277 (54&#xb7;9%) had symptomatic disease and 228 (45&#xb7;1%) had asymptomatic tuberculosis. There were no significant differences between symptomatic and asymptomatic individuals in phylogenetic clustering (213 [76&#xb7;9%] of 277 vs 195 [85&#xb7;5%] of 228; p=0&#xb7;37), THD (median 0&#xb7;39 [IQR 0&#xb7;06-0&#xb7;62] vs 0&#xb7;50 [0&#xb7;09-0&#xb7;65]; p=0&#xb7;12), or LBI (0&#xb7;00863 [0&#xb7;00810-0&#xb7;00988] vs 0&#xb7;00871 [0&#xb7;00829-0&#xb7;01020]; p=0&#xb7;088). Bayesian transmission trees showed no significant difference in the number of secondary infections inferred from symptomatic compared with asymptomatic individuals (p=0&#xb7;56). These findings were consistent across genomic clusters and robust to model assumptions. INTERPRETATION: We identified no differences in transmission between individuals who were symptomatic and those who were asymptomatic using multiple genomic measures. In this high-transmission setting, where systematic screening is implemented, our findings indicate that asymptomatic tuberculosis substantially contributes to tuberculosis transmission at the population level. These results suggest that symptom-based case detection alone is likely to be insufficient to interrupt transmission and highlight the importance of expanded screening strategies in high-risk populations. FUNDING: US National Institutes of Health and the Brazilian National Research Council (CNPq).

Humans↗