Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Clinical sequelae of gut microbiome development and disruption in hospitalized preterm infants.

Aberrant preterm infant gut microbiota assembly predisposes to early-life disorders and persistent health problems. Here, we characterize gut microbiome dynamics over the first 3 months of life in 236 preterm infants hospitalized in three neonatal intensive care units using shotgun metagenomics of 2,512 stools and metatranscriptomics of 1,381 stools. Strain tracking, taxonomic and functional profiling, and comprehensive clinical metadata identify Enterobacteriaceae, enterococci, and staphylococci as primarily exploiting available niches to populate the gut microbiome. Clostridioides difficile lineages persist between individuals in single centers, and Staphylococcus epidermidis lineages persist within and, unexpectedly, between centers. Collectively, antibiotic and non-antibiotic medications influence gut microbiome composition to greater extents than maternal or baseline variables. Finally, we identify a persistent low-diversity gut microbiome in neonates who develop necrotizing enterocolitis after day of life 40. Overall, we comprehensively describe gut microbiome dynamics in response to medical interventions in preterm, hospitalized neonates.

Humans↗

Whole genome sequence data set of methicillin-resistant Staphylococcus aureus isolated from a milkman associated with cows with subclinical mastitis in Kiruhura district, Uganda.

The whole-genome sequence data set for methicillin-resistant Staphylococcus aureus, which was isolated from a milkman associated with cows with subclinical mastitis in the Kiruhura district of Uganda, is presented here. The assembled genome size was 2822,509 bp, with a 33% GC, 2 Contigs, a Contig N50 of 2818,424, and 1 Contig L50. You can access the genome sequence and related metadata at https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_056782255.1/. This dataset can be used again for resistance gene mapping, comparing genomic analysis, and comprehending genetic diversity among MRSA isolates from Ugandan milkmen.

Antimicrobial-resistant genes Staphylococcus aureu↗

Global reach and sustained engagement of a structured digital education program in medical mycology: an observational analysis of the 2025 ESCMID-EFISG webinar series.

OBJECTIVES: Evaluate the 2025 European Society of Clinical Microbiology and Infectious Diseases-European Fungal Infection Study Group webinar series to assess digital education as a scalable, equitable model for global professional development in medical mycology. METHODS: This observational study analyzed Zoom metadata across 17 webinars (January-December 2025). Metrics included registration, unique viewers, peak concurrent views, attendance rate, and duration. RESULTS: The series recorded 4631 registrations and 1372 unique participants. Median live attendance was 199 (interquartile range [IQR] 138-269), with a 39.3% attendance rate (IQR 33.7-47.5%) and peak concurrent viewership of 165 (IQR 106-229). Median session duration was 108 minutes. Webinars engaged a median of 60 countries (range 26-89) simultaneously, spanning 128 countries globally. Faculty comprised 67 unique experts from 23 countries with balanced gender representation (52.8% men, 47.2% women), of whom 16.4% (n = 11/67) were affiliated with institutions in low- and middle-income countries. CONCLUSION: Structured digital programs achieve wide global reach and sustained engagement. Strong participation in long-form sessions supports implementing Continuing Medical Education accreditation and unrestricted on-demand access to enhance global health equity.

Antimicrobial resistance↗

Combining different standards and different approaches for health information retrieval in a quality-controlled gateway.

Internet as source of information is increasing in preeminence in numerous fields, including health. We describe in this paper the CISMeF project (acronym of Catalogue and Index of French-speaking Medical Sites) which has been designed to help the health information consumers and health professionals to find what they are looking for among the numerous health documents available online. The catalogue is founded on two standards: a set of metadata and a terminology based on the MeSH thesaurus which has the same structure and use as an ontology of the medical domain. The structure of the catalogue allows us to place the project at an overlap between the present Web, which is informal, and the forthcoming Semantic Web. Many features of information retrieval and navigation through the catalogue were developed. These features take into account the kind of the end-user (health professional, medical student, patient). The CISMeF-patients catalogue is a sub-catalogue of CISMeF and is dedicated to the patients and the general public. It shares the same model as CISMeF whereas MEDLINE and MedlinePlus do not. We also propose to couple two approaches (morphological processing and data mining) to help the users by correcting and refining their queries.

France↗

Analysing proteomic data.

The rapid growth of proteomics has been made possible by the development of reproducible 2D gels and biological mass spectrometry. However, despite technical improvements 2D gels are still less than perfectly reproducible and gels have to be aligned so spots for identical proteins appear in the same place. Gels can be warped by a variety of techniques to make them concordant. When gels are manipulated to improve registration, information is lost, so direct methods for gel registration which make use of all available data for spot matching are preferable to indirect ones. In order to identify proteins from gel spots a property or combination of properties that are unique to that protein are required. These can then be used to search databases for possible matches. Molecular mass, pI, amino acid composition and short sequence tags can all be used in database searches. Currently the method of choice for protein identification is mass spectrometry. Proteins are eluted from the gels and cleaved with specific endoproteases to produce a series of peptides of different molecular mass. In peptide mass fingerprinting, the peptide profile of the unknown protein is compared with theoretical peptide libraries generated from sequences in the different databases. Tandem mass spectroscopy (MS/MS) generates short amino acid sequence tags for the individual peptides. These partial sequences combined with the original peptide masses are then used for database searching, greatly improving specificity. Increasingly protein identification from MS/MS data is being fully or partially automated. When working with organisms, which do not have sequenced genomes (the case with most helminths), protein identification by database searching becomes problematical. A number of approaches to cross species protein identification have been suggested, but if the organism being studied is only distantly related to any organism with a sequenced genome then the likelihood of protein identification remains small. The dynamic nature of the proteome means that there really is no such thing as a single representative proteome and a complete set of metadata (data about the data) is going to be required if the full potential of database mining is to be realised in the future.

Animals↗

Comparing statistical and semantic approaches for identifying change from land cover datasets.

In this paper, we examine methods for integrating spatial data which apparently should be comparable because they are of the same data type or theme, but which are incompatible or discordant because the classes of that theme are different. For a variety of reasons including changes in methods, in understanding of the resource, and in policy initiatives in the commissioning of the survey, this problem is widespread in the results of natural resources surveys. We present two generic methods: one method is grounded in a statistical approach using discriminant analysis, and the other exploits the knowledge of experts. We use the context of land cover mapping of Great Britain to explore these approaches for integrating discordant data. We demonstrate that the expert-based approach gives very good levels of identification of locations with incompatible classifications at different times, and gives a much better rate of recognition of change. Some conclusions are made about the need to expand current metadata and data quality reporting to include descriptions of:- data conceptualisations, semantics and ontologies;- who decided and defined what the features of interest in a dataset are, and why. If the benefits of spatial data initiatives such as GRID, E-science and INSPIRE are to be fully realised then some method needs to be found to communicate that information most effectively to the potential user of the data.

Conservation of Natural Resources↗

Spatiotemporal patterns of Rift Valley fever virus in Africa: a retrospective genomic epidemiology and phylodynamic modelling study.

BACKGROUND: Rift Valley fever virus (RVFV) is a mosquito-borne zoonotic pathogen causing outbreaks in humans and ruminants across Africa and the Arabian Peninsula. Originally restricted to the Great Rift Valley, RVFV has expanded geographically, prompting its classification by WHO as a pathogen of pandemic potential. We investigated the evolutionary and spatial dynamics of RVFV across Africa. METHODS: We used genomic data generated at the International Livestock Research Institute Nairobi genomic laboratory (BioProject PRJNA1106221) and combined with publicly available datasets retrieved from the National Center for Biotechnology (NCBI) GenBank nucleotide database. In retrieving RVFV genome sequences from the NCBI GenBank, we applied the search terms "Rift Valley fever virus segment L AND 6404[SLEN]", "Rift Valley fever virus segment M AND 3885[SLEN]", and "Rift Valley fever virus segment S AND 1520:1690[SLEN]" for L (Large), M (Medium), and S (Small) segments, respectively. For sequences without additional spatiotemporal information, we searched PubMed to extract the associated sequence metadata. We performed molecular clock analysis, phylogenetic inference, phylodynamic modelling (continuous phylogeographic reconstruction), and landscape phylogeography on the three RVFV genome segments (L, M, and S). We aimed to assess evolutionary rates, dispersal patterns, and environmental drivers. Focus was placed on lineage C, the most widely distributed variant. FINDINGS: The global dataset used in this study consisted of large (n=236), medium (n=237), and small (n=247), which were further filtered to exclude potential reassortants and vaccine strains. Genome sequences retrieved from NCBI GenBank database comprised large (n=180), medium (n=184), and small (n=202). The genome sequences from retrospective human and livestock isolates comprised large (n=56), medium (n=53), and small (n=45) collected in Burundi (2018), Kenya (2007, 2018, 2019, 2021, and 2022), and Rwanda (2018 and 2022). Our dataset revealed that RVFV exhibited low overall genetic diversity. Lineage C, however, showed evidence of active evolution, with substitution rates ranging from 3·58 × 10-4 to 9·76 × 10-4 substitutions per site per year. This lineage probably originated in Zimbabwe in the mid-1970s and has since expanded across eastern and southern Africa. Phylogeographic reconstructions revealed rapid spread, with diffusion coefficients exceeding 50 000 km2 per year. INTERPRETATION: Lineage C appears capable of establishing endemic transmission in new regions, with ongoing diversification observed during interepidemic periods. These observations reinforce the value of continuous genomic surveillance, particularly during cryptic transmission phases when adaptive mutations might emerge. Although further evidence is needed, observed trends in climate variability and land-use change point to the potential benefit of targeted surveillance in settings that could be at increased risk, including urban centres and wetlands. FUNDING: This work was supported by the German Federal Ministry for Economic Cooperation and Development, the Rockefeller Foundation, and the Africa Centres for Disease Control and Prevention.

Rift Valley fever virus↗

Multidimensional prophage profiling of carbapenem-resistant Enterobacteriaceae in Thailand: a nationwide, multicentre, genomic study.

BACKGROUND: Prophages influence bacterial fitness, resistance, and evolution, yet their epidemiology remains poorly understood in carbapenem-resistant Enterobacteriaceae (CRE). In this nationwide study in Thailand, we aimed to describe prophage repertoires in clinical CRE isolates and to explore their potential relevance for molecular epidemiology. METHODS: We performed a nationwide, retrospective, genomic analysis of all CRE clinical isolates collected through our previous national surveillance study involving 11 hospitals in 11 provinces in Thailand between March 25, 2012, and Jul 21, 2017. Whole-genome sequencing data from 747 CRE isolates were analysed. Intact prophages were identified using PHAge Search Tool Enhanced Release (PHASTER) and clustered by nucleotide sequence similarity. Prophage profiles were compared across multilocus sequence types, carbapenemase genotypes, specimens, geography, and patient demographics (age and sex). FINDINGS: Of the included 747 CRE isolates, 170 (23%) were Escherichia coli and 577 (77%) were Klebsiella pneumoniae. 220 (29%) of 747 strains had been isolated from female patients and 264 (35%) from male patients; metadata on patient sex were missing for 263 (35%) isolates. The median patient age was 63 years (IQR 50-72). 71 (10%) of isolates were from blood, 283 (38%) from sputum, 284 (38%) from urine, and 109 (15%) from other specimens. 374 distinct prophage clusters were identified, with significantly more prophages per genome in K pneumoniae (mean 3&#xb7;01 [SD 1&#xb7;55]) than in E coli (1&#xb7;64 [1&#xb7;46]; p<0&#xb7;0001). Prophage repertoires largely mirrored bacterial multilocus sequence types. However, even within the highly clonal K pneumoniae sequence type 16 lineage, discrete prophage variation was identified, with common profiles observed in geographically dispersed patients. Respiratory K pneumoniae frequently carried a mosaic prophage with environmental signatures and a type VI secretion system, whereas blood-derived E coli harboured a prophage with immune-modulating genes. Distinct prophage clusters were observed across clinical specimens, age groups, carbapenemase genotype, and geographical region. Strains coharbouring blaNDM-1 plus blaOXA-232 (114 [15%] of 747) had the highest prophage loads. INTERPRETATION: The prophage content was shaped by the bacterial lineage, ecological niche, and temporal dynamics, providing an additional layer of epidemiological resolution beyond conventional genome typing. Integrating prophage profiling into molecular surveillance frameworks could help to identify transmission events, improve infectious source attribution, and enhance infection control strategies. FUNDING: Japan Agency for Medical Research and Development.

Female↗

Comparison of phylogenetic metrics of transmission between symptomatic and asymptomatic tuberculosis in individuals who were incarcerated in Brazil in 2008-24: a retrospective genomic epidemiology study.

BACKGROUND: Tuberculosis control efforts have traditionally targeted symptomatic individuals; however, the role of asymptomatic cases in sustaining transmission is increasingly recognised. We aimed to quantify the contribution of asymptomatic tuberculosis to recent transmission using genomic and epidemiological data from a high-transmission setting. METHODS: We conducted a retrospective genomic epidemiology study of Mycobacterium tuberculosis isolates collected in Mato Grosso do Sul, Brazil, between Aug 25, 2008, and March 19, 2024. Available isolates underwent whole-genome sequencing. Demographic, clinical, incarceration history, and laboratory metadata were obtained from surveillance records. From Jan 1, 2017, to March 19, 2024, active case finding was conducted in the state's three largest prisons (all male-only facilities), during which sputum samples were collected from individuals irrespective of symptoms and tested using GeneXpert and culture. Comparisons of transmission between individuals with and without symptoms were restricted to individuals who were incarcerated and were identified through active case finding and for whom high-quality, M tuberculosis lineage 4 genomes were available. Metrics of recent transmission included phylogenetic clustering, time-scaled haplotype density (THD), local branching index (LBI), and transmission probabilities inferred using Bayesian Reconstruction and Evolutionary Analysis of Transmission Histories. FINDINGS: 4448 tuberculosis cases were notified in Mato Grosso do Sul in 2008-24. After excluding cases for which M tuberculosis isolates were not available or had low sequencing quality, who had contaminated cultures or mixed infection, or who were infected with non-lineage 4 M tuberculosis, we included 2362 lineage 4 M tuberculosis isolates with high-quality genome sequences. 1849 (78&#xb7;3%) of 2362 isolates were part of a genomic cluster. Among 2362 individuals with tuberculosis, 1137 (48&#xb7;1%) were incarcerated at diagnosis. Of these individuals, 505 were identified through active case finding in three male-only prisons. The median age was 30 years (IQR 25-37); 304 (60&#xb7;2%) had mixed ethnicity, 90 (17&#xb7;8%) were White, 56 (11&#xb7;1%) were Black, 13 (2&#xb7;6%) were Indigenous, and six (1&#xb7;2%) were Asian. 277 (54&#xb7;9%) had symptomatic disease and 228 (45&#xb7;1%) had asymptomatic tuberculosis. There were no significant differences between symptomatic and asymptomatic individuals in phylogenetic clustering (213 [76&#xb7;9%] of 277 vs 195 [85&#xb7;5%] of 228; p=0&#xb7;37), THD (median 0&#xb7;39 [IQR 0&#xb7;06-0&#xb7;62] vs 0&#xb7;50 [0&#xb7;09-0&#xb7;65]; p=0&#xb7;12), or LBI (0&#xb7;00863 [0&#xb7;00810-0&#xb7;00988] vs 0&#xb7;00871 [0&#xb7;00829-0&#xb7;01020]; p=0&#xb7;088). Bayesian transmission trees showed no significant difference in the number of secondary infections inferred from symptomatic compared with asymptomatic individuals (p=0&#xb7;56). These findings were consistent across genomic clusters and robust to model assumptions. INTERPRETATION: We identified no differences in transmission between individuals who were symptomatic and those who were asymptomatic using multiple genomic measures. In this high-transmission setting, where systematic screening is implemented, our findings indicate that asymptomatic tuberculosis substantially contributes to tuberculosis transmission at the population level. These results suggest that symptom-based case detection alone is likely to be insufficient to interrupt transmission and highlight the importance of expanded screening strategies in high-risk populations. FUNDING: US National Institutes of Health and the Brazilian National Research Council (CNPq).

Humans↗

Data input module for Birth Defects Systems Manager.

The need for a computational bioinformatics infrastructure to manage the vast digital information from functional genomics and proteomics motivated us to develop Birth Defects Systems Manager (BDSM) as an open resource to facilitate analysis and discovery in developmental biology and developmental toxicity. This report describes the design, development and implementation of the data loading module of BDSM, referred to as LoadBDSM. It includes a shared data directory resource that can be granted various levels of security for different research groups or investigators to manage experimental datasets individually or in groups. LoadBDSM allows the upload of data and experiment details using controlled semantics for developmental exposure (toxicant, dosing scenario, intervention), biological sample (species, tissue, stage) and disease outcome (time, risk, phenotype). It adheres to existing controlled vocabulary plus rules of inference (ontologies) for experiment, data and metadata annotations. LoadBDSM extends the capabilities of BDSM to support the emergence of "embryo-formatics" defined here as the data, information and knowledge from genomic sciences applied to, or derived from, an embryological context. This includes, but is not limited to, delineating pathways and biological regulatory networks for specific chemicals or classes of developmental toxicants, developing novel biomarkers indicative of exposure and/or predictive of adverse effects, and integrating modern computing and information technology with data from molecular biology.

Abnormalities, Drug-Induced↗

Meningioma transcriptomic landscape demonstrates novel subtypes with regional associated biology and patient outcome.

Meningiomas, although mostly benign, can be recurrent and fatal. World Health Organization (WHO) grading of the tumor does not always identify high-risk meningioma, and better characterizations of their aggressive biology are needed. To approach this problem, we combined 13 bulk RNA sequencing (RNA-seq) datasets to create a dimension-reduced reference landscape of 1,298 meningiomas. The clinical and genomic metadata effectively correlated with landscape regions, which led to the identification of meningioma subtypes with specific biological signatures. The time to recurrence also correlated with the map location. Further, we developed an algorithm that maps new patients onto this landscape, where the nearest neighbors predict outcome. This study highlights the utility of combining bulk transcriptomic datasets to visualize the complexity of tumor populations. Further, we provide an interactive tool for understanding the disease and predicting patient outcomes. This resource is accessible via the online tool Oncoscape, where the scientific community can explore the meningioma landscape.

Meningioma↗

Protocol for histology-anchored macroscopic staging of gonadal maturity in exploited fishes.

Here, we present a protocol to assign gonadal maturity stages in commercially exploited fishes using a histology-anchored workflow. We describe steps for recording field metadata, photographing gonads, fixing central gonadal tissue, and processing paraffin sections. We then detail procedures for staining sections with hematoxylin and eosin, diagnosing gametogenic features, and assigning stages using a common reproductive-phase framework with species- and sex-specific reference descriptors. This protocol standardizes documentation and decision logic rather than proposing a new maturity scale.

Developmental biology↗

Continuing dental education on the World Wide Web.

Continuing dental education (CDE) courses delivered on the World Wide Web (Web CDE) offer numerous advantages over traditional CDE; however, two major issues--location of suitable courses and course quality--need resolution. Locating high-quality courses is difficult due to the lack of the standardized metadata that allows search engines to match courses to practitioners' needs. Web directories created by professional organizations are beginning to show promise, but require further development. Search engines and Web directories are discussed and improvements currently underway summarized. Course quality remains a highly significant concern. A national effort to create Web CDE course quality standards is underway that includes proposed standards. These proposed standards are summarized and used to comment on the current state of Web CDE courses. Examples are given when possible. Three emerging Web CDE technologies and a look to the future of Web CDE are discussed.

Computer-Assisted Instruction↗

A system for simultaneous multiple subject, multiple stimulus modality, and multiple channel collection and analysis of sensory evoked potentials.

A system has been developed for collecting sensory evoked potentials simultaneously from multiple channels for multiple subjects at up to 80 kHz sample rate per channel. Sample rates up to 200 kHz are available for four or less chambers and a single channel per chamber. A variety of visual, somatosensory, and auditory stimuli may be presented singly or simultaneously. Collected waveforms are associated with searchable text (metadata) to allow convenient selection from a relational database. Multiple waveforms can then be easily grouped for analysis and processed. Results can be exported to other software for further graphics or statistical processing. Scripting and event logging are available to provide automation and improve data confidence. Sample data are presented from control animals for each of the sensory modalities for comparison with historical data collected from other systems.

Animals↗

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa↗

ThermoData Engine (TDE): software implementation of the dynamic data evaluation concept.

The first full-scale software implementation of the dynamic data evaluation concept {ThermoData Engine (TDE)} is described for thermophysical property data. This concept requires the development of large electronic databases capable of storing essentially all experimental data known to date with detailed descriptions of relevant metadata and uncertainties. The combination of these electronic databases with expert-system software, designed to automatically generate recommended data based on available experimental data, leads to the ability to produce critically evaluated data dynamically or 'to order'. Six major design tasks are described with emphasis on the software architecture for automated critical evaluation including dynamic selection and application of prediction methods and enforcement of thermodynamic consistency. The direction of future enhancements is discussed.

Journal Article↗

Integrating multi-omics technologies to decipher microbiome functions.

Multi-omics approaches have revolutionized our understanding of microbial communities by enabling simultaneous interrogation of genomic, transcriptomic, proteomic, and metabolomic data. The systematic integration and analysis of these deep datasets help decipher the functional roles of microbiomes, providing critical insights into microbial activities, interactions, and dynamics across diverse environments. Biological complexity makes multi-omics analysis of a single, isolated organism demanding but highly informative, yet this complexity increases further when samples comprise hundreds to thousands of individual species. As microbiome research continues to expand into clinical, environmental, and engineered systems, standardized workflows, benchmarked datasets, and community-driven initiatives are essential to ensure reproducibility, standardization and interpretability. Establishing and disseminating best practices for experimental design, data processing, and integrative analyses will be critical for maximizing comparability and scientific rigor across studies. This perspective highlights recent advances in multi-omics microbiome research, outlines key obstacles in data integration and metadata harmonization, and proposes a collaborative roadmap for scalable, FAIR-compliant multi-omics investigations and potentially disruptive Artificial Intelligence (AI) advances comparable to those of AlphaFold in the field of microbiome science.

Multiomics↗

Structure-centric searching enables global mapping of the public metabolome.

Searching and learning from aggregated public metabolomics data spanning thousands of studies remained largely inaccessible. Here we present StructureMASST, a web-based application enabling scalable, structure-centric searches across public metabolomics repositories using molecule names or chemical representations. It queries a precomputed knowledgebase of 2.19 billion spectral matches and 420 million metadata links, supports modification-tolerant and mass-shift searches, and maps chemical structures across taxonomy, biological context and environmental conditions to accelerate discovery.

Journal Article↗