Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Whole genome sequence data set of methicillin-resistant Staphylococcus aureus isolated from a milkman associated with cows with subclinical mastitis in Kiruhura district, Uganda.

The whole-genome sequence data set for methicillin-resistant Staphylococcus aureus, which was isolated from a milkman associated with cows with subclinical mastitis in the Kiruhura district of Uganda, is presented here. The assembled genome size was 2822,509 bp, with a 33% GC, 2 Contigs, a Contig N50 of 2818,424, and 1 Contig L50. You can access the genome sequence and related metadata at https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_056782255.1/. This dataset can be used again for resistance gene mapping, comparing genomic analysis, and comprehending genetic diversity among MRSA isolates from Ugandan milkmen.

Antimicrobial-resistant genes Staphylococcus aureu↗

Global reach and sustained engagement of a structured digital education program in medical mycology: an observational analysis of the 2025 ESCMID-EFISG webinar series.

OBJECTIVES: Evaluate the 2025 European Society of Clinical Microbiology and Infectious Diseases-European Fungal Infection Study Group webinar series to assess digital education as a scalable, equitable model for global professional development in medical mycology. METHODS: This observational study analyzed Zoom metadata across 17 webinars (January-December 2025). Metrics included registration, unique viewers, peak concurrent views, attendance rate, and duration. RESULTS: The series recorded 4631 registrations and 1372 unique participants. Median live attendance was 199 (interquartile range [IQR] 138-269), with a 39.3% attendance rate (IQR 33.7-47.5%) and peak concurrent viewership of 165 (IQR 106-229). Median session duration was 108 minutes. Webinars engaged a median of 60 countries (range 26-89) simultaneously, spanning 128 countries globally. Faculty comprised 67 unique experts from 23 countries with balanced gender representation (52.8% men, 47.2% women), of whom 16.4% (n = 11/67) were affiliated with institutions in low- and middle-income countries. CONCLUSION: Structured digital programs achieve wide global reach and sustained engagement. Strong participation in long-form sessions supports implementing Continuing Medical Education accreditation and unrestricted on-demand access to enhance global health equity.

Antimicrobial resistance↗

Spatiotemporal patterns of Rift Valley fever virus in Africa: a retrospective genomic epidemiology and phylodynamic modelling study.

BACKGROUND: Rift Valley fever virus (RVFV) is a mosquito-borne zoonotic pathogen causing outbreaks in humans and ruminants across Africa and the Arabian Peninsula. Originally restricted to the Great Rift Valley, RVFV has expanded geographically, prompting its classification by WHO as a pathogen of pandemic potential. We investigated the evolutionary and spatial dynamics of RVFV across Africa. METHODS: We used genomic data generated at the International Livestock Research Institute Nairobi genomic laboratory (BioProject PRJNA1106221) and combined with publicly available datasets retrieved from the National Center for Biotechnology (NCBI) GenBank nucleotide database. In retrieving RVFV genome sequences from the NCBI GenBank, we applied the search terms "Rift Valley fever virus segment L AND 6404[SLEN]", "Rift Valley fever virus segment M AND 3885[SLEN]", and "Rift Valley fever virus segment S AND 1520:1690[SLEN]" for L (Large), M (Medium), and S (Small) segments, respectively. For sequences without additional spatiotemporal information, we searched PubMed to extract the associated sequence metadata. We performed molecular clock analysis, phylogenetic inference, phylodynamic modelling (continuous phylogeographic reconstruction), and landscape phylogeography on the three RVFV genome segments (L, M, and S). We aimed to assess evolutionary rates, dispersal patterns, and environmental drivers. Focus was placed on lineage C, the most widely distributed variant. FINDINGS: The global dataset used in this study consisted of large (n=236), medium (n=237), and small (n=247), which were further filtered to exclude potential reassortants and vaccine strains. Genome sequences retrieved from NCBI GenBank database comprised large (n=180), medium (n=184), and small (n=202). The genome sequences from retrospective human and livestock isolates comprised large (n=56), medium (n=53), and small (n=45) collected in Burundi (2018), Kenya (2007, 2018, 2019, 2021, and 2022), and Rwanda (2018 and 2022). Our dataset revealed that RVFV exhibited low overall genetic diversity. Lineage C, however, showed evidence of active evolution, with substitution rates ranging from 3·58 × 10-4 to 9·76 × 10-4 substitutions per site per year. This lineage probably originated in Zimbabwe in the mid-1970s and has since expanded across eastern and southern Africa. Phylogeographic reconstructions revealed rapid spread, with diffusion coefficients exceeding 50 000 km2 per year. INTERPRETATION: Lineage C appears capable of establishing endemic transmission in new regions, with ongoing diversification observed during interepidemic periods. These observations reinforce the value of continuous genomic surveillance, particularly during cryptic transmission phases when adaptive mutations might emerge. Although further evidence is needed, observed trends in climate variability and land-use change point to the potential benefit of targeted surveillance in settings that could be at increased risk, including urban centres and wetlands. FUNDING: This work was supported by the German Federal Ministry for Economic Cooperation and Development, the Rockefeller Foundation, and the Africa Centres for Disease Control and Prevention.

Rift Valley fever virus↗

Multidimensional prophage profiling of carbapenem-resistant Enterobacteriaceae in Thailand: a nationwide, multicentre, genomic study.

BACKGROUND: Prophages influence bacterial fitness, resistance, and evolution, yet their epidemiology remains poorly understood in carbapenem-resistant Enterobacteriaceae (CRE). In this nationwide study in Thailand, we aimed to describe prophage repertoires in clinical CRE isolates and to explore their potential relevance for molecular epidemiology. METHODS: We performed a nationwide, retrospective, genomic analysis of all CRE clinical isolates collected through our previous national surveillance study involving 11 hospitals in 11 provinces in Thailand between March 25, 2012, and Jul 21, 2017. Whole-genome sequencing data from 747 CRE isolates were analysed. Intact prophages were identified using PHAge Search Tool Enhanced Release (PHASTER) and clustered by nucleotide sequence similarity. Prophage profiles were compared across multilocus sequence types, carbapenemase genotypes, specimens, geography, and patient demographics (age and sex). FINDINGS: Of the included 747 CRE isolates, 170 (23%) were Escherichia coli and 577 (77%) were Klebsiella pneumoniae. 220 (29%) of 747 strains had been isolated from female patients and 264 (35%) from male patients; metadata on patient sex were missing for 263 (35%) isolates. The median patient age was 63 years (IQR 50-72). 71 (10%) of isolates were from blood, 283 (38%) from sputum, 284 (38%) from urine, and 109 (15%) from other specimens. 374 distinct prophage clusters were identified, with significantly more prophages per genome in K pneumoniae (mean 3&#xb7;01 [SD 1&#xb7;55]) than in E coli (1&#xb7;64 [1&#xb7;46]; p<0&#xb7;0001). Prophage repertoires largely mirrored bacterial multilocus sequence types. However, even within the highly clonal K pneumoniae sequence type 16 lineage, discrete prophage variation was identified, with common profiles observed in geographically dispersed patients. Respiratory K pneumoniae frequently carried a mosaic prophage with environmental signatures and a type VI secretion system, whereas blood-derived E coli harboured a prophage with immune-modulating genes. Distinct prophage clusters were observed across clinical specimens, age groups, carbapenemase genotype, and geographical region. Strains coharbouring blaNDM-1 plus blaOXA-232 (114 [15%] of 747) had the highest prophage loads. INTERPRETATION: The prophage content was shaped by the bacterial lineage, ecological niche, and temporal dynamics, providing an additional layer of epidemiological resolution beyond conventional genome typing. Integrating prophage profiling into molecular surveillance frameworks could help to identify transmission events, improve infectious source attribution, and enhance infection control strategies. FUNDING: Japan Agency for Medical Research and Development.

Female↗

Comparison of phylogenetic metrics of transmission between symptomatic and asymptomatic tuberculosis in individuals who were incarcerated in Brazil in 2008-24: a retrospective genomic epidemiology study.

BACKGROUND: Tuberculosis control efforts have traditionally targeted symptomatic individuals; however, the role of asymptomatic cases in sustaining transmission is increasingly recognised. We aimed to quantify the contribution of asymptomatic tuberculosis to recent transmission using genomic and epidemiological data from a high-transmission setting. METHODS: We conducted a retrospective genomic epidemiology study of Mycobacterium tuberculosis isolates collected in Mato Grosso do Sul, Brazil, between Aug 25, 2008, and March 19, 2024. Available isolates underwent whole-genome sequencing. Demographic, clinical, incarceration history, and laboratory metadata were obtained from surveillance records. From Jan 1, 2017, to March 19, 2024, active case finding was conducted in the state's three largest prisons (all male-only facilities), during which sputum samples were collected from individuals irrespective of symptoms and tested using GeneXpert and culture. Comparisons of transmission between individuals with and without symptoms were restricted to individuals who were incarcerated and were identified through active case finding and for whom high-quality, M tuberculosis lineage 4 genomes were available. Metrics of recent transmission included phylogenetic clustering, time-scaled haplotype density (THD), local branching index (LBI), and transmission probabilities inferred using Bayesian Reconstruction and Evolutionary Analysis of Transmission Histories. FINDINGS: 4448 tuberculosis cases were notified in Mato Grosso do Sul in 2008-24. After excluding cases for which M tuberculosis isolates were not available or had low sequencing quality, who had contaminated cultures or mixed infection, or who were infected with non-lineage 4 M tuberculosis, we included 2362 lineage 4 M tuberculosis isolates with high-quality genome sequences. 1849 (78&#xb7;3%) of 2362 isolates were part of a genomic cluster. Among 2362 individuals with tuberculosis, 1137 (48&#xb7;1%) were incarcerated at diagnosis. Of these individuals, 505 were identified through active case finding in three male-only prisons. The median age was 30 years (IQR 25-37); 304 (60&#xb7;2%) had mixed ethnicity, 90 (17&#xb7;8%) were White, 56 (11&#xb7;1%) were Black, 13 (2&#xb7;6%) were Indigenous, and six (1&#xb7;2%) were Asian. 277 (54&#xb7;9%) had symptomatic disease and 228 (45&#xb7;1%) had asymptomatic tuberculosis. There were no significant differences between symptomatic and asymptomatic individuals in phylogenetic clustering (213 [76&#xb7;9%] of 277 vs 195 [85&#xb7;5%] of 228; p=0&#xb7;37), THD (median 0&#xb7;39 [IQR 0&#xb7;06-0&#xb7;62] vs 0&#xb7;50 [0&#xb7;09-0&#xb7;65]; p=0&#xb7;12), or LBI (0&#xb7;00863 [0&#xb7;00810-0&#xb7;00988] vs 0&#xb7;00871 [0&#xb7;00829-0&#xb7;01020]; p=0&#xb7;088). Bayesian transmission trees showed no significant difference in the number of secondary infections inferred from symptomatic compared with asymptomatic individuals (p=0&#xb7;56). These findings were consistent across genomic clusters and robust to model assumptions. INTERPRETATION: We identified no differences in transmission between individuals who were symptomatic and those who were asymptomatic using multiple genomic measures. In this high-transmission setting, where systematic screening is implemented, our findings indicate that asymptomatic tuberculosis substantially contributes to tuberculosis transmission at the population level. These results suggest that symptom-based case detection alone is likely to be insufficient to interrupt transmission and highlight the importance of expanded screening strategies in high-risk populations. FUNDING: US National Institutes of Health and the Brazilian National Research Council (CNPq).

Humans↗

Meningioma transcriptomic landscape demonstrates novel subtypes with regional associated biology and patient outcome.

Meningiomas, although mostly benign, can be recurrent and fatal. World Health Organization (WHO) grading of the tumor does not always identify high-risk meningioma, and better characterizations of their aggressive biology are needed. To approach this problem, we combined 13 bulk RNA sequencing (RNA-seq) datasets to create a dimension-reduced reference landscape of 1,298 meningiomas. The clinical and genomic metadata effectively correlated with landscape regions, which led to the identification of meningioma subtypes with specific biological signatures. The time to recurrence also correlated with the map location. Further, we developed an algorithm that maps new patients onto this landscape, where the nearest neighbors predict outcome. This study highlights the utility of combining bulk transcriptomic datasets to visualize the complexity of tumor populations. Further, we provide an interactive tool for understanding the disease and predicting patient outcomes. This resource is accessible via the online tool Oncoscape, where the scientific community can explore the meningioma landscape.

Meningioma↗

Protocol for histology-anchored macroscopic staging of gonadal maturity in exploited fishes.

Here, we present a protocol to assign gonadal maturity stages in commercially exploited fishes using a histology-anchored workflow. We describe steps for recording field metadata, photographing gonads, fixing central gonadal tissue, and processing paraffin sections. We then detail procedures for staining sections with hematoxylin and eosin, diagnosing gametogenic features, and assigning stages using a common reproductive-phase framework with species- and sex-specific reference descriptors. This protocol standardizes documentation and decision logic rather than proposing a new maturity scale.

Developmental biology↗

Continuing dental education on the World Wide Web.

Continuing dental education (CDE) courses delivered on the World Wide Web (Web CDE) offer numerous advantages over traditional CDE; however, two major issues--location of suitable courses and course quality--need resolution. Locating high-quality courses is difficult due to the lack of the standardized metadata that allows search engines to match courses to practitioners' needs. Web directories created by professional organizations are beginning to show promise, but require further development. Search engines and Web directories are discussed and improvements currently underway summarized. Course quality remains a highly significant concern. A national effort to create Web CDE course quality standards is underway that includes proposed standards. These proposed standards are summarized and used to comment on the current state of Web CDE courses. Examples are given when possible. Three emerging Web CDE technologies and a look to the future of Web CDE are discussed.

Computer-Assisted Instruction↗

A system for simultaneous multiple subject, multiple stimulus modality, and multiple channel collection and analysis of sensory evoked potentials.

A system has been developed for collecting sensory evoked potentials simultaneously from multiple channels for multiple subjects at up to 80 kHz sample rate per channel. Sample rates up to 200 kHz are available for four or less chambers and a single channel per chamber. A variety of visual, somatosensory, and auditory stimuli may be presented singly or simultaneously. Collected waveforms are associated with searchable text (metadata) to allow convenient selection from a relational database. Multiple waveforms can then be easily grouped for analysis and processed. Results can be exported to other software for further graphics or statistical processing. Scripting and event logging are available to provide automation and improve data confidence. Sample data are presented from control animals for each of the sensory modalities for comparison with historical data collected from other systems.

Animals↗

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa↗

Integrating multi-omics technologies to decipher microbiome functions.

Multi-omics approaches have revolutionized our understanding of microbial communities by enabling simultaneous interrogation of genomic, transcriptomic, proteomic, and metabolomic data. The systematic integration and analysis of these deep datasets help decipher the functional roles of microbiomes, providing critical insights into microbial activities, interactions, and dynamics across diverse environments. Biological complexity makes multi-omics analysis of a single, isolated organism demanding but highly informative, yet this complexity increases further when samples comprise hundreds to thousands of individual species. As microbiome research continues to expand into clinical, environmental, and engineered systems, standardized workflows, benchmarked datasets, and community-driven initiatives are essential to ensure reproducibility, standardization and interpretability. Establishing and disseminating best practices for experimental design, data processing, and integrative analyses will be critical for maximizing comparability and scientific rigor across studies. This perspective highlights recent advances in multi-omics microbiome research, outlines key obstacles in data integration and metadata harmonization, and proposes a collaborative roadmap for scalable, FAIR-compliant multi-omics investigations and potentially disruptive Artificial Intelligence (AI) advances comparable to those of AlphaFold in the field of microbiome science.

Multiomics↗

Structure-centric searching enables global mapping of the public metabolome.

Searching and learning from aggregated public metabolomics data spanning thousands of studies remained largely inaccessible. Here we present StructureMASST, a web-based application enabling scalable, structure-centric searches across public metabolomics repositories using molecule names or chemical representations. It queries a precomputed knowledgebase of 2.19 billion spectral matches and 420 million metadata links, supports modification-tolerant and mass-shift searches, and maps chemical structures across taxonomy, biological context and environmental conditions to accelerate discovery.

Journal Article↗

Temporal stability and lack of variance in microbiome composition and functionality in fit recreational athletes.

Human gut microbiome composition and function is influenced by environmental and lifestyle factors, including exercise and fitness. We studied the composition and functionality of the faecal microbiome of recreational (non-elite) runners (n&#x2009;=&#x2009;62) with serial shotgun metagenomics, at 4 time points over a 7-week period. Gut microbiome composition and function was stable over time. Grouping of samples on the basis of their fitness level (fair, good, excellent, and superior) or habitual training (low (4-6&#xa0;h/week), medium (7-9&#xa0;h/week), high (10-12&#xa0;h/week), and extreme (13&#x2009;+&#x2009;hours/week)) revealed no significant microbiome-related differences. Overall, the species Faecalibacterium prausnitzii, Blautia wexlerae, and Prevotella copri were the most abundant members of the gut microbiome. Analysis of co-abundance groups (CAGs) revealed no significant relationship between CAGs and fitness levels or training subgroups. Functional pathways were similar across all samples and timepoints with no clustering based on associated metadata. The most abundant genes identified within samples corresponded to pathways for nucleoside and nucleotide biosynthesis, amino acid biosynthesis, and cell wall biosynthesis. Collectively, these results describe the microbiome of active recreational runners and note temporal stability amongst participants.

Humans↗

Analysis of molecular data of Arabidopsis thaliana (L.) Heynh. (Brassicaceae) with Geographical Information Systems (GIS).

A Geographical Information System (GIS) is used to analyse allelic information of 13 sequenced loci of natural populations of Arabidopsis thaliana and to identify geographical structures. GIS provides tools for visualization and analysis of geographical population structures using molecular data. The geographical distribution of the number of variable positions in the alignments, the distribution of recombinant sequence blocks, and the distribution of a newly defined measure, the differentiation index, are studied. The differentiation index is introduced to measure the sequence divergence among individual plants sampled from various geographical localities. The numbers of variable positions and the differentiation index are also used for a metadata analysis covering about 26 kb of the genome. This analysis reveals, for the first time, differences in DNA sequence structures of geographically different populations of A. thaliana. The broadly defined west Mediterranean region consists of accessions with the highest numbers of polymorphic positions followed by the west European region. The GIS technology Kriging is used to define Arabidopsis specific diversity zones in Europe. The highest genetic variability is observed along the Atlantic coast from the western Iberian Peninsula to southern Great Britain, while lowest variability is found in central Europe.

Arabidopsis↗

Validating existing data in the Environmental Technology Verification Program.

Establishing the credibility of existing data is an ongoing issue, particularly when the data sets are to be used for a secondary purpose, i.e., not the original reason for which they were collected. If the secondary purpose is similar to the primary purpose, the potential user may have little difficulty establishing credibility since the acceptance criteria for both purposes should be similar. If the secondary purpose is different, then data credibility may be more difficult to establish because the experiment generating the data may not have been conducted optimally for the secondary purpose and all of the necessary quality assurance data ("metadata") may not have been collected. In either case, a process will be required to determine the acceptability of the data. For this reason, at the time the U.S. Environmental Protection Agency (EPA) Environmental Technology Verification (ETV) program was established, similar certification and verification programs run by states or foreign countries routinely used existing data sets, for cost reasons, rather than generate new data by testing. The issue of whether existing data could be used in the ETV program immediately surfaced. In response, a policy and a process that addressed existing data were written and published in Appendix C of the ETV Quality and Management Plan (Hayes et al., 1998). This paper discusses how the ETV program determines the credibility of existing data used to verify the performance of environmental technologies.

Data Interpretation, Statistical↗

A search tool based on 'encapsulated' MeSH thesaurus to retrieve quality health resources on the internet.

In the year 2001, the Internet has become a major source of health information for the health professional and the Netizen. The objective of Doc' CISMeF (D'C) was to create a powerful generic search tool based on a structured information model which 'encapsulates' the MeSH thesaurus to index and retrieve quality health resources on the Internet. To index resources, D'C uses four sections in its information model: 'meta-term', keyword, subheading, and resource type. Two search options are available: simple and advanced. The simple search requires the end-user to input a single term or expression. If this term belongs to the D'C information structure model, it will be exploded. If not, a full-text search is performed. In the advanced search, complex searches are possible combining Boolean operators with meta-terms, keywords, subheadings and resource types. D'C uses two standard tools for organising information: the MeSH thesaurus and the Dublin Core metadata format. Resources included in D'C are described according to the following elements: title, author or creator, subject and keywords, description, publishers, date, resource type, format, identifier, and language.

Abstracting and Indexing↗

Implementing context and team based access control in healthcare intranets.

The establishment of an efficient access control system in healthcare intranets is a critical security issue directly related to the protection of patients' privacy. Our C-TMAC (Context and Team-based Access Control) model is an active security access control model that layers dynamic access control concepts on top of RBAC (Role-based) and TMAC (Team-based) access control models. It also extends them in the sense that contextual information concerning collaborative activities is associated with teams of users and user permissions are dynamically filtered during runtime. These features of C-TMAC meet the specific security requirements of healthcare applications. In this paper, an experimental implementation of the C-TMAC model is described. More specifically, we present the operational architecture of the system that is used to implement C-TMAC security components in a healthcare intranet. Based on the technological platform of an Oracle Data Base Management System and Application Server, the application logic is coded with stored PL/SQL procedures that include Dynamic SQL routines for runtime value binding purposes. The resulting active security system adapts to current need-to-know requirements of users during runtime and provides fine-grained permission granularity. Apart from identity certificates for authentication, it uses attribute certificates for communicating critical security metadata, such as role membership and team participation of users.

Computer Communication Networks↗

MDB: a database system utilizing automatic construction of modules and STAR-derived universal language.

MOTIVATION: The value of information greatly increases if stored in databases. The objective was to construct a multi-purpose database system primarily designed to store and provide access to three-dimensional structures of biological molecules including theoretical models. RESULTS: A dictionary defining data format and structure for three-dimensional models of biological molecules (MDB dictionary) was developed. The dictionary was written using universal, standardized data description language. This language can be applied to describe data with no restrictions on their origin or type, including metadata. Thus both the data definitions (format) and database descriptions are created using the uniform language and processed with universal software. A database and data design technique that allowed use of dictionaries to automatically construct relational databases was developed. This technique was employed to construct the MDB database system. Data design developed and applied in the MDB project makes it possible to carry out data curation utilizing the database engine to identify errors. It also allows storage and query of data at different levels of consistency with the standard format specifications, i.e. both the correctly formatted data, and data that requires further curation. AVAILABILITY: The MDB dictionary is available at http://www.gwer.ch/proteinstructure/mdb and as part of the PDB resources at http://pdb.rutgers.edu/mmcif/.

Computational Biology↗