Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Global reach and sustained engagement of a structured digital education program in medical mycology: an observational analysis of the 2025 ESCMID-EFISG webinar series.

OBJECTIVES: Evaluate the 2025 European Society of Clinical Microbiology and Infectious Diseases-European Fungal Infection Study Group webinar series to assess digital education as a scalable, equitable model for global professional development in medical mycology. METHODS: This observational study analyzed Zoom metadata across 17 webinars (January-December 2025). Metrics included registration, unique viewers, peak concurrent views, attendance rate, and duration. RESULTS: The series recorded 4631 registrations and 1372 unique participants. Median live attendance was 199 (interquartile range [IQR] 138-269), with a 39.3% attendance rate (IQR 33.7-47.5%) and peak concurrent viewership of 165 (IQR 106-229). Median session duration was 108 minutes. Webinars engaged a median of 60 countries (range 26-89) simultaneously, spanning 128 countries globally. Faculty comprised 67 unique experts from 23 countries with balanced gender representation (52.8% men, 47.2% women), of whom 16.4% (n = 11/67) were affiliated with institutions in low- and middle-income countries. CONCLUSION: Structured digital programs achieve wide global reach and sustained engagement. Strong participation in long-form sessions supports implementing Continuing Medical Education accreditation and unrestricted on-demand access to enhance global health equity.

Antimicrobial resistance↗

Combining different standards and different approaches for health information retrieval in a quality-controlled gateway.

Internet as source of information is increasing in preeminence in numerous fields, including health. We describe in this paper the CISMeF project (acronym of Catalogue and Index of French-speaking Medical Sites) which has been designed to help the health information consumers and health professionals to find what they are looking for among the numerous health documents available online. The catalogue is founded on two standards: a set of metadata and a terminology based on the MeSH thesaurus which has the same structure and use as an ontology of the medical domain. The structure of the catalogue allows us to place the project at an overlap between the present Web, which is informal, and the forthcoming Semantic Web. Many features of information retrieval and navigation through the catalogue were developed. These features take into account the kind of the end-user (health professional, medical student, patient). The CISMeF-patients catalogue is a sub-catalogue of CISMeF and is dedicated to the patients and the general public. It shares the same model as CISMeF whereas MEDLINE and MedlinePlus do not. We also propose to couple two approaches (morphological processing and data mining) to help the users by correcting and refining their queries.

France↗

Analysing proteomic data.

The rapid growth of proteomics has been made possible by the development of reproducible 2D gels and biological mass spectrometry. However, despite technical improvements 2D gels are still less than perfectly reproducible and gels have to be aligned so spots for identical proteins appear in the same place. Gels can be warped by a variety of techniques to make them concordant. When gels are manipulated to improve registration, information is lost, so direct methods for gel registration which make use of all available data for spot matching are preferable to indirect ones. In order to identify proteins from gel spots a property or combination of properties that are unique to that protein are required. These can then be used to search databases for possible matches. Molecular mass, pI, amino acid composition and short sequence tags can all be used in database searches. Currently the method of choice for protein identification is mass spectrometry. Proteins are eluted from the gels and cleaved with specific endoproteases to produce a series of peptides of different molecular mass. In peptide mass fingerprinting, the peptide profile of the unknown protein is compared with theoretical peptide libraries generated from sequences in the different databases. Tandem mass spectroscopy (MS/MS) generates short amino acid sequence tags for the individual peptides. These partial sequences combined with the original peptide masses are then used for database searching, greatly improving specificity. Increasingly protein identification from MS/MS data is being fully or partially automated. When working with organisms, which do not have sequenced genomes (the case with most helminths), protein identification by database searching becomes problematical. A number of approaches to cross species protein identification have been suggested, but if the organism being studied is only distantly related to any organism with a sequenced genome then the likelihood of protein identification remains small. The dynamic nature of the proteome means that there really is no such thing as a single representative proteome and a complete set of metadata (data about the data) is going to be required if the full potential of database mining is to be realised in the future.

Animals↗

Comparing statistical and semantic approaches for identifying change from land cover datasets.

In this paper, we examine methods for integrating spatial data which apparently should be comparable because they are of the same data type or theme, but which are incompatible or discordant because the classes of that theme are different. For a variety of reasons including changes in methods, in understanding of the resource, and in policy initiatives in the commissioning of the survey, this problem is widespread in the results of natural resources surveys. We present two generic methods: one method is grounded in a statistical approach using discriminant analysis, and the other exploits the knowledge of experts. We use the context of land cover mapping of Great Britain to explore these approaches for integrating discordant data. We demonstrate that the expert-based approach gives very good levels of identification of locations with incompatible classifications at different times, and gives a much better rate of recognition of change. Some conclusions are made about the need to expand current metadata and data quality reporting to include descriptions of:- data conceptualisations, semantics and ontologies;- who decided and defined what the features of interest in a dataset are, and why. If the benefits of spatial data initiatives such as GRID, E-science and INSPIRE are to be fully realised then some method needs to be found to communicate that information most effectively to the potential user of the data.

Conservation of Natural Resources↗

Spatiotemporal patterns of Rift Valley fever virus in Africa: a retrospective genomic epidemiology and phylodynamic modelling study.

BACKGROUND: Rift Valley fever virus (RVFV) is a mosquito-borne zoonotic pathogen causing outbreaks in humans and ruminants across Africa and the Arabian Peninsula. Originally restricted to the Great Rift Valley, RVFV has expanded geographically, prompting its classification by WHO as a pathogen of pandemic potential. We investigated the evolutionary and spatial dynamics of RVFV across Africa. METHODS: We used genomic data generated at the International Livestock Research Institute Nairobi genomic laboratory (BioProject PRJNA1106221) and combined with publicly available datasets retrieved from the National Center for Biotechnology (NCBI) GenBank nucleotide database. In retrieving RVFV genome sequences from the NCBI GenBank, we applied the search terms "Rift Valley fever virus segment L AND 6404[SLEN]", "Rift Valley fever virus segment M AND 3885[SLEN]", and "Rift Valley fever virus segment S AND 1520:1690[SLEN]" for L (Large), M (Medium), and S (Small) segments, respectively. For sequences without additional spatiotemporal information, we searched PubMed to extract the associated sequence metadata. We performed molecular clock analysis, phylogenetic inference, phylodynamic modelling (continuous phylogeographic reconstruction), and landscape phylogeography on the three RVFV genome segments (L, M, and S). We aimed to assess evolutionary rates, dispersal patterns, and environmental drivers. Focus was placed on lineage C, the most widely distributed variant. FINDINGS: The global dataset used in this study consisted of large (n=236), medium (n=237), and small (n=247), which were further filtered to exclude potential reassortants and vaccine strains. Genome sequences retrieved from NCBI GenBank database comprised large (n=180), medium (n=184), and small (n=202). The genome sequences from retrospective human and livestock isolates comprised large (n=56), medium (n=53), and small (n=45) collected in Burundi (2018), Kenya (2007, 2018, 2019, 2021, and 2022), and Rwanda (2018 and 2022). Our dataset revealed that RVFV exhibited low overall genetic diversity. Lineage C, however, showed evidence of active evolution, with substitution rates ranging from 3·58 × 10-4 to 9·76 × 10-4 substitutions per site per year. This lineage probably originated in Zimbabwe in the mid-1970s and has since expanded across eastern and southern Africa. Phylogeographic reconstructions revealed rapid spread, with diffusion coefficients exceeding 50 000 km2 per year. INTERPRETATION: Lineage C appears capable of establishing endemic transmission in new regions, with ongoing diversification observed during interepidemic periods. These observations reinforce the value of continuous genomic surveillance, particularly during cryptic transmission phases when adaptive mutations might emerge. Although further evidence is needed, observed trends in climate variability and land-use change point to the potential benefit of targeted surveillance in settings that could be at increased risk, including urban centres and wetlands. FUNDING: This work was supported by the German Federal Ministry for Economic Cooperation and Development, the Rockefeller Foundation, and the Africa Centres for Disease Control and Prevention.

Rift Valley fever virus↗

Multidimensional prophage profiling of carbapenem-resistant Enterobacteriaceae in Thailand: a nationwide, multicentre, genomic study.

BACKGROUND: Prophages influence bacterial fitness, resistance, and evolution, yet their epidemiology remains poorly understood in carbapenem-resistant Enterobacteriaceae (CRE). In this nationwide study in Thailand, we aimed to describe prophage repertoires in clinical CRE isolates and to explore their potential relevance for molecular epidemiology. METHODS: We performed a nationwide, retrospective, genomic analysis of all CRE clinical isolates collected through our previous national surveillance study involving 11 hospitals in 11 provinces in Thailand between March 25, 2012, and Jul 21, 2017. Whole-genome sequencing data from 747 CRE isolates were analysed. Intact prophages were identified using PHAge Search Tool Enhanced Release (PHASTER) and clustered by nucleotide sequence similarity. Prophage profiles were compared across multilocus sequence types, carbapenemase genotypes, specimens, geography, and patient demographics (age and sex). FINDINGS: Of the included 747 CRE isolates, 170 (23%) were Escherichia coli and 577 (77%) were Klebsiella pneumoniae. 220 (29%) of 747 strains had been isolated from female patients and 264 (35%) from male patients; metadata on patient sex were missing for 263 (35%) isolates. The median patient age was 63 years (IQR 50-72). 71 (10%) of isolates were from blood, 283 (38%) from sputum, 284 (38%) from urine, and 109 (15%) from other specimens. 374 distinct prophage clusters were identified, with significantly more prophages per genome in K pneumoniae (mean 3&#xb7;01 [SD 1&#xb7;55]) than in E coli (1&#xb7;64 [1&#xb7;46]; p<0&#xb7;0001). Prophage repertoires largely mirrored bacterial multilocus sequence types. However, even within the highly clonal K pneumoniae sequence type 16 lineage, discrete prophage variation was identified, with common profiles observed in geographically dispersed patients. Respiratory K pneumoniae frequently carried a mosaic prophage with environmental signatures and a type VI secretion system, whereas blood-derived E coli harboured a prophage with immune-modulating genes. Distinct prophage clusters were observed across clinical specimens, age groups, carbapenemase genotype, and geographical region. Strains coharbouring blaNDM-1 plus blaOXA-232 (114 [15%] of 747) had the highest prophage loads. INTERPRETATION: The prophage content was shaped by the bacterial lineage, ecological niche, and temporal dynamics, providing an additional layer of epidemiological resolution beyond conventional genome typing. Integrating prophage profiling into molecular surveillance frameworks could help to identify transmission events, improve infectious source attribution, and enhance infection control strategies. FUNDING: Japan Agency for Medical Research and Development.

Female↗

Comparison of phylogenetic metrics of transmission between symptomatic and asymptomatic tuberculosis in individuals who were incarcerated in Brazil in 2008-24: a retrospective genomic epidemiology study.

BACKGROUND: Tuberculosis control efforts have traditionally targeted symptomatic individuals; however, the role of asymptomatic cases in sustaining transmission is increasingly recognised. We aimed to quantify the contribution of asymptomatic tuberculosis to recent transmission using genomic and epidemiological data from a high-transmission setting. METHODS: We conducted a retrospective genomic epidemiology study of Mycobacterium tuberculosis isolates collected in Mato Grosso do Sul, Brazil, between Aug 25, 2008, and March 19, 2024. Available isolates underwent whole-genome sequencing. Demographic, clinical, incarceration history, and laboratory metadata were obtained from surveillance records. From Jan 1, 2017, to March 19, 2024, active case finding was conducted in the state's three largest prisons (all male-only facilities), during which sputum samples were collected from individuals irrespective of symptoms and tested using GeneXpert and culture. Comparisons of transmission between individuals with and without symptoms were restricted to individuals who were incarcerated and were identified through active case finding and for whom high-quality, M tuberculosis lineage 4 genomes were available. Metrics of recent transmission included phylogenetic clustering, time-scaled haplotype density (THD), local branching index (LBI), and transmission probabilities inferred using Bayesian Reconstruction and Evolutionary Analysis of Transmission Histories. FINDINGS: 4448 tuberculosis cases were notified in Mato Grosso do Sul in 2008-24. After excluding cases for which M tuberculosis isolates were not available or had low sequencing quality, who had contaminated cultures or mixed infection, or who were infected with non-lineage 4 M tuberculosis, we included 2362 lineage 4 M tuberculosis isolates with high-quality genome sequences. 1849 (78&#xb7;3%) of 2362 isolates were part of a genomic cluster. Among 2362 individuals with tuberculosis, 1137 (48&#xb7;1%) were incarcerated at diagnosis. Of these individuals, 505 were identified through active case finding in three male-only prisons. The median age was 30 years (IQR 25-37); 304 (60&#xb7;2%) had mixed ethnicity, 90 (17&#xb7;8%) were White, 56 (11&#xb7;1%) were Black, 13 (2&#xb7;6%) were Indigenous, and six (1&#xb7;2%) were Asian. 277 (54&#xb7;9%) had symptomatic disease and 228 (45&#xb7;1%) had asymptomatic tuberculosis. There were no significant differences between symptomatic and asymptomatic individuals in phylogenetic clustering (213 [76&#xb7;9%] of 277 vs 195 [85&#xb7;5%] of 228; p=0&#xb7;37), THD (median 0&#xb7;39 [IQR 0&#xb7;06-0&#xb7;62] vs 0&#xb7;50 [0&#xb7;09-0&#xb7;65]; p=0&#xb7;12), or LBI (0&#xb7;00863 [0&#xb7;00810-0&#xb7;00988] vs 0&#xb7;00871 [0&#xb7;00829-0&#xb7;01020]; p=0&#xb7;088). Bayesian transmission trees showed no significant difference in the number of secondary infections inferred from symptomatic compared with asymptomatic individuals (p=0&#xb7;56). These findings were consistent across genomic clusters and robust to model assumptions. INTERPRETATION: We identified no differences in transmission between individuals who were symptomatic and those who were asymptomatic using multiple genomic measures. In this high-transmission setting, where systematic screening is implemented, our findings indicate that asymptomatic tuberculosis substantially contributes to tuberculosis transmission at the population level. These results suggest that symptom-based case detection alone is likely to be insufficient to interrupt transmission and highlight the importance of expanded screening strategies in high-risk populations. FUNDING: US National Institutes of Health and the Brazilian National Research Council (CNPq).

Humans↗

Database and tools for analysis of topographic organization and map transformations in major projection systems of the brain.

Integration of dispersed and complicated information collected from the brain is needed to build new knowledge. But integration may be hampered by rigid presentation formats, diversity of data formats among laboratories, and lack of access to lower level data. We have addressed some of the fundamental issues related to this challenge at the level of anatomical data, by producing a coordinate based digital atlas and database application for a major projection system in the rat brain: the cerebro-ponto-cerebellar system. This application, Functional Anatomy of the Cerebro-Cerebellar System in rat (FACCS), is available via the Rodent Brain WorkBench (http://www.rbwb.org). The data included are x,y,z-coordinate lists describing exact distributions of tissue elements (axonal terminal fields of axons, or cell bodies) that are labeled with axonal tracing techniques. All data are translated to a common local coordinate system to facilitate across animal comparison. A search capability allows queries based on, e.g. location of tracer injection sites, tracer category, size of the injection sites, and contributing author. A graphic search tool allows the user to move a volume cursor inside a coordinate system to detect particular injection sites having connections to a specific tissue volume at chosen density levels. Tools for visualization and analysis of selected data are included, as well as an option to download individual data sets for further analysis. With this application, data and metadata from different experiments are mapped into the same information structure and made available for re-use and re-analysis in novel combinations. The application is prepared for future handling of data from other projection systems as well as other data categories.

Anatomy, Artistic↗

Data input module for Birth Defects Systems Manager.

The need for a computational bioinformatics infrastructure to manage the vast digital information from functional genomics and proteomics motivated us to develop Birth Defects Systems Manager (BDSM) as an open resource to facilitate analysis and discovery in developmental biology and developmental toxicity. This report describes the design, development and implementation of the data loading module of BDSM, referred to as LoadBDSM. It includes a shared data directory resource that can be granted various levels of security for different research groups or investigators to manage experimental datasets individually or in groups. LoadBDSM allows the upload of data and experiment details using controlled semantics for developmental exposure (toxicant, dosing scenario, intervention), biological sample (species, tissue, stage) and disease outcome (time, risk, phenotype). It adheres to existing controlled vocabulary plus rules of inference (ontologies) for experiment, data and metadata annotations. LoadBDSM extends the capabilities of BDSM to support the emergence of "embryo-formatics" defined here as the data, information and knowledge from genomic sciences applied to, or derived from, an embryological context. This includes, but is not limited to, delineating pathways and biological regulatory networks for specific chemicals or classes of developmental toxicants, developing novel biomarkers indicative of exposure and/or predictive of adverse effects, and integrating modern computing and information technology with data from molecular biology.

Abnormalities, Drug-Induced↗

Meningioma transcriptomic landscape demonstrates novel subtypes with regional associated biology and patient outcome.

Meningiomas, although mostly benign, can be recurrent and fatal. World Health Organization (WHO) grading of the tumor does not always identify high-risk meningioma, and better characterizations of their aggressive biology are needed. To approach this problem, we combined 13 bulk RNA sequencing (RNA-seq) datasets to create a dimension-reduced reference landscape of 1,298 meningiomas. The clinical and genomic metadata effectively correlated with landscape regions, which led to the identification of meningioma subtypes with specific biological signatures. The time to recurrence also correlated with the map location. Further, we developed an algorithm that maps new patients onto this landscape, where the nearest neighbors predict outcome. This study highlights the utility of combining bulk transcriptomic datasets to visualize the complexity of tumor populations. Further, we provide an interactive tool for understanding the disease and predicting patient outcomes. This resource is accessible via the online tool Oncoscape, where the scientific community can explore the meningioma landscape.

Meningioma↗

Protocol for histology-anchored macroscopic staging of gonadal maturity in exploited fishes.

Here, we present a protocol to assign gonadal maturity stages in commercially exploited fishes using a histology-anchored workflow. We describe steps for recording field metadata, photographing gonads, fixing central gonadal tissue, and processing paraffin sections. We then detail procedures for staining sections with hematoxylin and eosin, diagnosing gametogenic features, and assigning stages using a common reproductive-phase framework with species- and sex-specific reference descriptors. This protocol standardizes documentation and decision logic rather than proposing a new maturity scale.

Developmental biology↗

Continuing dental education on the World Wide Web.

Continuing dental education (CDE) courses delivered on the World Wide Web (Web CDE) offer numerous advantages over traditional CDE; however, two major issues--location of suitable courses and course quality--need resolution. Locating high-quality courses is difficult due to the lack of the standardized metadata that allows search engines to match courses to practitioners' needs. Web directories created by professional organizations are beginning to show promise, but require further development. Search engines and Web directories are discussed and improvements currently underway summarized. Course quality remains a highly significant concern. A national effort to create Web CDE course quality standards is underway that includes proposed standards. These proposed standards are summarized and used to comment on the current state of Web CDE courses. Examples are given when possible. Three emerging Web CDE technologies and a look to the future of Web CDE are discussed.

Computer-Assisted Instruction↗

A system for simultaneous multiple subject, multiple stimulus modality, and multiple channel collection and analysis of sensory evoked potentials.

A system has been developed for collecting sensory evoked potentials simultaneously from multiple channels for multiple subjects at up to 80 kHz sample rate per channel. Sample rates up to 200 kHz are available for four or less chambers and a single channel per chamber. A variety of visual, somatosensory, and auditory stimuli may be presented singly or simultaneously. Collected waveforms are associated with searchable text (metadata) to allow convenient selection from a relational database. Multiple waveforms can then be easily grouped for analysis and processed. Results can be exported to other software for further graphics or statistical processing. Scripting and event logging are available to provide automation and improve data confidence. Sample data are presented from control animals for each of the sensory modalities for comparison with historical data collected from other systems.

Animals↗

The impact of Life Science Identifier on informatics data.

Since the Life Science Identifier (LSID) data identification and access standard made its official debut in late 2004, several organizations have begun to use LSIDs to simplify the methods used to uniquely name, reference and retrieve distributed data objects and concepts. In this review, the authors build on introductory work that describes the LSID standard by documenting how five early adopters have incorporated the standard into their technology infrastructure and by outlining several common misconceptions and difficulties related to LSID use, including the impact of the byte identity requirement for LSID-identified objects and the opacity recommendation for use of the LSID syntax. The review describes several shortcomings of the LSID standard, such as the lack of a specific metadata standard, along with solutions that could be addressed in future revisions of the specification.

Computational Biology↗

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa↗

ThermoData Engine (TDE): software implementation of the dynamic data evaluation concept.

The first full-scale software implementation of the dynamic data evaluation concept {ThermoData Engine (TDE)} is described for thermophysical property data. This concept requires the development of large electronic databases capable of storing essentially all experimental data known to date with detailed descriptions of relevant metadata and uncertainties. The combination of these electronic databases with expert-system software, designed to automatically generate recommended data based on available experimental data, leads to the ability to produce critically evaluated data dynamically or 'to order'. Six major design tasks are described with emphasis on the software architecture for automated critical evaluation including dynamic selection and application of prediction methods and enforcement of thermodynamic consistency. The direction of future enhancements is discussed.

Journal Article↗

Bringing chemical data onto the Semantic Web.

Present chemical data storage methodologies place many restrictions on the use of the stored data. The absence of sufficient high-quality metadata prevents intelligent computer access to the data without human intervention. This creates barriers to the automation of data mining in activities such as quantitative structure-activity relationship modelling. The application of Semantic Web technologies to chemical data is shown to reduce these limitations. The use of unique identifiers and relationships (represented as uniform resource identifiers, URIs, and resource description framework, RDF) held in a triplestore provides for greater detail and flexibility in the sharing and storage of molecular structures and properties.

Journal Article↗

ChemSem: an extensible and scalable RSS-based seminar alerting system for scientific collaboration.

A seminar announcement system based on the extensive use of XML-based data structures, CML/MathML for carrying more domain-specific molecular content, and open source software components is described. The output is a resource description framework (RDF) site summary (RSS) feed, which potentially carries many advantages over conventional announcement mechanisms, including the ability to aggregate and then sort multiple and diverse RSS feeds on the basis of declared metadata and to feed into RDF-based mechanisms for establishing links between different subject areas.

Journal Article↗