Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

GenBank mining reveals novel insights into Rhizobium phylogeny: Identical 16S rRNA sequences are mainly uncoupled from species designation, host plant, and geographic origin: How this search suggested the definition of a direct 'microbial h-index'.

16S rDNA is the historical gold standard for bacterial identification, particularly in metabarcoding approaches reliant on sequence similarity thresholds. We analyzed 6,660 Rhizobium 16S rRNA gene sequences from GenBank to examine the relationship between sequence identity and three metadata: species name, host plant, and geographic origin. Using an iterative BLAST-based pipeline, we detected 116,069 pairwise matches and assessed concordance among sequences (average length 1,328 bp) sharing 100% identity. For those in which the organism name, host plant and country of isolation were present in the record, surprisingly, 66.59% of identical sequence pairs showed full discordance across all three metadata, while only 1.40% shared the same name, host, and country. The most widespread sequence, detected 371 times, was associated with over 56 different host plants across 25 countries and bore multiple species name designations. These results highlight a striking mismatch between the 16S barcode and the taxonomic, ecological, and phenotypic variability it is assumed to reflect, likely arising from the slow evolution of rRNA genes contrasted with the mobility of ecologically relevant genes via horizontal transfer on plasmids, transposons, and phages. Our findings further challenge the limitations of relying on 16S rRNA alone for fine-scale taxonomic and metadata-based inference in capturing the true functional and ecological diversity of bacteria, endorsing the critical importance of polyphasic taxonomic approaches that integrate genomic, phenotypic, and ecological data. An interesting byproduct of the analysis was to realize the possibility of treating these data as if they were 'citations.' The more one finds the same query sequence, the more that sequence can be considered biologically 'cited', i.e., re-proposed elsewhere in the world. Thus, one can also analyze the h-index of such a ranking. In our Rhizobium dataset, we calculated an h-index = 201, meaning the sequence ranked 201st had 202 identical homologues in GenBank. Although the research effort on given species is directly connected with it, this number provides a quantitative indicator of a taxon's sequence recurrence and distribution within public databases, independent of nomenclatural inconsistencies, offering a novel framework for assessing bacterial representation across global datasets.

RNA, Ribosomal, 16S↗

Network-based integration of metabolomics data from large-scale repositories.

INTRODUCTION: Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. OBJECTIVES: This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. METHODS: We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github.com/EloisaRL/Metabolomic-data-analysis-app/tree/main . RESULTS: As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. CONCLUSION: Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Metabolomics↗

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis↗

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software↗

The need for standardization and improved open (meta)data practices in metaproteomics.

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices. Video Abstract.

Proteomics↗

Community-driven advances in computational mass spectrometry: The perspective of EuBIC-MS members.

Advances in data acquisition, artificial intelligence, and integrative bioinformatics are driving the rapid evolution of computational mass spectrometry, and in turn, transforming modern proteomics, metabolomics, and lipidomics. These developments have greatly increased the scale and complexity of mass spectrometry data, underscoring the importance of evolving accurate, transparent, efficient and reproducible data processing workflows. Addressing these challenges requires collaborative innovation that brings together expertise in software engineering, statistics, and biology. The European Bioinformatics Community for Mass Spectrometry (EuBIC-MS), an initiative of the European Proteomics Association (EuPA), fosters a culture of open, community-driven development through its biennial Developers Meetings and Winter Schools. This commentary summarizes the scientific background and outcomes of the EuBIC-MS Developers Meeting 2025, which took place in Novacella, Italy. Three keynote presentations highlighted major frontiers in the field: deep proteome and phosphoproteome profiling, text mining for protein-protein interaction extraction, and scalable proteomics for AI-driven drug discovery. Seven community-selected hackathons addressed emerging challenges such as single-cell proteomics data analysis, FAIR metadata extraction, deep learning frameworks, R-Python interoperability, and DIA validation. Together, these efforts demonstrate the potential for scientific and technical innovation to arise from open collaboration, and highlight how community-driven initiatives can accelerate progress in computational mass spectrometry. SIGNIFICANCE: Modern proteomics increasingly depends on computational advances to translate complex, high-dimensional data into biological knowledge. The EuBIC-MS Developers Meeting 2025 exemplifies how community-driven collaboration can directly accelerate this process by bringing together experts from bioinformatics, statistics, and experimental proteomics to co-develop open, interoperable, and reproducible analytical tools. By fostering shared software frameworks, transparent benchmarking, and collaborative problem solving, the EuBIC-MS community helps ensure that technological innovation translates into reliable biological insights. This collaborative model strengthens the foundation for quantitative, system-level understanding of proteomes and establishes a sustainable path for integrating artificial intelligence and next-generation data acquisition into routine biological discovery. This commentary shows some current highlights in the field of computational mass spectrometry and community-based approaches undertaken during the most recent Developers Meeting to solve these challenges. The approaches discussed and initiated during the meeting - ranging from deep proteome profiling and phosphosite mapping to text mining, single-cell data analysis, and FAIR metadata extraction - address key bottlenecks that currently limit the biological interpretability and comparability of proteomics data.

Mass Spectrometry↗

Effect of repeated mass drug administration on the transmission of yaws: a retrospective genomic epidemiology study.

BACKGROUND: Yaws, a neglected tropical disease caused by Treponema pallidum subspecies pertenue (T p pertenue), has evaded eradication, in part due to a high proportion of asymptomatic cases. Repeated mass drug administration (MDA), whereby an entire population is repeatedly treated irrespective of disease, could provide a solution. Here, we aimed to investigate the effect of MDA on the genomic epidemiology of T p pertenue. METHODS: We conducted a retrospective genomic epidemiology study on samples collected during a cluster-randomised trial of mass administration of azithromycin for yaws eradication in the Namatanai District of Papua New Guinea. Participants were in 38 wards (administrative units encompassing several villages) in three local-level government areas (LLGs). The experimental group received an initial round of MDA followed by two further rounds 6 months and 12 months after the first round. The control group received one round of MDA followed by two rounds of treatment targeting clinical cases and contacts only, on the same schedule as the MDA in the experimental group. A follow-up survey on both groups was done 18 months after the first MDA round. Swab samples were collected at each round from ulcerative and nodular skin lesions, and blood was collected by finger-prick for serological testing at 18 months. Metadata on ulcer size (cm) and duration (days) were recorded at each round, and treponemal and non-treponemal antibodies were recorded at 18 months. Samples from swabs positive for T p pertenue underwent library preparation and whole-genome sequencing. We examined the phylogenetic relationships between genomes, linking them with geospatial and patient metadata to understand the impact of MDA on T p pertenue diversity and transmission. FINDINGS: Swabs collected from 297 individuals with active yaws from April 30, 2018, to Nov 2, 2019, yielded 222 good-quality Tp pertenue genomes. We identified 20 sublineages of T p pertenue in the control group and 21 in the experimental group at the beginning of the study. At the end of the study, there were 13 sublineages in the control group and three in the experimental group, of which two persisted in both groups. Three sublineages not detected at baseline were observed in the control group after commencing MDA. The two sublineages that persisted in both groups had non-synonymous mutations in penicillin-binding proteins. One of these sublineages evolved macrolide resistance in three individuals and was associated with lowered treponemal antibody (p=0&#xb7;0036) and longer ulcer duration (p=0&#xb7;015). Despite the study taking place within a small island, sublineages were geographically clustered, with pairs of samples from the same ward (odds ratio 7&#xb7;1, 95% CI 5&#xb7;7-8&#xb7;8; p<0&#xb7;0001) or neighbouring wards (4&#xb7;3, 3&#xb7;3-5&#xb7;4; p<0&#xb7;0001) more likely to share the same sublineages compared with pairs from different LLGs. Additionally, older individuals were more likely to share sublineages than were younger individuals (1&#xb7;5, 1&#xb7;2-1&#xb7;9; p<0&#xb7;0001). INTERPRETATION: Repeated MDA was successful in reducing and maintaining the genetic diversity of T p pertenue at a low level but was associated with the development of macrolide resistance. Yaws re-emergence after MDA was attributed to multiple sublineages, of which the majority were detected in the population before MDA. Participants within the same ward were more likely to share sublineages than those that were more widely geographically separated, suggesting that re-emergence was driven by local transmission. These findings could inform future yaws elimination strategies. FUNDING: European Research Council, EU, Provincial Deputation of Barcelona, Barber&#xe0; Solid&#xe0;ria Foundation, Wellcome, and Fundaci&#xf3; "la Caixa".

Adolescent↗

Transition of Staphylococcus aureus tetracycline resistance plasmid pT181 from independent multicopy replicon to predominantly integrated chromosomal element over 65 years.

Mobile genetic elements (MGEs), including plasmids, phages and genome islands, are major sources of bacterial genetic diversity. The small plasmid pT181 confers tetracycline resistance in bacterial pathogen Staphylococcus aureus via an efflux pump, TetK. pT181 was one of the earliest sequenced S. aureus plasmids, and has been isolated in both clinical and livestock-associated strains for decades, both as an independent replicon and integrated in the chromosome as part of staphylococcal cassette chromosome mec (SCCmec). Bacterial genome analysis tools and high-quality sequences with metadata are publicly available, but these resources remain underleveraged for examining historical data, especially when studying the spread of MGEs across a species and over time. Using publicly available reads and metadata, we explored the evolution of pT181 over almost seven decades of samples to identify temporal trends in sequence evolution, copy number changes, and spread across S. aureus and beyond. pT181 was prevalent across S. aureus (found in 9.5% of 83,366 genomes tested), with a conserved sequence outside of three hypervariable regions. The history of pT181 since 1954 is characterized by spread across strains, significant variation in plasmid copy number of the independent replicon, and increasing frequency of integration of the plasmid into the S. aureus chromosome. We have identified multiple chromosomal integration locations of the plasmid, including outside of the previously characterized SCCmec. We find that pT181 has been transferred across staphylococcaceae and into a Gram-negative species. The repeated integration of pT181 into the chromosome may indicate co-evolution of the plasmid and the host, potentially to facilitate increased antibiotic resistance.

Journal Article↗

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models↗

Organizing medical networked information (OMNI).

The Internet has become a major source of biomedical information over the last 5 years. Several projects have recently been established to help users find respectable information sources quickly. OMNI (Organizing Medical Networked Information) is one such filtering and indexing project. OMNI has focused on the quality of information and the application to Internet resources of standard tools for organizing information such as the National Library of Medicine's Medical Subject Headings and the Dublin Core metadata format. Now two years old, the OMNI project fulfils a valuable role for the UK biomedical community, through its gateway service (http:@omni.ac.uk), its printed resource guides and its training workshop programme. OMNI is also a focus for biomedical metadata activities in the UK. The gateway continues to grow in size and further work on information quality issues and integration is planned.

Abstracting and Indexing↗

Performance Profiles of Short DNA Barcode Segments for Family Level Detection of Asteraceae Within Asterales.

Short DNA barcodes may facilitate sequence recovery from degraded material, but their ability to retain target-family identity while excluding related taxa varies among genomic regions. We computationally evaluated 16 nuclear, plastid, and mitochondrial marker regions from 11 Asterales families using 279,956 NCBI locus-record matches and an accession-disjoint discovery/test design. Thirty-one candidate segments of 50-200 bp (mean, 98.55 bp) were screened in discovery data and evaluated for within-Asteraceae sequence recall, differentiation from non-Asteraceae Asterales, in silico primer behavior, phylogenetic placement, and exploratory matching across 808 metadata-defined metagenomic samples. Conserved regions such as matR and rbcL showed high within-Asteraceae identity, whereas ITS1, ITS, and trnH-psbA showed larger differences from related-family backgrounds; ITS2 and ycf1 showed intermediate profiles. Candidate segments were placed within or immediately adjacent to Asteraceae reference branches in segment-specific maximum-likelihood analyses, although support and topology varied among regions. Metadata-defined target-containing groups had higher mean query coverage and identity than background groups; because target presence was not independently verified and no classifier was fitted, these comparisons were descriptive and did not estimate diagnostic accuracy. Definitionally linked sequence statistics were interpreted as structural associations rather than evidence of causal evolutionary mechanisms. These results provide a family-level computational comparison of candidate short segments for Asteraceae detection within Asterales. Species identification, operational marker combinations, threshold robustness, and laboratory performance require validation using taxonomically dense, voucher-linked, and experimentally characterized datasets.

Asteraceae↗

Labeling and filtering of medical information on the Internet.

Internet information undergoes no quality controls and virtually anybody can publish anything. Because of this, it is difficult for searchers to take information retrieved from the Internet at face value. A related problem is the uncontrolled promotion of medical products on the Internet. A further problem of today's Internet is that authors use no uniform keywords and other descriptive labels, which deteriorates the quality of search results. A solution for all these problems could be widespread use of descriptive and evaluative metainformation associated with medical Internet information. Our concept is based on a recently established infrastructure for assigning metadata to Internet information, the so-called PICS Standard (Platform for Internet Content Selection). We prototyped a PICS-based rating vocabulary for medical information (med-PICS), containing descriptive and evaluative categories, to be used by the webauthor and third-party label services (such as medical associations), respectively. We propose an international effort to assign metadata to medical Internet information.

Humans↗

Neuronal database integration: the Senselab EAV data model.

We discuss an approach towards integrating heterogeneous nervous system data using an augmented Entity-Attribute-Value (EAV) schema design. This approach, widely used in implementing electronic patient record systems (EPRSs), allows the physical schema of the database to be relatively immune to changes in domain knowledge. This is because new kinds of facts are added as data (or as metadata) rather than hard-coded as the names of newly created tables or columns. Because the domain knowledge is stored as metadata, a framework developed in one scientific domain can be ported to another with only modest revision. We describe our progress in creating a code framework that handles browsing and hyperlinking of the different kinds of data.

Databases, Factual↗

Representation by standard terminologies of health status concepts contained in two health status assessment instruments used in rheumatic disease management.

Health and functional status data have been shown to have clinical utility in predicting outcome. Various metadata registries in the form of patient self-administered health assessment questionnaires have been incorporated into routine clinical care and clinical research of patients with rheumatic disease. Examples of such health assessment instruments are the Clinical Health Assessment Questionnaire (CLINHAQ) and the Modified Health Assessment Questionnaire (MHAQ). These instruments contain concepts that are an integral part of the health and functional status domain. Using an automated indexing tool we examined the clinical content coverage by SNOMED RT and the Unified Medical Language System (UMLS) Metathesaurus for health and functional status concepts identified in the MHAQ and CLINHAQ. Significant differences existed between the overall representational ability of SNOMED and UMLS for concepts identified in the MHAQ (49%, vs. 77% respectively, p < .005) and for concepts identified in the CLINHAQ (30% vs. 64% respectively p < .005). Representational capability by SNOMED-RT and UMLS for concepts in a given health assessment instrument was carried across four semantic classes of "attitudes", "symptoms", "activities", and "social attributes". The conceptual content coverage of health status assessment concepts contained in the MHAQ and CLINHAQ by SNOMED-RT and UMLS was incomplete but better for UMLS with its panoply of vocabulary sources. This observed overall improved representation by UMLS appeared to be due to better representation of concepts in "activities" and "social attributes" semantic classes. Representation of health or functional status concepts in a computerized medical record should be founded on a universally agreed concept model of that domain. Established functional and health status metadata registries can serve as important sources for concepts and candidate classes within that domain.

Health Status↗

IML: An image markup language.

Image Markup Language is an extensible markup language (XML) schema used to describe both image metadata and annotations. It describes both data pertaining to an entire image, and data that are tied to specific regions or features of the image. Developed for a specific domain in Medical Education, this pa-per describes extensions to take advantage of the Dublin Core metadata standard, and of an XML schema for vector graphics representation. We have developed a prototype system of open source tools implementing an authoring system, a client system, and an image annotation database which can be queried though the Web.

Diagnostic Imaging↗

MeSHmap: a text mining tool for MEDLINE.

Our research goal is to explore text mining from the metadata included in MEDLINE documents. We present MeSHmap our prototype text mining system that exploits the MeSH indexing accompanying MEDLINE records. MeSHmap supports searches via PubMed followed by user driven exploration of the MeSH terms and subheadings in the retrieved set. The potential of the system goes beyond text retrieval. It may also be used to compare entities of the same type such as pairs of drugs or pairs of procedures etc. In addition there is the potential to generate maps of entities (drugs or diseases etc.) such that the strength of the link between two entities in the map represents their similarity as expressed in the MeSH metadata of the MEDLINE documents. Higher level operators have been proposed to support these comparison and mapping functions. This paper motivates and describes MeSHmap. Future work will include user evaluations of the system.

Abstracting and Indexing↗

SQLGEN: a framework for rapid client-server database application development.

SQLGEN is a framework for rapid client-server relational database application development. It relies on an active data dictionary on the client machine that stores metadata on one or more database servers to which the client may be connected. The dictionary generates dynamic Structured Query Language (SQL) to perform common database operations; it also stores information about the access rights of the user at log-in time, which is used to partially self-configure the behavior of the client to disable inappropriate user actions. SQLGEN uses a microcomputer database as the client to store metadata in relational form, to transiently capture server data in tables, and to allow rapid application prototyping followed by porting to client-server mode with modest effort. SQLGEN is currently used in several production biomedical databases.

Computer Communication Networks↗

Automatic query mapping among genomic databases: a pilot exploration.

As databases in the human genome project proliferate, it is important for users of one genomic database to identify similar or inconsistent data in other autonomously developed genomic databases. To do so, the user needs to issue the same query across multiple databases. We describe an approach that allows a query issued against one database to be automatically mapped to an equivalent query against another structurally different database. Our approach features two components: 1) a database designed to capture knowledge (metadata) that describes the correspondences among individual database components and 2) a module that utilizes the metadata to perform query mappings. As a demonstration, we apply our query mapping approach to two chromosome map databases (DB/12 and GDB).

Algorithms↗