Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Achieving evolvable Web-database bioscience applications using the EAV/CR framework: recent advances.

The EAV/CR framework, designed for database support of rapidly evolving scientific domains, utilizes metadata to facilitate schema maintenance and automatic generation of Web-enabled browsing interfaces to the data. EAV/CR is used in SenseLab, a neuroscience database that is part of the national Human Brain Project. This report describes various enhancements to the framework. These include (1) the ability to create "portals" that present different subsets of the schema to users with a particular research focus, (2) a generic XML-based protocol to assist data extraction and population of the database by external agents, (3) a limited form of ad hoc data query, and (4) semantic descriptors for interclass relationships and links to controlled vocabularies such as the UMLS.

Database Management Systems↗

Exploring the portability of informatics capabilities from a clinical application to a bioscience application.

This report describes XDesc (eXperiment Description), a pilot project that serves as a case study exploring the degree to which an informatics capability developed in a clinical application can be ported for use in the biosciences. In particular, XDesc uses the Entity-Attribute-Value database implementation (including a great deal of metadata-based functionality) developed in TrialDB, a clinical research database, for use in describing the samples used in microarray experiments stored in the Yale Microarray Database (YMD). XDesc was linked successfully to both TrialDB and YMD, and was used to describe the data in three different microarray research projects involving Drosophila. In the process, a number of new desirable capabilities were identified in the bioscience domain. These were implemented on a pilot basis in XDesc, and subsequently "folded back" into TrialDB itself, enhancing its capabilities for dealing with clinical data. This case study provides a concrete example of how informatics research and development in clinical and bioscience domains has the potential for synergy and for cross-fertilization.

Clinical Medicine↗

Proteomic data exchange and storage: the need for common standards and public repositories.

The ever increasing volumes of proteomic data now being produced by laboratories across the world have resulted in major issues in data storage and accessibility. The further demands of multilaboratory initiatives has highlighted issues when collaborators cannot import data generated within the same project but generated by different hardware types and processed by laboratory-specific work flows and analyses packages. There is an increasing need for common data standards that will allow the interchange of data between different instrumentation, search engines, and between laboratory databases. This could then lead to the establishment of data repositories from where benchmark datasets could be accessed and reanalyzed. The Human Proteome Organization is currently supporting efforts to establish such standards. The work of the Proteomics Standards Initiative has lead to the development of the mzData XML interchange standard and is now broadening its scope to produce a spectral analysis output format, mzIdent. Accompanying controlled vocabularies allow the accurate, while systematic, representation of metadata throughout both schema.

Benchmarking↗

The Primary Care Electronic Library (PCEL) five years on: open source evaluation of usage.

BACKGROUND: The Primary Care Electronic Library (PCEL) is a collection of indexed and abstracted internet resources. PCEL contains a directory of quality-assured internet material with associated search facilities. PCEL has been indexed, using metadata and established taxonomies. Site development requires an understanding of usage; this paper reports the use of open source tools to evaluate usage. This evaluation was conducted during a six-month period of development of PCEL. OBJECTIVE: To use open source to evaluate changes in usage of an electronic library. METHOD: We defined data we needed for analysis; this included: page requests, visits, unique visitors, page requests per visit, geographical location of users, NHS users, chronological information about users and resources used. RESULTS: During the evaluation period, page requests increased from 3500 to 10,000; visits from 1250 to 2300; and unique visitors from 750 to 1500. Up to 83% of users come from the UK, 15% were NHS users. The page requests of NHS users are slowly increasing but not as fast as requests by other users in the UK. PCEL is primarily used Monday to Friday, 9 a.m. to 5 p.m. Monday is the busiest day with use lessening through the week. NHS users had a different list of top ten resources accessed than non-NHS users, with only four resources appearing in both. CONCLUSIONS: Open source tools provide useful data which can be used to evaluate online resources. Improving the functionality of PCEL has been associated with increased use.

Information Dissemination↗

Pithos - a scalable and secure data container for FAIR-compliant research data management in life sciences.

Modern research techniques have led to exponential growth in the volume and complexity of scientific data. Consequently, managing these volumes securely and efficiently has become a major challenge. While all research domains face these challenges, life science research is particularly affected because current approaches often rely on a large set of different file formats, with metadata stored in separated databases or spreadsheets. This leads to fragmented datasets, orphaned data, and compromised research reproducibility. Traditional solutions also force researchers to choose between security and accessibility, with encrypted files preventing selective access and indexed formats lacking adequate security for sensitive data. These limitations are particularly problematic in large-scale genomic studies where researchers must decompress multi-gigabyte files to access specific regions, creating computational bottlenecks and inefficient network usage when working with cloud-stored datasets. We introduce Pithos, a next-generation file format specifically designed for scientific data management in distributed cloud environments. The format uses content-defined chunking to enable efficient deduplication across distributed storage systems, thereby reducing storage costs and bandwidth requirements. The append-only structure ensures data immutability and allows for incremental updates without compromising content. Benchmark results show that Pithos outperforms existing solutions in read and write performance, with comparable or improved storage efficiency.

Biological Science Disciplines↗

Distributed heterogeneous inspecting system and its middleware-based solution.

There are many cases when an organization needs to monitor the data and operations of its supervised departments, especially those departments which are not owned by this organization and are managed by their own information systems. Distributed Heterogeneous Inspecting System (DHIS) is the system an organization uses to monitor its supervised departments by inspecting their information systems. In DHIS, the inspected systems are generally distributed, heterogeneous, and constructed by different companies. DHIS has three key processes-abstracting core data sets and core operation sets, collecting these sets, and inspecting these collected sets. In this paper, we present the concept and mathematical definition of DHIS, a metadata method for solving the interoperability, a security strategy for data transferring, and a middleware-based solution of DHIS. We also describe an example of the inspecting system at WENZHOU custom.

Computer Communication Networks↗

Improving the precision of the keyword-matching pornographic text filtering method using a hybrid model.

With the flooding of pornographic information on the Internet, how to keep people away from that offensive information is becoming one of the most important research areas in network information security. Some applications which can block or filter such information are used. Approaches in those systems can be roughly classified into two kinds: metadata based and content based. With the development of distributed technologies, content based filtering technologies will play a more and more important role in filtering systems. Keyword matching is a content based method used widely in harmful text filtering. Experiments to evaluate the recall and precision of the method showed that the precision of the method is not satisfactory, though the recall of the method is rather high. According to the results, a new pornographic text filtering model based on reconfirming is put forward. Experiments showed that the model is practical, has less loss of recall than the single keyword matching method, and has higher precision.

Algorithms↗

eEurope 2002: Quality Criteria for Health Related Websites.

BACKGROUND: A number of organisations have begun to provide specific tools for searching, rating, and grading this information, while others have set up codes of conduct by which site providers can attest to their high quality services. The aim of such tools is to assist individuals to sift through the mountains of information available so as to be better able to discern valid and reliable messages from those which are misleading or inaccurate. OBJECTIVE: Recognising that European citizens are avid consumers of health related information on the internet and recognising that they are already using the types of rating system described above, the European Council at Feira on June 19-20 2000 supported an initiative within eEurope 2002 to develop a core set of Quality Criteria for Health Related Websites. The specific aim was to draw up a commonly agreed set of simple quality criteria on which Member States, as well as public and private bodies, may draw in the development of quality initiatives for health related websites. These criteria should be applied in addition to relevant Community law. METHODS: A meeting was held during 2001 which drew together key players from Government departments, International Organisations, non-governmental organisations and industry, to explore current practices and experiments in this field. Some sixty invited participants from all the Member States, Norway, Switzerland, and the United States of America took part in the meeting of June 7-8, 2001: they included delegates from industrial, medical, and patient interest groups, delegates from Member States' governments, and key invited speakers from the field of health information ethics. These individuals, and many others, also took part in the web-based consultation which was open from august to November 2001. RESULTS: The broad headings for quality criteria identified include Transparency and Honesty, Authority, Privacy and data protection, Updating of information, Accountability, Responsible partnering, Editorial policy, Accessibility, the latter includes attention to guidelines on physical accessibility as well as general findability, searchability, readability, usability, etc. A metadata labelling system may be used to make health data more findable. Such a system may also be used in conjunction with quality criteria to give higher ranking by search engines to those sites or pages labelled as complying with defined quality criteria. CONCLUSIONS: The set of quality criteria is based upon a broad consensus among specialists in this field, health authorities, and prospective users. It is now to be expected that national and regional health authorities, relevant professional associations, and private medical website owners will 1) implement the Quality Criteria for Health Related Websites in a manner appropriate to their website and consumers; 2) develop information campaigns to educate site developers and citizens about minimum quality standards for health related websites; 3) draw on the wide range of health information offered across the European Union and localise such information for the benefit of citizens (translation and cultural adaptation); 4) exchange information and experience at European level about how quality standards are being implemented.

Europe↗

Taxonomic informatics tools for the electronic Nomenclator Zoologicus.

Given the current trends, it seems inevitable that all biological documents will eventually exist in a digital format and be distributed across the internet. New network services and tools need to be developed to increase retrieval rates for documents and to refine data recovery. Biological data have traditionally been well managed using taxonomic principles. As part of a larger initiative to build an array of names-based network services that emulate taxonomic principles for managing biological information, we undertook the digitization of a major taxonomic reference text, Nomenclator Zoologicus. The process involved replicating the text to a high level of fidelity, parsing the content for inclusion within a database, developing tools to enable expert input into the product, and integrating the metadata and factual content within taxonomic network services. The result is a high-quality and freely available web application (http://uio.mbl.edu/NomenclatorZoologicus/) capable of being exploited in an array of biological informatics services.

Animals↗

The European Health Data Space and the Secondary Use of Sensitive Health Data.

INTRODUCTION: The European Health Data Space (EHDS) is one of the European Union's most ambitious data-governance projects. It aims to create a common framework through which electronic health data can be accessed and reused across Member States for care, research, innovation, policy, and public-interest purposes. Its practical viability depends not only on digital infrastructure, but also on legal, ethical, and organisational harmonisation, particularly for genetic and genomic data. METHODS: This paper examines the EHDS with emphasis on the secondary use of health data. It reviews the EHDS institutional architecture, discusses Finland's Findata as a national model for structured access, and analyses challenges for data holders and data donors, including interoperability, governance burdens, privacy protection, residual re-identification risk, and genomic-data sensitivity. RESULTS: A cross-border cancer-genomics case study shows that the EHDS can streamline data discovery and the routing of access requests, but does not by itself eliminate legal fragmentation, heterogeneous ethics review, and consent-related barriers. DISCUSSION: Effective implementation will require harmonisation beyond infrastructure, including clearer consent standards, more consistent ethics procedures, interoperable metadata, and proportionate safeguards for genomic data.

Electronic Health Records↗

Genome-resolved analysis of colonization factor repertoires reveals ecological stratification in cervid gut microbiomes.

INTRODUCTION: Colonization factors (CFs) are important microbial traits associated with persistence and host adaptation in the gut, yet their large-scale organization in cervid gut microbiomes remains unclear. METHODS: A total of 3,311 non-redundant high-quality metagenome-assembled genomes (MAGs), derived from 688 cervid gut metagenomic samples across 15 publicly available projects and one in-house dataset, were analyzed. CF-associated genes were identified by comparison against the GHA CF database, and CF repertoires were characterized at genome, host-species, and gastrointestinal-segment levels. RESULTS: A total of 138,729 CF-associated genes spanning 71 CF families were identified. MAGs from Cervinae contained richer CF repertoires than those from Caprinae, and CF47 (Peptidase_C69), CF24_29 (QueH), and CF18 (Glycos_transf_2) were among the most prevalent families. CF repertoires were strongly structured by taxonomy, showed a moderate association with bacterial phylogenetic distance, and formed two recurrent genome-level configurations with distinct KEGG functional profiles. Integration of sample metadata further revealed differentiation of CF repertoires across host species and gastrointestinal segments, representing the major ecological dimensions examined in this study. Segment-associated CF variation was accompanied by redistribution of broader functional profiles, including enrichment of carbohydrate and lipid metabolism in the jejunum, membrane transport in the ileum, xenobiotics biodegradation in the cecum, and environmental adaptation in the rumen. DISCUSSION: These findings provide a genome-resolved view of CF repertoire organization in cervid gut microbiomes and demonstrate that colonization-associated functions are structured across microbial lineages and ecological contexts. This study highlights the importance of considering microbial taxonomy and host-associated environments when interpreting the distribution of CF repertoires in mammalian gut ecosystems.

Cervidae↗

Genomic and ecological systems-thinking framework for pathogenic Leptospira in Puerto Rico.

INTRODUCTION: Leptospirosis is a complex zoonotic disease requiring high-resolution surveillance. A systems-thinking framework was used to connect genomic and ecological data and map the geographic and host-based structuring of co-circulating pathogenic Leptospira lineages in Puerto Rico. METHODS: Forty-four core genomes of L. interrogans, L. borgpetersenii, and L. kirschneri from human, domestic, and wildlife hosts were analyzed. Spatiotemporal and landscape metadata were integrated using root-to-tip regression, isolation-by-distance profiling and calibrated single-nucleotide polymorphism (SNP) thresholds (≤1, ≤5, and ≤10 SNPs) to define transmission clusters. RESULTS: Leptospira species exhibited distinct ecological pathways partitioned by geography, explaining 56% of genomic variance for L. interrogans and 91% for L. borgpetersenii (PERMANOVA). L. interrogans displayed high landscape connectivity across multiple hosts, forming localized networks (≤1 to ≤10 SNPs) that capture active spillovers (human-to-rat linkages at ≤1 SNP) and resolved into rodent host-specific lineages (R2 = 0.34). Conversely, L. borgpetersenii showed spatial and temporal genomic homogeneity and a lack of host-associated structure within an unpartitioned transmission pool dominated by Mus musculus. As a result, fixed genomic thresholds yielded disparate outcomes: L. interrogans resolved into 4 to 5 discrete, expanding clusters, whereas L. borgpetersenii grouped into a single uniform population at the ≤10-SNP threshold. CONCLUSION: Co-circulating pathogenic leptospires occupy distinct ecological niches shaped by varying host restriction and environmental persistence. Fixed genomic thresholds lack universal applicability; effective genomic epidemiological surveillance must employ species-specific threshold calibration to accurately map transmission pathways.

Puerto Rico↗

Cruella: developing a scalable tissue microarray data management system.

CONTEXT: Compared with DNA microarray technology, relatively little information is available concerning the special requirements, design influences, and implementation strategies of data systems for tissue microarray technology. These issues include the requirement to accommodate new and different data elements for each new project as well as the need to interact with pre-existing models for clinical, biological, and specimen-related data. OBJECTIVE: To design and implement a flexible, scalable tissue microarray data storage and management system that could accommodate information regarding different disease types and different clinical investigators, and different clinical investigation questions, all of which could potentially contribute unforeseen data types that require dynamic integration with existing data. DESIGN: The unpredictability of the data elements combined with the novelty of automated analysis algorithms and controlled vocabulary standards in this area require flexible designs and practical decisions. Our design includes a custom Java-based persistence layer to mediate and facilitate interaction with an object-relational database model and a novel database schema. User interaction is provided through a Java Servlet-based Web interface. RESULTS: Cruella has become an indispensable resource and is used by dozens of researchers every day. The system stores millions of experimental values covering more than 300 biological markers and more than 30 disease types. The experimental data are merged with clinical data that has been aggregated from multiple sources and is available to the researchers for management, analysis, and export. CONCLUSION: Cruella addresses many of the special considerations for managing tissue microarray experimental data and the associated clinical information. A metadata-driven approach provides a practical solution to many of the unique issues inherent in tissue microarray research, and allows relatively straightforward interoperability with and accommodation of new data models.

Database Management Systems↗

OmicsPred as a centralised resource for genetic prediction of multi-omic traits.

Genetic prediction of multi-omic data has emerged as a cost-effective alternative to direct omics profiling, particularly useful for identifying molecular features associated with disease susceptibility. However, despite its popularity, multi-omic imputation models are fragmented across studies, hindering findability, accessibility, interoperability and re-use. To address this, we developed OmicsPred (https://www.omicspred.org), a centralised platform for the deposition and dissemination of genetic prediction models of multi-omic traits. OmicsPred unifies the most commonly used molecular imputation models (e.g. from PredictDB) and other published studies totalling 3,339,469 prediction models spanning transcriptomic, proteomic, and metabolomic traits (as of May 2026). Each model is accompanied by metadata describing score development and predictive performance, and distributed in formats compatible with popular analytic tools, such as PGS Catalog Calculator and MetaXcan. To demonstrate the utility of the resource for systematic target discovery, we perform a multi-omic phenome-wide association analysis in Million Veterans Program data.

Journal Article↗

Programmatic access to ICTV virus taxonomy through a public ontology API.

The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. To address this, we developed a public and sustainable solution leveraging ontology-based APIs. Successive ICTV Master Species List (MSL) releases were transformed into a structured ontology and deployed as a unified representation through the Ontology Lookup Service (OLS). The framework also provides ICTV-NCBI mappings and helper libraries for integration into downstream systems. This enables, for the first time, public programmatic retrieval of current and historical virological taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.

API↗

SimpleMicrobiome: An integrated web-based platform for streamlined microbiome data analysis and visualization.

Microbiome studies require multiple analytical steps after initial sequence processing. These steps commonly include data harmonization, preprocessing, taxonomic profiling, diversity analysis, differential abundance testing, predictive modeling, network inference, and preparation of publication-ready outputs. Although robust packages are available for many of these tasks, routine use often depends on command-line workflows, repeated data reformatting, and method-specific scripting. These requirements can limit accessibility for experimental researchers and complicate consistent analysis across interdisciplinary teams. We developed SimpleMicrobiome, a web-based R Shiny platform that integrates established microbiome analysis methods into a single interactive downstream workflow. The application accepts standard abundance, taxonomy, and metadata tables, supports interactive preprocessing and sample filtering, and provides modules for taxa profile visualization, alpha and beta diversity analysis, ANCOM-BC2 and MaAsLin2 differential abundance testing, Random Forest modeling with SHAP-based interpretation, microbial association network inference using SparCC and SPIEC-EASI through NetCoMi, correlation heatmaps, and dbRDA/CAP-style association biplots. The platform is implemented as a modular Shiny application so that preprocessing choices are propagated across downstream analyses, results can be exported as figures and tables, and the same application can be run through the public server, source-code installation, or a Docker image. SimpleMicrobiome consolidates major downstream microbiome analysis tasks in an accessible browser-based environment while retaining links to established analytical frameworks. The platform may reduce technical barriers for non-programming users, improve consistency across exploratory and reporting-oriented analyses, and support collaborative microbiome research. The public application is available at https://simplemicrobiome.mglab.org, the source code is available at https://github.com/yjcho2252/SimpleMicrobiome, and a Docker image for local deployment is available at https://hub.docker.com/r/mglab2252/simplemicrobiome.

differential abundance↗

[CISMeF: catalog and index of French-speaking medical sites].

The Internet has now become a major source of health information. The aim of CISMeF is to catalogue and index the main French-speaking sites and documents concerning health. This project was initiated by Rouen University Hospital. Its URL is http://www.chu-rouen.fr/cismef. CISMeF covers all areas of health care and medical sciences, and is indexed both alphabetically and according to subject. It was set up on a Sun workstation under the Sun UNIX operating system and is entirely based on static HTML. By May 1999, the number of sites and documents indexed was already over 6,500, with a mean of 75 new sites added each week. CISMeF is updated via a five-step process: resource collection, filtering, description, classification, and indexing. The Net Scoring criteria are used to assess the quality of health information on the Internet. These criteria concern eight categories: credibility, content, links, design, interactivity, quantitative aspects, ethics and accessibility. CISMeF uses two standard tools to organize information: the MeSH (medical subject heading) thesaurus from the Medline reference database (National Library of Medicine, USA) and the Dublin core metadata format. The sites and documents included in CISMeF are described using the following elements from the Dublin core project: title, author or creator, subject and keywords, description, publisher, date, resource type, format, identifier, and language.

Abstracting and Indexing↗

Maintaining a catalog of manually-indexed, clinically-oriented World Wide Web content.

With no quality controls and a highly distributed means of posting information, finding high-quality, clinically-oriented content on the World Wide Web can be difficult. Maintaining a catalog of such information can be equally challenging. CliniWeb is a catalog of quality-filtered and clinically-oriented content on the Web designed to enhance access to such information. This paper describes a group of semi-automated tools have been developed to maintain the CliniWeb database. One allows easier identification of content by utilizing Web crawling techniques from high-level pages. Another allows easier selection of content for inclusion and its indexing. A final one checks links to help keep the database current. These are augmented by general plans to adopt more detailed metadata and linkages into the medical literature.

Abstracting and Indexing↗