Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Validating existing data in the Environmental Technology Verification Program.

Establishing the credibility of existing data is an ongoing issue, particularly when the data sets are to be used for a secondary purpose, i.e., not the original reason for which they were collected. If the secondary purpose is similar to the primary purpose, the potential user may have little difficulty establishing credibility since the acceptance criteria for both purposes should be similar. If the secondary purpose is different, then data credibility may be more difficult to establish because the experiment generating the data may not have been conducted optimally for the secondary purpose and all of the necessary quality assurance data ("metadata") may not have been collected. In either case, a process will be required to determine the acceptability of the data. For this reason, at the time the U.S. Environmental Protection Agency (EPA) Environmental Technology Verification (ETV) program was established, similar certification and verification programs run by states or foreign countries routinely used existing data sets, for cost reasons, rather than generate new data by testing. The issue of whether existing data could be used in the ETV program immediately surfaced. In response, a policy and a process that addressed existing data were written and published in Appendix C of the ETV Quality and Management Plan (Hayes et al., 1998). This paper discusses how the ETV program determines the credibility of existing data used to verify the performance of environmental technologies.

Data Interpretation, Statistical↗

Where one size does not fit all: understanding the needs of potential users of a portal to breast cancer knowledge online.

The article argues that, although the Internet has great potential for assisting people to find information on breast cancer, at present that potential is not being realised. The literature shows considerable dissatisfaction with information provision for breast cancer, including on the Internet where appropriate information suited to particular needs often cannot be found. An Australian project (Breast Cancer Knowledge Online [BCKOnline]), in its first stage, set out to explore the needs for breast cancer information using an ethnographic method and a purposive sample of 77 participants, most of them women with breast cancer. A portal, which will enable users to tailor information to their particular needs, is at present being developed based on the results of the needs analysis. The process includes user-selected profiles, enabled through "user-centric" resource descriptions, and a metadata repository that links the profiles with specific information resources. The article presents limited results from the needs analysis-those highlighting the differences between younger and older women and the problems with present Internet information provision as seen by the sample. The final section discusses how the portal will both tailor information to needs and assist with the problems with the Internet revealed in the literature.

Adult↗

A search tool based on 'encapsulated' MeSH thesaurus to retrieve quality health resources on the internet.

In the year 2001, the Internet has become a major source of health information for the health professional and the Netizen. The objective of Doc' CISMeF (D'C) was to create a powerful generic search tool based on a structured information model which 'encapsulates' the MeSH thesaurus to index and retrieve quality health resources on the Internet. To index resources, D'C uses four sections in its information model: 'meta-term', keyword, subheading, and resource type. Two search options are available: simple and advanced. The simple search requires the end-user to input a single term or expression. If this term belongs to the D'C information structure model, it will be exploded. If not, a full-text search is performed. In the advanced search, complex searches are possible combining Boolean operators with meta-terms, keywords, subheadings and resource types. D'C uses two standard tools for organising information: the MeSH thesaurus and the Dublin Core metadata format. Resources included in D'C are described according to the following elements: title, author or creator, subject and keywords, description, publishers, date, resource type, format, identifier, and language.

Abstracting and Indexing↗

Implementing context and team based access control in healthcare intranets.

The establishment of an efficient access control system in healthcare intranets is a critical security issue directly related to the protection of patients' privacy. Our C-TMAC (Context and Team-based Access Control) model is an active security access control model that layers dynamic access control concepts on top of RBAC (Role-based) and TMAC (Team-based) access control models. It also extends them in the sense that contextual information concerning collaborative activities is associated with teams of users and user permissions are dynamically filtered during runtime. These features of C-TMAC meet the specific security requirements of healthcare applications. In this paper, an experimental implementation of the C-TMAC model is described. More specifically, we present the operational architecture of the system that is used to implement C-TMAC security components in a healthcare intranet. Based on the technological platform of an Oracle Data Base Management System and Application Server, the application logic is coded with stored PL/SQL procedures that include Dynamic SQL routines for runtime value binding purposes. The resulting active security system adapts to current need-to-know requirements of users during runtime and provides fine-grained permission granularity. Apart from identity certificates for authentication, it uses attribute certificates for communicating critical security metadata, such as role membership and team participation of users.

Computer Communication Networks↗

Integrating large-scale genotype and phenotype data.

With the completion of the Human Genome Project, a new emphasis is focusing on the sequence variation and the resulting phenotype. The number of data available from genomic studies addressing this relationship is rapidly growing. In order to analyze these data as a whole, they need to be integrated, aggregated and annotated in a timely manner. The Pharmacogenetics and Pharmacogenomics Knowledge Base PharmGKB; ( ) assembles and disseminates these data and their associated metadata that are needed for unambiguous identification and replication. Assembling these data in a timely manner is challenging, and the scalability of these data produce major challenges for a knowledge base such as PharmGKB. However, it is only through rapid global meta-annotation of these data that we will understand the relationship between specific genotype(s) and the related phenotype. PharmGKB has confronted these challenges, and these experiences and solutions can benefit all genome communities.

Animals↗

Computational knowledge integration in biopharmaceutical research.

An initiative to increase biopharmaceutical research productivity by capturing, sharing and computationally integrating proprietary scientific discoveries with public knowledge is described. This initiative involves both organisational process change and multiple interoperating software systems. The software components rely on mutually supporting integration techniques. These include a richly structured ontology, statistical analysis of experimental data against stored conclusions, natural language processing of public literature, secure document repositories with lightweight metadata, web services integration, enterprise web portals and relational databases. This approach has already begun to increase scientific productivity in our enterprise by creating an organisational memory (OM) of internal research findings, accessible on the web. Through bringing together these components it has also been possible to construct a very large and expanding repository of biological pathway information linked to this repository of findings which is extremely useful in analysis of DNA microarray data. This repository, in turn, enables our research paradigm to be shifted towards more comprehensive systems-based understandings of drug action.

Algorithms↗

MDB: a database system utilizing automatic construction of modules and STAR-derived universal language.

MOTIVATION: The value of information greatly increases if stored in databases. The objective was to construct a multi-purpose database system primarily designed to store and provide access to three-dimensional structures of biological molecules including theoretical models. RESULTS: A dictionary defining data format and structure for three-dimensional models of biological molecules (MDB dictionary) was developed. The dictionary was written using universal, standardized data description language. This language can be applied to describe data with no restrictions on their origin or type, including metadata. Thus both the data definitions (format) and database descriptions are created using the uniform language and processed with universal software. A database and data design technique that allowed use of dictionaries to automatically construct relational databases was developed. This technique was employed to construct the MDB database system. Data design developed and applied in the MDB project makes it possible to carry out data curation utilizing the database engine to identify errors. It also allows storage and query of data at different levels of consistency with the standard format specifications, i.e. both the correctly formatted data, and data that requires further curation. AVAILABILITY: The MDB dictionary is available at http://www.gwer.ch/proteinstructure/mdb and as part of the PDB resources at http://pdb.rutgers.edu/mmcif/.

Computational Biology↗

FuNTB: a functional network clustering tool for the analysis of genome-wide genetic variants in Mycobacterium tuberculosis.

MOTIVATION: Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mtb), still claims around 1.25 million lives each year. The growing threat of drug resistance-often driven by single‑nucleotide polymorphisms (SNPs) in Mtb genomes underscores the need for high‑quality genomic data and powerful bioinformatics tools. We present FuNTB, a python‑based pipeline that detects non‑synonymous SNPs in Mtb and builds functional network clusters to reveal genotype-phenotype relationships. RESULTS: FuNTB profiles non‑synonymous SNPs at the gene level across user‑defined phenotypes, pinpointing both shared and unique mutations. It ingests annotated Variant Call Format (VCF) files or MTBseq outputs and merges them with clinical metadata to produce network‑XML files compatible with Cytoscape and Gephi. When applied to the CRyPTIC Mtb collection, FuNTB rapidly recovered established resistance genes and surfaced novel candidates, validating its utility for mapping genotype-phenotype associations. AVAILABILITY AND IMPLEMENTATION: FuNTB is implemented in Python 3.8+ and is freely available under the MIT license at https://doi.org/10.5281/zenodo.15399917.

Mycobacterium tuberculosis↗

GRUMB: a genome-resolved metagenomic framework for monitoring urban microbiomes and diagnosing pathogen risk.

SUMMARY: Urban infrastructure hosts dynamic microbial communities that complicate biosurveillance and AMR monitoring. Existing tools rarely combine genome-resolved reconstruction with ecological modeling and batch-aware analytics tailored to infrastructure-scale studies. We present GRUMB (Genome-Resolved Urban Microbiome Biosurveillance), an open-source, SLURM-compatible pipeline that reconstructs high-quality metagenome-assembled genomes (MAGs) from shotgun sequencing reads and integrates taxonomic/functional annotation (CARD, VFDB), batch-aware normalization, ecological diagnostics and machine learning classification of environment types with uncertainty and risk scoring. GRUMB accepts either SRA project accessions or paired-end FASTQ files with metadata, and produces assemblies, MAGs, taxonomic and functional profiles, ecological outputs and risk-informed classification. Its modular design enables reproducible, infrastructure-scale biosurveillance across diverse environments. AVAILABILITY AND IMPLEMENTATION: GRUMB is freely available under the MIT License at: https://github.com/SuleimanAminu/genome-resolved-urban-microbiome-biosurveillance; Zenodo DOI: https://doi.org/10.5281/zenodo.15505402. Requirements: Linux (Ubuntu 20.04+), Python 3.11, R 4.2+, SLURM. Issues and feature requests are tracked on GitHub.

Microbiota↗

MetaFX: feature extraction from whole-genome metagenomic sequencing data.

MOTIVATION: Microbial communities consist of thousands of microorganisms and viruses and have a tight connection with an environment, such as gut microbiota modulation of host body metabolism. However, the direct relationship between the presence of certain microorganism and the host state often remains unknown. Toolkits using reference-based approaches are limited to microbes present in databases. Reference-free methods often require enormous resources for metagenomic assembly or results in many poorly interpretable features based on k-mers. RESULTS: Here we present MetaFX-an open-source library for feature extraction from whole-genome metagenomic sequencing data and classification of groups of samples. Using a large volume of metagenomic samples deposited in databases, MetaFX compares samples grouped by metadata criteria (e.g. disease, treatment, etc.) and constructs genomic features distinct for certain types of communities. Features constructed based on statistical k-mer analysis and de Bruijn graphs partition. Those features are used in machine learning models for classification of novel samples. Extracted features can be visualized on de Bruijn graphs and annotated for providing biological insights. We demonstrate the utility of MetaFX by building classification models for 590 human gut samples with inflammatory bowel disease. Our results outperform the previous research disease prediction accuracy up to 17%, and improves classification results compared to taxonomic analysis by 9±10% on average. AVAILABILITY AND IMPLEMENTATION: MetaFX is a feature extraction toolkit applicable for metagenomic datasets analysis and samples classification. The source code, test data, and relevant information for MetaFX are freely accessible at https://github.com/ctlab/metafx under the MIT License. Alternatively, MetaFX can be obtained via http://doi.org/10.5281/zenodo.16949369.

Metagenomics↗

moiraine: an R package to construct reproducible pipelines for the application and comparison of multi-omics integration methods.

MOTIVATION: In the past decades, many statistical methods for integrating multi-omics data have been developed. They have been implemented into software tools, which differ widely in their programming choices, such as the format required for data input, or the format of the generated integration results. This lack of standards renders cumbersome and time-intensive the application and comparison of different integration tools to the same multi-omics dataset. RESULTS: We have developed the moiraine R package for constructing reproducible multi-omics integration pipelines, which enables users to apply one or more statistical methods for multi-omics integration to their own multi-omics dataset. moiraine facilitates the preprocessing of the omics datasets and automates their formatting for the integration step. It simplifies the interpretation and evaluation of the integration results through the construction of visualizations in which metadata about samples and features can easily be included. Crucially, it enables the comparison of results obtained with different integration tools, allowing users to assess the robustness of their results. AVAILABILITY AND IMPLEMENTATION: The moiraine R package is publicly available at https://github.com/Plant-Food-Research-Open/moiraine; an archival snapshot of the package is available on Zenodo at https://doi.org/10.5281/zenodo.17172718. A detailed tutorial is available at https://plant-food-research-open.github.io/moiraine-manual/.

Software↗

treestructure: an R package to detect population structure in phylogenetic trees.

MOTIVATION: How population structure can shape genetic diversity is a longstanding problem in population genetics. While the use of geographic locations, when available, can help answer some of these questions, it is still difficult to determine population structure when such metadata are not available or when the potential population structure is not easily observed. Here, we present an updated version of treestructure, an R package that implements a statistical test based on coalescent theory to detect unobserved population structure in a time-scaled phylogenetic tree. AVAILABILITY: treestructure is available at CRAN at https://cloud.r-project.org/web/packages/treestructure/ and at https://emvolz-phylodynamics.github.io/treestructure/.

Phylogeny↗

Using cancer profiles to identify synthetic lethal therapeutic targets and predictive biomarkers in cancer gene dependency data.

MOTIVATION: Large scale loss-of-function screens utilising CRISPR or siRNA can provide profound insights into the importance of individual genes for the survival of a cancer cell and can drive the identification of therapeutic targets and biomarkers, and the development of targeted drugs. However, the analysis of these data and the substantial bodies of metadata that relate to them, is technically challenging and typically requires substantial expertise in data science and computer coding. RESULTS: To facilitate the analysis of cancer gene dependency data by cancer biologists and clinical scientists, we have developed DepMine-a computational toolkit providing a powerful system for framing complex queries relating cancer gene dependency to the underlying genetic changes that occur in cancer cells. DepMine identifies synthetic lethal relationships between putative target genes and complex 'cancer profiles' built from user-specified combinations of mutations, copy-number variation, and expression levels, and can refine these to optimal biomarker definitions for target dependency. AVAILABILITY: The Python implementation of DepMine and associated data files can be obtained at https://github.com/UOSbioinformaticslab/depmine and is free to academics and Not-For-Profit organisations. The DepMine release referenced in this paper is archived as DOI: 10.5281/zenodo.19570601.

Humans↗

Interactive exploration of biobank-scale ancestral recombination graphs with Lorax.

MOTIVATION: Ancestral Recombination Graphs (ARGs) provide a comprehensive representation of genetic ancestry and underpin analyses of natural selection, disease association, and population history. However, existing visualization tools are limited in scalability and interactivity, making ARGs difficult to explore at biobank scale. RESULTS: We introduce Lorax, a GPU-accelerated, web-native platform for real-time visualization of population-scale ARGs. Lorax integrates genomic position, coalescent time, local genealogy, and metadata, enabling interactive exploration of ancestry and variant inheritance in biobank-scale datasets. AVAILABILITY AND IMPLEMENTATION: Lorax is freely available as a live demo at https://lorax.ucsc.edu/ and as a Python package "lorax-arg" on PyPI. The source code and documentation are available on GitHub at https://github.com/pratikkatte/lorax.

Software↗

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

MOTIVATION: An unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. RESULTS: Within 49 h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Machine Learning↗

WxS-QC-a quality control pipeline for human germline short-variant Whole-Genome and Whole-Exome cohorts for population-scale analyses.

SUMMARY: Whole-exome (WES) and whole-genome (WGS) sequencing are rapidly becoming preferred methods for population-scale analysis of the human genetic landscape. However, there are currently no standardized quality control (QC) pipelines for human WES and WGS datasets. In this paper, we present WxS-QC, a powerful, scalable, and convenient pipeline for the QC of human germline short-variant WGS and WES cohorts for population-scale analyses. Our pipeline is suitable for both rare-variant discovery and common-variant association studies. It is based on deeply refactored gnomAD v3 and v4 quality control pipelines, contains several methods we have developed de novo, and is aligned with current best practices in WGS/WES germline cohort QC. We provide all methods in a single codebase, aligned to work together and controlled via a single YAML config, with automatic export of resulting graphs and summary tables, excellent performance and scalability, and comprehensive documentation. The pipeline can run in any UNIX-like environment and can efficiently process cohorts of up to 200 000 whole-exome samples, with the potential to handle bigger datasets. AVAILABILITY AND IMPLEMENTATION: The pipeline code is written in Python using the Hail library and is freely available under the BSD-3 license here: https://github.com/wtsi-hgi/wxs-qc. The detailed description of the pipeline is available in the pipeline documentation: https://github.com/wtsi-hgi/wxs-qc/blob/main/README.md. We also provide an open dataset with all required metadata, which is available at https://wxs-qc-data.cog.sanger.ac.uk/wxs-qc_public_dataset_v3.tar. An example of test dataset analysis is available in the supplementary materials.

Humans↗

SEMEDA: ontology based semantic integration of biological databases.

MOTIVATION: Many molecular biological databases are implemented on relational Database Management Systems, which provide standard interfaces like JDBC and ODBC for data and metadata exchange. By using these interfaces, many technical problems of database integration vanish and issues related to semantics remain, e.g. the use of different terms for the same things, different names for equivalent database attributes and missing links between relevant entries in different databases. RESULTS: In this publication, principles and methods that were used to implement SEMEDA (Semantic Meta Database) are described. Database owners can use SEMEDA to provide semantically integrated access to their databases as well as to collaboratively edit and maintain ontologies and controlled vocabularies. Biologists can use SEMEDA to query the integrated databases in real time without having to know the structure or any technical details of the underlying databases. AVAILABILITY: SEMEDA is available at http://www-bm.ipk-gatersleben.de/semeda/. Database providers who intend to grant access to their databases via SEMEDA are encouraged to contact the authors.

Computational Biology↗

PDBML: the representation of archival macromolecular structure data in XML.

SUMMARY: The Protein Data Bank (PDB) has recently released versions of the PDB Exchange dictionary and the PDB archival data files in XML format collectively named PDBML. The automated generation of these XML files is driven by the data dictionary infrastructure in use at the PDB. The correspondences between the PDB dictionary and the XML schema metadata are described as well as the XML representations of PDB dictionaries and data files.

Amino Acid Sequence↗