Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

YeastHub: a semantic web use case for integrating data in the life sciences domain.

MOTIVATION: As the semantic web technology is maturing and the need for life sciences data integration over the web is growing, it is important to explore how data integration needs can be addressed by the semantic web. The main problem that we face in data integration is a lack of widely-accepted standards for expressing the syntax and semantics of the data. We address this problem by exploring the use of semantic web technologies-including resource description framework (RDF), RDF site summary (RSS), relational-database-to-RDF mapping (D2RQ) and native RDF data repository-to represent, store and query both metadata and data across life sciences datasets. RESULTS: As many biological datasets are presently available in tabular format, we introduce an RDF structure into which they can be converted. Also, we develop a prototype web-based application called YeastHub that demonstrates how a life sciences data warehouse can be built using a native RDF data store (Sesame). This data warehouse allows integration of different types of yeast genome data provided by different resources in different formats including the tabular and RDF formats. Once the data are loaded into the data warehouse, RDF-based queries can be formulated to retrieve and query the data in an integrated fashion. AVAILABILITY: The YeastHub website is accessible via the following URL: http://yeasthub.gersteinlab.org.

Biology↗

An integrated culturomic and genomic database and analysis platform for methanogenic archaea.

Methanogenic archaea research is challenged by limited strain resources, fragmented genomic data, inconsistent genome quality, substantial uncultured lineages, and difficulties in laboratory culturing, hindering advances in biogas production, climate mitigation, and microbial ecology. These archaea play crucial roles in global carbon cycling and anaerobic environments, yet scattered data and unculturable strains limit systematic studies and applications. To address this, we created MethArDB (Methanogenic Archaeal Genome Database), a specialized database for methanogenic archaea, compiling 3919 genomes, 87 host-associated plasmids, and 42 phages, with standardized quality classifications (complete, scaffold, draft), protein sequences, and metadata on geography, habitats, metabolism, and inheritable elements. Integrated MethArCT (Methanogenic Archaeal Culturomics Toolkit) employs a dual-threshold orthologous/paralogous protein analysis to evaluate metabolic pathway completeness, predicting cultivation parameters and suggesting candidate cultivation strategies, including potential medium formulations and conditions, to support strain isolation. Overall, MethArDB and MethArCT form an integrated platform combining genomics and culturomics to facilitate methanogenic archaea research. Database URL:  http://methardb.cn.

Genome, Archaeal↗

Global spread of Streptococcus pyogenes A genomics-supported narrative review.

Group A Streptococcus (GAS) has recently reemerged as a leading cause of both mild and severe invasive infections worldwide, with recent upsurges in invasive disease among children and adults. Notwithstanding a partial synchronicity with the COVID-19 pandemic, this rapid global dissemination of more virulent GAS lineages has been promptly detected, as well as the molecular shifts underlying the observed changes in clinical patterns. Whole-genome sequencing (WGS)-based genomic epidemiology allowed us to gain relevant insights into this upsurge as it was happening. This review integrates the canonical research publication-based approach with genomic data and metadata and identifies a subset of genomic clusters playing a major role in invasive GAS (iGAS) infections worldwide, which were named as Global Pathogenic Lineages (GPLs). The four GPLs broadly coincide with five sequence types (STs): GPL1 with ST28, GPL2 with ST15 and ST315, GPL3 with ST52, and GPL4 with ST39. While non-GPLs clusters maintain a baseline reservoir of antimicrobial-resistance and virulence genes, GPLs show varying but noteworthy resistance profiles and are frequent causes of iGAS. The integration of WGS into routine diagnostics procedures is a forthcoming improvement, aimed not only at informing tailored therapy and implementing infection control strategies, but also to perform continuous surveillance. Ongoing WGS in clinical microbiology, as a matter of fact, will provide unparalleled insights into lineage emergence, transmission dynamics, and the geographic clustering of virulence and resistance determinants.

Streptococcus pyogenes↗

Programmatic access to ICTV virus taxonomy through a public ontology API.

BACKGROUND: The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. FINDINGS: To address this, we developed a public and sustainable solution leveraging ontology-based APIs. All available ICTV Master Species List (MSL) releases, from MSL1 to MSL41, were transformed into a unified, semantically structured ontology comprising more than 195,000 current and historical entities and deployed through the Ontology Lookup Service (OLS). The ontology is automatically rebuilt and republished whenever a new MSL release becomes available. Complementary ICTV-NCBI mappings and helper libraries support integration into downstream systems. CONCLUSIONS: Together, these resources enable, for the first time, public programmatic retrieval of current and historical ICTV taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints, including resolution of former taxonomic terms to their current accepted taxon or taxa and retrieval of taxon histories across releases. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.

API↗

The life sciences Global Image Database (GID).

Although a vast amount of life sciences data is generated in the form of images, most scientists still store images on extremely diverse and often incompatible storage media, without any type of metadata structure, and thus with no standard facility with which to conduct searches or analyses. Here we present a solution to unlock the value of scientific images. The Global Image Database (GID) is a web-based (http://www.gwer.ch/qv/gid/gid.ht m ) structured central repository for scientific annotated images. The GID was designed to manage images from a wide spectrum of imaging domains ranging from microscopy to automated screening. The annotations in the GID define the source experiment of the images by describing who the authors of the experiment are, when the images were created, the biological origin of the experimental sample and how the sample was processed for visualization. A collection of experimental imaging protocols provides details of the sample preparation, and labeling, or visualization procedures. In addition, the entries in the GID reference these imaging protocols with the probe sequences or antibody names used in labeling experiments. The GID annotations are searchable by field or globally. The query results are first shown as image thumbnail previews, enabling quick browsing prior to original-sized annotated image retrieval. The development of the GID continues, aiming at facilitating the management and exchange of image data in the scientific community, and at creating new query tools for mining image data.

Biological Science Disciplines↗

EMPIAR: the Electron Microscopy Public Image Archive.

Public archiving in structural biology is well established with the Protein Data Bank (PDB; wwPDB.org) catering for atomic models and the Electron Microscopy Data Bank (EMDB; emdb-empiar.org) for 3D reconstructions from cryo-EM experiments. Even before the recent rapid growth in cryo-EM, there was an expressed community need for a public archive of image data from cryo-EM experiments for validation, software development, testing and training. Concomitantly, the proliferation of 3D imaging techniques for cells, tissues and organisms using volume EM (vEM) and X-ray tomography (XT) led to calls from these communities to publicly archive such data as well. EMPIAR (empiar.org) was developed as a public archive for raw cryo-EM image data and for 3D reconstructions from vEM and XT experiments and now comprises over a thousand entries totalling over 2 petabytes of data. EMPIAR resources include a deposition system, entry pages, facilities to search, visualize and download datasets, and a REST API for programmatic access to entry metadata. The success of EMPIAR also poses significant challenges for the future in dealing with the very fast growth in the volume of data and in enhancing its reusability.

Imaging, Three-Dimensional↗

The Vertebrate Genome Annotation (Vega) database.

The Vertebrate Genome Annotation (Vega) database (http://vega.sanger.ac.uk) has been designed to be a community resource for browsing manual annotation of finished sequences from a variety of vertebrate genomes. Its core database is based on an Ensembl-style schema, extended to incorporate curation-specific metadata. In collaboration with the genome sequencing centres, Vega attempts to present consistent high-quality annotation of the published human chromosome sequences. In addition, it is also possible to view various finished regions from other vertebrates, including mouse and zebrafish. Vega displays only manually annotated gene structures built using transcriptional evidence, which can be examined in the browser. Attempts have been made to standardize the annotation procedure across each vertebrate genome, which should aid comparative analysis of orthologues across the different finished regions.

Animals↗

HubMed: a web-based biomedical literature search interface.

HubMed is an alternative search interface to the PubMed database of biomedical literature, incorporating external web services and providing functions to improve the efficiency of literature search, browsing and retrieval. Users can create and visualize clusters of related articles, export citation data in multiple formats, receive daily updates of publications in their areas of interest, navigate links to full text and other related resources, retrieve data from formatted bibliography lists, navigate citation links and store annotated metadata for articles of interest. HubMed is freely available at http://www.hubmed.org/.

Internet↗

Pathbase: a new reference resource and database for laboratory mouse pathology.

Pathbase (http://www.pathbase.net) is a web accessible database of histopathological images of laboratory mice, developed as a resource for the coding and archiving of data derived from the analysis of mutant or genetically engineered mice and their background strains. The metadata for the images, which allows retrieval and interoperability with other databases, is derived from a series of orthogonal ontologies and controlled vocabularies. One of these controlled vocabularies, MPATH, was developed by the Pathbase Consortium as a formal description of the content of mouse histopathological images. The database currently has over 1000 images on-line with 2000 more under curation and presents a paradigm for the development of future databases dedicated to aspects of experimental biology.

Animals↗

Building messaging substrates for Web and Grid applications.

Grid application frameworks have increasingly aligned themselves with the developments in Web services. Web services are currently the most popular infrastructure based on service-oriented architecture (SOA) paradigm. There are three core areas within the SOA framework: (i) a set of capabilities that are remotely accessible, (ii) communications using messages and (iii) metadata pertaining to the aforementioned capabilities. In this paper, we focus on issues related to the messaging substrate hosting these services; we base these discussions on the NARADABROKERING system. We outline strategies to leverage capabilities available within the substrate without the need to make any changes to the service implementations themselves. We also identify the set of services needed to build Grids of Grids. Finally, we discuss another technology, HPSEARCH, which facilitates the administration of the substrate and the deployment of applications via a scripting interface. These issues have direct relevance to scientific Grid applications, which need to go beyond remote procedure calls in client-server interactions to support integrated distributed applications that couple databases, high performance computing codes and visualization codes.

Computer Simulation↗

Identification of Sample Processing Errors in Microbiome Studies Using Host Genetic Profiles.

In microbiome studies, sample processing errors are frequent and difficult to detect, especially in large studies involving multiple sites, personnel, and sample types. We present two complementary approaches to identify such errors using host DNA profiled via metagenomic sequencing of microbiome samples. The first approach compares host SNPs inferred from metagenomics to independently obtained genotypes (e.g., microarray genotypes) to match samples to their donors, while the second method compares metagenomics-inferred SNPs between samples to identify samples supplied by the same donor. Furthermore, we demonstrate that combining these methods with experimental metadata provides greater confidence in the identification of errors. Analyzing a longitudinal vaginal microbiome dataset, we demonstrate the ability of our approach to identify mislabeled samples. Using subsampling, we further show that our methods are robust to low sequencing coverage. Overall, our analysis highlights the frequency of processing errors in microbiome studies. We therefore recommend applying error-detection methods in all studies with suitable data.

Journal Article↗

The globalization of crystallographic knowledge.

The rapid growth of the World Wide Web provides major new opportunities for distributed databases, especially in macromolecular science. A new generation of technology, based on structured documents (SD), is being developed which will integrate documents and data in a seamless manner. This offers experimentalists the chance to publish and archive high-quality data from any discipline. Data and documents from different disciplines can be combined and searched using technology such as eXtensible Markup Language (XML) and its associated support for hypermedia (XLL), metadata (RDF) and stylesheets (XSL). Opportunities in crystallography and related disciplines are described.

Crystallography↗

OILing the way to machine understandable bioinformatics resources.

The complex questions and analyses posed by biologists, as well as the diverse data resources they develop, require the fusion of evidence from different, independently developed, and heterogeneous resources. The web, as an enabler for interoperability, has been an excellent mechanism for data publication and transportation. Successful exchange and integration of information, however, depends on a shared language for communication (a terminology) and a shared understanding of what the data means (an ontology). Without this kind of understanding, semantic heterogeneity remains a problem for both humans and machines. One means of dealing with heterogeneity in bioinformatics resources is through terminology founded upon an ontology. Bioinformatics resources tend to be rich in human readable and understandable annotation, with each resource using its own terminology. These resources are machine readable, but not machine understandable. Ontologies have a role in increasing this machine understanding, reducing the semantic heterogeneity between resources and thus promoting the flexible and reliable interoperation of bioinformatics resources. This paper describes a solution derived from the semantic web [a machine understandable world-wide web (WWW)], the ontology inference layer (OIL), as a solution for semantic bioinformatics resources. The nature of the heterogeneity problems are presented along with a description of how metadata from domain ontologies can be used to alleviate this problem. A companion paper in this issue gives an example of the development of a bio-ontology using OIL.

Algorithms↗

Building a bioinformatics ontology using OIL.

This paper describes the initial stages of building an ontology of bioinformatics and molecular biology. The conceptualization is encoded using the ontology inference layer (OIL), a knowledge representation language that combines the modeling style of frame-based systems with the expressiveness and reasoning power of description logics (DLs). This paper is the second of a pair in this special issue. The first described the core of the OIL language and the need to use ontologies to deliver semantic bioinformatics resources. In this paper, the early stages of building an ontology component of a bioinformatics resource querying application are described. This ontology (TaO) holds the information about molecular biology represented in bioinformatics resources and the bioinformatics tasks performed over these resources. It, therefore, represents the metadata of the resources the application can query. It also manages the terminologies used in constructing the query plans used to retrieve instances from those external resources. The methodology used in this task capitalizes upon features of OIL-The conceptualization afforded by the frame-based view of OIL's syntax; the expressive power and reasoning of the logical formalism; and the ability to encode both handcrafted, hierarchies of concepts, as well as defining concepts in terms of their properties, which can then be used to establish a classification and infer relationships not encoded by the ontologist. This ability forms the basis of the methodology described here: For each portion of the TaO, a basic framework of concepts is asserted by the ontologist. Then, the properties of these concepts are defined by the ontologist and the logic's reasoning power used to reclassify and infer further relationships. This cycle of elaboration and refinement is iterated on each portion of the ontology until a satisfactory ontology has been created.

Algorithms↗

The emergency department triage of community-acquired pneumonia project data and documentation systems: a model for multicenter clinical trials.

Multicenter clinical trials are complex undertakings that require significant resources to ensure efficient, high quality research. This paper describes the goals, design, and implementation of a multicenter clinical trial database management system to support this aim. A large number of study sites or patients, and the goal of automatically generating large portions of data management infrastructure from common metadata, motivated the development of the system. This paper also describes extensions for a generalized project documentation system, and discusses plans for further extensions and improvements based on observed strengths, limitations, and anticipated technological change.

Community-Acquired Infections↗

Font adaptive word indexing of modern printed documents.

We propose an approach for the word-level indexing of modern printed documents which are difficult to recognize using current OCR engines. By means of word-level indexing, it is possible to retrieve the position of words in a document, enabling queries involving proximity of terms. Web search engines implement this kind of indexing, allowing users to retrieve Web pages on the basis of their textual content. Nowadays, digital libraries hold collections of digitized documents that can be retrieved either by browsing the document images or relying on appropriate metadata assembled by domain experts. Word indexing tools would therefore increase the access to these collections. The proposed system is designed to index homogeneous document collections by automatically adapting to different languages and font styles without relying on OCR engines for character recognition. The approach is based on three main ideas: the use of Self Organizing Maps (SOM) to perform unsupervised character clustering, the definition of one suitable vector-based word representation whose size depends on the word aspect-ratio, and the run-time alignment of the query word with indexed words to deal with broken and touching characters. The most appropriate applications are for processing modern printed documents (17th to 19th centuries) where current OCR engines are less accurate. Our experimental analysis addresses six data sets containing documents ranging from books of the 17th century to contemporary journals.

Abstracting and Indexing↗

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques↗

Potential meets reality: GIS and public health research in Australia.

Geographical Information Systems-computerised systems for the capture, storage, retrieval, analysis and display of spatial data-have recently been promoted as important tools for the study of public health. Attention must also be given to the issues involved in this relatively new application, especially in Australian conditions. These include the coarse spatial resolution of most health and social data, the propagation of error through the need to use estimates and concordance tables to handle data in mismatched official spatial boundaries, the inflexible analytical capacity of most GIS for the needs of epidemiology, and difficulties in access to data, which are compounded by the absence of a good metadata register. The conflict between the need for spatial precision in GIS and preserving the confidentiality of health data is a salient issue. Medical geographers and public health researchers using GIS must recognise these issues in order to work together and toward extending the use of GIS technology beyond broad ecological and accessibility studies.

Australia↗