Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Dysregulated Sheddase Signalling as a Molecular Driver of Plaque Instability Revealed by Integrative Transcriptomics.

Atherosclerosis is a major cause of mortality due to chronic and progressive low-grade inflammation and fibroproliferative remodelling of the intima of arteries. Comprehensive understanding of the interplay between plaque biology and the mechanisms underlying plaque vulnerability and rupture is essential. Here, we aimed to investigate the transcriptomic profiles of stable and unstable atherosclerotic plaques using RNA sequencing data from human carotid atherosclerotic plaque samples based on next-generation knowledge discovery (NGKD) methods. High-throughput RNA-seq data from plaques dissected in stable and unstable regions of four patients were obtained from the Gene Expression Omnibus (GEO) database. GEO RNA-seq Experiments Interactive Navigator (GREIN) software was used to obtain raw gene-level counts and filtered metadata for this dataset. The data were further filtered and normalized using Express analyst to derive differentially expressed genes (DEGs) in unstable plaques compared to stable plaques. The DEGs were further analysed using WebGestalt, STRING DB, preranked gene set enrichment analysis (GSEA), and Ingenuity Pathway Analysis (IPA) software. We identified 4792 DEGs in unstable plaques based on a p-value cutoff of <&#x2009;0.05. NGKD analysis revealed that the sheddase pathway, collagen degradation, activation of matrix metalloproteinases (MMPs), and extracellular matrix (ECM) degradation ranked among the top five upregulated pathways, whereas the inhibition of MMPs and smooth muscle contraction pathways were identified as the most prominent downregulated pathways in unstable plaques. We found that the sheddase pathway was one of the most significantly upregulated canonical pathways in unstable plaques and this finding opens new avenues for potential therapeutic interventions in patients with atherosclerosis.

Humans↗

Data extraction and ad hoc query of an entity-attribute-value database.

Entity--attribute--value (EAV) tables form the major component of several mainstream electronic patient record systems (EPRSs). Such systems have been optimized for real-time retrieval of individual patient data. Data warehousing, on the other hand, involves cross-patient data retrieval based on values of patient attributes, with a focus on ad hoc query. Attribute-centric query is inherently more difficult when data are stored in EAV form than when they are stored conventionally. The authors illustrate their approach to the attribute-centric query problem with ACT/DB, a database for managing clinical trials data. This approach is based on metadata supporting a query front end that essentially hides the EAV/non-EAV nature of individual attributes from the user. The authors' work does not close the query problem, and they identify several complex subproblems that are still to be solved.

Clinical Trials as Topic↗

Common data model for neuroscience data and data model exchange.

OBJECTIVE: Generalizing the data models underlying two prototype neurophysiology databases, the authors describe and propose the Common Data Model (CDM) as a framework for federating a broad spectrum of disparate neuroscience information resources. DESIGN: Each component of the CDM derives from one of five superclasses-data, site, method, model, and reference-or from relations defined between them. A hierarchic attribute-value scheme for metadata enables interoperability with variable tree depth to serve specific intra- or broad inter-domain queries. To mediate data exchange between disparate systems, the authors propose a set of XML-derived schema for describing not only data sets but data models. These include biophysical description markup language (BDML), which mediates interoperability between data resources by providing a meta-description for the CDM. RESULTS: The set of superclasses potentially spans data needs of contemporary neuroscience. Data elements abstracted from neurophysiology time series and histogram data represent data sets that differ in dimension and concordance. Site elements transcend neurons to describe subcellular compartments, circuits, regions, or slices; non-neuroanatomic sites include sequences to patients. Methods and models are highly domain-dependent. CONCLUSIONS: True federation of data resources requires explicit public description, in a metalanguage, of the contents, query methods, data formats, and data models of each data resource. Any data model that can be derived from the defined superclasses is potentially conformant and interoperability can be enabled by recognition of BDML-described compatibilities. Such metadescriptions can buffer technologic changes.

Animals↗

How the past teaches the future: ACMI distinguished lecture.

More than 30 years of experience in developing a computer-based patient record system, The Medical Record (TMR), in multiple settings, in multiple specialty groups, and at multiple sites has taught us many lessons. Lessons related to computer-based patient records include the importance of a data model in which input, storage, and planned use are independent; separation of patient-specific data from metadata; a modular design to localize the program code that deals with a set of data; redundant storage to optimize tasks and response time; and integration of decision support into work process. Lessons related to medical informatics include the importance of a clinical-technical partnership, control of tools at the leading edge, and rapid prototyping in the real world. Finally, changes in technology move the challenges but do not eliminate them.

History, 20th Century↗

Techniques for optimization of queries on integrated biological resources.

Today, scientific data are inevitably digitized, stored in a wide variety of formats, and are accessible over the Internet. Scientific discovery increasingly involves accessing multiple heterogeneous data sources, integrating the results of complex queries, and applying further analysis and visualization applications in order to collect datasets of interest. Building a scientific integration platform to support these critical tasks requires accessing and manipulating data extracted from flat files or databases, documents retrieved from the Web, as well as data that are locally materialized in warehouses or generated by software. The lack of efficiency of existing approaches can significantly affect the process with lengthy delays while accessing critical resources or with the failure of the system to report any results. Some queries take so much time to be answered that their results are returned via email, making their integration with other results a tedious task. This paper presents several issues that need to be addressed to provide seamless and efficient integration of biomolecular data. Identified challenges include: capturing and representing various domain specific computational capabilities supported by a source including sequence or text search engines and traditional query processing; developing a methodology to acquire and represent semantic knowledge and metadata about source contents, overlap in source contents, and access costs; developing cost and semantics based decision support tools to select sources and capabilities, and to generate efficient query evaluation plans.

Algorithms↗

Public health, GIS, and the internet.

Internet access and use of georeferenced public health information for GIS application will be an important and exciting development for the nation's Department of Health and Human Services and other health agencies in this new millennium. Technological progress toward public health geospatial data integration, analysis, and visualization of space-time events using the Web portends eventual robust use of GIS by public health and other sectors of the economy. Increasing Web resources from distributed spatial data portals and global geospatial libraries, and a growing suite of Web integration tools, will provide new opportunities to advance disease surveillance, control, and prevention, and insure public access and community empowerment in public health decision making. Emerging supercomputing, data mining, compression, and transmission technologies will play increasingly critical roles in national emergency, catastrophic planning and response, and risk management. Web-enabled public health GIS will be guided by Federal Geographic Data Committee spatial metadata, OpenGIS Web interoperability, and GML/XML geospatial Web content standards. Public health will become a responsive and integral part of the National Spatial Data Infrastructure.

Geographic Information Systems↗

Image file formats: past, present, and future.

Despite the rapid growth of the Internet for storage and display of World Wide Web-based teaching files, the available image file formats have remained relatively limited. The recently developed portable networks graphics (PNG) format is versatile and offers several advantages over the older Internet standard image file formats that make it an attractive option for digital teaching files. With the PNG format, it is possible to repeatedly open, edit, and save files with lossless compression along with gamma and chromicity correction. The two-dimensional interlacing capabilities of PNG allow an image to fill in from top to bottom and from right to left, making retrieval faster than with other formats. In addition, images can be viewed closer to the original settings, and metadata (ie, information about data) can be incorporated into files. The PNG format provides a network-friendly, patent-free, lossless compression scheme that is truly cross-platform and has many new features that are useful for multimedia and Web-based radiologic teaching. The widespread acceptance of PNG by the World Wide Web Consortium and by the most popular Web browsers and graphic manipulation software companies suggests an expanding role in the future of multimedia teaching file development.

Computer-Assisted Instruction↗

Satellog: a database for the identification and prioritization of satellite repeats in disease association studies.

BACKGROUND: To date, 35 human diseases, some of which also exhibit anticipation, have been associated with unstable repeats. Anticipation has been reported in a number of diseases in which repeat expansion may have a role in etiology. Despite the growing importance of unstable repeats in disease, currently no resource exists for the prioritization of repeats. Here we present Satellog, a database that catalogs all pure 1-16 repeat unit satellite repeats in the human genome along with supplementary data. Satellog analyzes each pure repeat in UniGene clusters for evidence of repeat polymorphism. RESULTS: A total of 5,546 such repeats were identified, providing the first indication of many novel polymorphic sites in the genome. Overall, polymorphic repeats were over-represented within 3'-UTR sequence relative to 5'-UTR and coding sequence. Interestingly, we observed that repeat polymorphism within coding sequence is restricted to trinucleotide repeats whereas UTR sequence tolerated a wider range of repeat period polymorphisms. For each pure repeat we also calculate its repeat length percentile rank, its location either within or adjacent to EnsEMBL genes, and its expression profile in normal tissues according to the GeneNote database. CONCLUSION: Satellog provides the ability to dynamically prioritize repeats based on any of their characteristics (i.e. repeat unit, class, period, length, repeat length percentile rank, genomic co-ordinates), polymorphism profile within UniGene, proximity to or presence within gene regions (i.e. cds, UTR, 15 kb upstream etc.), metadata of the genes they are detected within and gene expression profiles within normal human tissues. Unstable repeats associated with 31 diseases were analyzed in Satellog to evaluate their common repeat properties. The utility of Satellog was highlighted by prioritizing repeats for Huntington's disease and schizophrenia. Satellog is available online at http://satellog.bcgsc.ca.

3' Untranslated Regions↗

A Taxonomic Search Engine: federating taxonomic databases using web services.

BACKGROUND: The taxonomic name of an organism is a key link between different databases that store information on that organism. However, in the absence of a single, comprehensive database of organism names, individual databases lack an easy means of checking the correctness of a name. Furthermore, the same organism may have more than one name, and the same name may apply to more than one organism. RESULTS: The Taxonomic Search Engine (TSE) is a web application written in PHP that queries multiple taxonomic databases (ITIS, Index Fungorum, IPNI, NCBI, and uBIO) and summarises the results in a consistent format. It supports "drill-down" queries to retrieve a specific record. The TSE can optionally suggest alternative spellings the user can try. It also acts as a Life Science Identifier (LSID) authority for the source taxonomic databases, providing globally unique identifiers (and associated metadata) for each name. CONCLUSION: The Taxonomic Search Engine is available at http://darwin.zoology.gla.ac.uk/~rpage/portal/ and provides a simple demonstration of the potential of the federated approach to providing access to taxonomic names.

Classification↗

Workflows in bioinformatics: meta-analysis and prototype implementation of a workflow generator.

BACKGROUND: Computational methods for problem solving need to interleave information access and algorithm execution in a problem-specific workflow. The structures of these workflows are defined by a scaffold of syntactic, semantic and algebraic objects capable of representing them. Despite the proliferation of GUIs (Graphic User Interfaces) in bioinformatics, only some of them provide workflow capabilities; surprisingly, no meta-analysis of workflow operators and components in bioinformatics has been reported. RESULTS: We present a set of syntactic components and algebraic operators capable of representing analytical workflows in bioinformatics. Iteration, recursion, the use of conditional statements, and management of suspend/resume tasks have traditionally been implemented on an ad hoc basis and hard-coded; by having these operators properly defined it is possible to use and parameterize them as generic re-usable components. To illustrate how these operations can be orchestrated, we present GPIPE, a prototype graphic pipeline generator for PISE that allows the definition of a pipeline, parameterization of its component methods, and storage of metadata in XML formats. This implementation goes beyond the macro capacities currently in PISE. As the entire analysis protocol is defined in XML, a complete bioinformatic experiment (linked sets of methods, parameters and results) can be reproduced or shared among users. AVAILABILITY: http://if-web1.imb.uq.edu.au/Pise/5.a/gpipe.html (interactive), ftp://ftp.pasteur.fr/pub/GenSoft/unix/misc/Pise/ (download). CONCLUSION: From our meta-analysis we have identified syntactic structures and algebraic operators common to many workflows in bioinformatics. The workflow components and algebraic operators can be assimilated into re-usable software components. GPIPE, a prototype implementation of this framework, provides a GUI builder to facilitate the generation of workflows and integration of heterogeneous analytical tools.

Algorithms↗

MeMo: a hybrid SQL/XML approach to metabolomic data management for functional genomics.

BACKGROUND: The genome sequencing projects have shown our limited knowledge regarding gene function, e.g. S. cerevisiae has 5-6,000 genes of which nearly 1,000 have an uncertain function. Their gross influence on the behaviour of the cell can be observed using large-scale metabolomic studies. The metabolomic data produced need to be structured and annotated in a machine-usable form to facilitate the exploration of the hidden links between the genes and their functions. DESCRIPTION: MeMo is a formal model for representing metabolomic data and the associated metadata. Two predominant platforms (SQL and XML) are used to encode the model. MeMo has been implemented as a relational database using a hybrid approach combining the advantages of the two technologies. It represents a practical solution for handling the sheer volume and complexity of the metabolomic data effectively and efficiently. The MeMo model and the associated software are available at http://dbkgroup.org/memo/. CONCLUSION: The maturity of relational database technology is used to support efficient data processing. The scalability and self-descriptiveness of XML are used to simplify the relational schema and facilitate the extensibility of the model necessitated by the creation of new experimental techniques. Special consideration is given to data integration issues as part of the systems biology agenda. MeMo has been physically integrated and cross-linked to related metabolomic and genomic databases. Semantic integration with other relevant databases has been supported through ontological annotation. Compatibility with other data formats is supported by automatic conversion.

Computer Simulation↗

HeatMapper: powerful combined visualization of gene expression profile correlations, genotypes, phenotypes and sample characteristics.

BACKGROUND: Accurate interpretation of data obtained by unsupervised analysis of large scale expression profiling studies is currently frequently performed by visually combining sample-gene heatmaps and sample characteristics. This method is not optimal for comparing individual samples or groups of samples. Here, we describe an approach to visually integrate the results of unsupervised and supervised cluster analysis using a correlation plot and additional sample metadata. RESULTS: We have developed a tool called the HeatMapper that provides such visualizations in a dynamic and flexible manner and is available from http://www.erasmusmc.nl/hematologie/heatmapper/. CONCLUSION: The HeatMapper allows an accessible and comprehensive visualization of the results of gene expression profiling and cluster analysis.

Antigens, CD34↗

MultiSeq: unifying sequence and structure data for evolutionary analysis.

BACKGROUND: Since the publication of the first draft of the human genome in 2000, bioinformatic data have been accumulating at an overwhelming pace. Currently, more than 3 million sequences and 35 thousand structures of proteins and nucleic acids are available in public databases. Finding correlations in and between these data to answer critical research questions is extremely challenging. This problem needs to be approached from several directions: information science to organize and search the data; information visualization to assist in recognizing correlations; mathematics to formulate statistical inferences; and biology to analyze chemical and physical properties in terms of sequence and structure changes. RESULTS: Here we present MultiSeq, a unified bioinformatics analysis environment that allows one to organize, display, align and analyze both sequence and structure data for proteins and nucleic acids. While special emphasis is placed on analyzing the data within the framework of evolutionary biology, the environment is also flexible enough to accommodate other usage patterns. The evolutionary approach is supported by the use of predefined metadata, adherence to standard ontological mappings, and the ability for the user to adjust these classifications using an electronic notebook. MultiSeq contains a new algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of a homologous group of distantly related proteins. The method, based on the multidimensional QR factorization of multiple sequence and structure alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. CONCLUSION: MultiSeq is a major extension of the Multiple Alignment tool that is provided as part of VMD, a structural visualization program for analyzing molecular dynamics simulations. Both are freely distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics and MultiSeq is included with VMD starting with version 1.8.5. The MultiSeq website has details on how to download and use the software: http://www.scs.uiuc.edu/~schulten/multiseq/

Algorithms↗

A database application for pre-processing, storage and comparison of mass spectra derived from patients and controls.

BACKGROUND: Statistical comparison of peptide profiles in biomarker discovery requires fast, user-friendly software for high throughput data analysis. Important features are flexibility in changing input variables and statistical analysis of peptides that are differentially expressed between patient and control groups. In addition, integration the mass spectrometry data with the results of other experiments, such as microarray analysis, and information from other databases requires a central storage of the profile matrix, where protein id's can be added to peptide masses of interest. RESULTS: A new database application is presented, to detect and identify significantly differentially expressed peptides in peptide profiles obtained from body fluids of patient and control groups. The presented modular software is capable of central storage of mass spectra and results in fast analysis. The software architecture consists of 4 pillars, 1) a Graphical User Interface written in Java, 2) a MySQL database, which contains all metadata, such as experiment numbers and sample codes, 3) a FTP (File Transport Protocol) server to store all raw mass spectrometry files and processed data, and 4) the software package R, which is used for modular statistical calculations, such as the Wilcoxon-Mann-Whitney rank sum test. Statistic analysis by the Wilcoxon-Mann-Whitney test in R demonstrates that peptide-profiles of two patient groups 1) breast cancer patients with leptomeningeal metastases and 2) prostate cancer patients in end stage disease can be distinguished from those of control groups. CONCLUSION: The database application is capable to distinguish patient Matrix Assisted Laser Desorption Ionization (MALDI-TOF) peptide profiles from control groups using large size datasets. The modular architecture of the application makes it possible to adapt the application to handle also large sized data from MS/MS- and Fourier Transform Ion Cyclotron Resonance (FT-ICR) mass spectrometry experiments. It is expected that the higher resolution and mass accuracy of the FT-ICR mass spectrometry prevents the clustering of peaks of different peptides and allows the identification of differentially expressed proteins from the peptide profiles.

Algorithms↗

An online database for brain disease research.

BACKGROUND: The Stanley Medical Research Institute online genomics database (SMRIDB) is a comprehensive web-based system for understanding the genetic effects of human brain disease (i.e. bipolar, schizophrenia, and depression). This database contains fully annotated clinical metadata and gene expression patterns generated within 12 controlled studies across 6 different microarray platforms. DESCRIPTION: A thorough collection of gene expression summaries are provided, inclusive of patient demographics, disease subclasses, regulated biological pathways, and functional classifications. CONCLUSION: The combination of database content, structure, and query speed offers researchers an efficient tool for data mining of brain disease complete with information such as: cross-platform comparisons, biomarkers elucidation for target discovery, and lifestyle/demographic associations to brain diseases.

Bipolar Disorder↗

Location-based health information services: a new paradigm in personalised information delivery.

Brute health information delivery to various devices can be easily achieved these days, making health information instantly available whenever it is needed and nearly anywhere. However, brute health information delivery risks overloading users with unnecessary information that does not answer their actual needs, and might even act as noise, masking any other useful and relevant information delivered with it. Users' profiles and needs are definitely affected by where they are, and this should be taken into consideration when personalising and delivering information to users in different locations. The main goal of location-based health information services is to allow better presentation of the distribution of health and healthcare needs and Internet resources answering them across a geographical area, with the aim to provide users with better support for informed decision-making. Personalised information delivery requires the acquisition of high quality metadata about not only information resources, but also information service users, their geographical location and their devices. Throughout this review, experience from a related online health information service, HealthCyberMap http://healthcybermap.semanticweb.org/, is referred to as a model that can be easily adapted to other similar services. HealthCyberMap is a Web-based directory service of medical/health Internet resources exploring new means to organise and present these resources based on consumer and provider locations, as well as the geographical coverage or scope of indexed resources. The paper also provides a concise review of location-based services, technologies for detecting user location (including IP geolocation), and their potential applications in health and healthcare.

Journal Article↗

The diagnostic path, a useful visualisation tool in virtual microscopy.

BACKGROUND: The Virtual Microscopy based on completely digitalised histological slide. Concerning this digitalisation many new features in mircoscopy can be processed by the computer. New applications are possible or old, well known techniques of image analyses can be adapted for routine use. AIMS: A so called diagnostic path observes in the way of a professional sees through a histological virtual slide combined with the text information of the dictation process. This feature can be used for image retrieval, quality assurance or for educational purpose. MATERIALS AND METHODS: The diagnostic path implements a metadata structure of image information. It stores and processes the different images seen by a pathologist during his "slide viewing" and the obtained image sequence ("observation path"). Contemporary, the structural details of the pathology reports were analysed. The results were transferred into an XML structure. Based on this structure, a report editor and a search function were implemented. The report editor compiles the "diagnostic path", which is the connection from the image viewing sequence ("observation path") and the oral report sequence of the findings ("dictation path"). The time set ups of speech and image viewing serve for the link between the two sequences. The search tool uses the obtained diagnostic path. It allows the user to search for particular histological hallmarks in pathology reports and in the corresponding images. RESULTS: The new algorithm was tested on 50 pathology reports and 74 attached histological images. The creation of a new individual diagnostic path is automatically performed during the routine diagnostic process. The test prototype experienced an insignificant prolongation of the diagnosis procedure (oral case description and stated diagnosis by the pathologist) and a fast and reliable retrieval, especially useful for continuous education and quality control of case description and diagnostic work. DISCUSSION: The Digital Virtual Microscope has been designed to handle 1000 images per day in the daily routine work of a pathology institution. It implies the necessity of an automatic mechanism of image meta dating. The non - deterministic correlation between the oral statements (case report) and image information content guides the image meta dating. The presented software opens up new possibilities for a content oriented search in a virtual slide, and can successfully support medical education and diagnostic quality assurance.

Journal Article↗

Host clustering of Campylobacter species and enteric pathogens in a longitudinal cohort of infants, family members and livestock in rural Eastern Ethiopia.

BACKGROUND: Livestock are recognized as major reservoirs for Campylobacter species and other enteric pathogens, posing infection risks to humans. High prevalence of Campylobacter during early childhood has been linked to environmental enteric dysfunction and stunting, particularly in low-resource settings. METHODS: A total of 280 samples from Campylobacter positive households with complete metadata were analyzed by shotgun metagenomic sequencing followed by bioinformatic analysis via the CZ-ID metagenomic pipeline (Illumina mNGS Pipeline v7.1). Further statistical analyses in JMP PRO 16 explored the microbiome, emphasizing Campylobacter and other enteric pathogens. Two-way hierarchical clustering and split k-mer analysis examined host structuring, patterns of co-infections and genetic relationships. Principal component analysis was used to characterize microbiome composition across the seven sample types. RESULTS: The study identified that microbiome composition was strongly host-driven, with more than 3844 genera detected, and two principal components explaining 62% of the total variation. Twenty-one dominant (based on relative abundance) Campylobacter species showed distinct clustering patterns for humans, ruminants, and broad hosts. The broad-host cluster included the most prevalent species, C. jejuni, C. concisus, and C. coli, present across sample types&#xa0;and a sub-cluster within C. jejuni involving humans, chickens, and ruminants. Campylobacter species from chickens showed strong positive correlations with mothers (r&#x2009;=&#x2009;0.76), siblings (r&#x2009;=&#x2009;0.61) and infants (r&#x2009;=&#x2009;0.54), while co-occurrence analysis found a higher likelihood (Pr&#x2009;>&#x2009;0.5) of pairs such as C. jejuni with C. coli, C. concisus, and C. showae. Analysis of the top 50 most abundant microbial taxa showed a distinct cluster uniquely present in human stool and absent in all livestock. The study also found frequent co-occurrence of C. jejuni with other enteric pathogens such as Salmonella, and Shigella, particularly in human and chicken. Additionally, instances of Candidatus Campylobacter infans (C. infans) were identified co-occurring with Salmonella and Shigella species in stool samples from infants, mothers, and siblings. CONCLUSIONS: A comprehensive analysis of Campylobacter diversity in humans and livestock in a low-resource setting revealed that infants can be exposed to multiple Campylobacter species early in life. C. jejuni is the dominant species with a propensity for co-occurrence with other notable enteric bacterial pathogens, including Salmonella, and Shigella, especially among infants. Video Abstract.

Animals↗