Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metadata”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Meningioma transcriptomic landscape demonstrates novel subtypes with regional associated biology and patient outcome.

Meningiomas, although mostly benign, can be recurrent and fatal. World Health Organization (WHO) grading of the tumor does not always identify high-risk meningioma, and better characterizations of their aggressive biology are needed. To approach this problem, we combined 13 bulk RNA sequencing (RNA-seq) datasets to create a dimension-reduced reference landscape of 1,298 meningiomas. The clinical and genomic metadata effectively correlated with landscape regions, which led to the identification of meningioma subtypes with specific biological signatures. The time to recurrence also correlated with the map location. Further, we developed an algorithm that maps new patients onto this landscape, where the nearest neighbors predict outcome. This study highlights the utility of combining bulk transcriptomic datasets to visualize the complexity of tumor populations. Further, we provide an interactive tool for understanding the disease and predicting patient outcomes. This resource is accessible via the online tool Oncoscape, where the scientific community can explore the meningioma landscape.

Meningioma↗

Protocol for histology-anchored macroscopic staging of gonadal maturity in exploited fishes.

Here, we present a protocol to assign gonadal maturity stages in commercially exploited fishes using a histology-anchored workflow. We describe steps for recording field metadata, photographing gonads, fixing central gonadal tissue, and processing paraffin sections. We then detail procedures for staining sections with hematoxylin and eosin, diagnosing gametogenic features, and assigning stages using a common reproductive-phase framework with species- and sex-specific reference descriptors. This protocol standardizes documentation and decision logic rather than proposing a new maturity scale.

Developmental biology↗

Continuing dental education on the World Wide Web.

Continuing dental education (CDE) courses delivered on the World Wide Web (Web CDE) offer numerous advantages over traditional CDE; however, two major issues--location of suitable courses and course quality--need resolution. Locating high-quality courses is difficult due to the lack of the standardized metadata that allows search engines to match courses to practitioners' needs. Web directories created by professional organizations are beginning to show promise, but require further development. Search engines and Web directories are discussed and improvements currently underway summarized. Course quality remains a highly significant concern. A national effort to create Web CDE course quality standards is underway that includes proposed standards. These proposed standards are summarized and used to comment on the current state of Web CDE courses. Examples are given when possible. Three emerging Web CDE technologies and a look to the future of Web CDE are discussed.

Computer-Assisted Instruction↗

Bioconductor: an open source framework for bioinformatics and computational biology.

This chapter describes the Bioconductor project and details of its open source facilities for analysis of microarray and other high-throughput biological experiments. Particular attention is paid to concepts of container and workflow design, connections of biological metadata to statistical analysis products, support for statistical quality assessment, and calibration of inference uncertainty measures when tens of thousands of simultaneous statistical tests are performed.

Animals↗

A system for simultaneous multiple subject, multiple stimulus modality, and multiple channel collection and analysis of sensory evoked potentials.

A system has been developed for collecting sensory evoked potentials simultaneously from multiple channels for multiple subjects at up to 80 kHz sample rate per channel. Sample rates up to 200 kHz are available for four or less chambers and a single channel per chamber. A variety of visual, somatosensory, and auditory stimuli may be presented singly or simultaneously. Collected waveforms are associated with searchable text (metadata) to allow convenient selection from a relational database. Multiple waveforms can then be easily grouped for analysis and processed. Results can be exported to other software for further graphics or statistical processing. Scripting and event logging are available to provide automation and improve data confidence. Sample data are presented from control animals for each of the sensory modalities for comparison with historical data collected from other systems.

Animals↗

The impact of Life Science Identifier on informatics data.

Since the Life Science Identifier (LSID) data identification and access standard made its official debut in late 2004, several organizations have begun to use LSIDs to simplify the methods used to uniquely name, reference and retrieve distributed data objects and concepts. In this review, the authors build on introductory work that describes the LSID standard by documenting how five early adopters have incorporated the standard into their technology infrastructure and by outlining several common misconceptions and difficulties related to LSID use, including the impact of the byte identity requirement for LSID-identified objects and the opacity recommendation for use of the LSID syntax. The review describes several shortcomings of the LSID standard, such as the lack of a specific metadata standard, along with solutions that could be addressed in future revisions of the specification.

Computational Biology↗

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa↗

ThermoData Engine (TDE): software implementation of the dynamic data evaluation concept.

The first full-scale software implementation of the dynamic data evaluation concept {ThermoData Engine (TDE)} is described for thermophysical property data. This concept requires the development of large electronic databases capable of storing essentially all experimental data known to date with detailed descriptions of relevant metadata and uncertainties. The combination of these electronic databases with expert-system software, designed to automatically generate recommended data based on available experimental data, leads to the ability to produce critically evaluated data dynamically or 'to order'. Six major design tasks are described with emphasis on the software architecture for automated critical evaluation including dynamic selection and application of prediction methods and enforcement of thermodynamic consistency. The direction of future enhancements is discussed.

Journal Article↗

Bringing chemical data onto the Semantic Web.

Present chemical data storage methodologies place many restrictions on the use of the stored data. The absence of sufficient high-quality metadata prevents intelligent computer access to the data without human intervention. This creates barriers to the automation of data mining in activities such as quantitative structure-activity relationship modelling. The application of Semantic Web technologies to chemical data is shown to reduce these limitations. The use of unique identifiers and relationships (represented as uniform resource identifiers, URIs, and resource description framework, RDF) held in a triplestore provides for greater detail and flexibility in the sharing and storage of molecular structures and properties.

Journal Article↗

ChemSem: an extensible and scalable RSS-based seminar alerting system for scientific collaboration.

A seminar announcement system based on the extensive use of XML-based data structures, CML/MathML for carrying more domain-specific molecular content, and open source software components is described. The output is a resource description framework (RDF) site summary (RSS) feed, which potentially carries many advantages over conventional announcement mechanisms, including the ability to aggregate and then sort multiple and diverse RSS feeds on the basis of declared metadata and to feed into RDF-based mechanisms for establishing links between different subject areas.

Journal Article↗

Integrating multi-omics technologies to decipher microbiome functions.

Multi-omics approaches have revolutionized our understanding of microbial communities by enabling simultaneous interrogation of genomic, transcriptomic, proteomic, and metabolomic data. The systematic integration and analysis of these deep datasets help decipher the functional roles of microbiomes, providing critical insights into microbial activities, interactions, and dynamics across diverse environments. Biological complexity makes multi-omics analysis of a single, isolated organism demanding but highly informative, yet this complexity increases further when samples comprise hundreds to thousands of individual species. As microbiome research continues to expand into clinical, environmental, and engineered systems, standardized workflows, benchmarked datasets, and community-driven initiatives are essential to ensure reproducibility, standardization and interpretability. Establishing and disseminating best practices for experimental design, data processing, and integrative analyses will be critical for maximizing comparability and scientific rigor across studies. This perspective highlights recent advances in multi-omics microbiome research, outlines key obstacles in data integration and metadata harmonization, and proposes a collaborative roadmap for scalable, FAIR-compliant multi-omics investigations and potentially disruptive Artificial Intelligence (AI) advances comparable to those of AlphaFold in the field of microbiome science.

Multiomics↗

Structure-centric searching enables global mapping of the public metabolome.

Searching and learning from aggregated public metabolomics data spanning thousands of studies remained largely inaccessible. Here we present StructureMASST, a web-based application enabling scalable, structure-centric searches across public metabolomics repositories using molecule names or chemical representations. It queries a precomputed knowledgebase of 2.19 billion spectral matches and 420 million metadata links, supports modification-tolerant and mass-shift searches, and maps chemical structures across taxonomy, biological context and environmental conditions to accelerate discovery.

Journal Article↗

Temporal stability and lack of variance in microbiome composition and functionality in fit recreational athletes.

Human gut microbiome composition and function is influenced by environmental and lifestyle factors, including exercise and fitness. We studied the composition and functionality of the faecal microbiome of recreational (non-elite) runners (n = 62) with serial shotgun metagenomics, at 4 time points over a 7-week period. Gut microbiome composition and function was stable over time. Grouping of samples on the basis of their fitness level (fair, good, excellent, and superior) or habitual training (low (4-6 h/week), medium (7-9 h/week), high (10-12 h/week), and extreme (13 + hours/week)) revealed no significant microbiome-related differences. Overall, the species Faecalibacterium prausnitzii, Blautia wexlerae, and Prevotella copri were the most abundant members of the gut microbiome. Analysis of co-abundance groups (CAGs) revealed no significant relationship between CAGs and fitness levels or training subgroups. Functional pathways were similar across all samples and timepoints with no clustering based on associated metadata. The most abundant genes identified within samples corresponded to pathways for nucleoside and nucleotide biosynthesis, amino acid biosynthesis, and cell wall biosynthesis. Collectively, these results describe the microbiome of active recreational runners and note temporal stability amongst participants.

Humans↗

Analysis of molecular data of Arabidopsis thaliana (L.) Heynh. (Brassicaceae) with Geographical Information Systems (GIS).

A Geographical Information System (GIS) is used to analyse allelic information of 13 sequenced loci of natural populations of Arabidopsis thaliana and to identify geographical structures. GIS provides tools for visualization and analysis of geographical population structures using molecular data. The geographical distribution of the number of variable positions in the alignments, the distribution of recombinant sequence blocks, and the distribution of a newly defined measure, the differentiation index, are studied. The differentiation index is introduced to measure the sequence divergence among individual plants sampled from various geographical localities. The numbers of variable positions and the differentiation index are also used for a metadata analysis covering about 26 kb of the genome. This analysis reveals, for the first time, differences in DNA sequence structures of geographically different populations of A. thaliana. The broadly defined west Mediterranean region consists of accessions with the highest numbers of polymorphic positions followed by the west European region. The GIS technology Kriging is used to define Arabidopsis specific diversity zones in Europe. The highest genetic variability is observed along the Atlantic coast from the western Iberian Peninsula to southern Great Britain, while lowest variability is found in central Europe.

Arabidopsis↗

Validating existing data in the Environmental Technology Verification Program.

Establishing the credibility of existing data is an ongoing issue, particularly when the data sets are to be used for a secondary purpose, i.e., not the original reason for which they were collected. If the secondary purpose is similar to the primary purpose, the potential user may have little difficulty establishing credibility since the acceptance criteria for both purposes should be similar. If the secondary purpose is different, then data credibility may be more difficult to establish because the experiment generating the data may not have been conducted optimally for the secondary purpose and all of the necessary quality assurance data ("metadata") may not have been collected. In either case, a process will be required to determine the acceptability of the data. For this reason, at the time the U.S. Environmental Protection Agency (EPA) Environmental Technology Verification (ETV) program was established, similar certification and verification programs run by states or foreign countries routinely used existing data sets, for cost reasons, rather than generate new data by testing. The issue of whether existing data could be used in the ETV program immediately surfaced. In response, a policy and a process that addressed existing data were written and published in Appendix C of the ETV Quality and Management Plan (Hayes et al., 1998). This paper discusses how the ETV program determines the credibility of existing data used to verify the performance of environmental technologies.

Data Interpretation, Statistical↗

Where one size does not fit all: understanding the needs of potential users of a portal to breast cancer knowledge online.

The article argues that, although the Internet has great potential for assisting people to find information on breast cancer, at present that potential is not being realised. The literature shows considerable dissatisfaction with information provision for breast cancer, including on the Internet where appropriate information suited to particular needs often cannot be found. An Australian project (Breast Cancer Knowledge Online [BCKOnline]), in its first stage, set out to explore the needs for breast cancer information using an ethnographic method and a purposive sample of 77 participants, most of them women with breast cancer. A portal, which will enable users to tailor information to their particular needs, is at present being developed based on the results of the needs analysis. The process includes user-selected profiles, enabled through "user-centric" resource descriptions, and a metadata repository that links the profiles with specific information resources. The article presents limited results from the needs analysis-those highlighting the differences between younger and older women and the problems with present Internet information provision as seen by the sample. The final section discusses how the portal will both tailor information to needs and assist with the problems with the Internet revealed in the literature.

Adult↗

A search tool based on 'encapsulated' MeSH thesaurus to retrieve quality health resources on the internet.

In the year 2001, the Internet has become a major source of health information for the health professional and the Netizen. The objective of Doc' CISMeF (D'C) was to create a powerful generic search tool based on a structured information model which 'encapsulates' the MeSH thesaurus to index and retrieve quality health resources on the Internet. To index resources, D'C uses four sections in its information model: 'meta-term', keyword, subheading, and resource type. Two search options are available: simple and advanced. The simple search requires the end-user to input a single term or expression. If this term belongs to the D'C information structure model, it will be exploded. If not, a full-text search is performed. In the advanced search, complex searches are possible combining Boolean operators with meta-terms, keywords, subheadings and resource types. D'C uses two standard tools for organising information: the MeSH thesaurus and the Dublin Core metadata format. Resources included in D'C are described according to the following elements: title, author or creator, subject and keywords, description, publishers, date, resource type, format, identifier, and language.

Abstracting and Indexing↗

Implementing context and team based access control in healthcare intranets.

The establishment of an efficient access control system in healthcare intranets is a critical security issue directly related to the protection of patients' privacy. Our C-TMAC (Context and Team-based Access Control) model is an active security access control model that layers dynamic access control concepts on top of RBAC (Role-based) and TMAC (Team-based) access control models. It also extends them in the sense that contextual information concerning collaborative activities is associated with teams of users and user permissions are dynamically filtered during runtime. These features of C-TMAC meet the specific security requirements of healthcare applications. In this paper, an experimental implementation of the C-TMAC model is described. More specifically, we present the operational architecture of the system that is used to implement C-TMAC security components in a healthcare intranet. Based on the technological platform of an Oracle Data Base Management System and Application Server, the application logic is coded with stored PL/SQL procedures that include Dynamic SQL routines for runtime value binding purposes. The resulting active security system adapts to current need-to-know requirements of users during runtime and provides fine-grained permission granularity. Apart from identity certificates for authentication, it uses attribute certificates for communicating critical security metadata, such as role membership and team participation of users.

Computer Communication Networks↗