Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Argonaute--a database for gene regulation by mammalian microRNAs.

MicroRNAs (miRNAs) constitute a recently discovered class of small non-coding RNAs that regulate expression of target genes either by decreasing the stability of the target mRNA or by translational inhibition. They are involved in diverse processes, including cellular differentiation, proliferation and apoptosis. Recent evidence also suggests their importance for cancerogenesis. By far the most important model systems in cancer research are mammalian organisms. Thus, we decided to compile comprehensive information on mammalian miRNAs, their origin and regulated target genes in an exhaustive, curated database called Argonaute (http://www.ma.uni-heidelberg.de/apps/zmf/argonaute/interface). Argonaute collects latest information from both literature and other databases. In contrast to current databases on miRNAs like miRBase::Sequences, NONCODE or RNAdb, Argonaute hosts additional information on the origin of an miRNA, i.e. in which host gene it is encoded, its expression in different tissues and its known or proposed function, its potential target genes including Gene Ontology annotation, as well as miRNA families and proteins known to be involved in miRNA processing. Additionally, target genes are linked to an information retrieval system that provides comprehensive information from sequence databases and a simultaneous search of MEDLINE with all synonyms of a given gene. The web interface allows the user to get information for a single or multiple miRNAs, either selected or uploaded through a text file. Argonaute currently has information on 839 miRNAs from human, mouse and rat.

Animals↗

SYSTOMONAS--an integrated database for systems biology analysis of Pseudomonas.

To provide an integrated bioinformatics platform for a systems biology approach to the biology of pseudomonads in infection and biotechnology the database SYSTOMONAS (SYSTems biology of pseudOMONAS) was established. Besides our own experimental metabolome, proteome and transcriptome data, various additional predictions of cellular processes, such as gene-regulatory networks were stored. Reconstruction of metabolic networks in SYSTOMONAS was achieved via comparative genomics. Broad data integration is realized using SOAP interfaces for the well established databases BRENDA, KEGG and PRODORIC. Several tools for the analysis of stored data and for the visualization of the corresponding results are provided, enabling a quick understanding of metabolic pathways, genomic arrangements or promoter structures of interest. The focus of SYSTOMONAS is on pseudomonads and in particular Pseudomonas aeruginosa, an opportunistic human pathogen. With this database we would like to encourage the Pseudomonas community to elucidate cellular processes of interest using an integrated systems biology strategy. The database is accessible at http://www.systomonas.de.

Bacterial Proteins↗

Mouse Tumor Biology Database (MTB): status update and future directions.

The Mouse Tumor Biology (MTB) database provides access to data about endogenously arising tumors (both spontaneous and induced) in genetically defined mice (inbred, hybrid, mutant and genetically engineered mice). Data include information on the frequency and latency of mouse tumors, pathology reports and images, genomic changes occurring in the tumors, genetic (strain) background and literature or contributor citations. Data are curated from the primary literature or submitted directly from researchers. MTB is accessed via the Mouse Genome Informatics web site (http://www.informatics.jax.org). Integrated searches of MTB are enabled through use of multiple controlled vocabularies and by adherence to standardized nomenclature, when available. Recently MTB has been redesigned and its database infrastructure replaced with a robust relational database management system (RDMS). Web interface improvements include a new advanced query form and enhancements to already existing search capabilities. The Tumor Frequency Grid has been revised to enhance interactivity, providing an overview of reported tumor incidence across mouse strains and an entrée into the database. A new pathology data submission tool allows users to submit, edit and release data to the MTB system.

Animals↗

Databases in use at the individual monitoring service of ITN-DPRSN.

In this work, the databases developed for routine use at the Individual Monitoring for External Radiation Service (IMS) of the Radiological Protection and Nuclear Safety Department (DPRSN) at the Nuclear and Technological Institute (ITN) in Portugal are presented. At the IMS there are two dosimetry systems running simultaneously, one based on film and the other one on thermoluminescent detectors (TLD). Two databases were initially and independently home-developed in order to meet each service's needs. A few modifications were introduced and while each service's requirements were maintained where needed, the databases were adapted in order to store the same type of information relative to the facilities and monitored workers, as well as to produce similar shaped reports and technical information. The necessary administrative features of the services were considered in the database development, made user-friendly and welcomed by the ordinary users. The improvements allowed a more direct analysis of the annual doses and an easy identification of professions and practices associated with higher dose values.

Academies and Institutes↗

The mammalian protein-protein interaction database and its viewing system that is linked to the main FANTOM2 viewer.

Here, we describe the development of a mammalian protein-protein interaction (PPI) database and of a PPI Viewer application to display protein interaction networks (http://fantom21.gsc.riken.go.jp/PPI/). In the database, we stored the mammalian PPIs identified through our PPI assays (internal PPIs), as well as those we extracted and processed (external PPIs) from publicly available data sources, the DIP and BIND databases and MEDLINE abstracts by using FACTS, a new functional inference and curation system. We integrated the internal and external PPIs into the PPI database, which is linked to the main FANTOM2 viewer. In addition, we incorporated into the PPI Viewer information regarding the luciferase reporter activity of internal PPIs and the data confidence of external PPIs; these data enable visualization and evaluation of the reliability of each interaction. Using the described system, we successfully identified several interactions of biological significance. Therefore, the PPI Viewer is a useful tool for exploring FANTOM2 clone-related protein interactions and their potential effects on signaling and cellular communication.

Animals↗

Iditis: protein structure database.

The validation, enrichment and organization of the data stored in PDB files is essential for those data to be used accurately and efficiently for modelling, experimental design and the determination of molecular interactions. The Iditis protein structure database has been designed to allow the widest possible range of queries to be performed across all available protein structures. The Iditis database is the most comprehensive protein structure resource currently available, and contains over 500 fields of information describing all publicly deposited protein structures. A custom-written database engine and graphical user interface provide a natural and simple environment for the construction of searches for complex sequence- and structure-based motifs. Extensions and specialized interfaces allow the data generated by the database to used in conjunction with a wide range of applications.

Database Management Systems↗

SCOP, Structural Classification of Proteins database: applications to evaluation of the effectiveness of sequence alignment methods and statistics of protein structural data.

The Structural Classification of Proteins (SCOP) database provides a detailed and comprehensive description of the relationships of all known protein structures. The classification is on hierarchical levels: the first two levels, family and superfamily, describe near and far evolutionary relationships; the third, fold, describes geometrical relationships. The distinction between evolutionary relationships and those that arise from the physics and chemistry of proteins is a feature that is unique to this database, so far. The database can be used as a source of data to calibrate sequence search algorithms and for the generation of population statistics on protein structures. The database and its associated files are freely accessible from a number of WWW sites mirrored from URL http://scop. mrc-lmb.cam.ac.uk/scop/.

Algorithms↗

Computational method for temporal pattern discovery in biomedical genomic databases.

With the rapid growth of biomedical research databases, opportunities for scientific inquiry have expanded quickly and led to a demand for computational methods that can extract biologically relevant patterns among vast amounts of data. A significant challenge is identifying temporal relationships among genotypic and clinical (phenotypic) data. Few software tools are available for such pattern matching, and they are not interoperable with existing databases. We are developing and validating a novel software method for temporal pattern discovery in biomedical genomics. In this paper, we present an efficient and flexible query algorithm (called TEMF) to extract statistical patterns from time-oriented relational databases. We show that TEMF - as an extension to our modular temporal querying application (Chronus II) - can express a wide range of complex temporal aggregations without the need for data processing in a statistical software package. We show the expressivity of TEMF using example queries from the Stanford HIV Database.

Artificial Intelligence↗

Architecture of a mediator for a bioinformatics database federation.

Developments in our ability to integrate and analyze data held in existing heterogeneous data resources can lead to an increase in our understanding of biological function at all levels. However, supporting ad hoc queries across multiple data resources and correlating data retrieved from these is still difficult. To address this, we are building a mediator based on the functional data model database, P/FDM, which integrates access to heterogeneous distributed biological databases. Our architecture makes use of the existing search capabilities and indexes of the underlying databases, without infringing on their autonomy. Central to our design philosophy is the use of schemas. We have adopted a federated architecture with a five-level schema, arising from the use of the ANSI-SPARC three-level schema to describe both the existing autonomous data resources and the mediator itself. We describe the use of mapping functions and list comprehensions in query splitting, producing execution plans, code generation, and result fusion. We give an example of cross-database querying involving data held locally in P/FDM systems and external data in SRS.

Algorithms↗

Active concept learning in image databases.

Concept learning in content-based image retrieval systems is a challenging task. This paper presents an active concept learning approach based on the mixture model to deal with the two basic aspects of a database system: the changing (image insertion or removal) nature of a database and user queries. To achieve concept learning, we a) propose a new user directed semi-supervised expectation-maximization algorithm for mixture parameter estimation, and b) develop a novel model selection method based on Bayesian analysis that evaluates the consistency of hypothesized models with the available information. The analysis of exploitation versus exploration in the search space helps to find the optimal model efficiently. Our concept knowledge transduction approach is able to deal with the cases of image insertion and query images being outside the database. The system handles the situation where users may mislabel images during relevance feedback. Experimental results on Corel database show the efficacy of our active concept learning approach and the improvement in retrieval performance by concept transduction.

Algorithms↗

PlantPan: A comprehensive multi-species plant pan-genome database.

The pan-genome represents the complete genomic diversity of specific species, serving as a valuable resource for studying species evolution, crop domestication, and guiding crop breeding and improvement. While there are several single-species-specific plant pan-genome databases, the availability of multi-species pan-genome databases is limited. Additionally, variations in methods and data types used for plant pan-genome analysis across different databases hinder the comparison and integration of pan-genome information from various projects at multi-species or single-species levels. To tackle this challenge, we introduce PlantPan, a comprehensive database housing the results of pan-genome analysis for 195 genomes from 11 plant species. PlantPan aims to provide extensive information, including gene-centric and sequence-centric pan-genome information, graph-based pan-genome, pan-genome openness profiles, gene functions and its variation characteristics, homologous genes, and gene clusters across different species. Statistically, PlantPan incorporates 9 163 011 genes, 694 191 gene clusters, 526 973 370 genome variations, and 1 616 089 non-redundant genome variation groups at the species level, 33 455,098 genome synteny, and 177 827 non-redundant genome synteny groups at the species level. Regarding functional genes, PlantPan contains 5 222 720 genes related to transcription factors, 395 247 literature-reported resistance genes, 455 748 predicted microbial/disease resistance genes, and 1 612 112 genes related to molecular pathways. In summary, PlantPan is a vital platform for advancing the application of pan-genomes in molecular breeding for crops and evolutionary research for plants.

Genome, Plant↗

AdOnco: a database for clinical and scientific documentation of head and neck oncology.

OBJECTIVES: The aim of this project was to design, develop, and implement a head and neck cancer computer database for clinical and scientific use. METHODS: A relational database based on Filemaker Pro 6.0 was developed and integrated into our local network. A precise and easy-to-handle interface should allow for a quick overview of the patient's oncological history and for optimized data acquisition. An automatic report function was integrated to enhance the quality of daily patient care. For evaluation purposes, statistical analysis functions were implemented. RESULTS: Over a 14-month period, 410 patient records were available through the local network. The automated report function and the well-organized screen desktop resulted in time-efficient and accurate patient care. Additionally, the quality of information presented to referring physicians increased notably. The statistical analysis data provided by the database were reliable and easy to export. CONCLUSIONS: We developed an oncology database for both clinical and scientific purposes and integrated it successfully into our patient documentation system. The combination of clinical and scientific features proved to be very effective in daily routine and research.

Algorithms↗

d-matrix - database exploration, visualization and analysis.

BACKGROUND: Motivated by a biomedical database set up by our group, we aimed to develop a generic database front-end with embedded knowledge discovery and analysis features. A major focus was the human-oriented representation of the data and the enabling of a closed circle of data query, exploration, visualization and analysis. RESULTS: We introduce a non-task-specific database front-end with a new visualization strategy and built-in analysis features, so called d-matrix. d-matrix is web-based and compatible with a broad range of database management systems. The graphical outcome consists of boxes whose colors show the quality of the underlying information and, as the name suggests, they are arranged in matrices. The granularity of the data display allows consequent drill-down. Furthermore, d-matrix offers context-sensitive categorization, hierarchical sorting and statistical analysis. CONCLUSIONS: d-matrix enables data mining, with a high level of interactivity between humans and computer as a primary factor. We believe that the presented strategy can be very effective in general and especially useful for the integration of distinct data types such as phenotypical and molecular data.

Cardiovascular Diseases↗

SCOWLP: a web-based database for detailed characterization and visualization of protein interfaces.

BACKGROUND: Currently there is a strong need for methods that help to obtain an accurate description of protein interfaces in order to be able to understand the principles that govern molecular recognition and protein function. Many of the recent efforts to computationally identify and characterize protein networks extract protein interaction information at atomic resolution from the PDB. However, they pay none or little attention to small protein ligands and solvent. They are key components and mediators of protein interactions and fundamental for a complete description of protein interfaces. Interactome profiling requires the development of computational tools to extract and analyze protein-protein, protein-ligand and detailed solvent interaction information from the PDB in an automatic and comparative fashion. Adding this information to the existing one on protein-protein interactions will allow us to better understand protein interaction networks and protein function. DESCRIPTION: SCOWLP (Structural Characterization Of Water, Ligands and Proteins) is a user-friendly and publicly accessible web-based relational database for detailed characterization and visualization of the PDB protein interfaces. The SCOWLP database includes proteins, peptidic-ligands and interface water molecules as descriptors of protein interfaces. It contains currently 74,907 protein interfaces and 2,093,976 residue-residue interactions formed by 60,664 structural units (protein domains and peptidic-ligands) and their interacting solvent. The SCOWLP web-server allows detailed structural analysis and comparisons of protein interfaces at atomic level by text query of PDB codes and/or by navigating a SCOP-based tree. It includes a visualization tool to interactively display the interfaces and label interacting residues and interface solvent by atomic physicochemical properties. SCOWLP is automatically updated with every SCOP release. CONCLUSION: SCOWLP enriches substantially the description of protein interfaces by adding detailed interface information of peptidic-ligands and solvent to the existing protein-protein interaction databases. SCOWLP may be of interest to many structural bioinformaticians. It provides a platform for automatic global mapping of protein interfaces at atomic level, representing a useful tool for classification of protein interfaces, protein binding comparative studies, reconstruction of protein complexes and understanding protein networks. The web-server with the database and its additional summary tables used for our analysis are available at http://www.scowlp.org.

Algorithms↗

ProtRepeatsDB: a database of amino acid repeats in genomes.

BACKGROUND: Genome wide and cross species comparisons of amino acid repeats is an intriguing problem in biology mainly due to the highly polymorphic nature and diverse functions of amino acid repeats. Innate protein repeats constitute vital functional and structural regions in proteins. Repeats are of great consequence in evolution of proteins, as evident from analysis of repeats in different organisms. In the post genomic era, availability of protein sequences encoded in different genomes provides a unique opportunity to perform large scale comparative studies of amino acid repeats. ProtRepeatsDB http://bioinfo.icgeb.res.in/repeats/ is a relational database of perfect and mismatch repeats, access to which is designed as a resource and collection of tools for detection and cross species comparisons of different types of amino acid repeats. DESCRIPTION: ProtRepeatsDB (v1.2) consists of perfect as well as mismatch amino acid repeats in the protein sequences of 141 organisms, the genomes of which are now available. The web interface of ProtRepeatsDB consists of different tools to perform repeat s; based on protein IDs, organism name, repeat sequences, and keywords as in FASTA headers, size, frequency, gene ontology (GO) annotation IDs and regular expressions (REGEXP) describing repeats. These tools also allow formulation of a variety of simple, complex and logical queries to facilitate mining and large-scale cross-species comparisons of amino acid repeats. In addition to this, the database also contains sequence analysis tools to determine repeats in user input sequences. CONCLUSION: ProtRepeatsDB is a multi-organism database of different types of amino acid repeats present in proteins. It integrates useful tools to perform genome wide queries for rapid screening and identification of amino acid repeats and facilitates comparative and evolutionary studies of the repeats. The database is useful for identification of species or organism specific repeat markers, interspecies variations and polymorphism.

Amino Acid Sequence↗

A database application for pre-processing, storage and comparison of mass spectra derived from patients and controls.

BACKGROUND: Statistical comparison of peptide profiles in biomarker discovery requires fast, user-friendly software for high throughput data analysis. Important features are flexibility in changing input variables and statistical analysis of peptides that are differentially expressed between patient and control groups. In addition, integration the mass spectrometry data with the results of other experiments, such as microarray analysis, and information from other databases requires a central storage of the profile matrix, where protein id's can be added to peptide masses of interest. RESULTS: A new database application is presented, to detect and identify significantly differentially expressed peptides in peptide profiles obtained from body fluids of patient and control groups. The presented modular software is capable of central storage of mass spectra and results in fast analysis. The software architecture consists of 4 pillars, 1) a Graphical User Interface written in Java, 2) a MySQL database, which contains all metadata, such as experiment numbers and sample codes, 3) a FTP (File Transport Protocol) server to store all raw mass spectrometry files and processed data, and 4) the software package R, which is used for modular statistical calculations, such as the Wilcoxon-Mann-Whitney rank sum test. Statistic analysis by the Wilcoxon-Mann-Whitney test in R demonstrates that peptide-profiles of two patient groups 1) breast cancer patients with leptomeningeal metastases and 2) prostate cancer patients in end stage disease can be distinguished from those of control groups. CONCLUSION: The database application is capable to distinguish patient Matrix Assisted Laser Desorption Ionization (MALDI-TOF) peptide profiles from control groups using large size datasets. The modular architecture of the application makes it possible to adapt the application to handle also large sized data from MS/MS- and Fourier Transform Ion Cyclotron Resonance (FT-ICR) mass spectrometry experiments. It is expected that the higher resolution and mass accuracy of the FT-ICR mass spectrometry prevents the clustering of peaks of different peptides and allows the identification of differentially expressed proteins from the peptide profiles.

Algorithms↗

Searching biomedical databases on complementary medicine: the use of controlled vocabulary among authors, indexers and investigators.

BACKGROUND: The optimal retrieval of a literature search in biomedicine depends on the appropriate use of Medical Subject Headings (MeSH), descriptors and keywords among authors and indexers. We hypothesized that authors, investigators and indexers in four biomedical databases are not consistent in their use of terminology in Complementary and Alternative Medicine (CAM). METHODS: Based on a research question addressing the validity of spinal palpation for the diagnosis of neuromuscular dysfunction, we developed four search concepts with their respective controlled vocabulary and key terms. We calculated the frequency of MeSH, descriptors, and keywords used by authors in titles and abstracts in comparison to standard practices in semantic and analytic indexing in MEDLINE, MANTIS, CINAHL, and Web of Science. RESULTS: Multiple searches resulted in the final selection of 38 relevant studies that were indexed at least in one of the four selected databases. Of the four search concepts, validity showed the greatest inconsistency in terminology among authors, indexers and investigators. The use of spinal terms showed the greatest consistency. Of the 22 neuromuscular dysfunction terms provided by the investigators, 11 were not contained in the controlled vocabulary and six were never used by authors or indexers. Most authors did not seem familiar with the controlled vocabulary for validity in the area of neuromuscular dysfunction. Recently, standard glossaries have been developed to assist in the research development of manual medicine. CONCLUSIONS: Searching biomedical databases for CAM is challenging due to inconsistent use of controlled vocabulary and indexing procedures in different databases. A standard terminology should be used by investigators in conducting their search strategies and authors when writing titles, abstracts and submitting keywords for publications.

Abstracting and Indexing↗

Structure-activity relations: maximizing the usefulness of mutagenicity and carcinogenicity databases.

The most important criteria for the development and analysis of databases for elucidating the structural bases of toxicological activity include the integrity of the databases with respect to uniformity of the experimental protocol and interpretation of the test results and inclusion of chemicals representing different chemical classes and differing mechanisms of action. Within these criteria, it is demonstrated that when the chemicals are chosen at random, the larger the database, the better the predictivity of chemicals not included in the learning set. It is shown however, that when chemicals are selected on the basis of structural features, that a learning set of approximately 180 chemicals is as informative as a database consisting of 800 chemicals chosen at random.

Animals↗