Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

MeRNA: a database of metal ion binding sites in RNA structures.

Metal ions are essential for the folding of RNA into stable tertiary structures and for the catalytic activity of some RNA enzymes. To aid in the study of the roles of metal ions in RNA structural biology, we have created MeRNA (Metals in RNA), a comprehensive compilation of all metal binding sites identified in RNA 3D structures available from the PDB and Nucleic Acid Database. Currently, our database contains information relating to binding of 9764 metal ions corresponding to 23 distinct elements, in 256 RNA structures. The metal ion locations were confirmed and ligands characterized using original literature references. MeRNA includes eight manually identified metal-ion binding motifs, which are described in the literature. MeRNA is searchable by PDB identifier, metal ion, method of structure determination, resolution and R-values for X-ray structure and distance from metal to any RNA atom or to water. New structures with their respective binding motifs will be added to the database as they become available. The MeRNA database will further our understanding of the roles of metal ions in RNA folding and catalysis and have applications in structural and functional analysis, RNA design and engineering. The MeRNA database is accessible at http://merna.lbl.gov.

Binding Sites↗

PepSeeker: a database of proteome peptide identifications for investigating fragmentation patterns.

Proteome science relies on bioinformatics tools to characterize proteins via their proteolytic peptides which are identified via characteristic mass spectra generated after their ions undergo fragmentation in the gas phase within the mass spectrometer. The resulting secondary ion mass spectra are compared with protein sequence databases in order to identify the amino acid sequence. Although these search tools (e.g. SEQUEST, Mascot, X!Tandem, Phenyx) are frequently successful, much is still not understood about the amino acid sequence patterns which promote/protect particular fragmentation pathways, and hence lead to the presence/absence of particular ions from different ion series. In order to advance this area, we have developed a database, PepSeeker (http://nwsr.smith.man.ac.uk/pepseeker), which captures this peptide identification and ion information from proteome experiments. The database currently contains >185,000 peptides and associated database search information. Users may query this resource to retrieve peptide, protein and spectral information based on protein or peptide information, including the amino acid sequence itself represented by regular expressions coupled with ion series information. We believe this database will be useful to proteome researchers wishing to understand gas phase peptide ion chemistry in order to improve peptide identification strategies. Questions can be addressed to j.selley@manchester.ac.uk.

Databases, Protein↗

Megx.net--database resources for marine ecological genomics.

Marine microbial genomics and metagenomics is an emerging field in environmental research. Since the completion of the first marine bacterial genome in 2003, the number of fully sequenced marine bacteria has grown rapidly. Concurrently, marine metagenomics studies are performed on a regular basis, and the resulting number of sequences is growing exponentially. To address environmentally relevant questions like organismal adaptations to oceanic provinces and regional differences in the microbial cycling of nutrients, it is necessary to couple sequence data with geographical information and supplement them with contextual information like physical, chemical and biological data. Therefore, new specialized databases are needed to organize and standardize data storage as well as centralize data access and interpretation. We introduce Megx.net, a set of databases and tools that handle genomic and metagenomic sequences in their environmental contexts. Megx.net includes (i) a geographic information system to systematically store and analyse marine genomic and metagenomic data in conjunction with contextual information; (ii) an environmental genome browser with fast search functionalities; (iii) a database with precomputed analyses for selected complete genomes; and (iv) a database and tool to classify metagenomic fragments based on oligonucleotide signatures. These integrative databases and webserver will help researchers to generate a better understanding of the functioning of marine ecosystems. All resources are freely accessible at http://www.megx.net.

Databases, Genetic↗

SGDB: a database of synthetic genes re-designed for optimizing protein over-expression.

Here we present the Synthetic Gene Database (SGDB): a relational database that houses sequences and associated experimental information on synthetic (artificially engineered) genes from all peer-reviewed studies published to date. At present, the database comprises information from more than 200 published experiments. This resource not only provides reference material to guide experimentalists in designing new genes that improve protein expression, but also offers a dataset for analysis by bioinformaticians who seek to test ideas regarding the underlying factors that influence gene expression. The SGDB was built under MySQL database management system. We also offer an XML schema for standardized data description of synthetic genes. Users can access the database at http://www.evolvingcode.net/codon/sgdb/index.php, or batch downloads all information through XML files. Moreover, users may visually compare the coding sequences of a synthetic gene and its natural counterpart with an integrated web tool at http://www.evolvingcode.net/codon/sgdb/aligner.php, and discuss questions, findings and related information on an associated e-forum at http://www.evolvingcode.net/forum/viewforum.php?f=27.

Databases, Nucleic Acid↗

TassDB: a database of alternative tandem splice sites.

Subtle alternative splice events at tandem splice sites are frequent in eukaryotes and substantially increase the complexity of transcriptomes and proteomes. We have developed a relational database, TassDB (TAndem Splice Site DataBase), which stores extensive data about alternative splice events at GYNGYN donors and NAGNAG acceptors. These splice events are of subtle nature since they mostly result in the insertion/deletion of a single amino acid or the substitution of one amino acid by two others. Currently, TassDB contains 114 554 tandem splice sites of eight species, 5209 of which have EST/mRNA evidence for alternative splicing. In addition, human SNPs that affect NAGNAG acceptors are annotated. The database provides a user-friendly interface to search for specific genes or for genes containing tandem splice sites with specific features as well as the possibility to download large datasets. This database should facilitate further experimental studies and large-scale bioinformatics analyses of tandem splice sites. The database is available at http://helios.informatik.uni-freiburg.de/TassDB/.

Alternative Splicing↗

REDIdb: the RNA editing database.

The RNA Editing Database (REDIdb) is an interactive, web-based database created and designed with the aim to allocate RNA editing events such as substitutions, insertions and deletions occurring in a wide range of organisms. The database contains both fully and partially sequenced DNA molecules for which editing information is available either by experimental inspection (in vitro) or by computational detection (in silico). Each record of REDIdb is organized in a specific flat-file containing a description of the main characteristics of the entry, a feature table with the editing events and related details and a sequence zone with both the genomic sequence and the corresponding edited transcript. REDIdb is a relational database in which the browsing and identification of editing sites has been simplified by means of two facilities to either graphically display genomic or cDNA sequences or to show the corresponding alignment. In both cases, all editing sites are highlighted in colour and their relative positions are detailed by mousing over. New editing positions can be directly submitted to REDIdb after a user-specific registration to obtain authorized secure access. This first version of REDIdb database stores 9964 editing events and can be freely queried at http://biologia.unical.it/py_script/search.html.

DNA, Complementary↗

Eukaryotic genome size databases.

Three independent databases of eukaryotic genome size information have been launched or re-released in updated form since 2005: the Plant DNA C-values Database (www.kew.org/genomesize/homepage.html), the Animal Genome Size Database (www.genomesize.com) and the Fungal Genome Size Database (www.zbi.ee/fungal-genomesize/). In total, these databases provide freely accessible genome size data for >10,000 species of eukaryotes assembled from more than 50 years' worth of literature. Such data are of significant importance to the genomics and broader scientific community as fundamental features of genome structure, for genomics-based comparative biodiversity studies, and as direct estimators of the cost of complete sequencing programs.

Animals↗

SNAPPI-DB: a database and API of Structures, iNterfaces and Alignments for Protein-Protein Interactions.

SNAPPI-DB, a high performance database of Structures, iNterfaces and Alignments of Protein-Protein Interactions, and its associated Java Application Programming Interface (API) is described. SNAPPI-DB contains structural data, down to the level of atom co-ordinates, for each structure in the Protein Data Bank (PDB) together with associated data including SCOP, CATH, Pfam, SWISSPROT, InterPro, GO terms, Protein Quaternary Structures (PQS) and secondary structure information. Domain-domain interactions are stored for multiple domain definitions and are classified by their Superfamily/Family pair and interaction interface. Each set of classified domain-domain interactions has an associated multiple structure alignment for each partner. The API facilitates data access via PDB entries, domains and domain-domain interactions. Rapid development, fast database access and the ability to perform advanced queries without the requirement for complex SQL statements are provided via an object oriented database and the Java Data Objects (JDO) API. SNAPPI-DB contains many features which are not available in other databases of structural protein-protein interactions. It has been applied in three studies on the properties of protein-protein interactions and is currently being employed to train a protein-protein interaction predictor and a functional residue predictor. The database, API and manual are available for download at: http://www.compbio.dundee.ac.uk/SNAPPI/downloads.jsp.

Cluster Analysis↗

EMBL Nucleotide Sequence Database in 2006.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl) at the EMBL European Bioinformatics Institute, UK, offers a large and freely accessible collection of nucleotide sequences and accompanying annotation. The database is maintained in collaboration with DDBJ and GenBank. Data are exchanged between the collaborating databases on a daily basis to achieve optimal synchrony. Webin is the preferred tool for individual submissions of nucleotide sequences, including Third Party Annotation, alignments and bulk data. Automated procedures are provided for submissions from large-scale sequencing projects and data from the European Patent Office. In 2006, the volume of data has continued to grow exponentially. Access to the data is provided via SRS, ftp and variety of other methods. Extensive external and internal cross-references enable users to search for related information across other databases and within the database. All available resources can be accessed via the EBI home page at http://www.ebi.ac.uk/. Changes over the past year include changes to the file format, further development of the EMBLCDS dataset and developments to the XML format.

Base Sequence↗

euHCVdb: the European hepatitis C virus database.

The hepatitis C virus (HCV) genome shows remarkable sequence variability, leading to the classification of at least six major genotypes, numerous subtypes and a myriad of quasispecies within a given host. A database allowing researchers to investigate the genetic and structural variability of all available HCV sequences is an essential tool for studies on the molecular virology and pathogenesis of hepatitis C as well as drug design and vaccine development. We describe here the European Hepatitis C Virus Database (euHCVdb, http://euhcvdb.ibcp.fr), a collection of computer-annotated sequences based on reference genomes. The annotations include genome mapping of sequences, use of recommended nomenclature, subtyping as well as three-dimensional (3D) molecular models of proteins. A WWW interface has been developed to facilitate database searches and the export of data for sequence and structure analyses. As part of an international collaborative effort with the US and Japanese databases, the European HCV Database (euHCVdb) is mainly dedicated to HCV protein sequences, 3D structures and functional analyses.

Databases, Protein↗

Misclassification model for person-time analysis of automated medical care databases.

A misclassification model is presented for the assessment of bias in rate ratios estimated by person-time analyses of automated medical care databases. The model allows for misclassification of events and person-time and applies to both differential and nondifferential errors. The focus is on medical care exposures that occur at discrete points in time (e.g., vaccinations) and on adverse events that are closely associated in time. Bias corrections for rate ratios and binomial tests of equality of event rates during exposed and unexposed person-time are developed and illustrated. For nondifferential under- or over-ascertainment of events, the observed rate ratio (r) is unbiased at the null hypothesis (true rate ratio R = 1), negatively biased when R > 1, and positively biased when R < 1 (i.e., biased toward the null). Differential under-ascertainment of unexposed events and differential over-ascertainment of exposed events positively bias r when R = 1. Differential event sensitivities cause larger biases in rate ratios than differential false event rates. False positive exposures bias observed event rate ratios more than false negative exposures. Biases are small when event sensitivities are nondifferential and when less than 10% of database exposures and events are false. The usefulness of the model for critical sensitivity analysis is illustrated by an example from a linked database study of childhood vaccine safety. Greater dissemination of data quality assessments, sensitivity analyses, and methods used to supplement automated databases are needed to further our understanding of the appropriate role of medical care databases in epidemiologic research.

Asthma↗

Patients with diagnosed diabetes mellitus can be accurately identified in an Indian Health Service patient registration database.

OBJECTIVE: The computerized patient registration databases maintained by the Indian Health Service (IHS) represent a potentially important source of data about the epidemic of diabetes among American Indian and Alaskan Native people. The purpose of this study is to determine the accuracy of this data source, and to identify the optimal search criteria to identify patients with a diagnosis of diabetes in an IHS patient registration database. METHODS: The authors compared the results of a series of computerized searches to a "gold standard" sample of 465 manually reviewed charts from a large IHS facility. RESULTS: Among patients ages 15 years and older, the best criterion for identifying patients diagnosed with diabetes was the presence of at least one purpose of visit narrative identified by a 250.00 to 250.93 ICD-9 code. The presence of a single computerized code for diabetes identified patients with diagnosed diabetes with a sensitivity of 92% (95% confidence interval [CI] 81, 97), a specificity of 99% (95% CI 98, 99), and a calculated positive predictive value of 94% (95% CI 85, 99). In a separate chart review of 462 charts of patients who had at least one 250.00 to 250.93 ICD-9 code recorded in the database, 435 had a diagnosis of diabetes for an observed positive predictive value of 94%. Because the prevalence of diabetes varies by age of the patient, the positive predictive value of the ability to identify patients with diabetes also varies by age. CONCLUSION: A computerized search of an IHS patient database can identify patients with a diagnosis of diabetes with an accuracy that is similar to the reported accuracy from other health care system databases.

Abstracting and Indexing↗

Comparison of systematic search and database methods for constructing segments of protein structure.

Two principal methods of determining the conformation of short pieces of polypeptide backbone in proteins have been developed: using a database of known structures and systematically generating all conformations. In this paper, we compare the effectiveness of these two techniques. The completeness of the database for segments of different lengths is examined and it is found to contain most conformations for segments seven residues long, but to deteriorate rapidly for longer regions. When the database segment is to be incorporated into the rest of a structure, at least seven residues are required to build four new residues, because of the need to position the segment relative to the rest of the structure. It is found that such positioning using flanking residues results in large errors in the inserted region. We conclude that the database method is currently not effective for comparative modeling, even for short segments. The systematic search procedure is found to generate almost all structures of short segments found in proteins. In contrast to the database method, low root mean square error structures are obtained for a set of trial segments embedded in the rest of a protein structure. Thus, it should be considered the method of choice.

Computer Simulation↗

Identifying quality outliers in a large, multiple-institution database by using customized versions of the Simplified Acute Physiology Score II and the Mortality Probability Model II0.

OBJECTIVE: To assess whether customized versions of the Simplified Acute Physiology Score (SAPS) II and the Mortality Probability Model (MPM) II0 agree on the identity of intensive care unit quality outliers within a multiple-center database. DESIGN: Retrospective database analysis. SETTING AND PATIENTS: Patient subset of the Project IMPACT database consisting of 39,617 adult patients admitted to surgical, medical, and mixed surgical-medical intensive care units at 54 hospitals between 1995 and 1999 who met inclusion criteria for SAPS II and MPM II0. INTERVENTIONS: Customized versions of SAPS II and MPM II0 were obtained by fitting new logistic regressions to the data by using the risk score as the independent variable and outcome at hospital discharge as the dependent variable. The data set was divided randomly into a training set and a validation set. Each model was customized by using the training set; model performance was then assessed in the validation set by using the area under the receiver operating characteristic curve and the Hosmer-Lemeshow statistic. The final models were based on the entire data set. The level of agreement between the customized models on the identity of quality outliers was evaluated by using kappa analysis. MEASUREMENTS AND MAIN RESULTS: Both customized models exhibited good discrimination and good calibration in this database. The area under the receiver operating characteristic curve was 0.83 for MPM II0 and 0.872 for SAPS II following model customization. The Hosmer-Lemeshow statistic was 12.3 ( >.14) for MPM II0, and 8.17 (p >.42) for SAPS II, after customization. Kappa analysis showed only fair agreement between the two customized models with regard to the identity of the quality outliers: kappa = 0.44 (95% confidence interval, 0.24, 0.65). CONCLUSIONS: Customization of SAPS II and MPM II0 to the Project IMPACT database resulted in well-calibrated models. Despite this, the models exhibited only a moderate level of agreement in which hospitals were designated as quality outliers. Seventeen of the 54 hospitals were categorized differently depending on which of the two scoring systems was used. Therefore, the rating of quality of care appears, in part, to be a function of the prediction model used.

APACHE↗

Multimedia computer database for neurosurgery.

OBJECTIVE: There is a need for an efficient mechanism of storing and analyzing neurosurgical clinical, imaging, and operative data to facilitate clinical audit, research, education, and preparation of scientific presentations. METHODS: A computer database was developed to meet this need. The recorded data include diagnoses, digitized neuroimaging studies, operative details (with intraoperative video clips), transcranial Doppler studies, outcomes, complications, admissions, and clinic visits. The anatomy, pathology, and clinical presentation are recorded for each diagnosis. RESULTS: The database provides an audit of neurosurgery cases, which includes admission Glasgow Coma Scale score, length of intensive care and hospital stays, Glasgow Outcome Scale score, and complications. Clinical research is facilitated by flexible search strategies based on the anatomy, pathology, or clinical presentation of diseases, or any of the recorded intraoperative or outcome factors. The system can be used to assess the influence on outcome of factors, such as transcranial Doppler velocity, intraoperative blood pressure, and the use of ventricular drainage, intraoperative angiography, or temporary clipping. The database can be used to track patients with untreated or partially treated conditions, such as incidental or incompletely coiled aneurysms. The recorded images and video clips are used for teaching and producing multimedia presentations and reports. The database is designed to enable secure Internet connections among institutions so that outcomes and complications can be compared among surgeons and institutions. CONCLUSION: This multimedia computer database facilitates clinical audit, research, teaching, and presentation activities.

Databases, Factual↗

Assessment of a race-specific normative HRT-III database to differentiate glaucomatous from normal eyes.

PURPOSE: To determine if a new, normative, race-specific database enhances the ability of confocal scanning laser ophthalmoscopy to differentiate normal from glaucomatous eyes. METHODS: One eye of eligible normal and glaucoma patients was enrolled. All subjects underwent a complete ophthalmologic examination, standard achromatic perimetry (SITA-SAP, 24-2), and confocal scanning laser ophthalmoscopy [Heidelberg retinal tomograph (HRT-II)] within 1 month of enrollment. Racial groups were defined by self-report. Glaucoma was defined by the existence of reproducible SAP loss (pattern standard deviation <5% and/or Glaucoma Hemifield Test outside normal limits) on 2 consecutive fields. Normal subjects had 2 normal visual fields (pattern standard deviation >5% and Glaucoma Hemifield Test within 97% normal limits) and a normal clinical examination. HRT-II examinations were exported to the HRT-III software, which includes a large race-specific normative database consisting of 733 white and 215 black eyes. Moorfields regression analysis (MRA) for the most abnormal optic disc sector was compared between the HRT-II (MRA2) and the HRT-III software before (MRA3-B) and after (MRA3-A) adjustment for race. Sectors outside the 99.9% confidence interval limits ("outside normal limits") were determined to be abnormal. RESULTS: We enrolled 124 black (52 glaucoma, 72 normal) and 96 white (32 glaucoma, 64 normal) subjects. Mean age was 51+/-13 years and 50+/-16 years for blacks and whites, respectively (P = 0.45). Visual field mean deviation was -7.3+/-6.7 db for glaucomatous eyes and -0.4+/-1.1 db for normal eyes (P < 0.001). Sensitivity and specificity for the HRT-II was 71.9% and 95.3%, respectively, for white subjects and 50.0% and 98.6%, respectively, for black subjects. Using the expanded HRT-III database, analysis yielded a sensitivity of 81.3% and specificity of 93.8% for whites and a sensitivity of 71.2% and specificity of 86.1% for blacks. After an adjustment for black ethnicity was made in the HRT-III program, the sensitivity and specificity for blacks was 65.4% and 90.3%, respectively. CONCLUSIONS: A new, larger, race-specific HRT-III database increases sensitivity while maintaining specificity for whites and increases sensitivity but decreases specificity for blacks. New software and databases based on race require careful scrutiny before use in clinical practice.

Adult↗

A graphical anatomical database of neural connectivity.

We describe a graphical anatomical database program, called XANAT (so named because it was developed under the X window system in UNIX), that allows the results of numerous studies on neuroanatomical connections to be stored, compared and analysed in a standardized format. Data are entered into the database by drawing injection and label sites from a particular tracer study directly onto canonical representations of the neuroanatomical structures of interest, along with providing descriptive text information. Searches may then be performed on the data by querying the database graphically, for example by specifying a region of interest within the brain for which connectivity information is desired, or via text information, such as keywords describing a particular brain region, or an author name or reference. Analyses may also be performed by accumulating data across multiple studies and displaying a colour-coded map that graphically represents the total evidence for connectivity between regions. Thus, data may be studied and compared free of areal boundaries (which often vary from one laboratory to the next), and instead with respect to standard landmarks, such as the position relative to well-known neuro-anatomical substrates or stereotaxic coordinates. If desired, areal boundaries may also be defined by the user to facilitate the interpretation of results. We demonstrate the application of the database to the analysis of pulvinar-cortical connections in the macaque monkey, for which the results of over 120 neuro-anatomical experiments were entered into the database. We show how these techniques can be used to elucidate connectivity trends and patterns that may otherwise go unnoticed.

Animals↗

SubtiList: a relational database for the Bacillus subtilis genome.

In the framework of the international collaborative project aiming to sequence the whole Bacillus subtilis chromosome, we have created a relational database for managing and analysing information associated with the molecular genetics of this bacterium: SubtiList. It allows recovery of non-redundant DNA sequences of the B. subtilis genome, as well as related information, i.e. genes, proteins, etc. A logical structure has been designed with appropriate links between the different objects, and a set of procedures has been implemented for data updating and management. The database is organized around a core constituted by all known contigs of B. subtilis, i.e. sets of non-redundant sequences created from original entries in the EMBL data library. A user-friendly interface has been developed to make the database easy to consult. Sequence analysis tools have been integrated into the database, such as a program for rapid similarity searching of protein data banks, and a powerful DNA pattern searching program. Thanks to the consistency of SubtiList, we have performed a codon usage analysis by Factorial Correspondence Analysis, and a study of the distribution of the isoelectric points of known proteins of B. subtilis. The SubtiList database is available through anonymous ftp (address 'ftp.pasteur.fr' or IP number 157.99.64.12, directory '/pub/GenomeDB/SubtiList').

Algorithms↗