Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

A dynamic two-dimensional polyacrylamide gel electrophoresis database: the mycobacterial proteome via Internet.

Proteome analysis by two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) and mass spectrometry, in combination with protein chemical methods, is a powerful approach for the analysis of the protein composition of complex biological samples. Data organization is imperative for efficient handling of the vast amount of information generated. Thus we have constructed a 2-D PAGE database to store and compare protein patterns of cell-associated and culture-supernatant proteins of different mycobacterial strains. In accordance with the guidelines for federated 2-DE databases, we developed a program that generates a dynamic 2-D PAGE database for the World-Wide-Web to organise and publish, via the internet, our results from proteome analysis of different Mycobacterium tuberculosis as well as Mycobacterium bovis BCG strains. The uniform resource locator for the database is http://www.mpiib-berlin.mpg.de/2D-PAGE and can be read with a Java compatible browser. The interactive hypertext markup language documents displayed are generated dynamically in each individual session from a rational data file, a 2-D gel image file and a map file describing the protein spots as polygons. The program consists of common gateway interface scripts written in PERL, minimizing the administrative workload of the database. Furthermore, the database facilitates not only interactive use, but also worldwide active participation of other scientific groups with their own data, requiring only minimal computer hardware and knowledge of information technology.

Bacterial Proteins↗

The structure and dipole moment of globular proteins in solution and crystalline states: use of NMR and X-ray databases for the numerical calculation of dipole moment.

The large dipole moment of globular proteins has been well known because of the detailed studies using dielectric relaxation and electro-optical methods. The search for the origin of these dipolemoments, however, must be based on the detailed knowledge on protein structure with atomic resolutions. At present, we have two sources of information on the structure of protein molecules: (1) x-ray databases obtained in crystalline state; (2) NMR databases obtained in solution state. While x-ray databases consist of only one model, NMR databases, because of the fluctuation of the protein folding in solution, consist of a number of models, thus enabling the computation of dipole moment repeated for all these models. The aim of this work, using these databases, is the detailed investigation on the interdependence between the structure and dipole moment of protein molecules. The dipole moment of protein molecules has roughly two components: one dipole moment is due to surface charges and the other, core dipole moment, is due to polar groups such as N--H and C==O bonds. The computation of surface charge dipole moment consists of two steps: (A) calculation of the pK shifts of charged groups for electrostatic interactions and (B) calculation of the dipole moment using the pK corrected for electrostatic shifts. The dipole moments of several proteins were computed using both NMR and x-ray databases. The dipole moments of these two sets of calculations are, with a few exceptions, in good agreement with one another and also with measured dipole moments.

Animals↗

Use of mass spectrometric molecular weight information to identify proteins in sequence databases.

During the last decade new ionization techniques have made it possible to measure the molecular weight of many intact proteins by mass spectrometry, and they have made it much easier to obtain a mass spectrometric peptide map of a protein. At the same time advances in protein and DNA sequencing technology are resulting in an exponential increase in the number of sequences deposited in databases. Here we investigate the possibility to use mass spectrometric data to identify proteins in databases. Searching a database by total molecular weight is found to be an easy and sometimes sufficient approach. For more specificity and for error tolerance in both the mass spectrometric data and the database information we search by partial mass spectrometric peptide map of the protein. In general, just four to six proteolytic peptides measured with a mass accuracy between 0.1 and 0.01% allow a useful search of databases such as the Protein Identification Resource (PIR). As the size of DNA and protein sequence databases grows, protein identification by partial mass spectrometric peptide maps should become increasingly powerful and may become a general method to identify and characterize proteins.

Amino Acid Sequence↗

Microcomputer-assisted filing system of cardiac catheterization records using a relational database management system.

To efficiently store and retrieve cardiac catheterization records, we have developed a computer-assisted database, which comprises a 16-bit microcomputer with dual floppy disk drives, a 20 MB random-access memory, hard disk drive, and a line printer. All programmings were accomplished using a relational database management system (R:base 5000, Microrim, Inc.). Data inquiry procedures could be performed with direct operational commands of the system as well as with preprogrammed command files, and final results of searches were printed out with a line printer. The major advantages of the present system described in this report include: (1) the relatively easy and rapid creation of the database, (2) ease of modification of the database structures even after the system design is finished, (3) operational commands in combination with conditional operator(s) are flexible and powerful enough to allow the end user to retrieve data based on various kinds of criteria, (4) a high-level programming language provided by the R:base automates a series of database procedures with relative ease, (5) relational capabilities of the database management system can enhance the possibility of reconstruction of a new data file from a single or several preexisting data files, and (6) the system can be realized at reasonable cost.

Cardiac Catheterization↗

The gene-protein database of Escherichia coli: edition 4.

The gene-protein database of Escherichia coli has as its core an index that links each of the protein spots from a two-dimensional polyacrylamide gel to the gene that encodes the protein. Additional information about each protein and its gene is generated from two-dimensional gel analysis or collated from the literature to form the database. Earlier editions of the database have provided periodic updates of information. The current edition does this, but also introduces a new reference gel image produced by an electrophoresis system recently adopted in this laboratory. The new gel system was chosen because it offers an improved opportunity for other investigations to produce close replicas of the reference gel pattern, thereby allowing easier access to the information of the database and encouraging independent contribution to the database. The new gel format also is larger and hence more compatible with computer assisted image analysis, which has become essential for a project of this magnitude. This edition continues the use of the former reference gel images, but adds a reference image of an equilibrium gel of E. coli strain W3110 produced by the new standardized gel system. At this time, 55% of the protein spots annotated on the previous equilibrium reference gel for this organism have been located on the new reference image, and these identifications are included in the tables of the database.

Bacterial Proteins↗

Mouse liver protein database: a catalog of proteins detected by two-dimensional gel electrophoresis.

Alterations in the abundance or structure of mouse liver proteins are being studied using two-dimensional gel electrophoresis (2-DE) to build a database of protein changes correlating with exposure to ionizing radiation or toxic chemicals. Thus far, studies have included the analysis of proteins from the offspring of exposed parents or from the exposed individuals themselves. In order to characterize and identify proteins found altered by such exposures, sex- and strain-related differences in protein patterns have been analyzed, and the subcellular locations of a large portion of the mapped proteins have been determined. As part of these studies, data are collected and stored using a variety of computer hardware and software tools that allow the accumulation of information on the origin of samples, gel identification, experiment description, and protein similarities and differences. This accumulation of information constitutes the mouse liver protein database. Relational database software is used to tie the different facets of the database together so that the results of a variety of experiments can be compared and interrelated. The database optimizes the information obtained from 2-DE gel sets by allowing use of the data for many purposes, including monitoring of gel resolution to ensure the collection of high quality data and correlation of protein effects induced by different agents. This first edition of the Argonne National Laboratory mouse liver protein database lays the foundation for future work and communication that should elucidate the significance of observed protein effects as possible markers of exposure to toxic agents.

Animals↗

HSC-2DPAGE and the two-dimensional gel electrophoresis database of dog heart proteins.

A two-dimensional gel electrophoresis database of dog (Canis familiaris) proteins is presented. The database contains 1212 protein spots which have been characterised in terms of their pI and Mr. This database has been integrated into the HSC-2DPAGE database which is accessible on the Internet via the World Wide Web with the uniform resource location (URL): (http://www.harefield.nthames.nhs.uk/nhli/ protein/index.html). Identifications for 80 of the protein spots have been obtained by visual cross-matching with the human heart protein database in HSC-2DPAGE (42 spots), N-terminal microsequence analysis (25 spots) and peptide mass fingerprinting (20 spots). This database is being used in studies of alterations in protein expression in models of heart failure and heart disease.

Amino Acid Sequence↗

Mining the human proteome: experience with the human lymphoid protein database.

We have undertaken an effort in the past five years aimed at developing a database of lymphoid proteins detectable by two-dimensional (2-D) polyacrylamide gel electrophoresis. The database contains 2-D patterns and derived information pertaining to: (i) polypeptide constituents of unstimulated and stimulated mature T cells and immature thymocytes; (ii) cultured T cells and cell lines that have been manipulated by transfection with a variety of constructs or by treatment with specific agents; (iii) single cell-derived T and B cell clones; (iv) cells obtained from patients with lymphoproliferative disorders and leukemia; and (v) a variety of other relevant cell populations. The database has experienced a substantial expansion in 2-D patterns it contains, numbering currently 9167 individual 2-D patterns. This number represents a fraction of the 30,682 2-D patterns maintained in our databases. The capacity to design and undertake experiments, produce high-quality 2-D patterns, and to undertake simple or rudimentary analyses of 2-D patterns to meet the basic needs of the experiments for which the 2-D gels were produced has exceeded the capacity to fully and uniformly integrate information generated from any gel image or experiment, across all images and experiments. While only a fraction of the information in the 2-D patterns contained in the lymphoid database has been mined, novel findings derived from querying the database point to the merits of this protein based approach. Additional resources have recently been put into place to mine more effectively data pertaining to protein expression in lymphoid cells.

Cell Cycle↗

Infevers: an evolving mutation database for auto-inflammatory syndromes.

The Infevers database (http://fmf.igh.cnrs.fr/infevers/) was established in 2002 to provide investigators with access to a central source of information about all sequence variants associated with periodic fevers: Familial Mediterranean fever (FMF), TNF Receptor Associated Periodic Syndrome (TRAPS), Hyper IgD Syndrome (HIDS), Familial Cold Autoinflammatory Syndrome/Muckle-Wells Syndrome/Chronic Infantile Neurological Cutaneous and Articular Syndrome (FCAS/MWS/CINCA). The prototype of this group of disorders is FMF, a recessive disease characterized by recurrent bouts of unexplained inflammation. FMF is the pivotal member of an expanding family of autoinflammatory disorders, a new term coined to describe illnesses resulting from a defect of the innate immune response. Therefore, we decided to extend the Infevers database to genes connected with autoinflammatory diseases. We present here the biological content of the Infevers database, including the introduction of two new entries: Crohn/Blau and Pyogenic sterile arthritis, pyoderma gangrenosum and acne (PAPA syndrome). Infevers has a range of query capabilities, allowing for simple or complex interrogation of the database. Currently, the database contains 291 sequence variants in related genes (MEFV, TNFRSF1A, MVK, CARD15, PSTPIP1, and CIAS1), consisting of published data and personal communications, which has revealed or refined the preferential mutational sites for each gene. This database will continue to evolve in its content and to improve in its presentation.

Arthritis↗

HAEdb: a novel interactive, locus-specific mutation database for the C1 inhibitor gene.

Hereditary angioneurotic edema (HAE) is an autosomal dominant disorder characterized by episodic local subcutaneous and submucosal edema and is caused by the deficiency of the activated C1 esterase inhibitor protein (C1-INH or C1INH; approved gene symbol SERPING1). Published C1-INH mutations are represented in large universal databases (e.g., OMIM, HGMD), but these databases update their data rather infrequently, they are not interactive, and they do not allow searches according to different criteria. The HAEdb, a C1-INH gene mutation database (http://hae.biomembrane.hu) was created to contribute to the following expectations: 1) help the comprehensive collection of information on genetic alterations of the C1-INH gene; 2) create a database in which data can be searched and compared according to several flexible criteria; and 3) provide additional help in new mutation identification. The website uses MySQL, an open-source, multithreaded, relational database management system. The user-friendly graphical interface was written in the PHP web programming language. The website consists of two main parts, the freely browsable search function, and the password-protected data deposition function. Mutations of the C1-INH gene are divided in two parts: gross mutations involving DNA fragments >1 kb, and micro mutations encompassing all non-gross mutations. Several attributes (e.g., affected exon, molecular consequence, family history) are collected for each mutation in a standardized form. This database may facilitate future comprehensive analyses of C1-INH mutations and also provide regular help for molecular diagnostic testing of HAE patients in different centers.

Angioedema↗

PolyMAPr: programs for polymorphism database mining, annotation, and functional analysis.

Pharmacogenomic and disease-association studies rely on identifying a comprehensive set of polymorphisms within candidate genes. Public SNP databases are a rich source of polymorphism data, but mining them effectively requires overcoming at least four challenges: ensuring accurate annotations for genes and polymorphisms, eliminating both inter- and intra-database redundancy, integrating data from multiple public sources with data generated locally, and prioritizing the variants for further study. PolyMAPr (Polymorphism Mining and Annotation Programs)' was developed to overcome these challenges and to improve the efficiency of database mining and polymorphism annotation. PolyMAPr takes as input a file containing a list of genes to be processed and files containing each annotated gene sequence. Polymorphic sequences obtained from public databases (dbSNP, CGAP, and JSNP) or through local SNP discovery efforts, as well as oligonucleotide sequences (e.g., PCR primers), are mapped to the annotated gene sequences and named according to suggested nomenclature guidelines. The functional effects of nonsynonymous coding-region SNPs (cSNPs) and any variants that might alter exon splicing enhancer (ESE) sites, putative transcription factor binding sites, or intron-exon splice sites are predicted. The output files are accessible though a browser interface. In addition, the results are also provided in Extensible Markup Language (XML) format to facilitate uploading them into a local relational database. PolyMAPr increases the efficiency of mining public databases for genetic variants within candidate genes and provides a mechanism by which data from multiple sources (both public and private) can be uniformly integrated, thereby significantly reducing the effort required to obtain a comprehensive set of polymorphisms for pharmacogenomic and disease-association studies. PolyMAPr can be obtained from http://pharmacogenomics.wustl.edu.

Databases, Nucleic Acid↗

The human TBX5 gene mutation database.

Germline mutations of the TBX5 gene were identified as the primary cause in up to 70% of patients with Holt-Oram syndrome (HOS), an autosomal dominant disorder characterized by malformations of the upper limbs and cardiac defects. Furthermore, somatic mutations of the TBX5 gene have been described in diseased heart tissues of patients with congenital heart defects of different cause. The relationship between genotype and phenotype remains unclear and the underlying mechanism of the pathogenic effect is not solved. In this report, we introduce the 'TBX5 Gene Mutation Database,' an online locus specific database containing germline and somatic mutations of the TBX5 gene. The permanently updated data collection includes all reported mutations beginning with the first description of the gene in 1997. With our database we complement the existing resources by: 1) giving a complete review of the so far reported mutation spectrum in TBX5 considering the clinical relevance; 2) linkage of the mutational data to the corresponding gene location and PubMed-Abstracts; and 3) additional links to other related resources like SNP database, sequences and literature references. The usage of our database will help to quickly find informations about genetic variations within the TBX5 gene. Here we describe the database structure, content, and potential applications (http://www.uni-leipzig.de/~genetik/TBX5).

Databases, Genetic↗

The NIDDK liver transplantation database.

UNLABELLED: The NIDDK Liver Transplantation Database was established to prospectively investigate questions related to the experience of patients evaluated for and undergoing liver transplantation. This article presents the study design, methods, and quality of data collection, along with some of the overall results. METHODS: An initial 4-year planning phase was used to develop data collection instruments and quality control procedures regarding assessment for transplantation, liver donors, and the recipients' pre-, peri- and postoperative course. During the 1990-1995 implementation phase, three clinical centers refined the data collection instruments and enrolled and followed consecutive liver transplant candidates who consented to be included in the protocol. RESULTS: The Database contains more than 49,000 data forms from 1563 candidates, 1002 donors, and 916 transplant recipients followed up to 5 years after transplantation. Overall, 95% of protocol forms were completed. The Database includes uniformly defined histology results of liver biopsies performed per protocol and for complications throughout follow-up. In addition, the Database maintains an inventory of available sera for the Serum Bank. All test results of studies performed on the sera are added to the Database. Of 1563 evaluated patients, 59% were deemed eligible for liver transplantation. Of the others who were too well or had contraindications, 15% became eligible later. Characteristics of patients in this study were generally comparable to those of patients nationally. CONCLUSIONS: The NIDDK Liver Transplantation Database has yielded comprehensive and high quality data and is a rich resource for extensive analysis about many important clinical aspects of liver transplantation.

Adolescent↗

Potential limitations of electronic database studies of prescription non-aspirin non-steroidal anti-inflammatory drugs (NANSAIDs) and risk of myocardial infarction (MI).

PURPOSE: To determine whether specific limitations in electronic database studies may lead to biased estimates of the association between prescription, non-selective non-aspirin non-steroidal anti-inflammatory drugs (NANSAIDs) and myocardial infarction (MI) METHODS: Using our case-control study of NANSAIDs and first, non-fatal MI, we determined the odds ratio (OR) for prescription NANSAIDs and MI. In the 'Replicating Electronic Database Analysis,' we considered non-prescription NANSAID users to be 'non-users,' did not stratify by aspirin use, and did not adjust for confounders typically unavailable or incomplete in existing databases. In the 'Misclassification Assessment Analysis,' we removed non-prescription NANSAIDs from the 'non-user' category. In the 'Confounding Assessment Analysis #1,' we additionally adjusted for smoking, family history, and years of education. In the 'Confounding Assessment Analysis #2,' we also adjusted for body mass index (BMI) and physical activity. In the 'Interaction Assessment Analysis,' we stratified on aspirin use and repeated the latter analysis. RESULTS: The prevalence of current NANSAID and aspirin use was higher in our controls than in electronic database studies, consistent with the fact that non-prescription NANSAIDs accounted for 81% of all NANSAID use. Education, physical activity, and BMI also were associated with prescription NANSAID use. When each potential source of bias was removed, the OR for NANSAIDs moved further from 1.0 (i.e., toward a protective association with MI): 'Replicating Electronic Database' analysis (OR 1.00, 95% confidence interval [CI]: 0.78--1.28); 'Misclassification Assessment Analysis' (OR 0.89, 95%CI: 0.70--1.14); 'Confounding Assessment Analysis #1' (OR 0.85, 95%CI: 0.66--1.10); 'Confounding Assessment Analysis #2' (OR 0.78, 95%CI: 0.60--1.01); 'Interaction Assessment Analysis' (OR 0.69, 95%CI: 0.51--0.95). CONCLUSIONS: Limitations in electronic databases may be responsible for the lack of association of NANSAIDs on lower MI risk noted in these studies. Further studies-preferably randomized trials-are needed to address the risk-benefit ratio of NANSAID use.

Adult↗

Validity of asthma diagnoses recorded in the Medical Services database of Quebec.

The goal of this study was to evaluate the validity of asthma diagnoses recorded in the Medical Services (physician billing) database of the Canadian province of Quebec. The predictive positive value (PPV) and predictive negative value (PNV) of two operational definitions of asthma based on diagnoses recorded in the database were evaluated. Patients 16-80 years old treated by a respiratory or a family physician in 2002 were selected from the database. The diagnosis derived from the Medical Services database was compared to the diagnosis written in the patient's medical chart. The PPV and PNV of the first operational definition based on one asthma diagnosis or more recorded in the database over a 1-year period were found to be 0.75 and 0.96 for respiratory physicians and 0.67 and 0.99 for family physicians, for patients 16-44 years old. The PPV increased to 0.78 for family physicians and to 0.77 for respiratory physicians when the second operational definition based on two diagnoses of asthma or more was used. Results tended to be lower for 45-80 years old patients. We conclude that diagnoses recorded in the Medical Services database of Quebec are valid to identify patients with asthma.

Adolescent↗

Refill adherence for patients with asthma and COPD: comparison of a pharmacy record database with manually collected repeat prescriptions.

PURPOSE: To compare refill adherence data based on two different methods of data capturing, that is, manually collected repeat prescriptions and a pharmacy record database. METHODS: The study comprised a comparison of adherence data from manually collected repeat prescriptions of asthma and chronic obstructive pulmonary disease (COPD) drugs with fixed dosages dispensed in 2002 and the corresponding data from a pharmacy record database. Data were collected in the county of Jämtland in Sweden. Refill adherence was calculated for the different collection methods. RESULTS: Data from 285 manually collected repeat prescriptions for asthma/COPD drugs for 2002 showed that 35% of the prescribings had been satisfactory refilled, while 42% showed an undersupply and 23% an oversupply. The pharmacy record database had 490 prescribings for asthma/COPD drugs registered in 2002, 28% of these had a satisfactory refill adherence, while 43% showed an undersupply, and 29% an oversupply. Based on the database it could be shown that 11% of the individuals had used more than one repeat prescription of the same medicine during 2002. Based on the pharmacy record database for 1999-2002, it was shown that 29% of the prescribings had been satisfactory refilled whereas undersupply increased (53%) and oversupply decreased (18%) as compared to the 1-year data. CONCLUSIONS: Refill adherence determined from manually collected repeat prescriptions and from a pharmacy record database did not differ for a 1-year period. Four-year data might give a better overview of patients' refill adherence than 1-year data.

Anti-Asthmatic Agents↗

Presence of pharmacoepidemiology in three bibliographic databases: Medline, IPA and SCI.

OBJECTIVE: The objective of this study is to make a comparative description of the evolution and distribution of international research into pharmacoepidemiology, using three bibliographic databases, in order to select the most appropriate for future bibliometric studies. METHODS: Bibliographic searches were performed using the following databases: Medline (1966-99), IPA (1970-99) and SCI (1990-99), using the term 'pharmacoepidemiology'. On the basis of these searches, the number of original articles per year and per journal title were noted. The growth of the output of scientific writing was found to fit Price's law. RESULTS: A total of 845 original articles were recovered: 467 from IPA, 219 from Medline and 159 from SCI. The highest mean number of original articles per year (33.4) was obtained with the IPA database. Price's exponential growth pattern was observed among all three databases. The total numbers of journals in which the original articles were published were 102 in Medline, 65 in IPA and 60 in SCI. The journals providing a single original article comprised 65% of the Medline titles and 61% of those in IPA and SCI. CONCLUSIONS: International research into pharmacoepidemiology presents an exponential growth pattern, in accordance with Price's law. There is a large degree of publishing dispersion. IPA was found to be the bibliographic database that recovered the greatest number of original articles, nearly half of which were published in Pharmacoepidemiology and Drug Safety. We therefore consider the latter database appropriate for bibliometric studies in the field of pharmacoepidemiology.

Bibliometrics↗

YAKUSA: a fast structural database scanning method.

YAKUSA is a program designed for rapid scanning of a structural database with a query protein structure. It searches for the longest common substructures called SHSPs (structural high-scoring pairs) existing between a query structure and every structure in the structural database. It makes use of protein backbone internal coordinates (alpha angles) in order to describe protein structures as sequences of symbols. The structural similarities are established in 5 steps, the first 3 being analogous to those used in BLAST: (1) building up a deterministic finite automaton describing all patterns identical or similar to those in the query structure; (2) searching for all these patterns in every structure in the database; (3) extending the patterns to longer matching substructures (i.e., SHSPs); (4) selecting compatible SHSPs for each query-database structure pair; and (5) ranking the query-database structure pairs using 3 scores based on SHSP similarity, on SHSP probabilities, and on spatial compatibility of SHSPs. Structural fragment probabilities are estimated according to a mixture transition distribution model, which is an approximation of a high-order Markov chain model. With regard to sensitivity and selectivity of the structural matches, YAKUSA compares well to the best related programs, although it is by far faster: A typical database scan takes about 40 s CPU time on a desktop personal computer. It has also been implemented on a Web server for real-time searches.

Algorithms↗