Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Pfam: a comprehensive database of protein domain families based on seed alignments.

Databases of multiple sequence alignments are a valuable aid to protein sequence classification and analysis. One of the main challenges when constructing such a database is to simultaneously satisfy the conflicting demands of completeness on the one hand and quality of alignment and domain definitions on the other. The latter properties are best dealt with by manual approaches, whereas completeness in practice is only amenable to automatic methods. Herein we present a database based on hidden Markov model profiles (HMMs), which combines high quality and completeness. Our database, Pfam, consists of parts A and B. Pfam-A is curated and contains well-characterized protein domain families with high quality alignments, which are maintained by using manually checked seed alignments and HMMs to find and align all members. Pfam-B contains sequence families that were generated automatically by applying the Domainer algorithm to cluster and align the remaining protein sequences after removal of Pfam-A domains. By using Pfam, a large number of previously unannotated proteins from the Caenorhabditis elegans genome project were classified. We have also identified many novel family memberships in known proteins, including new kazal, Fibronectin type III, and response regulator receiver domains. Pfam-A families have permanent accession numbers and form a library of HMMs available for searching and automatic annotation of new protein sequences.

Amino Acid Sequence↗

Delayed extraction improves specificity in database searches by matrix-assisted laser desorption/ionization peptide maps.

Peptide mass maps obtained by matrix-assisted laser desorption ionization (MALDI) are an attractive means to identify proteins by searches in sequence databases. Here we demonstrate that the recently introduced delayed ion-extraction technique, when coupled to reflectron MALDI time-of-flight mass spectrometry, leads to dramatically improved search specificity. Routine resolution in the range of 6,000 to 12,000 allows assignment of monoisotopic masses throughout the peptide mass range. Database searches can be performed with high precision by use of a mass accuracy which is currently better than 30 ppm over a wide mass range and better than 5 ppm for a narrow mass range. This high performance makes it possible to identify proteins with fewer peptide masses than before. Additional low intensity peaks can be assigned after a search because of the improved signal-to-noise ratio of delayed-extraction peptide mass spectra, increasing sequence coverage of matched proteins. The improvements in database search specificity can be used to identify the components of simple protein mixtures. In combination with advanced sample preparation and automation techniques, delayed-extraction MALDI time-of-flight mass spectrometry is now an extremely powerful tool for the database identification of proteins.

Amino Acid Sequence↗

Balanced centralized and distributed database design in a clinical research environment.

Clinical research databases can meet both research and clinical needs, but this ideal is seldom achieved. Priorities often differ for those who collect and ultimately use the data and those who develop data systems. Traditional database designs also create logistical barriers that hamper communication. The Michigan Alzheimer's Disease Research Center has developed a secure, distributed data system with centralized data entry that provides an intuitive, individually customized interface for investigators in their clinics, laboratories and offices. Data are kept in a form that can be readily understood without reference to a code-book. Investigators can modify and query their own copies of the database without knowledge of programming languages. Balancing centralized and distributed designs for research databases enhance the accuracy and completeness of data collection and increases the use of data for research and clinical care.

Aged↗

A database model for medical consultation.

The database model presented in this article is suitable for applications in which queries may require noncrisp references to certain attributes. The data item (attribute) values may be crisp or fuzzy. For instance, such adjectives as "high" or "normal" may be attribute values for the attribute "blood pressure." A disease or a condition can be described by a number of symptoms which may be crisp alphanumeric values or fuzzy terms such as "high" or "normal." A query into this database can retrieve diseases which have "similar" symptoms. The similarity or "indistinguishability" is a measure defined by the database user on the relations that describe a family of diseases. This database system in conjunction with a rule base can provide the framework for a medical consultation system.

Computer Simulation↗

Update of the androgen receptor gene mutations database.

The current version of the androgen receptor (AR) gene mutations database is described. The total number of reported mutations has risen from 309 to 374 during the past year. We have expanded the database by adding information on AR-interacting proteins; and we have improved the database by identifying those mutation entries that have been updated. Mutations of unknown significance have now been reported in both the 5' and 3' untranslated regions of the AR gene, and in individuals who are somatic mosaics constitutionally. In addition, single nucleotide polymorphisms, including silent mutations, have been discovered in normal individuals and in individuals with male infertility. A mutation hotspot associated with prostatic cancer has been identified in exon 5. The database is available on the internet (http://www.mcgill.ca/androgendb/), from EMBL-European Bioinformatics Institute (ftp.ebi.ac.uk/pub/databases/androgen), or as a Macintosh FilemakerPro or Word file (MC33@musica.mcgill.ca).

3' Untranslated Regions↗

Coping with change: intellectual property rights, new legislation, and the human mutation database initiative.

In 1996, the European Union issued a directive requiring member states to protect databases against unauthorized copying. Similar legislation is currently being considered and will probably be enacted in the US. Such database legislation 1) will almost certainly increase existing pressures on human mutation databases to commercialize, and 2) could inadvertently make human mutation data harder to acquire and use. Strategies for minimizing these difficulties are discussed. Alternatively, a nonprofit, community-wide "depository" could probably support itself by selling sophisticated bioinformatic products to the private sector. The proposed depository would offer substantially similar databases to academic and government users at little or no cost.

Databases, Factual↗

The Kidney Development Database.

The Kidney Development Database is a bioinformatics resource dedicated to providing easily accessible information on gene expression during the development of the pro-, meso-, and metanephroi of a range of vertebrates. It also contains data on mutant phenotypes and on the effects of experimental manipulation of kidneys developing in culture. The database is searchable by gene name or by expression pattern. It is now being used as a test bed for more "advanced" search strategies that measure hypotheses of gene interactions against expression data to test for any clashes that would make the hypotheses untenable and that scan the database for potentially interesting correlations between changes in gene expression. The Kidney Development Database can be accessed free of charge via the World Wide Web at either of the following uniform resource locators (URLs); http://golgi.ana.ed.ac.uk/kidhome.html, and http://www.ana.ed.ac.uk/anatomy/database/kidbase/ kidhome.html.

Animals↗

The microbial proteome database--an automated laboratory catalogue for monitoring protein expression in bacteria.

Laboratories devoted to high-throughput characterisation of purified proteins arrayed via two-dimensional (2-D) gel electrophoresis face an arduous task in maintaining a centralised and constantly evolving record of information relating to the characterisation of proteins and their responses following biological challenges. The Microbial Proteome Database (MPD) has been conceived as an in-house resource for complementing the plethora of genomic databases available for such organisms. The database utilises commercially available software to provide an electronic 'lab book' of information obtained daily from 2-D electrophoresis gels, image analysis packages, protein characterisation methodologies, and biological experimentation. The MPD begins from a single 2-D gel image (a 2-D 'reference map') with clickable spots that link to a 'protein catalogue' (ProtCat) with spot information including protein identity, changes in expression determined under experimental conditions, cellular location, mass, and pI. The entry for each protein then contains further links to gel images corresponding to the presence of the particular protein within different subproteomes (as defined by the pH of narrow- and wide-range immobilised pH gradients or from differential extraction methods used to determine the location of the protein within a functional cell). The database currently contains information from strains of three microbial species (Escherichia coil, Pseudomonas aeruginosa and Staphylococcus aureus) and 32 master gel images. The rapid accessibility of information obtained from microbial proteomes is an essential step towards the integrated analysis of these organisms at the gene, transcript, protein and functional levels and will aid in reducing turnaround times between sample preparation and the discovery of molecules of biological significance.

Automation↗

BOLD--a biological O-linked glycan database.

Glycans can be O-linked to proteins via the hydroxyl group of serine, threonine, tyrosine, hydroxylysine or hydroxyproline. Sometimes the glycan is O-linked to the hydroxyl group via a phosphodiester bond. The core monosaccharide residue may be N-acetylgalactosamine, N-acetylglucosamine, galactose, glucose, fucose, mannose, xylose or arabinose. These O-linked glycans can remain as a monosaccharide, but often a complex structure is built up by stepwise addition of monosaccharides. Monosaccharides known to be added include galactose, N-acetylglucosamine, fucose, N-acetylneuraminic acid, N-glycolylneuraminic acid and 2-keto-3-deoxynonulosonic acid. O-linked glycans can also contain sulfate and phosphate residues. This leads to the possibility of the existence of numerous O-glycan structures. The biological O-linked database (BOLD) is a relational database that contains information on O-linked glycan structures, their biological sources (with a link to the SWISS-PROT protein database), the references in which the glycan was described (with a link to MEDLINE), and the methods used to determine the glycan structure. The database provides a valuable resource for glycobiology researchers interested in O-linked oligosaccharide structures that have been previously described on proteins from different species and tissues.

Carbohydrate Conformation↗

An evaluation of New Jersey's hospital discharge database for surveillance of severe occupational injuries.

Computerized population-based hospital discharge data in New Jersey offer new opportunities for surveillance of serious work-related injuries. This database was evaluated for its potential in identifying selected injuries that occurred at work during 1985 and 1986. Hospital discharge data were compared with data collected by telephone interview of discharged patients. A total of 1,575 unique hospital discharge records for the selected injuries included finger amputation (1,041), thumb amputation (209), crush injury of the lower limb (208), toxic effects of heavy metals (69), and eye burns (48). Of 809 study subjects sent letters, 445 (55%) could be contacted and 289 (36%) were interviewed for the study. Sixty-one percent (175) said their injury was work related. A comparison was made between self-reported injury at work, and the presence of workers' compensation payer codes on the discharge database. The agreement beyond chance (Kappa) was 0.78 (95% CI = 0.67, 0.89). The sensitivity of this indicator of work relatedness was 83%; specificity was 98%. These data suggest that workers' compensation payment on the hospital discharge database may be a good to excellent proxy indicator of the work relatedness of these injuries. However, this proxy indicator will underestimate the number of work-related injuries by about 20%. Only 11% of hospital discharge records had external cause of injury codes (E-codes), which reduces the utility of the database for understanding the causal mechanisms of work-related injuries.

Accidents, Occupational↗

Anabaptist genealogy database.

In late 1996 we set out to build a computer-searchable genealogy of the Old Order Amish of Lancaster County, Pennsylvania, for use by geneticists. The goals of the project included: 1) using the genealogy to expedite the mapping of genes mutated in three rare recessive disorders under study at the National Institutes of Health (NIH); 2) building a freely available software package, PedHunter, to answer genetically relevant queries on our database and other similar databases; and 3) providing genealogy assistance to researchers outside NIH. All of these scientific goals had to be accomplished while maintaining the confidentiality of the persons in the database and the confidentiality of preliminary research results. We expanded the project to include complementary data sources that contained many individuals who were Anabaptist, but not Amish, and many individuals who never lived in Lancaster County. For this reason, the project was renamed Anabaptist Genealogy Database (AGDB). All of the initial goals of the project have been accomplished, and we recently marked the 5-year anniversary of answering the first of over 100 queries by researchers outside NIH. Thus, it is an opportune time to review the construction of AGDB, summarize its usage to date, and speculate on future projects it might stimulate and facilitate.

Databases, Genetic↗

Completeness of state administrative databases for surveillance of congenital heart disease.

BACKGROUND: Tracking birth prevalence of cardiac defects is essential to determining time and space clusters, and identifying potential associated factors. Resource limitations on state birth defects surveillance programs sometimes require that databases already available be used for ascertaining such defects. This study evaluated the data quality of state administrative databases for ascertaining congenital heart defects (CHD) and specific diagnoses of CHD. METHODS: Children's Hospital of Wisconsin (CHW) medical records for infants born 1997-1999 and treated for CHD (n = 373) were abstracted and each case assigned CHD diagnoses based on definitive diagnostic reports (echocardiograms, catheterizations, surgical or autopsy reports). These data were linked to state birth and death records, and birth and postnatal (< 1 year of age) hospital discharge summaries at the Wisconsin Bureau of Health Information (WBHI). Presence of any code/checkbox indicating CHD (generic CHD) and exact matches to abstracted diagnoses were evaluated. RESULTS: Fifty-eight percent of cases with generic CHD were identified by state databases. Postnatal hospital discharge summaries identified 48%, birth hospital discharge summaries 27%, birth certificates 9% and death records 4% of these cases. Exact matches were found for 52% of 633 specific diagnoses. Postnatal hospital discharge summaries provided most matches. CONCLUSION: State databases identified 60% of generic CHD and exactly matched about half of specific CHD diagnoses. The postnatal hospital discharge summaries performed best in both in identifying generic CHD and matching specific CHD diagnoses. Vital records had limited value in ascertaining CHD.

Birth Certificates↗

Development, databases and the Internet.

There is now a rapidly expanding population of interlinked developmental biology databases on the World Wide Web that can be readily accessed from a desk-top PC using programs such as Netscape or Mosaic. These databases cover popular organisms (Arabidopsis, Caenorhabditis, Drosophila, zebrafish, mouse, etc.) and include gene and protein sequences, lists of mutants, information on resources and techniques, and teaching aids. More complex are databases relating domains of gene expression to embryonic anatomy and these range from existing text-based systems for specific organs such as kidney, to a massive project under development, that will cover gene expression during the whole of mouse embryogenesis. In this brief article, we review selected examples of databases currently available, look forward to what will be available soon, and explain how to gain access to the World Wide Web.

Animals↗

The MRC-5 human embryonal lung fibroblast two-dimensional gel cellular protein database: quantitative identification of polypeptides whose relative abundance differs between quiescent, proliferating and SV40 transformed cells.

A new version of the MRC-5 two-dimensional gel cellular protein database (Celis et al., Electrophoresis 1989, 10, 76-115) is presented. Gels were scanned with a Molecular Dynamics laser scanner and processed by the PDQUEST II software. A total of 1895 [35S]methionine-labeled cellular polypeptides (1323 with isoelectric focusing and 572 with nonequilibrium pH gradient electrophoresis) are recorded in this database, containing quantitative and qualitative data on the relative abundance of cellular proteins synthesized by quiescent, proliferating and SV40 transformed MRC-5 fibroblasts. Of the 592 proteins quantitated so far, the levels of 138 were up- or down-regulated (51 and 87, respectively) by two times or more in the transformed cells as compared to their normal proliferating counterparts, while only 14 behaved similarly in quiescent cells. Seven MRC-5 SV40 proteins, including plastin and two interferon-induced proteins, were not detected in the master MRC-5 images. The identity of 36 of the transformation-sensitive proteins whose levels are up or down regulated by two times or more was determined and additional information can be transferred from the master transformed human epithelial amnion cells (AMA) database (Celis et al., Electrophoresis 1990, 11, 989-1071) for those polypeptides of known and unknown identity that have been matched to AMA polypeptides. As more information is gathered in this and other laboratories, including data on oncogene proteins and transcription factors, this comprehensive database will outline an integrated picture of the expression levels and properties of the thousands of protein components of organelles, pathways and cytoskeletal systems that may be directly or indirectly involved in properties associated with the transformed state.

Cell Transformation, Viral↗

The rat liver epithelial (RLE) cell nuclear protein database.

The master two-dimensional computer database of rat liver epithelial (RLE) cellular proteins (Wirth et al., Electrophoresis 1991, 12, 931-954) has been expanded to include detailed information concerning 1100 nucleoplasmic (cytosolic) and 850 particulate associated [35S]methionine labeled as well as 215 nucleoplasmic and 269 particulate associated [32P]orthophosphate labeled RLE nuclear polypeptides, respectively. The RLE nuclear protein database developed using the Elsie 5 gel analysis system contains both qualitative and quantitative annotations including polypeptide identification number, protein name (if known), molecular weight and pI information, quantitation and polypeptide spot shape, subcellular location, as well as specific information regarding transformation (chemical and spontaneous) and growth-related characteristics. Microsequencing of polypeptides directly from two-dimensional (2-D) blotted membranes has recently been established in our laboratory and provides a highly efficient and rapid means of polypeptide identification in the absence of specific antibodies. At present the RLE protein database is still in the developmental stage and is continually being updated as additional information is obtained. Nonetheless, it is anticipated that knowledge obtained concerning the identification and characterization of specific transformation and/or growth regulatory proteins in the RLE in vitro cell system will not only have direct application to other rodent and human 2-D protein databases currently under development but will also complement them.

Amino Acid Sequence↗

The human myocardial two-dimensional gel protein database: update 1994.

An updated human heart protein two-dimensional electrophoresis (2-DE) database is presented. The database, which contains some 1388 protein spots characterised in terms of M(r) and pI, has been analysed further by Western immunoblotting and protein sequencing. From a total of 103 protein spots analysed, 49 have been identified by immunoblotting and 32 have been identified by protein sequencing. A further six proteins have tentatively been assigned by comparison with the human heart 2-DE protein database of Jungblut et al. (Electrophoresis) 1994, 15, 685-607). This database is being used in studies of alterations in protein expression in the diseased and transplanted human heart.

Amino Acid Sequence↗

A qualitative and quantitative protein database approach identifies individual and groups of functionally related proteins that are differentially regulated in simian virus 40 (SV40) transformed human keratinocytes: an overview of the functional changes associated with the transformed phenotype.

A qualitative and quantitative two-dimensional (2-D) gel database approach has been used to identify individual and groups of proteins that are differentially regulated in simian virus 40 (SV40) transformed human keratinocytes (K14). Five hundred and sixty [35S]methionine-labeled proteins (462 isoelectric focusing, IEF; 98 nonequilibrium pH gradient electrophoresis, NEPHGE), out of the 3038 recorded in the master keratinocyte database, were excised from dry, silver-stained gels of normal proliferating primary keratinocytes and K14 cells and the radioactivity was determined by liquid scintillation counting. Two hundred and thirty five proteins were found to be either up- (177) or down-regulated (58) in the transformed cells by 50% or more, and of these, 115 corresponded to known proteins in the keratinocyte database (J.E. Celis et al., Electrophoresis 1993, 14, 1091-1198). The lowest abundance acidic protein quantitated was present in about 60,000 molecules per cell, assuming a value of 10(8) molecules per cell for total actin. The results identified individual, and groups of functionally related proteins that are differentially regulated in K14 keratinocytes and that play a role in a variety of cellular activities that include general metabolism, the cytoskeleton, DNA replication and cell proliferation, transcription and translation, protein folding, assembly, repair and turnover, membrane traffic, signal transduction, and differentiation. In addition, the results revealed several transformation sensitive proteins of unknown identity in the database as well as known proteins of yet undefined functions. Within the latter group, members of the S100 protein family--whose genes are clustered on human chromosome 1q21--were among the highest down-regulated proteins in K14 keratinocytes. Visual inspection of films exposed for different periods of time revealed only one new protein in the transformed K14 keratinocytes and this corresponded to keratin 18, a cytokeratin expressed mainly by simple epithelia. Besides providing with the first global overview of the functional changes associated with the transformed phenotype of human keratinocytes, the data strengthened previous evidence indicating that transformation results in the abnormal expression of normal genes rather than in the expression of new ones.

Autoradiography↗

Identification of transformation sensitive proteins recorded in human two-dimensional gel protein databases by mass spectrometric peptide mapping alone and in combination with microsequencing.

A comprehensive human keratinocyte two-dimensional (2-D) gel protein database has been established to study the expression levels and properties of the thousands of proteins that orchestrate various keratinocyte functions both in health and disease, cancer included. A major task in establishing such a database is to identify known proteins in the 2-D gel patterns as well as to reveal hitherto unknown proteins. To date, protein identification has been performed by one or a combination of the following methods: (i) comigration with known proteins, (ii) Western blotting using specific antibodies, (iii) microsequencing and (iv) vaccinia virus expression of full length cDNAs. Recently, the systematic identification of proteins has gained a new dimension with the advent of computer programs for searching peptide molecular mass databases with experimentally obtained peptide mass maps. Here we investigate this approach to identify proteins that are highly up- or down-regulated in simian virus SV40 transformed human keratinocytes (K14). Peptide mass maps of several proteins, including keratins 7, 8, 18 and 19 were obtained either by plasma desorption mass spectrometry (PDMS) analysis of high performance liquid chromatography (HPLC) purified peptides or by matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS) of total digests. The results demonstrated that peptide mass maps can be used for a rapid and sensitive protein identification allowing fast screening of proteins recorded in 2-D gel databases. The mass spectrometric approach when combined with microsequencing strengthened identification, and added the possibility of full characterization of post-translational modifications and sequence variations.

Amino Acid Sequence↗