Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,765 records · Page 98Linked to original sources

The evolution of structural databases.

Starting with the Protein Data Bank (PDB) as a common ancestor, the evolution of structural databases has been driven by the rapprochement of the structural world and the practical applications. The result is an impressive number of secondary structural databases that is welcomed by structural biologists and bioinformaticians but runs the risk of producing an embarrassment of riches among non-specialist users. Given that any profit depends on the number of customers, efficient interfaces between many structural data banks must be available to make their contents easily accessible. Increasing the information content of central structural repositories might be the best way to guide users through the many, sometimes overlapping databases.

Computer Communication Networks↗

Pandit: a database of protein and associated nucleotide domains with inferred trees.

MOTIVATION: A large, high-quality database of homologous sequence alignments with good estimates of their corresponding phylogenetic trees will be a valuable resource to those studying phylogenetics. It will allow researchers to compare current and new models of sequence evolution across a large variety of sequences. The large quantity of data may provide inspiration for new models and methodology to study sequence evolution and may allow general statements about the relative effect of different molecular processes on evolution. RESULTS: The Pandit 7.6 database contains 4341 families of sequences derived from the seed alignments of the Pfam database of amino acid alignments of families of homologous protein domains (Bateman et al., 2002). Each family in Pandit includes an alignment of amino acid sequences that matches the corresponding Pfam family seed alignment, an alignment of DNA sequences that contain the coding sequence of the Pfam alignment when they can be recovered (overall, 82.9% of sequences taken from Pfam) and the alignment of amino acid sequences restricted to only those sequences for which a DNA sequence could be recovered. Each of the alignments has an estimate of the phylogenetic tree associated with it. The tree topologies were obtained using the neighbor joining method based on maximum likelihood estimates of the evolutionary distances, with branch lengths then calculated using a standard maximum likelihood approach.

Algorithms↗

CVD: the intestinal crypt/villus in situ hybridization database.

UNLABELLED: The intestinal crypt/villus in situ hybridization database (CVD) query interface is a web-based tool to search for genes with similar relative expression patterns along the crypt/villus axis of the mammalian intestine. The CVD is an online database holding information for relative gene expression patterns in the mammalian intestine and is based on the scoring of in situ hybridization experiments reported in the literature. CVD contains expression data for 88 different genes collected from 156 different in situ hybridization profiles. The web-based query interface allows execution of both single gene queries and pattern searches. The query results provide links to the most relevant public gene databases. AVAILABILITY: http://pc113.imbg.ku.dk/ps/

Animals↗

Markov model recognition and classification of DNA/protein sequences within large text databases.

MOTIVATION: Short sequence patterns frequently define regions of biological interest (binding sites, immune epitopes, primers, etc.), yet a large fraction of this information exists only within the scientific literature and is thus difficult to locate via conventional means (e.g. keyword queries or manual searches). We describe herein a system to accurately identify and classify sequence patterns from within large corpora using an n-gram Markov model (MM). RESULTS: As expected, on test sets we found that identification of sequences with limited alphabets and/or regular structures such as nucleic acids (non-ambiguous) and peptide abbreviations (3-letter) was highly accurate, whereas classification of symbolic (1-letter) peptide strings with more complex alphabets was more problematic. The MM was used to analyze two very large, sequence-containing corpora: over 7.75 million Medline abstracts and 9000 full-text articles from Journal of Virology. Performance was benchmarked by comparing the results with Journal of Virology entries in two existing manually curated databases: VirOligo and the HLA Ligand Database. Performance estimates were 98 +/- 2% precision/84% recall for primer identification and classification and 67 +/- 6% precision/85% recall for peptide epitopes. We also find a dramatic difference between the amounts of sequence-related data reported in abstracts versus full text. Our results suggest that automated extraction and classification of sequence elements is a promising, low-cost means of sequence database curation and annotation. AVAILABILITY: MM routine and datasets are available upon request.

Abstracting and Indexing↗

circASbase: A Comprehensive Database of Alternative Splicing Events in circRNAs.

Although extensive evidence has underscored the critical role of alternative splicing (AS) in generating mature circular RNA (circRNA) isoforms and augmenting their functional diversity, a significant gap remains in the availability of specialized databases housing circRNA AS events. To bridge this gap, we develop circASbase, a pioneering and comprehensive database that catalogs 452,129 AS events in 884,047 full-length circRNAs from 581 samples across 13 species, and provides rich annotations to facilitate understanding the splicing regulation of circRNA. Our findings reveal substantial differences between circRNAs and linear transcripts regarding the distribution and occurrence of AS events, highlighting the unique regulatory landscape of circRNAs. These special splicing events result in functional differences of circRNAs by affecting internal ribosome entry sites, N6-methyladenosine sites, open reading frames, protein features, microRNA targets, and more. In summary, circASbase not only meets the urgent need of the research community for data repositories, but also represents a significant advancement in our understanding of circRNA biology. With its user-friendly interfaces and web-based visualization tools, circASbase is poised to become an indispensable resource for researchers exploring the regulatory mechanisms and functional roles of AS events in circRNAs. This database will continuously drive new insights and discoveries in the field, setting the stage for further advancements in circRNA research. circASbase is freely available at http://reprod.njmu.edu.cn/cgi-bin/circASbase/.

Alternative Splicing↗

S/MARt DB: a database on scaffold/matrix attached regions.

S/MARt DB, the S/MAR transaction database, is a relational database covering scaffold/matrix attached regions (S/MARs) and nuclear matrix proteins that are involved in the chromosomal attachment to the nuclear scaffold. The data are mainly extracted from original publications, but a World Wide Web interface for direct submissions is also available. S/MARt DB is closely linked to the TRANSFAC database on transcription factors and their binding sites. It is freely accessible through the World Wide Web (http://transfac.gbf.de/SMARtDB/) for non-profit research.

Animals↗

ORFer--retrieval of protein sequences and open reading frames from GenBank and storage into relational databases or text files.

BACKGROUND: Functional genomics involves the parallel experimentation with large sets of proteins. This requires management of large sets of open reading frames as a prerequisite of the cloning and recombinant expression of these proteins. RESULTS: A Java program was developed for retrieval of protein and nucleic acid sequences and annotations from NCBI GenBank, using the XML sequence format. Annotations retrieved by ORFer include sequence name, organism and also the completeness of the sequence. The program has a graphical user interface, although it can be used in a non-interactive mode. For protein sequences, the program also extracts the open reading frame sequence, if available, and checks its correct translation. ORFer accepts user input in the form of single or lists of GenBank GI identifiers or accession numbers. It can be used to extract complete sets of open reading frames and protein sequences from any kind of GenBank sequence entry, including complete genomes or chromosomes. Sequences are either stored with their features in a relational database or can be exported as text files in Fasta or tabulator delimited format. The ORFer program is freely available at http://www.proteinstrukturfabrik.de/orfer. CONCLUSION: The ORFer program allows for fast retrieval of DNA sequences, protein sequences and their open reading frames and sequence annotations from GenBank. Furthermore, storage of sequences and features in a relational database is supported. Such a database can supplement a laboratory information system (LIMS) with appropriate sequence information.

Animals↗

Tests of methods for evaluating bibliographic databases: an analysis of the National Library of Medicine's handling of literatures in the medical behavioral sciences.

This article reports on five separate studies designed for the National Library of Medicine (NLM) to develop and test methodologies for evaluating the products of large databases. The methodologies were tested on literatures of the medical behavioral sciences (MBS). One of these studies examined how well NLM covered MBS monographic literature using CATLINE and OCLC. Another examined MBS journal and serial literature coverage in MEDLINE and other MBS-related databases available through DIALOG. These two studies used 1010 items derived from the reference lists of sixty-one journals, and tested for gaps and overlaps in coverage in the various databases. A third study examined the quality of the indexing NLM provides to MBS literatures and developed a measure of indexing as a system component. The final two studies explored how well MEDLINE retrieved documents on topics submitted by MBS professionals and how online searchers viewed MEDLINE (and other systems and databases) in handling MBS topics. The five studies yielded both broad research outcomes and specific recommendations to NLM.

Abstracting and Indexing↗

TOXLINE: evolution of an online interactive bibliographic database.

The National Library of Medicine has offered TOXLINE, an online interactive bibliographic database of biomedical (toxicology) information since 1972. Files from 11 secondary sources comprise the TOXLINE database. The sources supplied bibliographic records in different formats and data structures. Data from each supplier's format had to be converted into a format suitable for TOXLINE. Three different, successive retrieval systems were used for the TOXLINE database which required reformatting of the data. Algorithms for generating terms for inverted file search methods were tested. Special characters peculiar to the scientific literature were evaluated during search term generation. Developing search term algorithms for chemical names in the scientific literature required techniques different from those used for nonscientific literature. Problems with replication of bibliographic records from multiple secondary sources are described. Some observations about online interactive databases since TOXLINE was first offered are noted.

MEDLARS↗

Report of workshop on cellular protein databases derived from two-dimensional polyacrylamide gel electrophoresis.

A workshop entitled Cellular Protein Databases from Two-Dimensional Gel Electrophoresis was held in Atlanta, Georgia, 28 February-1 March 1987. Its purpose was to assess the status of two-dimensional gel electrophoresis of proteins as a research methodology in biological systems, particularly in the generation of cellular protein databases. The workshop participants summarized current studies on a variety of biological systems, both prokaryotic and eukaryotic. Analysis of the progress being made led to the conclusion that electrophoretic techniques, supported by automatic scanning of gel images and computer-assisted processing, analysis and matching of gel images, are now capable of generating databases of great potential value. Factors affecting the reproducibility of protein spot patterns on gels were identified, and the extent to which gel pattern variability causes difficulties in communicating results and in integrating information from different laboratories was assessed. Measures were suggested to help overcome obstacles to the generation of comprehensive cellular protein databases from the electrophoretic resolution of total cellular proteins.

Animals↗

Genomically linked cellular protein databases derived from two-dimensional polyacrylamide gel electrophoresis.

In its most useful form a cellular protein database should be genomically based, because it is the genome which determines both the total number of proteins a cell can make and the particular ones that will be made under any given condition. Such a database should trace each protein back to its structural gene, and should account for every structural gene of a cell. Recent advances in molecular biology greatly facilitate the construction of such gene-protein databases. The mapping of genes of unidentified proteins resolved from total cell extracts on two-dimensional gels can now be accomplished by largely biochemical methods, without the necessity of isolating mutants or performing genetic crosses. Other techniques permit one to search gels for the product of any newly discovered gene (or open reading frame) suspected of encoding a protein. Consequently, gene-protein indices can be built independently and simultaneously from either direction--deducing the genetic map from the protein pattern, or finding the protein pattern from information encoded in the genome. A database of this sort is being constructed for the bacterium, Escherichia coli. Given the current pace of DNA nucleotide sequencing, the development of total gene-protein indices for a variety of cells can be anticipated in the near future.

Amino Acids↗

A two-dimensional gel protein database of noncultured total normal human epidermal keratinocytes: identification of proteins strongly up-regulated in psoriatic epidermis.

A two-dimensional (2-D) gel database of proteins from noncultured total normal human epidermal keratinocytes has been established. A total of 1449 [35S]methionine labelled proteins (1112 isoelectric focusing, 337 nonequilibrium pH gradient electrophoresis) were resolved and recorded using computer assisted (PDQ-SCAN and PDQUEST software) 2-D gel electrophoresis. By matching the protein patterns of total keratinocytes and transformed human amnion cells (master database; Celis et al., Leukemia 1988, 2, 561-602) as well as by 2-D immunoblotting and microsequencing of keratinocyte proteins, it was possible to identify 72 polypeptides in the keratinocyte database. The database also includes data on polypeptides that are synthesized at a higher level by keratinocytes enriched in basal cells, and on six secreted proteins which are produced, albeit at a reduced rate, by normal keratinocytes and that are strongly up-regulated in psoriatic epidermis (Celis et al., FEBS Letters, in press).

Electrophoresis, Gel, Two-Dimensional↗

Construction and analysis of a database representing a neural map.

We describe the development and analysis of a quantitative database representing the global structural and functional organization of an entire sensory map. The database was derived from measurements of anatomical characteristics of a statistical sample of typical mechanosensory afferents in the cricket cercal sensory system. Anatomical characteristics of the neurons were measured quantitatively in three dimensions using a computer reconstruction system. The reconstructions of all neurons were aligned and scaled to a common standard set of dimensions, according to a highly reproducible set of intrinsic fiducial marks. The database therefore preserves accurate information about spatial relationships between the neurons within the ensemble. Algorithms were implemented to allow the integration of electrophysiological data about the stimulus/response characteristics of the reconstructed neurons into the database. The algorithms essentially map a physiological function onto a "field" representing the continuous distribution of synaptic terminals throughout the neural structure. Subsequent analysis allowed quantitative predictions of several important functional characteristics of the sensory map that emerge from its global organization. First, quantitative and testable predictions were made about ensemble response patterns within the map. The predicted patterns are presented as graphical images, similar to images that might be observed with activity-dependent dyes in the real neural system. Second, the synaptic innervation patterns from the sensory afferent map onto the dendrites of a postsynaptic target interneuron were predicted by calculating the overlap between the interneuron's dendrites with the afferent map. By doing so, several aspects of the stimulus/response properties of the interneuron were accurately predicted.

Animals↗

Design considerations for small, special-system developmental databases

Small developmental databases, devoted to special systems and run by a normal research laboratory, differ qualitatively as well as quantitatively from larger databases described elsewhere in this volume. Resource limits set different optimum balances between computing elegance and ease of construction, while flexibility of design and semantics often assume greatest importance because specialist research databases are most useful when understanding of the biology is partial and still evolving. In this article we discuss the main considerations when building a small database-scope, access, data input, storage, searching and semantics-and illustrate them with examples of real working databases.Copyright 1997 Academic Press Limited Copyright 1997Academic Press Limited

Journal Article↗

The mouse gene expression database GXD

The gene expression database (GXD) is being developed to store and integrate expression information for mouse development. GXD addresses many issues that apply to gene expression databases in general, and its data structures and supporting software tools are generalized in design and thus readily adaptable to other life stages and species. Integration of GXD with the mouse genome database (MGD) and interconnections with other relevant databases will place the gene expression data into the larger biological and analytical context. Here, we describe the design and implementation of GXD and illustrate, in particular, the gene expression annotator, an electronic system for submitting expression data to the database.Copyright 1997 Academic Press Limited Copyright 1997Academic Press Limited

Journal Article↗

Enhancing medical database security.

A methodology for the enhancement of database security in a hospital environment is presented in this paper which is based on both the discretionary and the mandatory database security policies. In this way the advantages of both approaches are combined to enhance medical database security. An appropriate classification of the different types of users according to their different needs and roles and a User Role Definition Hierarchy has been used. The experience obtained from the experimental implementation of the proposed methodology in a major general hospital is briefly discussed. The implementation has shown that the combined discretionary and mandatory security enforcement effectively limits the unauthorized access to the medical database, without severely restricting the capabilities of the system.

Computer Security↗

The Human Papillomavirus Database.

Papillomaviruses are responsible for a variety of diseases in humans and animals, ranging from harmless skin warts to lethal cancers. They also make up one of the most genetically diversified families of viruses known, and could represent a model system of DNA-virus evolution. A specialized genetic sequences database, The Human Papillomavirus Database and Analysis Project, was recently established in an effort to provide database services that are specific to papillomaviruses to the research community and to perform a variety of sequence-based analyses. This review is intended to present the scope of the information currently contained in the database and to outline some of the analyses that have been performed on the genetic sequences. These analyses will address issues including phylogenetic relationships, recombination events, selective pressures on different genes and the possibility of cross-species transmission in the case of the papillomaviruses. Copyright 1995 S. Karger AG, Basel

Journal Article↗

Winter Flounder Expressed Sequence Tags: Establishment of an EST Database and Identification of Novel Fish Genes.

: An EST database of more than 900 sequences has been constructed from complementary DNAs from six different tissues (stomach, intestine, pyloric cecum, liver, spleen, and ovary) of the winter flounder Pleuronectes americanus. Template preparation and automated sequencing were optimized to generate high-quality information in an economic fashion. Using computer scripts developed in our laboratory, the sequences were automatically compared with sequences in the databases via a Web-browser interface, and significant returns were recorded and organized on user-friendly HTML pages. Half (453) of the ESTs had significant matches to database sequences of known function, 33 matched ESTs from other organisms, 34 matched ribosomal RNAs, and 24 matched hypothetical open reading frames of unknown function. Forty-one percent (374) of the ESTs had no matches to sequences in the database and presumably represent previously unidentified cDNAs. Several sequences are the first isolated from teleost fish, and should be of interest for gene mapping and studies of developmental biology.

Journal Article↗