Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

EyeSite: a semi-automated database of protein families in the eye.

The EyeSite is a web-based database of protein families for proteins that function in the eye and their homologous sequences. The resource clusters proteins at different levels of homology in order to facilitate functional annotation of sequences and modelling of proteins from structural homologues. Eye proteins are organized into the tissue types in which they function and are clustered into homologous families using a novel protocol employing the TribeMCL algorithm. Homologous families are further subdivided into sequence clusters for which multiple sequence alignments are generated. Structural annotations from the CATH domain database are provided for nearly 90% of the sequences, and protein family annotations from the Pfam database for approximately 86%. Homology models have also been generated where appropriate. The EyeSite is stored in a relational database and is extensively linked to other online bioinformatics resources to help relate allelic variants, annotations and clinical details to the derived data in the database. The EyeSite is available for online search, sequence information and model retrieval at http://eyesite.cryst.bbk.ac.uk/.

Amino Acid Sequence↗

Organizing, storing, and analysing qualitative research information in a computer database.

OBJECTIVE: To describe a method to handle data in qualitative research. DESIGN: Description of the data handling process in a personal computer database. SETTING: The method was used in a Danish qualitative study with more than 100 interviews on quality of care. OUTCOME MEASURES: The author's own description and assessment of the system. RESULTS: Storing information in a database makes it possible to store and retrieve information from qualitative research divided in themes and subthemes. The themes and subthemes are based on an interview guide and the succeeding development of theory and new themes. The comprehensive sorting systems in a database enable the researcher to keep an overview of his data while analysing and reporting. CONCLUSION: Storing, analysing, and retrieving data from qualitative research demands careful planning. A computer database may be a good help in this process if the researcher will not use more complex and comprehensive software for his analysis.

Denmark↗

Integrating protein structures and precomputed genealogies in the Magnum database: examples with cellular retinoid binding proteins.

BACKGROUND: When accurate models for the divergent evolution of protein sequences are integrated with complementary biological information, such as folded protein structures, analyses of the combined data often lead to new hypotheses about molecular physiology. This represents an excellent example of how bioinformatics can be used to guide experimental research. However, progress in this direction has been slowed by the lack of a publicly available resource suitable for general use. RESULTS: The precomputed Magnum database offers a solution to this problem for ca. 1,800 full-length protein families with at least one crystal structure. The Magnum deliverables include 1) multiple sequence alignments, 2) mapping of alignment sites to crystal structure sites, 3) phylogenetic trees, 4) inferred ancestral sequences at internal tree nodes, and 5) amino acid replacements along tree branches. Comprehensive evaluations revealed that the automated procedures used to construct Magnum produced accurate models of how proteins divergently evolve, or genealogies, and correctly integrated these with the structural data. To demonstrate Magnum's capabilities, we asked for amino acid replacements requiring three nucleotide substitutions, located at internal protein structure sites, and occurring on short phylogenetic tree branches. In the cellular retinoid binding protein family a site that potentially modulates ligand binding affinity was discovered. Recruitment of cellular retinol binding protein to function as a lens crystallin in the diurnal gecko afforded another opportunity to showcase the predictive value of a browsable database containing branch replacement patterns integrated with protein structures. CONCLUSION: We integrated two areas of protein science, evolution and structure, on a large scale and created a precomputed database, known as Magnum, which is the first freely available resource of its kind. Magnum provides evolutionary and structural bioinformatics resources that are useful for identifying experimentally testable hypotheses about the molecular basis of protein behaviors and functions, as illustrated with the examples from the cellular retinoid binding proteins.

Amino Acid Sequence↗

A multinomial modeling analysis of the recognition-failure paradigm.

The recognition-failure paradigm has received much theoretical consideration, especially the Tulving-Wiseman function and its exceptions. We show that the Tulving-Wiseman function does a poor job of accounting for the data, both when its fit is measured with a model-based, goodness-of-fit statistic and when a logically equivalent reformulation of the function is compared with data. We then present a simple multinomial model based on retrieval-independence theory that is capable of measuring storage and retrieval processes in recognition failure. The model is used to conduct a meta-analysis of the recognition-failure paradigm, and shows that violations of the Tulving-Wiseman function occur under conditions in which weak storage is coupled with strong retrieval. In addition, if storage and retrieval are assumed to be positively correlated across conditions, the model produces a theoretically motivated, alternative equation to the Tulving-Wiseman function that provides a virtually identical fit to the data.

Humans↗

Clinical application of a magneto-optical disk image filing system/image save and carry (ISAC) system.

We propose the utilization of portable magneto optical disks for image filing. The main problems in PACS are the need for a high-speed local area network (LAN) and a large mass storage device. An image filing system--image save and carry (ISAC)--is one solution for these problems in present PACS and requires minimal additional hardware and cost for the installation. Whenever a patient is examined, the clerk carries the medical record and the ISAC magneto-optical disk for recording image data, and after inspection records the image data into the magneto-optical disk. We investigated the number of image retrievals done for inpatients and outpatients in 1988 at our radiation therapy department. The data storage requirements were on the average 18.5 MB for outpatients and 173.9 MB for inpatients. An ISAC display console needs also an easy-to-use man/machine interface for specifying images and image display characteristics in order to realize the ISAC system.

Computer Systems↗

FLASH: a fast look-up algorithm for string homology.

A key issue in managing today's large amounts of genetic data is the availability of efficient, accurate, and selective techniques for detecting homologies (similarities) between newly discovered and already stored sequences. A common characteristic of today's most advanced algorithms, such as FASTA, BLAST, and BLAZE is the need to scan the contents of the entire database, in order to find one or more matches. This design decision results in either excessively long search times or, as is the case of BLAST, in a sharp trade-off between the achieved accuracy and the required amount of computation. The homology detection algorithm presented in this paper, on the other hand, is based on a probabilistic indexing framework. The algorithm requires minimal access to the database in order to determine matches. This minimal requirement is achieved by using the sequences of interest to generate a highly redundant number of very descriptive tuples; these tuples are subsequently used as indices in a table look-up paradigm. In addition to the description of the algorithm, theoretical and experimental results on the sensitivity and accuracy of the suggested approach are provided. The storage and computational requirements are described and the probability of correct matches and false alarms is derived. Sensitivity and accuracy are shown to be close to those of dynamic programming techniques. A prototype system has been implemented using the described ideas. It contains the full Swiss-Prot database rel 25 (10 MR) and the genome of E. Coli (2 MR). The system is currently being expanded to include the complete Genbank database.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

Image Engine: an object-oriented multimedia database for storing, retrieving and sharing medical images and text.

This paper describes Image Engine, an object-oriented, microcomputer-based, multimedia database designed to facilitate the storage and retrieval of digitized biomedical still images, video, and text using inexpensive desktop computers. The current prototype runs on Apple Macintosh computers and allows network database access via peer to peer file sharing protocols. Image Engine supports both free text and controlled vocabulary indexing of multimedia objects. The latter is implemented using the TView thesaurus model developed by the author. The current prototype of Image Engine uses the National Library of Medicine's Medical Subject Headings (MeSH) vocabulary (with UMLS Meta-1 extensions) as its indexing thesaurus.

Abstracting and Indexing↗

Iterated sequence databank search methods.

Iterated sequence databank search methods were assessed from the viewpoint of someone with the sequence of a novel gene product wishing to find distant relatives to their protein and, with the specific searches against the PDB, also hoping to find a relative of known structure. We examined three methods in detail, spanning a range from simple pattern-matching to sophisticated weighted profiles. Rather than apply these methods 'blindly' (with default parameters) to a large number of test queries, we have concentrated on the globins, so allowing a more detailed investigation of each method on different data subsets with different parameter settings. Despite their widespread use, regular-expression matching proved to be very limited-seldom extending beyond the sub-family from which the pattern was derived. To attain any generality, the patterns had to be 'stripped-down' to include only the most highly conserved parts. The QUEST program avoided these problems by introducing a more flexible (weighted) matching. On the PDB sequences this was highly effective, missing only a few globins with probes based on each sub-family or even a single representative from each sub-family. In addition, very few false-positives were encountered, and those that did match, often only did so for a few cycles before being lost again. On the larger sequence collection, however, QUEST encountered problems with maintaining (or achieving) the alignment of the full globin family. psi-BLAST also recognised almost all the globins when matching against the PDB sequences, typically, missing three or four of the most distantly related sequences while picking-up a few false-positives. In contrast to QUEST, psi-BLAST performed very well on the larger databank, getting almost a full collection of globins although still retaining the same proportion of false-positives. SAM applied to the PDB sequences performed reasonably well with the myoglobin and hemoglobin families as probes, missing, typically several of the more difficult proteins but performed poorly with the leghemoglobin probe. Only with the full family range as a probe did it produce results comparable to psi-BLAST and QUEST. With the larger databank, SAM produced a good result but, again, this was only achieved using the full range of sequence variation with the default regulariser and use of Dirichlet mixtures completely failed in this situation.

Amino Acid Sequence↗

Specific interoperability problems of security infrastructure services.

Communication and co-operation in healthcare and welfare require a well-defined set of security services based on a standards-based interoperable security infrastructure and provided by a Trusted Third Party. Generally, the services describe status and relation of communicating principals, corresponding keys and attributes, and the access rights to both applications and data. Legal, social, behavioral and ethical requirements demand securely stored patient information and well-established access tools and tokens. Electronic signatures as means for securing integrity of messages and files, certified time stamps and time signatures are important for accessing and storing data in Electronic Health Record Systems. The key for all these services is a secure and reliable procedure for authentication (identification and verification). While mentioning technical problems (e.g. lifetime of the storage devices, migration of retrieval and presentation software), this paper aims at identifying harmonization and interoperability requirements of securing data items, files, messages, sets of archived items or documents, and life-long Electronic Health Records based on a secure certificate-based identification. It's commonly known that just relying on existing and emerging security standards does not necessarily guarantee interoperability of different security infrastructure approaches. So certificate separation can be a key to modern interoperable security infrastructure services.

Authorship↗

The validity of the EASE expert system for inhalation exposures.

Estimation and Assessment of Substance Exposure (EASE) is a computerized expert system developed by the UK Health and Safety Executive to facilitate exposure assessments in the absence of exposure measurements. The system uses a number of rules to predict a range of likely exposures or an 'end-point' for a given work situation. The purpose of this study was to identify a number of inhalation exposure measurements covering a wide range of end-points in the EASE system to compare with the predicted exposures. Occupational exposure data sets were identified from previous research projects or from consultancy work. Available information for each set of measurements was retrieved from archive storage and reviewed to ensure that it was adequate to enable EASE (version 2) predictions to be obtained. Exposure measurements and other relevant contextual data were abstracted and entered into a computer spreadsheet. EASE predictions were then obtained for each task or job and entered into the spreadsheet. In addition, we generated a random exposure range for each data set for comparison with the EASE predictions. Finally, we produced exposure assessments for a subset of the data using a structured subjective assessment method. We were able to identify approximately 4000 inhalation exposure measurements covering 52 different scenarios and 28 EASE end-points. The data included measurements of solvent vapours, non-fibrous dusts and fibres. In 62% of the end-points the EASE predictions were generally greater than the exposure measurements and in 30% of the end-points the EASE estimates were comparable with the measurements. The random allocation of exposure ranges was, as expected, less reliable than EASE, although there were still about one-third of the cases where the randomly generated exposure ranges generally agreed with the measurements. The structured subjective assessments undertaken by a human expert produced exposure estimates in better agreement with the measurements with about two-thirds of the end-points derived from these assessments in good agreement with the data. We argue that the inhalation exposure estimates from EASE could be improved by incorporating some of the parameters included in the structured subjective assessment methodology.

Air Pollutants, Occupational↗

Gleaning non-trivial structural, functional and evolutionary information about proteins by iterative database searches.

Using a number of diverse protein families as test cases, we investigate the ability of the recently developed iterative sequence database search method, PSI-BLAST, to identify subtle relationships between proteins that originally have been deemed detectable only at the level of structure-structure comparison. We show that PSI-BLAST can detect many, though not all, of such relationships, but the success critically depends on the optimal choice of the query sequence used to initiate the search. Generally, there is a correlation between the diversity of the sequences detected in the first pass of database screening and the ability of a given query to detect subtle relationships in subsequent iterations. Accordingly, a thorough analysis of protein superfamilies at the sequence level is necessary in order to maximize the chances of gleaning non-trivial structural and functional inferences, as opposed to a single search, initiated, for example, with the sequence of a protein whose structure is available. This strategy is illustrated by several findings, each of which involves an unexpected structural prediction: (i) a number of previously undetected proteins with the HSP70-actin fold are identified, including a highly conserved and nearly ubiquitous family of metal-dependent proteases (typified by bacterial O-sialoglycoprotease) that represent an adaptation of this fold to a new type of enzymatic activity; (ii) we show that, contrary to the previous conclusions, ATP-dependent and NAD-dependent DNA ligases are confidently predicted to possess the same fold; (iii) the C-terminal domain of 3-phosphoglycerate dehydrogenase, which binds serine and is involved in allosteric regulation of the enzyme activity, is shown to typify a new superfamily of ligand-binding, regulatory domains found primarily in enzymes and regulators of amino acid and purine metabolism; (iv) the immunoglobulin-like DNA-binding domain previously identified in the structures of transcription factors NFkappaB and NFAT is shown to be a member of a distinct superfamily of intracellular and extracellular domains with the immunoglobulin fold; and (v) the Rag-2 subunit of the V-D-J recombinase is shown to contain a kelch-type beta-propeller domain which rules out its evolutionary relationship with bacterial transposases.

Actins↗

CAVEAT: a program to facilitate the design of organic molecules.

A frequently encountered problem in the design of enzyme inhibitors and other biologically active molecules is the identification of molecular frameworks to serve as templates or linking units that can position functional groups in specific relative orientations. The program CAVEAT was designed to address this problem by searching 3D databases for such molecular fragments. Key innovations introduced in CAVEAT are a focus on relationships between bonds and the provision of automated methods to identify and classify structural frameworks. Performance has been a particular concern in formulating CAVEAT, since it is intended to be used in an interactive manner. The focus in this report is the design and implementation of the principal algorithms and the performance achieved.

Amino Acid Sequence↗

OoClamp: an IBM-compatible software system for electrophysiologic receptor studies in Xenopus oocytes.

A software system for IBM-compatible microcomputers running MS-DOS or Microsoft Windows 3.1 is described which facilitates the acquisition, analysis and storage of data from electrophysiologic studies of receptors expressed in Xenopus laevis oocytes. The system is designed to provide standardization of test conditions, automation of all routine functions, rapid, on-line analysis of data, and self-documentation and compact storage of data files. All system settings are optimized for use with the Xenopus expression system, but can be adapted to other large cells. An example application, expression of muscarinic acetylcholine receptors in Xenopus oocytes, is described.

Animals↗