Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

GenBank.

The GenBank sequence database incorporates publicly available DNA sequences of more than 105 000 different organisms, primarily through direct submission of sequence data from individual laboratories and large-scale sequencing projects. Most submissions are made using the BankIt (web) or Sequin programs and accession numbers are assigned by GenBank staff upon receipt. Data exchange with the EMBL Data Library and the DNA Data Bank of Japan helps ensure comprehensive worldwide coverage. GenBank data is accessible through NCBI's integrated retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome, mapping, protein structure and domain information, and the biomedical literature via PubMed. Sequence similarity searching is provided by the BLAST family of programs. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. NCBI also offers a wide range of World Wide Web retrieval and analysis services based on GenBank data. The GenBank database and related resources are freely accessible via the NCBI home page at http://www.ncbi.nlm.nih.gov.

Animals↗

The information needs and behaviour of clinical researchers: a user-needs analysis.

AIMS: As part of the strategy to set up a new information service, including a physical Resource Centre, the analysis of information needs of clinical research professionals involved with clinical research and development in the UK and Europe was required. It also aimed to identify differences in requirements between the various roles of professionals and establish what information resources are currently used. METHODS: A user-needs survey online of the members of The Institute. Group discussions with specialist subcommittees of members. RESULTS: Two hundred and ninety members responded to the online survey of 20 questions. This makes it a response rate of 7.9%. Members expressed a lack of information in their particular professional area, and lack the skills to retrieve and appraise information. DISCUSSION: The results of the survey are discussed in more detail, giving indications of what the information service should collect, what types of materials should be provided to members and what services should be on offer. RECOMMENDATION: These were developed from the results of the needs analysis and submitted to management for approval. Issues of concern, such as financial constraint and staff constraints are also discussed. CONCLUSIONS: There is an opportunity to build a unique collection of clinical research material, which will promote The Institute not only to members, but also to the wider health sector. Members stated that the most physical medical libraries don't provide what they need, but the main finding through the survey and discussions is that it's pointless to set up 'yet another medical library'.

Academies and Institutes↗

The use of structure information to increase alignment accuracy does not aid homologue detection with profile HMMs.

MOTIVATION: The best quality multiple sequence alignments are generally considered to derive from structural superposition. However, no previous work has studied the relative performance of profile hidden Markov models (HMMs) derived from such alignments. Therefore several alignment methods have been used to generate multiple sequence alignments from 348 structurally aligned families in the HOMSTRAD database. The performance of profile HMMs derived from the structural and sequence-based alignments has been assessed for homologue detection. RESULTS: The best alignment methods studied here correctly align nearly 80% of residues with respect to structure alignments. Alignment quality and model sensitivity are found to be dependent on average number, length, and identity of sequences in the alignment. The striking conclusion is that, although structural data may improve the quality of multiple sequence alignments, this does not add to the ability of the derived profile HMMs to find sequence homologues. SUPPLEMENTARY INFORMATION: A list of HOMSTRAD families used in this study and the corresponding Pfam families is available at http://www.sanger.ac.uk/Users/sgj/alignments/map.html CONTACT: sgj@sanger.ac.uk

Amino Acid Sequence↗

L/D Protein Ligand Database (PLD): additional understanding of the nature and specificity of protein-ligand complexes.

SUMMARY: The Protein Ligand Database (PLD) is a publicly available web-based database that aims to provide further understanding of protein-ligand interactions. The PLD contains biomolecular data including calculated binding energies, Tanimoto ligand similarity scores and protein percentage sequence similarities. The database has potential for application as a tool in molecular design. AVAILABILITY: http://www-mitchell.ch.cam.ac.uk/pld/

Amino Acid Sequence↗

An automated PACS image acquisition and recovery scheme for image integrity based on the DICOM standard.

The data quality and completeness of acquired images, which we refer to as integrity, is considered as the most important requirement in the image acquisition design of the Picture Archiving and Communication System (PACS). The Digital Imaging and Communications in Medicine (DICOM) standard significantly simplifies the task of acquiring radiological images from a DICOM compliant imaging system into the PACS. However, human interaction with the imaging system by changing the DICOM communication settings can result in missing images during the PACS image acquisition. A scheme based on the DICOM Query and Retrieve (Q/R) service class was developed to automatically identify and recover missing images. In addition, grouping sequential scanned images such as a CT and MR image series is another potential process that can miss images because of no indication of the end of series. Two methods are presented for determining the end of series and the pros and cons of each method are discussed in detail. Two experiments in a real clinical environment were conducted; one with and one without the Q/R implementation. The statistical results indicate two highlights from this work. First, the Q/R scheme faithfully recovered all missing images caused by human interaction with the DICOM compliant imaging system. Second, there was no single image slice missed when grouping slices into a series using the presented grouping algorithm in the two experimental periods.

Algorithms↗

Parasite genome databases and web-based resources.

In the last decade, high-throughput genome sequencing and complementary techniques such as microarray and proteomics have generated, and will continue to generate, ever-increasing amounts of data. These technologies of gene discovery, expression, and functional analysis have been applied to a vast array of organisms, including parasites. In most instances, the data are freely available via the Internet, and researchers are becoming increasingly reliant on up-to-date, centralized data repositories to complement wet bench science. This chapter presents an overview of resources relevant to researchers with an interest in para-site genomics and biology. After briefly touching on some of the publicly available nucleotide and protein sequence as well as domain databases, the focus turns to parasite genome projects and associated Web-based resources. A list of parasite sequencing projects current at the time of writing, including relevant Web site addresses, is provided. The available resources range from network sites and project pages at sequencing institutes to databases that integrate and curate sequence data and associated annotation with diverse biological datasets. Particular attention is given to three databases, GeneDB (http://www.genedb.org/), PlasmoDB (http://plasmodb. org/), and tigr db, detailing the scope of each database and the tools available for data querying and retrieval.

Animals↗

Decoding of auditory cortex signals with a LAMSTAR neural network.

OBJECTIVES: Each neuron has a specific set of stimuli, which it preferentially responds to (the receptive field of the neuron). For implantable cortical prosthetic devices specific points of the cortex (or groups of neurons) have to be stimulated to create perceptions of sensory stimulus with specific attributes (such as frequency, temporal characteristics, etc). Such applications would need real time decoding of signals. Previously mathematical techniques, such as computing the receptive field (using electrophysiology data) and artificial neural networks (Kohonen network or SOM and back propagation network) have been used to decode neural signals. METHODS: A Large Adaptive Memory Storage and Retrieval (LAMSTAR) neural-network-based decoder was designed to decode responses recorded from neurons in the auditory cortex. It was designed to identify the frequency of the tonal stimuli that elicited a particular discharge rate pattern recorded on two channels of a tungsten wire electrode array. RESULTS: The network functioned efficiently as a decoder with 100% accuracy for the small sample of stimulus-response data used. DISCUSSION: The results show that the network is effective in studying the functional organization of the auditory cortex and other sensory systems. Depending on the input sub-word, information about the kind of stimuli that activates particular parts of the sensory cortex can be studied.

Acoustic Stimulation↗

Self-contained patient data in ORCA to cope with an evolving vocabulary.

Because of the benefits of standardization in healthcare data for research, decision support, and quality assessment, much research effort focuses on collection of structured patient data. Many strategies to obtain such data are based on controlled vocabularies to guide data entry in a far more flexible way than a fixed-form approach. Medical controlled vocabularies evolve, but change is difficult to reconcile with standardization. Retrieval of data, collected with different versions of vocabularies, is not straightforward and has consequences for patient care and research. There are several strategies to cope with these problems: keep each version, keep a record of changes, or conversion of previously collected data. Each of these strategies has pros and cons regarding storage consumption, performance during patient care, and research. The approach in ORCA (Open Record for Care) is based on self-contained patient data and combines the strengths of these strategies.

Humans↗

An HMM model for coiled-coil domains and a comparison with PSSM-based predictions.

MOTIVATION: Large-scale sequence data require methods for the automated annotation of protein domains. Many of the predictive methods are based either on a Position Specific Scoring Matrix (PSSM) of fixed length or on a window-less Hidden Markov Model (HMM). The performance of the two approaches is tested for Coiled-Coil Domains (CCDs). The prediction of CCDs is used frequently, and its optimization seems worthwhile. RESULTS: We have conceived MARCOIL, an HMM for the recognition of proteins with a CCD on a genomic scale. A cross-validated study suggests that MARCOIL improves predictions compared to the traditional PSSM algorithm, especially for some protein families and for short CCDs. The study was designed to reveal differences inherent in the two methods. Potential confounding factors such as differences in the dimension of parameter space and in the parameter values were avoided by using the same amino acid propensities and by keeping the transition probabilities of the HMM constant during cross-validation. AVAILABILTY: The prediction program and the databases are available at http://www.wehi.edu.au/bioweb/Mauro/Marcoil

Algorithms↗

Transparent image access in a distributed picture archiving and communications system: the Master Database broker.

A distributed design is the most cost-effective system for small-to medium-scale picture archiving and communications systems (PACS) implementations. However, the design presents an interesting challenge to developers and implementers: to make stored image data, distributed throughout the PACS network, appear to be centralized with a single access point for users. A key component for the distributed system is a central or master database, containing all the studies that have been scanned into the PACS. Each study includes a list of one or more locations for that particular dataset so that applications can easily find it. Non-Digital Imaging and Communications in Medicine (DICOM) clients, such as our worldwide web (WWW)-based PACS browser, query the master database directly to find the images, then jump to the most appropriate location via a distributed web-based viewing system. The Master Database Broker provides DICOM clients with the same functionality by translating DICOM queries to master database searches and distributing retrieval requests transparently to the appropriate source. The Broker also acts as a storage service class provider, allowing users to store selected image subsets and reformatted images with the original study, without having to know on which server the original data are stored.

CD-ROM↗

Evaluation of the quality of information retrieval of clinical findings from a computerized patient database using a semantic terminological model.

OBJECTIVES: To measure the strength of agreement between the concepts and records retrieved from a computerized patient database, in response to physician-derived questions, using a semantic terminological model for clinical findings with those concepts and records excerpted clinically by manual identification. The performance of the semantic terminological model is also compared with the more established retrieval methods of free-text search, ICD-10, and hierarchic retrieval. DESIGN: A clinical database (Diabeta) of 106,000 patient problem record entries containing 2,625 unique concepts in an clinical academic department was used to compare semantic, free-text, ICD-10, and hierarchic data retrieval against a gold standard in response to a battery of 47 clinical questions. MEASUREMENTS: The performance of concept and record retrieval expressed as mean detection rate, positive predictive value, Yates corrected and Mantel-Haenszel chi-squared values, and Cohen kappa value, with significance estimated using the Mann-Whitney test. RESULTS: The semantic terminological model used to retrieve clinically useful concepts from a patient database performed well and better than other methods, with a mean detection rate of 0.86, a positive predictive value of 0.96, a Yates corrected chi-squared value of 1,537, a Mantel-Haenszel chi-squared value of 19,302, and a Cohen kappa of 0.88. Results for record retrieval were even better, with a mean record detection rate of 0.94, a positive predictive value of 0.99, a Yates corrected chi-squared value of 94, 774, a Mantel-Haenszel chi-squared value of 1,550,356, and a Cohen kappa value of 0.94. The mean detection rate, Yates corrected chi-squared value, and Cohen kappa value for semantic retrieval were significantly better than for the other methods. CONCLUSION: The use of a semantic terminological model in this test scenario provides an effective framework for representing clinical finding concepts and their relationships. Although currently incomplete, the model supports improved information retrieval from a patient database in response to clinically relevant questions, when compared with alternative methods of analysis.

Data Interpretation, Statistical↗

Applications of the MEGADATS database system in medical genetics.

The MEGADATS relational database system has many useful applications in the field of medical genetics. Some of these applications include storage, retrieval, and display of pedigree information; retrieval of sets of individuals, sibships, or families who meet given criteria; storage of necessary information for mailing lists, clinic data, etc; and combination of pedigree information and genotype information into the format needed for linkage analysis packages.

Genetics, Medical↗

Lessons learned from data logging in a multicenter clinical trial using a late-generation implantable cardioverter-defibrillator. The Guardian ATP 4210 Multicenter Investigators Group.

OBJECTIVES: This study examined patterns of implantable cardioverter-defibrillator use as documented by data logging. BACKGROUND: Implantable cardioverter-defibrillators are accepted therapy for malignant ventricular tachyarrhythmias; however, relatively little is known about their patterns of use. Incorporation of data-storage capacities into these devices provides insight into long-term defibrillator function. METHODS: Stored data-logging information was retrieved from 401 implanted cardioverter-defibrillators in 393 patients over an average of 303 days of follow-up. RESULTS: A total of 91,443 detections were recorded in 299 patients. One hundred-six patients (26%) had detections due to supraventricular tachycardias, electrical noise or other causes, resulting in inappropriate therapy delivery to 92 patients (23%). Two hundred eighty-one patients recorded 66,276 episodes of ventricular tachycardia or ventricular fibrillation. Of these, 74.4% episodes terminated spontaneously without any delivered therapy, 22.1% terminated after antitachycardia pacing, and 1.7% terminated after shock therapy. Antitachycardia pacing was activated without formal testing in 47% of all patients receiving this therapy and was successful in 96% of all episodes receiving this therapy. Acceleration of tachycardia to shock therapy occurred in 1.3% of all episodes and in 30.5% of patients receiving antitachycardia pacing. Thirty-four patients (8.7%) died during follow-up. Mortality was associated with patient age, heart failure functional class at implantation and frequency of shocks received during follow-up (all p < or = 0.05). CONCLUSIONS: Most ventricular tachyarrhythmia detections by this noncommitted implantable cardioverter-defibrillator resolve spontaneously, whereas the majority receiving therapy can be treated with antitachycardia pacing. Mortality after implantable cardioverter-defibrillator implantation is associated with age, heart failure class and frequency of shocks received during follow-up. Data-logging capabilities provide valuable insights into the patterns of defibrillator use.

Algorithms↗

Tower of Hanoi: evidence for the cost of goal retrieval.

Past research on the Tower of Hanoi problem has provided clear evidence for the importance of goal-subgoal structures in problem solving. However, the nature of the traditional Tower of Hanoi problem makes it impossible to determine whether there is any special cost associated with storing or retrieving goals. A variation of the Tower of Hanoi problem is described that allows one to determine separately if there is an effect of how long a goal has to be retained on storage time or how long ago it was formed on retrieval time. This paradigm provides evidence for an effect of retention interval on retrieval time and not on storage time. An ACT-R (Adaptive Control of Thought-Rational) simulation of these data is described, which treats goal memory as no different from other memories.

Adult↗

PathMaster: content-based cell image retrieval using automated feature extraction.

OBJECTIVE: Currently, when cytopathology images are archived, they are typically stored with a limited text-based description of their content. Such a description inherently fails to quantify the properties of an image and refers to an extremely small fraction of its information content. This paper describes a method for automatically indexing images of individual cells and their associated diagnoses by computationally derived cell descriptors. This methodology may serve to better index data contained in digital image databases, thereby enabling cytologists and pathologists to cross-reference cells of unknown etiology or nature. DESIGN: The indexing method, implemented in a program called PathMaster, uses a series of computer-based feature extraction routines. Descriptors of individual cell characteristics generated by these routines are employed as indexes of cell morphology, texture, color, and spatial orientation. MEASUREMENTS: The indexing fidelity of the program was tested after populating its database with images of 152 lymphocytes/lymphoma cells captured from lymph node touch preparations stained with hematoxylin and eosin. Images of "unknown" lymphoid cells, previously unprocessed, were then submitted for feature extraction and diagnostic cross-referencing analysis. RESULTS: PathMaster listed the correct diagnosis as its first differential in 94 percent of recognition trials. In the remaining 6 percent of trials, PathMaster listed the correct diagnosis within the first three "differentials." CONCLUSION: PathMaster is a pilot cell image indexing program/search engine that creates an indexed reference of images. Use of such a reference may provide assistance in the diagnostic/prognostic process by furnishing a prioritized list of possible identifications for a cell of uncertain etiology.

Abstracting and Indexing↗

The European Bioinformatics Institute's data resources.

As the amount of biological data grows, so does the need for biologists to store and access this information in central repositories in a free and unambiguous manner. The European Bioinformatics Institute (EBI) hosts six core databases, which store information on DNA sequences (EMBL-Bank), protein sequences (SWISS-PROT and TrEMBL), protein structure (MSD), whole genomes (Ensembl) and gene expression (ArrayExpress). But just as a cell would be useless if it couldn't transcribe DNA or translate RNA, our resources would be compromised if each existed in isolation. We have therefore developed a range of tools that not only facilitate the deposition and retrieval of biological information, but also allow users to carry out searches that reflect the interconnectedness of biological information. The EBI's databases and tools are all available on our website at www.ebi.ac.uk.

Animals↗