Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Whole-proteome interaction mining.

MOTIVATION: A major post-genomic scientific and technological pursuit is to describe the functions performed by the proteins encoded by the genome. One strategy is to first identify the protein-protein interactions in a proteome, then determine pathways and overall structure relating these interactions, and finally to statistically infer functional roles of individual proteins. Although huge amounts of genomic data are at hand, current experimental protein interaction assays must overcome technical problems to scale-up for high-throughput analysis. In the meantime, bioinformatics approaches may help bridge the information gap required for inference of protein function. In this paper, a previously described data mining approach to prediction of protein-protein interactions (Bock and Gough, 2001, Bioinformatics, 17, 455-460) is extended to interaction mining on a proteome-wide scale. An algorithm (the phylogenetic bootstrap) is introduced, which suggests traversal of a phenogram, interleaving rounds of computation and experiment, to develop a knowledge base of protein interactions in genetically-similar organisms. RESULTS: The interaction mining approach was demonstrated by building a learning system based on 1,039 experimentally validated protein-protein interactions in the human gastric bacterium Helicobacter pylori. An estimate of the generalization performance of the classifier was derived from 10-fold cross-validation, which indicated expected upper bounds on precision of 80% and sensitivity of 69% when applied to related organisms. One such organism is the enteric pathogen Campylobacter jejuni, in which comprehensive machine learning prediction of all possible pairwise protein-protein interactions was performed. The resulting network of interactions shares an average protein connectivity characteristic in common with previous investigations reported in the literature, offering strong evidence supporting the biological feasibility of the hypothesized map. For inferences about complete proteomes in which the number of pairwise non-interactions is expected to be much larger than the number of actual interactions, we anticipate that the sensitivity will remain the same but precision may decrease. We present specific biological examples of two subnetworks of protein-protein interactions in C. jejuni resulting from the application of this approach, including elements of a two-component signal transduction systems for thermoregulation, and a ferritin uptake network.

Algorithms↗

Controlled vocabulary and design of laboratory results displays.

Traditional data-review displays are driven by the ancillary systems that produced the data. A different paradigm is being used at Columbia-Presbyterian Medical Center (CPMC) where a controlled medical vocabulary-the Medical Entities Dictionary (MED) is the driving force behind laboratory data-review displays. Using hierarchical and semantic networks the authors have constructed a Web-based tool that considerably simplifies the MED-editing task required to create new displays. The tool uses knowledge in the MED to extract contextually relevant hierarchic and semantic sub-nets from the MED. The tool has a sensitivity of 92.2% and a relevance of 94.7% for retrieval of terms from the MED. Based on these results and given sufficient domains' structure within controlled vocabularies, we conclude that similar algorithms will enable applications to design and generate customized displays on-the-fly.

Algorithms↗

Comprehensive cardiac pacemaker information system: basis for a regional follow-up network.

The capabilities of a mini-computer-oriented permanent pacemaker information system are described. Extensive patient and pacer functional data are maintained in readily accessible files which may be displayed on a CRT terminal or printed out. Selective sorting of stored information may be accomplished according to any desired set of inclusive or exclusive criteria. Intelligent, comprehensive patient follow-up has been greatly facilitated through application of the system. In view of the rapidly expanding pacemaker population, it is suggested that cooperative regional networks operating with comparable information storage and retrieval structures will provide the only means for adequate patient surveillance and compilation of necessary pacemaker data.

Cardiac Pacing, Artificial↗

Measuring long-term memory storage and retrieval in children.

Given the successful use of selective reminding measures of learning and memory in experimental research, initial normative and psychometric data was collected to assess the potential clinical utility of a selective reminding measure with children. Sixty-six 5- to 8-year-old children were administered counterbalanced alternate forms of the selective reminding measure at two test periods separated by about four hours. Two of the three alternate forms tested were found to be of comparable difficulty, with the third being associated with slightly lower levels of performance. Statistical comparisons across the two test periods indicated that little practice effect occurs with the measure. Test-retest reliability coefficients were comparable to those found on some more established neuropsychological measures. It is hoped that the promise shown by this initial data will stimulate further clinical use and development of the selective reminding measure.

Age Factors↗

Using hit curves to compare search algorithm performance.

Databases continue to grow but the metrics available to evaluate information retrieval systems have not changed. Large collections such as MEDLINE and the World Wide Web contain many relevant documents for common queries. Ranking is therefore increasingly important and successful information retrieval systems, such as Google, have emphasized ranking. However, existing evaluation metrics such as precision and recall, do not directly account for ranking. This paper describes a novel way of measuring information retrieval performance using weighted hit curves adapted from the field of statistical detection to reflect multiple desirable characteristics such as relevance, importance, and methodologic quality. In statistical detection, hit curves have been proposed to represent occurrence of interesting events during a detection process. Similarly, hit curves can be used to study the position of relevant documents within large result sets. We describe hit curves in light of a formal model of information retrieval, show how hit curves represent system performance including ranking, and define ways to statistically compare performance of multiple systems using hit curves. We provide example scenarios where traditional measures are less suitable than hit curves and conclude that hit curves may be useful for evaluating retrieval from large collections where ranking performance is crucial.

Algorithms↗

Method to correlate tandem mass spectra of modified peptides to amino acid sequences in the protein database.

A method to correlate uninterpreted tandem mass spectra of modified peptides, produced under low-energy (10-50 eV) collision conditions, with amino acid sequences in a protein database has been developed. The fragmentation patterns observed in the tandem mass spectra of peptides containing covalent modifications is used to directly search and fit linear amino acid sequences in the database. Specific information relevant to sites of modification is not contained in the character-based sequence information of the databases. The search method considers each putative modification site as both modified and unmodified in one pass through the database and simultaneously considers up to three different sites of modification. The search method will identify the correct sequence if the tandem mass spectrum did not represent a modified peptide. This approach is demonstrated with peptides containing modifications such as S-carboxymethylated cysteine, oxidized methionine, phosphoserine, phosphothreonine, or phosphotyrosine. In addition, a scanning approach is used in which neutral loss scans are used to initiate the acquisition of product ion MS/MS spectra of doubly charged phosphorylated peptides during a single chromatographic run for data analysis with the database-searching algorithm. The approach described in this paper provides a convenient method to match the nascent tandem mass spectra of modified peptides to sequences in a protein database and thereby identify previously unknown sites of modification.

Algorithms↗

Harvesting chemical information from the Internet using a distributed approach: ChemXtreme.

The Internet is a comprehensive resource of chemical information which is at the same time largely unstructured. It provides a wealth of scientific information such as experimental data and requires a suitable automated data mining and analysis tool for its meaningful exploration. The Java based software presented here, ChemXtreme, is developed for harvesting chemical information from the Internet employing the Google API in combination with a distributed client/server text analysis architecture based on JavaRMI. It represents the first and until now the only toolkit for automated structured data retrieval from the Internet which is itself open source. ChemXtreme employs the "search the search engine" strategy, where the URLs returned from the search engine are analyzed further via textual pattern analysis. This process resembles the manual analysis of the hit list, where relevant data are captured and, by means of human intervention, are mined into a format suitable for further analysis. ChemXtreme on the other hand transforms chemical information automatically into a structured format suitable for storage in databases and further analysis and also provides links to the original information source. The query data retrieved from the search engine by the server is encoded, encrypted, and compressed and then sent to all the participating active clients in the network for parsing. Relevant information identified by the clients on the retrieved Web sites is sent back to the server, verified, and added to the database for data mining and further analysis. The distributed further analysis of URLs in a client/server architecture scales very favorably, thus producing only minimal overhead.

Chemistry↗

A case for automated tape in clinical imaging.

Electronic archiving of radiology images over many years will require many terabytes of storage with a need for rapid retrieval of these images. As more large PACS installations are installed and implemented, a data crisis occurs. The ability to store this large amount of data using the traditional method of optical jukeboxes or online disk alone becomes an unworkable solution. The amount of floor space number of optical jukeboxes, and off-line shelf storage required to store the images becomes unmanageable. With the recent advances in tape and tape drives, the use of tape for long term storage of PACS data has become the preferred alternative. A PACS system consisting of a centrally managed system of RAID disk, software and at the heart of the system, tape, presents a solution that for the first time solves the problems of multi-modality high end PACS, non-DICOM image, electronic medical record and ADT data storage. This paper will examine the installation of the University of Utah, Department of Radiology PACS system and the integration of automated tape archive. The tape archive is also capable of storing data other than traditional PACS data. The implementation of an automated data archive to serve the many other needs of a large hospital will also be discussed. This will include the integration of a filmless cardiology department and the backup/archival needs of a traditional MIS department. The need for high bandwidth to tape with a large RAID cache will be examined and how with an interface to a RIS pre-fetch engine, tape can be a superior solution to optical platters or other archival solutions. The data management software will be discussed in detail. The performance and cost of RAID disk cache and automated tape compared to a solution that includes optical will be examined.

Costs and Cost Analysis↗

Geographic information systems (GIS): new perspectives in understanding human health and environmental relationships.

Geographic information systems (GIS) and digital computer technology will advance the mission of the Centers for Disease Control and Prevention (CDC) and Agency for Toxic Substances and Disease Registry (ATSDR) to protect public health. Geographic positioning, topology, and planar and surface measurements are basic GIS properties which enable highly precise locational referencing of spatial phenomena. The growing uses of remotely sensed imagery and satellite facilitated global positioning systems are contributing to unprecedented surveillance of the environment and greater understanding of known and suspected environmental disease associations with human and animal health. Earth science and public health monitoring GIS databases offer new analytic opportunities for disease assessment and prevention.

CD-ROM↗

Protein identification in DNA databases by peptide mass fingerprinting.

Proteins can be identified using a set of peptide fragment weights produced by a specific digestion to search a protein database in which sequences have been replaced by fragment weights calculated for various cleavage methods. We present a method using multidimensional searches that greatly increases the confidence level for identification, allowing DNA sequence databases to be examined. This method provides a link between 2-dimensional gel electrophoresis protein databases and genome sequencing projects. Moreover, the increased confidence level allows unknown proteins to be matched to expressed sequence tags, potentially eliminating the need to obtain sequence information for cloning. Database searching from a mass profile is offered as a free service by an automatic server at the ETH, Zürich. For information, send an electronic message to the address cbrg/inf.ethz.ch with the line: help mass search, or help all.

Animals↗

Interfacing the radiology information system to the modality: an integrated approach.

The radiology information system (RIS) provides patient and examination information that is used in setting up and performing a radiologic procedure. In a digital imaging environment, information from the RIS can also be used to populate fields in the Digital Imaging and Communications in Medicine (DICOM) image header. Ideally, information from the RIS should be available at the modality at the time of the examination, and automatically be attached to the image in the appropriate DICOM fields before storage in the picture archiving and communications system (PACS). We have designed a highly integrated RIS interface for a digital radiography (DR) system. This interface employs browser technology to make RIS information conveniently available at the modality, and DICOM modality performed procedure step (MPPS) for RIS/DR information exchange. A novel feature of our approach is that a single display screen at the modality is used to alternatively display either the modality control window or the RIS window. Full access to RIS capabilities is available at the modality, including worklists and prior reports.

Computer Systems↗

[Rheumatology online: A survey among the members of the German Rheumatology Society].

OBJECTIVE: On behalf of the "Systemic Inflammatory Rheumatic Diseases Network" comprehensive, nationwide horizontal and vertical cross-linking of research and care is to be developed for the first time. The quality of scientific work and patient care is to be increased in the medium term through this improved communication and co-operation. Our objective was to determine what hardware and software are avail- able to the physicians involved, with a view to the Internet being used as a basis for communication and documentation within the network. METHODS: A survey was carried out among 723 active members of the German Rheumatology Society (DGRh). Data on the hardware and software used and on Internet access were collected using a unilateral questionnaire. RESULTS: The response rate among the addressed rheumatologists was 55.3%, with 64.1% of members in private practice replying. Of those responding 85% have Internet access, with rheumatologists in private practice using the Internet significantly less frequently at work than those working at a hospital (42% vs 80%). The latter accordingly reported a higher proportion of medical Internet usage (69% vs 52%, p<0.001). The survey demonstrated that software for private practices and hospitals shows a very variable picture with a multiplicity of systems being used. CONCLUSION: Use of the Internet for communication in the "Systemic Inflammatory Rheumatic Diseases Network" is practicable in hospitals but clearly restricted in the private practice sector. The widely varying software used in hospitals and private practices underlines the need for standardized, comprehensive documentation systems to be developed. To ensure acceptance and broadly based application, they need to be integrated into the existing computer infrastructure. In this context, Internetbased applications offer new opportunities through the use of system-independent file formats.

Data Collection↗

Tools for managing image flow in the modality to clinical-image-review chain.

Web-based clinical-image viewing is commonplace in large medical centers. As demands for product and performance escalate, physicians, sold on the concept of "any image, anytime, anywhere," fret when image studies cannot be viewed in a time frame to which they are accustomed. Image delivery pathways in large medical centers are oftentimes complicated by multiple networks, multiple picture archiving and communication systems (PACS), and multiple groups responsible for image acquisition and delivery to multiple destinations. When studies are delayed, it may be difficult to rapidly pinpoint bottlenecks. Described here are the tools used to monitor likely failure points in our modality to clinical-image-viewing chain and tools for reporting volume and throughput trends. Though perhaps unique to our environment, we believe that tools of this type are essential for understanding and monitoring image-study flow, re-configuring resources to achieve better throughput, and planning for anticipated growth. Without such tools, quality clinical-image delivery may not be what it should.

Data Display↗

Flying blind: using a digital dashboard to navigate a complex PACS environment.

Radiology workflows have become more distributed and complicated, and fewer tangible cues are available to the radiologist to help optimize task prioritization and selection. Additionally, faster scanners, more detailed exams, and increased demand for imaging services have precipitated a potential image overload for today's radiologists who are pressured to provide efficient, quality service in less time. Radiologists are faced with the task of operating within complex systems but are lacking tools to efficiently and effectively monitor these systems in real time. Dashboard technology can help address this deficiency in radiology and facilitate informed, optimized decisions about workflow. Possible areas of application include workflow consolidation, workload distribution, and urgency evaluation. Dashboards should be optimized, context-sensitive, customizable, and workflow-integrated. Further research is needed to identify the most important dashboard metrics, determine their optimal display, and validate their utility.

Data Display↗

The Regenstrief Medical Record System: a quarter century experience.

Entrusted with the records for more than 1.5 million patients, the Regenstrief Medical Record System (RMRS) has evolved into a fast and comprehensive data repository used extensively at three hospitals on the Indiana University Medical Center campus and more than 30 Indianapolis clinics. The RMRS routinely captures laboratory results, narrative reports, orders, medications, radiology reports, registration information, nursing assessments, vital signs, EKGs and other clinical data. In this paper, we describe the RMRS data model, file structures and architecture, as well as recent necessary changes to these as we coordinate a collaborative effort among all major Indianapolis hospital systems, improving patient care by capturing city-wide laboratory and encounter data. We believe that our success represents persistent efforts to build interfaces directly to multiple independent instruments and other data collection systems, using medical standards such as HL7, LOINC, and DICOM. Inpatient and outpatient order entry systems, instruments for visit notes and on-line questionnaires that replace hardcopy forms, and intelligent use of coded data entry supplement the RMRS. Physicians happily enter orders, problems, allergies, visit notes, and discharge summaries into our locally developed Gopher order entry system, as we provide them with convenient output forms, choice lists, defaults, templates, reminders, drug interaction information, charge information, and on-line articles and textbooks. To prepare for the future, we have begun wrapping our system in Web browser technology, testing voice dictation and understanding, and employing wireless technology.

Computer Terminals↗

Using annotated peptide mass spectrum libraries for protein identification.

A system for creating a library of tandem mass spectra annotated with corresponding peptide sequences was described. This system was based on the annotated spectra currently available in the Global Proteome Machine Database (GPMDB). The library spectra were created by averaging together spectra that were annotated with the same peptide sequence, sequence modifications, and parent ion charge. The library was constructed so that experimental peptide tandem mass spectra could be compared with those in the library, resulting in a peptide sequence identification based on scoring the similarity of the experimental spectrum with the contents of the library. A software implementation that performs this type of library search was constructed and successfully used to obtain sequence identifications. The annotated tandem mass spectrum libraries for the Homo sapiens, Mus musculus, and Saccharomyces cerevisiae proteomes and search software were made available for download and use by other groups.

Amino Acid Sequence↗