Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Automated and computer-assisted pathology support for a large chronic study.

Pathology support for the 24,192 mouse toxicological ED01 study at the National Center for Toxicological Research was provided by the University of Arkansas Pathology Services Project through a contract between NCTR and the University of Arkansas. To aid in the collection, storage, retrieval, and analysis of this large amount of data an automated computer-assisted pathology system was developed. Use of this system has resulted in accurate data, the ability to handle large amounts of data, and low-cost analysis of the data.

Animals↗

Ambiguity resolution while mapping free text to the UMLS Metathesaurus.

We propose a method for resolving ambiguities encountered when mapping free text to the UMLS Metathesaurus. Much of the research in medical informatics involves the manipulation of free text. The Metathesaurus contains extensive information which supports solutions to problems encountered while processing such text. After discussing the process of mapping free text to the Metathesaurus and describing the ambiguities which are often the result of such mapping, we provide examples of rules designed to eliminate mapping ambiguities. These rules refer to the context in which the ambiguity occurs and crucially depend on semantic types obtained from the Metathesaurus. We have conducted a preliminary test of the methodology and the results obtained indicate that the rules successfully resolve ambiguity around 80% of the time.

Abstracting and Indexing↗

Pivot/Remote: a distributed database for remote data entry in multi-center clinical trials.

1. INTRODUCTION. Data collection is a critical component of multi-center clinical trials. Clinical trials conducted in intensive care units (ICU) are even more difficult because the acute nature of illnesses in ICU settings requires that masses of data be collected in a short time. More than a thousand data points are routinely collected for each study patient. The majority of clinical trials are still "paper-based," even if a remote data entry (RDE) system is utilized. The typical RDE system consists of a computer housed in the CC office and connected by modem to a centralized data coordinating center (DCC). Study data must first be recorded on a paper case report form (CRF), transcribed into the RDE system, and transmitted to the DCC. This approach requires additional monitoring since both the paper CRF and study database must be verified. The paper-based RDE system cannot take full advantage of automatic data checking routines. Much of the effort (and expense) of a clinical trial is ensuring that study data matches the original patient data. 2. METHODS. We have developed an RDE system, Pivot/Remote, that eliminates the need for paper-based CRFs. It creates an innovative, distributed database. The database resides partially at the study clinical centers (CC) and at the DCC. Pivot/Remote is descended from technology introduced with Pivot [1]. Study data is collected at the bedside with laptop computers. A graphical user interface (GUI) allows the display of electronic CRFs that closely mimic the normal paper-based forms. Data entry time is the same as for paper CRFs. Pull-down menus, displaying the possible responses, simplify the process of entering data. Edit checks are performed on most data items. For example, entered dates must conform to some temporal logic imposed by the study. Data must conform to some acceptable range of values. Calculations, such as computing the subject's age or the APACHE II score, are automatically made as the data is entered. Data that is collected serially (BP, HR, etc.) can be displayed graphically in a trend form along with other related variables. An audit trail is created that automatically tracks all changes to the original data, making it possible to reconstruct the CRF to any point in time. On-line help provides information on the study protocol as well as assistance with the use of the system. Electronic security makes it possible to lock certain parts of the CRF once it has been monitored. Completed CRFs are transmitted to the DCC via electronic mail where it is reviewed and merged into the study database. Questions about subject data are transmitted back to the CC via electronic mail. This approach to maintaining the study database is unique in that the study data files are distributed among the CC and DCC. Until a subject's CRF is monitored (verified against the original patient data residing in the hospital record), it logically resides at the CC where it was collected. Copies are transmitted to the DCC and are only read there. Any pre-monitoring changes must be made to the data at the CC. Once the subject's CRF is monitored, it logically moves to the DCC, and any subsequent changes are made at the DCC with copies of the CRF flowing back to the CC. 3. DISCUSSION. Pivot/Remote eliminates the need for paper forms by utilizing portable computers that can be used at the patient bedside. A GUI makes it possible to quickly enter data. Because the user gets instant feedback on possible error conditions, time is saved because the original data is close at hand. The ability to display trended data or variables in the context of other data allows detection of erroneous conditions beyond simple range checks. The logical construction of the database minimizes the problem of managing dual databases (at the CC and DCC) and keeps CC personnel in the loop until all changes are made.

Computer Communication Networks↗

Structured data entry in ORCA: the strengths of two models combined.

The capture of patient data in a structured format receives increasing attention. Data can be extracted from free text using natural language processing techniques, but it can also be collected in a structured fashion at the time of data entry. The latter has the advantage that completeness and unambiguity can be promoted by offering predefined terms and options for description of findings. The paper discusses two models for supporting structured data entry. In the direct model, there is an immediate relationship between the terms and options for data entry and the structure of the underlying database. In the indirect model, terms and options for data entry are based on a controlled vocabulary and not directly related to the structure in which actual data is represented. Both models have been utilized by ORCA (Open Record for CAre). We discuss the pros and cons of these two models in relation to the type of patient data and the task involved. It is concluded that a strategic combination of both models has more strengths and less weaknesses than the use of each model only.

Data Collection↗

Database on the structure of small ribosomal subunit RNA.

About 8600 complete or nearly complete sequences are now available from the Antwerp database on small ribosomal subunit RNA. All these sequences are aligned with one another on the basis of the adopted secondary structure model, which is corroborated by the observation of compensating substitutions in the alignment. Literature references, accession numbers and detailed taxonomic information are also compiled. The database can be consulted via the World Wide Web at URL http://rrna.uia.ac.be/ssu/

Base Sequence↗

Security infrastructure services for electronic archives and electronic health records.

Communication and co-operation in the domain of healthcare and welfare require a well-defined set of security services based on a Public Key Infrastructure and provided by a Trusted Third Party (TTP). These services describe both status and relation of communicating principals, corresponding keys and attributes, and the access rights to applications and data. Additional services are needed to provide trustworthy information about dynamic issues of communication and co-operation such as time and location of processes, workflow relations, and system behaviour. Legal, social, behavioural and ethical requirements demand securely stored patient information and well-established access tools and tokens. Electronic (and more specifically digital) signatures--as important means for securing the integrity of a message or file--along with certified time stamps or time signatures are especially important for purposes of data storage in electronic archives and electronic health records (EHR). While just mentioning technical storage problems (e.g. lifetime of the storage devices, interoperability of retrieval and presentation software), this paper identifies mechanisms of securing data items, files, messages, sets of archived items or documents, electronic archive structures, and life-long electronic health records. Other workshop contributions will demonstrate related aspects of policies, patient privacy, and privilege management.

Access to Information↗

Number of nodes examined and staging accuracy in colorectal carcinoma.

PURPOSE: The objective of this study was to determine the number of nodes that need to be examined to accurately reflect the histology of the regional lymphatics in colorectal carcinoma. PATIENTS AND METHODS: Patients undergoing curative resection for T2 and T3 colorectal cancer between 1992 and 1996 were reviewed. Pathologic data from these patients were entered into a computerized database for storage, retrieval, and analysis. The major outcome measured was the number of nodes that need to be examined to achieve a node-positive rate consistent with that reported in the National Cancer Data Base (NCDB) report. RESULTS: The number of nodes examined ranged from 0 to 78 (mean, 17 nodes). Node-negative patients had fewer nodes examined (mean, 14 nodes) than node-positive patients (mean, 20 nodes; P =.003). The entire sample had a node-positive rate of 38.8% (95% confidence interval [CI], 32% to 45.5%), not statistically different from that in the NCDB report. When at least 14 nodes were examined, the percent of patients with at least one positive node was 33.3% (95% CI, 24.6% to 42.3%), not statistically different from the NCDB report. CONCLUSION: In a sample of patients statistically similar to the sample in the NCDB report, the examination of at least 14 nodes after resection of T2 or T3 carcinoma of the colon and rectum will accurately stage the lymphatic basin.

Colonic Neoplasms↗

Linking patient information systems to bibliographic resources.

Medical informatics researchers have explored a number of ways to integrate medical information resources into patient care systems. Particular attention has been given to the integration of on-line bibliographic resources. This paper presents an information model which breaks down the integration task into three components, each of which answers a question: what is the user's question?, where can the answer be found?, and how is the retrieval strategy composed? Twelve experimental systems are reviewed and their methods for addressing one or more of these questions are described.

Computer Systems↗

[The automated ECG laboratory: equipment and operational problems (author's transl)].

In the light of recent advances in technology, the basic equipment of an automated ECG laboratory is described. The main features of data acquisition terminals, data receiver/controller units, A/D converters, computers, visual displays and systems for storage and retrieval of tracings, are briefly discussed. Three major alternatives are open for computer-aided ECG interpretation today: 1) complete, dedicated system in the hospital; 2) ECG data collection system with offline analysis by hospital business computer; 3) ECG service center outside of the hospital. Advantages and possible limitations of these methods are discussed. At the Ospedale Civile Regionale of Udine we have choosen the first method. An HP 1530 ECG interpretative system and the 12-lead ECG analysis program developed by Caceres-USPHS are used. Analog tracing and interpretative printout are available in the laboratory and/or at the patient location in about one minute. Our system has been working for less than one year. At present, 150-200 ECG are processed daily. Such an ECG processing system has proven to yield considerable savings in time and manpower. Some operational problems related to shifting from manual to computer work have been gradually overcome and will be discussed.

Computers↗

Providing concept-oriented views for clinical data using a knowledge-based system: an evaluation.

OBJECTIVE: Clinical information systems typically present patient data in chronologic order, organized by the source of the information (e.g., laboratory, radiology). This study evaluates the functionality and utility of a knowledge-based system that generates concept-oriented views (organized around clinical concepts such as disease or organ system) of clinical data. DESIGN: The authors have developed a system that uses a knowledge base of interrelationships between medical concepts to infer relationships between data in electronic medical records. They use these inferences to produce summaries, or views, of the data that are relevant to a specific concept of interest. They evaluated the ability of the system to select relevant information, reduce information overload, and support physician information retrieval. MEASUREMENTS: The sensitivity and specificity of the system for identifying relevant patient information were calculated. Effect on information overload was assessed by comparing the amount of information in each view with the amount of information in the entire record. Information retrieval accuracy and cost (time) were used to measure the effect of using concept-oriented views on the efficiency and effectiveness of retrievals. RESULTS: The sensitivity and specificity of the system for identifying relevant clinical information were generally in the range of 70 to 80 percent. Concept-oriented views are effective in reducing the amount of information retrieved (over 80 percent reduction) and, compared with source-oriented views, are able to improve physician retrieval accuracy (p=0.04). CONCLUSION: Computer-generated, concept-oriented views can be used to reduce clinician information overload and improve the accuracy of clinical data retrieval.

Artificial Intelligence↗

Pithos - a scalable and secure data container for FAIR-compliant research data management in life sciences.

Modern research techniques have led to exponential growth in the volume and complexity of scientific data. Consequently, managing these volumes securely and efficiently has become a major challenge. While all research domains face these challenges, life science research is particularly affected because current approaches often rely on a large set of different file formats, with metadata stored in separated databases or spreadsheets. This leads to fragmented datasets, orphaned data, and compromised research reproducibility. Traditional solutions also force researchers to choose between security and accessibility, with encrypted files preventing selective access and indexed formats lacking adequate security for sensitive data. These limitations are particularly problematic in large-scale genomic studies where researchers must decompress multi-gigabyte files to access specific regions, creating computational bottlenecks and inefficient network usage when working with cloud-stored datasets. We introduce Pithos, a next-generation file format specifically designed for scientific data management in distributed cloud environments. The format uses content-defined chunking to enable efficient deduplication across distributed storage systems, thereby reducing storage costs and bandwidth requirements. The append-only structure ensures data immutability and allows for incremental updates without compromising content. Benchmark results show that Pithos outperforms existing solutions in read and write performance, with comparable or improved storage efficiency.

Biological Science Disciplines↗

AthaMap: an online resource for in silico transcription factor binding sites in the Arabidopsis thaliana genome.

Gene expression is controlled mainly by the binding of transcription factors to regulatory sequences. To generate a genomic map for regulatory sequences, the Arabidopsis thaliana genome was screened for putative transcription factor binding sites. Using publicly available data from the TRANSFAC database and from publications, alignment matrices for 23 transcription factors of 13 different factor families were used with the pattern search program Patser to determine the genomic positions of more than 2.4 x 10(6) putative binding sites. Due to the dense clustering of genes and the observation that regulatory sequences are not restricted to upstream regions, the prediction of binding sites was performed for the whole genome. The genomic positions and the underlying data were imported into the newly developed AthaMap database. This data can be accessed by positional information or the Arabidopsis Genome Initiative identification number. Putative binding sites are displayed in the defined region. Data on the matrices used and on the thresholds applied in these screens are given in the database. Considering the high density of sites it will be a valuable resource for generating models on gene expression regulation. The data are available at http://www.athamap.de.

Arabidopsis↗

Experiences with a distributed virtual patient record system.

TeleMed is a distributed diagnosis and analysis system, which permits physicians who are not collocated to consult on the status of a patient. The patient's record is dynamically constructed from data that may reside at several sites but which can be quickly assembled for viewing by pointing to the patient's name. Then, a graphical patient record appears, through which consulting physicians can retrieve textual and radiographic data with a single mouse click. TeleMed uses modern distributed object technology and emerging telecollaboration tools.

Computer Communication Networks↗

Development of an audiometry data processing system.

The development was reported of an electronic data processing system of computerized automatic audiometry and of Bekesy audiometry. This system using a microcomputer has the following features: (1) The large amount of data obtained through (a) automatic audiometry, (b) Bekesy audiometry (measurement at fixed or continuous frequency plus test practice, and (c) the Temporal Tone-Decay test can be stored on-line in real time, and are processed and displayed off-line at any time. (2) When retrieved, the output format of these audiometric data is of the same pattern as the conventional audiogram. (3) The output pattern on the CRT graphic display can be hard-copied whenever desired. (4) The system can be operated manually in the same way as the conventional method. (5) The SISI and DL tests can be executed manually and the resultant data keyed into computer storage. And (6) The operation from the data input to the retrieval output is performed interpretively through the CRT display, so that anyone even without special computer or audiological training can operate it at any time.

Audiometry↗

Electronic processing of medical visual information.

The potential data base for electronic processing has recently been enlarged significantly to include static and dynamic visual information. Emergence of the videodisk and other technology now enables computer-assisted instruction to be audiovisual and audiovisual communication to be interactive. One consequence, now that audiovisual display has been wedded to computer systems, is that computer-assisted instruction (CAI) in many medical subjects should no longer be considered acceptable unless it includes pictorial data. Therefore, the computer scientist must now master the principles and become knowledgeable in the skills of audiovisual communication to cope successfully with the impending revolution in education and to make good use of the newly available technology. Similarly, in planning storage and retrieval systems, images must be considered as part of the data. This presentation traces the development of interactive multimedial self-instruction by describing the innovative functions that have appeared in equipment during the past 13 years culminating in the QuadraSync. The low cost of some of these devices encourages their adoption.

Audiovisual Aids↗

The micro- or home computer as a teaching aid in postgraduate medical study.

The home computer 'magic' portrayed by the advertising media, supposed to bring computer science to the fingertips of every man, woman and child, is assessed by a practical medical lecturer. Reviewable data banks from which data can be retrieved almost instantaneously, the elimination of filing cabinets and stockpiles of notes, and the efficient storage of clinical patient data together with the relatively low cost and ease of operation make the computer one of the most exciting additions to modern teaching aids.

Computers↗