Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

GenBank.

GenBank (R) is a comprehensive sequence database that contains publicly available DNA sequences for more than 119 000 different organisms, obtained primarily through the submission of sequence data from individual laboratories and batch submissions from large-scale sequencing projects. Most submissions are made using the BankIt (web) or Sequin programs and accession numbers are assigned by GenBank staff upon receipt. Daily data exchange with the EMBL Data Library in the UK and the DNA Data Bank of Japan helps ensure worldwide coverage. GenBank is accessible through NCBI's retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome, mapping, protein structure and domain information, and the biomedical journal literature via PubMed. BLAST provides sequence similarity searches of GenBank and other sequence databases. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. To access GenBank and its related retrieval and analysis services, go to the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Animals↗

MARS: microarray analysis, retrieval, and storage system.

BACKGROUND: Microarray analysis has become a widely used technique for the study of gene-expression patterns on a genomic scale. As more and more laboratories are adopting microarray technology, there is a need for powerful and easy to use microarray databases facilitating array fabrication, labeling, hybridization, and data analysis. The wealth of data generated by this high throughput approach renders adequate database and analysis tools crucial for the pursuit of insights into the transcriptomic behavior of cells. RESULTS: MARS (Microarray Analysis and Retrieval System) provides a comprehensive MIAME supportive suite for storing, retrieving, and analyzing multi color microarray data. The system comprises a laboratory information management system (LIMS), a quality control management, as well as a sophisticated user management system. MARS is fully integrated into an analytical pipeline of microarray image analysis, normalization, gene expression clustering, and mapping of gene expression data onto biological pathways. The incorporation of ontologies and the use of MAGE-ML enables an export of studies stored in MARS to public repositories and other databases accepting these documents. CONCLUSION: We have developed an integrated system tailored to serve the specific needs of microarray based research projects using a unique fusion of Web based and standalone applications connected to the latest J2EE application server technology. The presented system is freely available for academic and non-profit institutions. More information can be found at http://genome.tugraz.at.

Algorithms↗

Reference master: a microcomputer-based storage and retrieval system for bibliographic references.

A complete system for housekeeping and retrieval of bibliographic references managing individual reprint collections is described. By the use of special hardware and individual data base software even large reprint collections in the range up to 65,000 papers are handled economically. A fast 8-bit microprocessor (HD 64180) in combination with a Winchester hard disk drive serves as the basis for rapid access to the desired information. An efficient string search algorithm written in assembly language guarantees a fast operation with a search speed of more than 6,000 entries/minute. The system cannot only prepare reference lists and reference files, but also incorporates an editor and maintains the control whether reprints are already on file or requested. The implementation of back-up schemes assure against data losses. Using a state of the art design single board computer and the most recent mass storage device technology, the system is as well small and cost effective, and thus suitable for personal use. In addition, some general questions and pitfalls concerning the management of scientific literature collections are touched upon.

Bibliographies as Topic↗

Collection and retrieval of structured clinical data from electronic patient records in general practice. A first-phase study to create a health care database for research and quality assessment.

OBJECTIVE: To evaluate prerequisites, practicalities, attitudes and limitations related to the collection of structured clinical data in everyday general practice for use in the future establishment of a national registration network. DESIGN: Prospective study. SETTING: Primary health care centres in south-western Sweden. SUBJECTS: Fourteen participating general practitioners in five primary health care centres. MAIN OUTCOME MEASURES: Feasibility and workload involved in structured data entry and in the retrieval of data from different record systems. The accuracy of clinical data in terms of clinical variables, correctness and representativeness. RESULTS: All four record systems could deliver basic data on the patient population. One centre had to be excluded from further data retrieval because of limitations in the data retrieval export format. Collecting data in everyday practice was feasible with acceptable data accuracy and moderate workload. CONCLUSION: It was feasible to collect, retrieve and store structured clinical data with respect to accuracy and extra workload. Interest in a national registration network and an increasing demand for information about primary health care in order to optimise clinical practices and support research, creates prerequisites for establishing a valid and reliable database. However, developmental work focusing on classification limitations, coding tools and routines for data retrieval is necessary.

Data Collection↗

Text-based knowledge discovery: search and mining of life-sciences documents.

Text literature is playing an increasingly important role in biomedical discovery. The challenge is to manage the increasing volume, complexity and specialization of knowledge expressed in this literature. Although information retrieval or text searching is useful, it is not sufficient to find specific facts and relations. Information extraction methods are evolving to extract automatically specific, fine-grained terms corresponding to the names of entities referred to in the text, and the relationships that connect these terms. Information extraction is, in turn, a means to an end, and knowledge discovery methods are evolving for the discovery of still more-complex structures and connections among facts. These methods provide an interpretive context for understanding the meaning of biological data.

Biological Science Disciplines↗

Selective dissemination and indexing of scientific information.

Selective dissemination of information to individuals provides a new and promising method for keeping abreast of current scientific information. Since SDI services are directed to the information needs of each individual, they are a significant step beyond grouporiented services and products, which require considerable expenditure of effort by each user as he sorts useful information from trash. However, SDI systems do require a high degree of precision in matching scientists against documents. They must operate more efficiently and economically than many current systems which occasionally provide a useful item of information to users. To meet these stringent requirements for quality, precision, efficiency, and economy, more research must be devoted to comparing and improving indexing methods, which are the basic component of all information storage and retrieval systems. It is incredible that so much money has been spent on the development and operation of scientific information systems before basic data on the comparative performance of various indexing methods have been gathered, analyzed, and confirmed by multiple investigators. The design of an effective information system would seem to require this type of basic knowledge, just as basic properties of alternative materials must be known before an engineer can design a building, bridge, or factory. Yet, except for the few studies mentioned in the previous section, research on indexing methods has been greatly neglected. Bourne's comment about studies of indexing languages is still an appropriate description of the situation: "In almost all the experimental reports, the investigator worked with an indexing language different than that of other experimenters. Consequently, no one has ever had his test results verified, or expanded, or made more precise by another experimenter" (47). Most existing information systems are based on keyword indexing, with concepts broken into isolated terms during input operations and recombined to synthesize the original concept during search and retrieval. Such systems tend to involve imprecise indexing, with a high level of "noise" in retrieved documents, difficult search strategy involving extensive post-coordination, and lengthy, complex computer manipulations. This situation reflects the fact that many producers of indexed data originally focused the design of their systems on the production of a published product with entries printed under short, concise index headings. Production of magnetic tapes as a by-product of the publication process, and their use for retrospective searching or for SDI services, was a much later development, almost an afterthought. Yet use of these tapes is growing so rapidly that it may be time to redesign the tape-producing systems, with ease of tape use for SDI services and retrospective searching as the primary consideration, and with publication of abstract and index bulletins or title listings relegated to secondary importance (49). The use of keywords to index documents creates a high degree of disorganization in information search and retrieval operations: Information is scattered under the many different terms that can be used to index different aspects of a concept. If the large-scale, comprehensive abstracting and indexing services were based on enumerative classifications with assignment of documents to logical hierarchical categories at the time of initial indexing, then many of the specialized information centers (50) and the 1300 abstracting and indexing services (3) would be unnecessary, and much of the reindexing and reprocessing of documents, the repackaging and reworking of abstracts and index data, and the resulting overlap and duplication characteristic of current information processing could be terminated. Partly because of the disorganization resulting from keyword indexing, the cost of a 5-year retrospective search of information on just one data base on magnetic tapes is a major investment (16). The effort and cost required to find a few items of useful information scattered among 1,285,000 abstracts indexed on 116 full reels of magnetic tape (11 million characters per reel) which will be needed for the 5-year Eighth Collective Index to Chemical Abstracts (1967-1971) (51) staggers the imagination. In contrast, when HICLASS systems based on enumerative hierarchical classifications are used, concepts that might be useful for later retrieval are identified and related items of information are grouped together during the indexing process. These enumerative classifications, with single-hit matching, make it possible to index and retrieve ideas as intact units and to perform simple sequential searches of the very small segment of a file that deals with a given topic (31). The experiments at both the Science Information Exchange and the National Cancer Institute, as described in this article, demonstrate that automated HICLASS systems are feasible and can operate at a very satisfactory level of performance. Although considerable effort may be required for the development and constant updating of detailed enumerative classifications, HICLASS categories may facilitate organization of data at the time of input, improve the precision of matching documents with users, and greatly simplify search logic and computer manipulations. If so, then output savings and performance would more than justify input costs, and the development and use of enumerative classifications would be a better solution to information problems than the current keyword-and-coordination approach. It is time to think beyond the ease of the single input step in information systems and to take a hard look at ways of easing retrieval problems for the multitude of information systems that process the indexed data (52). Indexing effort is expended only once, whereas search and retrieval effort is required by every user of a system. If information were better analyzed and organized during input operations, if more basic research were devoted to the effect of indexing methods on the performance of information systems, and if more emphasis were placed on the quality and usefulness of retrieved information, then the magnitude of problems related to the storage and retrieval of scientific information might be considerably reduced.

Abstracting and Indexing↗

Informatics and multiplexing of intact protein identification in bacteria and the archaea.

Although direct fragmentation of protein ions in a mass spectrometer is far more efficient than exhaustive mapping of 1-3 kDa peptides for complete characterization of primary structures predicted from sequenced genomes, the development of this approach is still in its infancy. Here we describe a statistical model (good to within approximately 5%) that shows that the database search specificity of this method requires only three of four fragment ions to match (at +/-0.1 Da) for a 99.8% probability of being correct in a database of 5,000 protein forms. Software developed for automated processing of protein ion fragmentation data and for probability-based retrieval of whole proteins is illustrated by identification of 18 archaeal and bacterial proteins with simultaneous mass-spectrometric (MS) mapping of their entire primary structures. Dissociation of two or three proteins at once for such identifications in parallel is also demonstrated, along with retention and exact localization of a phosphorylated serine residue through the fragmentation process. These conceptual and technical advances should assist future processing of whole proteins in a higher throughput format for more robust detection of co- and post-translational modifications.

Algorithms↗

The Medline/full-text research project.

This project was designed to test the relative efficacy of index terms and full-text for the retrieval of documents in those MEDLINE journals for which full-text searching was also available. The full-text files used were MEDIS from Mead Data Central and CCML from BRS Information Technologies. One hundred clinical medical topics were searched in these two files as well as the MEDLINE file to accumulate the necessary data. It was found that full-text identified significantly more relevant articles than did the indexed file, MEDLINE. The full-text searches, however, lacked the precision of searches done in the indexed file. Most relevant items missed in the full-text files, but identified in MEDLINE, were missed because the searcher failed to account for some aspect of natural language, used a logical or positional operator that was too restrictive, or included a concept which was implied, but not expressed in the natural language. Very few of the unique relevant full-text citations would have been retrieved by title or abstract alone. Finally, as of July, 1990 the more current issue of a journal was just as likely to appear in MEDLINE as in one of the full-text files.

Abstracting and Indexing↗

Integration of database capabilities into a patient reporting system.

A database design is described which automatically archives computer-generated patient imaging and radioassay reports. Selected phrases are condensed so that data storage will be efficient without sacrificing a prose style of report. An indexed file structure has been used to facilitate rapid record retrieval even when several hundred thousand records are stored. Personnel time is considerably reduced for recalling patient records, preparing periodic summaries of studies completed, and performing administrative functions such as billing and keeping track of checked out images. Complex queries, such as "list all the patients between the ages of 50 and 60 on digitalis referred for a stress cardiac study, with left ventricular ejection fraction less than 40% and apical dyskinesis," become feasible. A system for data backup is described to protect against catastrophic data loss.

Computers↗

Maintaining continuity of clinical operations while implementing large-scale filmless operations.

Texas Children's Hospital is a pediatric tertiary care facility in the Texas Medical Center with a large-scale, Digital Imaging and Communications in Medicine (DICOM)-compliant picture archival and communications system (PACS) installation. As our PACS has grown from an ultrasound niche PACS into a full-scale, multimodality operation, assuring continuity of clinical operations has become the number one task of the PACS staff. As new equipment is acquired and incorporated into the PACS, workflow processes, responsibilities, and job descriptions must be revised to accommodate filmless operations. Round-the-clock clinical operations must be supported with round-the-clock service, including three shifts, weekends, and holidays. To avoid unnecessary interruptions in clinical service, this requirement includes properly trained operators and users, as well as service personnel. Redundancy is a cornerstone in assuring continuity of clinical operations. This includes all PACS components such as acquisition, network interfaces, gateways, archive, and display. Where redundancy is not feasible, spare parts must be readily available. The need for redundancy also includes trained personnel. Procedures for contingency operations in the event of equipment failures must be devised, documented, and rehearsed. Contingency operations might be required in the event of scheduled as well as unscheduled service events, power outages, network outages, or interruption of the radiology information system (RIS) interface. Methods must be developed and implemented for reporting and documenting problems. We have a Trouble Call service that records a voice message and automatically pages the PACS Console Operator on duty. We also have developed a Maintenance Module on our RIS system where service calls are recorded by technologists and service actions are recorded and monitored by PACS support personnel. In a filmless environment, responsibility for the delivery of images to the radiologist and referring physician must be accepted by each imaging supervisor. Thus, each supervisor must initiate processes to verify correct patient and examination identification and the correct count and routing of images with each examination.

Computer Communication Networks↗

iScout: an intelligent scout for accessing and navigating large image sets in a PACS.

A new software tool for PACS, called iScout (intelligent scout), has been developed and optimized for a radiology workstation. The purpose of iScout is to display an overview of a large image series, allowing the user to select images for priority downloading from a PACS server to a PACS workstation. This allows radiologists to reduce the delays that are associated with downloading hundreds or even thousands of images. Several schemes that semiautomatically manage the download process are presented along with tests to measure performance. The results of the tests confirm that priority downloading provides faster access to images in large image series and that the time savings increase in proportion to the study size.

Computer Systems↗

Endosperm-preferred expression of maize genes as revealed by transcriptome-wide analysis of expressed sequence tags.

The transcriptome-wide endosperm-preferred expression of maize genes was addressed by analyzing a large database of expressed sequence tags (ESTs). We generated 30,531 high quality sequence-reads from the 5'-ends of cDNA libraries from maize endosperm harvested at 10, 15, and 20 days after pollination. A further 196,900 maize sequence-reads retrieved from public databases were added to this endosperm collection to generate MAIZEST, a database with tools for data storage and analysis. MAIZEST contains 227,431 ESTs, one third of which represents developing endosperm and the remaining two-thirds represent transcripts from 49 cDNA libraries constructed from different organs and tissues. Assembling the MAIZEST ESTs generated 29,206 putative transcripts, of which a set of 4032 assembled sequences was composed exclusively of sequences derived from endosperm cDNA libraries. After sequence analysis using overlapping parameters, a sub-set of 2403 assembled sequences was functionally annotated and revealed a wide variety of putative new genes involved in endosperm development and metabolism.

Expressed Sequence Tags↗

Clinical evaluation of newly developed CRT viewing station: CT reading and observer's performance.

The clinical performance of the new viewing station with six CRT monitors (17-inch, 1,024 x 1,280) was evaluated. In the primary interpretation of CT images, time measurements were carried out for eight radiologists. No significant differences in reading time existed between CRT and film in 3 of 4 readers in head CT series, and in 2 of 6 readers in body CT series. Compared with the previous system, the new prototype system achieved an approximately 30% decrease in reading time in both head and body CT studies and could reduce mental and eye fatigue.

Computer Systems↗

Survival estimates of a prognostic classification depended more on year of treatment than on imputation of missing values.

BACKGROUND AND OBJECTIVE: The International Germ Cell Consensus (IGCC) classification defines good, intermediate, and poor prognosis groups among patients with nonseminomatous germ cell cancer. In the database used to develop the IGCC classification (n = 5,202), >40% of patients were excluded because of missing values (n = 2,154). We looked for effects of this exclusion on survival estimates in the three IGCC prognosis groups. STUDY DESIGN AND SETTING: We imputed missing values using a multiple imputation procedure. The IGCC classification was applied to patients with complete data (n = 3,048) and with imputed data (n = 2,154), and 5-year survival was calculated for each prognosis group. RESULTS: Patients with missing values had a lower 5-year survival than those without missing values: 76% vs. 82%. Five-year survival in the complete and imputed data samples was 92% and 87% for the good prognosis groups and 80% and 70% for the intermediate prognosis groups, whereas 5-year survival for the poor prognosis groups in both samples was similar (50% and 47%, respectively). This difference in survival was largely explained by a higher proportion of missing values among patients treated before 1985, who had a worse survival than patients treated after 1985. CONCLUSION: Multiple imputation of the missing values led to lower survival estimates across the IGCC prognosis groups, compared with estimates based on the complete data. Although imputation of missing values gives statistically better survival estimates, adjustments for year of treatment are necessary to make the estimates applicable to currently diagnosed patients with testicular cancer.

Classification↗

National consensus on data elements for nurse managed health centers.

This report presents a summary of the findings from the National Network for Nurse Managed Health Centers Data Consensus Conference. Nationally, nurse-managed health centers are increasingly offering communities another option for access to high-quality primary care. The lack of agreed upon, standardized data elements for these centers has limited the ability to present clear information about their contributions as well as to inform policy related to their support and development. Fifty-three national invitees came to consensus in Washington, DC on the critical data elements for a national database for nurse-managed health centers. This database includes both clinical and financial/business practices elements. Consensus was not reached around some clinical areas. These areas are briefly discussed as well as the plans for next stages of data collection.

Community Health Centers↗