Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

Microarray databases: standards and ontologies.

A single microarray can provide information on the expression of tens of thousands of genes. The amount of information generated by a microarray-based experiment is sufficiently large that no single study can be expected to mine each nugget of scientific information. As a consequence, the scale and complexity of microarray experiments require that computer software programs do much of the data processing, storage, visualization, analysis and transfer. The adoption of common standards and ontologies for the management and sharing of microarray data is essential and will provide immediate benefit to the research community.

Database Management Systems↗

Options available--from start to finish--for obtaining data from DNA microarrays II.

Microarray technology has undergone a rapid evolution. With widespread interest in large-scale genomic research, an abundance of equipment and reagents have now become available and affordable to a large cross section of the scientific community. As protocols become more refined, careful investigators are able to obtain good quality microarray data quickly. In most recent times, however, perhaps one of the biggest obstacles researchers face is not the manufacture and use of microarrays at the bench, but storage and analysis of the array data. This review discusses the most recent equipment, reagents and protocols available to the researcher, as well as describing data analysis and storage options available from the evolving field of microarray informatics.

DNA↗

Elements of a systematic review.

This paper examines the subject of systematic reviews from a nursing viewpoint. The history of the evidence-based healthcare movement and the major differences between systematic reviews and traditional literature reviews are discussed. The steps of the process used by those conducting reviews are examined in detail. These include structuring a research question, searching and appraising the literature, data extraction, analysis and synthesis, and reporting the results. It is this process that ensures reviews can be considered as a legitimate form of nursing research.

Data Collection↗

Internationalization and localization: evaluating and testing a Website for Asian users.

The objective of this study was to combine internationalization and localization of Websites and improvement of Website usability with user-centred design methods. This study designed for internationalization and localization of Websites for Asian users, and implemented usability engineering into every phase of Website usability testing, based on the internationalization and localization perspectives of the honeywell.com/your home Website. The first step was to develop the usage scenarios. Three Asian usability specialists carried out one heuristic evaluation session for the current honeywell.com/your home Website. The usability problems were analysed and possible solutions to these problems were discussed. In the next phase, cluster analysis was utilized to test current information architecture. The results provided options for future information architecture development for this Website. Finally, a performance measurement test was conducted to investigate the performance for Asian users. Based on the results, suggestions for improving the Website usability from the localization perspective were provided. The results demonstrate the user-centred design (UCD) approach and stress international and local issues in Website development to Website designers.

Adult↗

Complementary roles for toxicologic pathology and mathematics in toxicogenomics, with special reference to data interpretation and oscillatory dynamics.

Toxicogenomics is an emerging multidisciplinary science that will profoundly impact the practice of toxicology. New generations of biologists, using evolving toxicogenomics tools, will generate massive data sets in need of interpretation. Mathematical tools are necessary to cluster and otherwise find meaningful structure in such data. The linking of this structure to gene functions and disease processes, and finally the generation of useful data interpretation remains a significant challenge. The training and background of pathologists make them ideally suited to contribute to the field of toxicogenomics, from experimental design to data interpretation. Toxicologic pathology, a discipline based on pattern recognition, requires familiarity with the dynamics of disease processes and interactions between organs, tissues, and cell populations. Optimal involvement of toxicologic pathologists in toxicogenomics requires that they communicate effectively with the many other scientists critical for the effective application of this complex discipline to societal problems. As noted by Petricoin III et al (Nature Genetics 32, 474-479, 2002), cooperation among regulators, sponsors and experts will be essential for realizing the potential of microarrays for public health. Following a brief introduction to the role of mathematics in toxicogenomics, "data interpretation" from the perspective of a pathologist is briefly discussed. Based on oscillatory behavior in the liver, the importance of an understanding of mathematics is addressed, and an approach to learning mathematics "later in life" is provided. An understanding of pathology by mathematicians involved in toxicogenomics is equally critical, as both mathematics and pathology are essential for transforming toxicogenomics data sets into useful knowledge.

Animals↗

Performance improvement of the SPIHT coder based on statistics of medical ultrasound images in the wavelet domain.

This paper proposes some modifications to the state-of-the-art Set Partitioning In Hierarchical Trees (SPIHT) image coder based on statistical analysis of the wavelet coefficients across various subbands and scales, in a medical ultrasound (US) image. The original SPIHT algorithm codes all the subbands with same precision irrespective of their significance, whereas the modified algorithm processes significant subbands with more precision and ignores the least significant subbands. The statistical analysis shows that most of the image energy in ultrasound images lies in the coefficients of vertical detail subbands while diagonal subbands contribute negligibly towards total image energy. Based on these statistical observations, this work presents a new modified SPIHT algorithm, which codes the vertical subbands with more precision while neglecting the diagonal subbands. This modification speeds up the coding/decoding process as well as improving the quality of the reconstructed medical image at low bit rates. The experimental results show that the proposed method outperforms the original SPIHT on average by 1.4 dB at the matching bit rates when tested on a series of medical ultrasound images. Further, the proposed algorithm needs 33% less memory as compared to the original SPIHT algorithm.

Algorithms↗

Data mining of molecular dynamics trajectories of nucleic acids.

Analysis, storage, and transfer of molecular dynamic trajectories are becoming the bottleneck of computer simulations. In this paper we discuss different approaches for data mining and data processing of huge trajectory files generated from molecular dynamic simulations of nucleic acids.

Computer Simulation↗

XML-based visualization of design and completeness in medical databases.

PURPOSE: mdplot (medical database plot) visualizes both structure and quality of data in medical databases by means of a summary representation of design and completeness in XML format. The goal is to identify attributes suitable for evaluation and to aid in creating open data models. METHODS: A three-stage visualization approach is applied. First, an overview of all classes in a database, second a detailed view of a specific class and third an analysis of individual attributes. Missing data is identified to enable specific efforts to improve data quality prior to analysis. For each class number of patients, attributes, and records per patient are provided. A condensed bar chart for each category of attributes (categorical, numerical, text and other) visualizes available content: The abscissa corresponds to the sequence of attributes; the ordinate represents completeness per attribute. By selection of a specific class, a detailed description is provided including mean completeness in each category as well as completeness per attribute. To analyse attributes that are collected at several time points per patient, a frequency distribution of records per patient can be generated. RESULTS: The new methodology was applied to two clinical research databases consisting of 292 attributes (955 patients) and 224 attributes (610 patients), respectively, and resulted in major restructuring of the systems. A public website is provided for generation of mdplots.

Data Collection↗

Electronic availability of EUROCARE-3 data: a tool for further analysis.

The EUROCARE-3 CD-ROM has been developed to provide more detailed data with respect to those published in the monograph. The CD-ROM provides estimates of age-specific and age-standardised survival figures, cumulative and interval-specific survival, observed and relative survival for 47 cancer sites or combinations of sites, based on >4 million adult cancer patients diagnosed from 1983 to 1994 and reported from 56 European cancer registries. In addition, the CD-ROM provides observed survival proportions for 25 childhood cancer entities based on 23,000 young patients diagnosed from 1990 to 1994. Survival indicators, corresponding standard errors and confidence intervals can be selected according to cancer site, registry or country, sex, age class and disease duration. Basic graphical display and export facilities have also been provided. As an example of how to use this CD-ROM, this paper will report a descriptive analysis of relative survival patterns for all cancers combined, by age, sex and country. The EUROCARE-3 CD-ROM can be ordered free of charge or directly downloaded at http://www.eurocare.it.

Adolescent↗

Imagene: an integrated computer environment for sequence annotation and analysis.

MOTIVATION: To be fully and efficiently exploited, data coming from sequencing projects together with specific sequence analysis tools need to be integrated within reliable data management systems. Systems designed to manage genome data and analysis tend to give a greater importance either to the data storage or to the methodological aspect, but lack a complete integration of both components. RESULTS: This paper presents a co-operative computer environment (called Imagenetrade mark) dedicated to genomic sequence analysis and annotation. Imagene has been developed by using an object-based model. Thanks to this representation, the user can directly manipulate familiar data objects through icons or lists. Imagene also incorporates a solving engine in order to manage analysis tasks. A global task is solved by successive divisions into smaller sub-tasks. During program execution, these sub-tasks are graphically displayed to the user and may be further re-started at any point after task completion. In this sense, Imagene is more transparent to the user than a traditional menu-driven package. Imagene also provides a user interface to display, on the same screen, the results produced by several tasks, together with the capability to annotate these results easily. In its current form, Imagene has been designed particularly for use in microbial sequencing projects. AVAILABILITY: Imagene best runs on SGI (Irix 6.3 or higher) workstations. It is distributed free of charge on a CD-ROM, but requires some Ilog licensed software to run. Some modules also require separate license agreements. Please contact the authors for specific academic conditions and other Unix platforms. CONTACT: imagene home page: http://wwwabi.snv.jussieu.fr/imagene

Bacillus subtilis↗

Pooled library tissue tags for EST-based gene discovery.

MOTIVATION: In gene discovery projects based on EST sequencing, effective post-sequencing identification methods are important in determining tissue sources of ESTs within pooled cDNA libraries. In the past, such identification efforts have been characterized by higher than necessary failure rates due to the presence of errors within the subsequence containing the oligo tag intended to define the tissue source for each EST. RESULTS: A large-scale EST-based gene discovery program at The University of Iowa has led to the creation of a unique software method named UITagCreator usable in the creation of large sets of synthetic tissue identification tags. The identification tags provide error detection and correction capability and, in conjunction with automated annotation software, result in a substantial improvement in the accurate identification of the tissue source in the presence of sequencing and base-calling errors. These identification rates are favorable, relative to past paradigms. AVAILABILITY: The UITagCreator source code and installation instructions, along with detection software usable in concert with created tag sets, is freely available at http://genome.uiowa.edu/pubsoft/software.html CONTACT: tomc@eng.uiowa.edu

Algorithms↗

Visualization and analysis of protein interactions.

SUMMARY: We have developed a new program called InterViewer for drawing large-scale protein interaction networks in three-dimensional space. Unique features of InterViewer include (1) it is much faster than other recent implementations of drawing algorithms; (2) it can be used not only for visualizing protein interactions but also for analyzing them interactively; and (3) it provides an integrated framework for querying protein interaction databases and directly visualizes the query results. AVAILABILITY: http://wilab.inha.ac.kr/protein/

Algorithms↗

PSI: indexing protein structures for fast similarity search.

MOTIVATION: We consider the problem of finding similarities in protein structure databases. Current techniques sequentially compare the given query protein to all of the proteins in the database to find similarities. Therefore, the cost of similarity queries increases linearly as the volume of the protein databases increase. As the sizes of experimentally determined and theoretically estimated protein structure databases grow, there is a need for scalable searching techniques. RESULTS: Our techniques extract feature vectors on triplets of SSEs (Secondary Structure Elements). Later, these feature vectors are indexed using a multidimensional index structure. For a given query protein, this index structure is used to quickly prune away unpromising proteins in the database. The remaining proteins are then aligned using a popular alignment tool such as VAST. We also develop a novel statistical model to estimate the goodness of a match using the SSEs. Experimental results show that our techniques improve the pruning time of VAST 3 to 3.5 times while maintaining similar sensitivity.

Algorithms↗

BRAGI: linking and visualization of database information in a 3D viewer and modeling tool.

BRAGI is a well-established package for viewing and modeling of three-dimensional (3D) structures of biological macromolecules. A new version of BRAGI has been developed that is supported on Windows, Linux and SGI. The user interface has been rewritten to give the standard 'look and feel' of the chosen operating system and to provide a more intuitive, easier usage. A large number of new features have been added. Information from public databases such as SWISS-PROT, InterPro, DALI and OMIM can be displayed in the 3D viewer. Structures can be searched for homologous sequences using the NCBI BLAST server.

Amino Acid Sequence↗

ATID: a web-oriented database for collection of publicly available alternative translational initiation events.

SUMMARY: Alternative translational initiation is an important cellular mechanism contributing to the diversity of protein products and functions. We develop a database that provides a comprehensive collection of alternative translational initiation events. The purpose of this alternative translational initiation database (ATID) is to facilitate the systematic study of alternative translational initiation of genes. The current version of database contains 300 genes from Homo sapiens, Mus musculus and other species. Each of the genes has two or more isoforms due to alternative translational initiation. Resources in ATID, including gene information, alternative products of genes and domain structures of isoforms, are provided through a user-friendly web interface. AVAILABILITY: The ATID database is available for public use at http://bioinfo.au.tsinghua.edu.cn/atie/.

Amino Acid Sequence↗

ClaNC: point-and-click software for classifying microarrays to nearest centroids.

SUMMARY: ClaNC (classification to nearest centroids) is a simple and an accurate method for classifying microarrays. This document introduces a point-and-click interface to the ClaNC methodology. The software is available as an R package. AVAILABILITY: ClaNC is freely available from http://students.washington.edu/adabney/clanc

Algorithms↗

What should be expected from feature selection in small-sample settings.

MOTIVATION: High-throughput technologies for rapid measurement of vast numbers of biological variables offer the potential for highly discriminatory diagnosis and prognosis; however, high dimensionality together with small samples creates the need for feature selection, while at the same time making feature-selection algorithms less reliable. Feature selection must typically be carried out from among thousands of gene-expression features and in the context of a small sample (small number of microarrays). Two basic questions arise: (1) Can one expect feature selection to yield a feature set whose error is close to that of an optimal feature set? (2) If a good feature set is not found, should it be expected that good feature sets do not exist? RESULTS: The two questions translate quantitatively into questions concerning conditional expectation. (1) Given the error of an optimal feature set, what is the conditionally expected error of the selected feature set? (2) Given the error of the selected feature set, what is the conditionally expected error of the optimal feature set? We address these questions using three classification rules (linear discriminant analysis, linear support vector machine and k-nearest-neighbor classification) and feature selection via sequential floating forward search and the t-test. We consider three feature-label models and patient data from a study concerning survival prognosis for breast cancer. With regard to the two focus questions, there is similarity across all experiments: (1) One cannot expect to find a feature set whose error is close to optimal, and (2) the inability to find a good feature set should not lead to the conclusion that good feature sets do not exist. In practice, the latter conclusion may be more immediately relevant, since when faced with the common occurrence that a feature set discovered from the data does not give satisfactory results, the experimenter can draw no conclusions regarding the existence or nonexistence of suitable feature sets. AVAILABILITY: http://ee.tamu.edu/~edward/feature_regression/

Artificial Intelligence↗