Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

IPPRED: server for proteins interactions inference.

SUMMARY: IPPRED is a web based server to infer protein-protein interactions through homology search between candidate proteins and those described as interacting. This simple inference allows to propose or to validate potential interactions. AVAILABILITY: IPPRED is freely available at http://cbi.labri.fr/outils/ippred/.

Algorithms↗

Data-adaptive test statistics for microarray data.

MOTIVATION: An important task in microarray data analysis is the selection of genes that are differentially expressed between different tissue samples, such as healthy and diseased. However, microarray data contain an enormous number of dimensions (genes) and very few samples (arrays), a mismatch which poses fundamental statistical problems for the selection process that have defied easy resolution. RESULTS: In this paper, we present a novel approach to the selection of differentially expressed genes in which test statistics are learned from data using a simple notion of reproducibility in selection results as the learning criterion. Reproducibility, as we define it, can be computed without any knowledge of the 'ground-truth', but takes advantage of certain properties of microarray data to provide an asymptotically valid guide to expected loss under the true data-generating distribution. We are therefore able to indirectly minimize expected loss, and obtain results substantially more robust than conventional methods. We apply our method to simulated and oligonucleotide array data. AVAILABILITY: By request to the corresponding author.

Algorithms↗

Do you do text?

Explore the source record for details and available documents.

Algorithms↗

Improving missing value estimation in microarray data with gene ontology.

MOTIVATION: Gene expression microarray experiments produce datasets with frequent missing expression values. Accurate estimation of missing values is an important prerequisite for efficient data analysis as many statistical and machine learning techniques either require a complete dataset or their results are significantly dependent on the quality of such estimates. A limitation of the existing estimation methods for microarray data is that they use no external information but the estimation is based solely on the expression data. We hypothesized that utilizing a priori information on functional similarities available from public databases facilitates the missing value estimation. RESULTS: We investigated whether semantic similarity originating from gene ontology (GO) annotations could improve the selection of relevant genes for missing value estimation. The relative contribution of each information source was automatically estimated from the data using an adaptive weight selection procedure. Our experimental results in yeast cDNA microarray datasets indicated that by considering GO information in the k-nearest neighbor algorithm we can enhance its performance considerably, especially when the number of experimental conditions is small and the percentage of missing values is high. The increase of performance was less evident with a more sophisticated estimation method. We conclude that even a small proportion of annotated genes can provide improvements in data quality significant for the eventual interpretation of the microarray experiments. AVAILABILITY: Java and Matlab codes are available on request from the authors. SUPPLEMENTARY MATERIAL: Available online at http://users.utu.fi/jotatu/GOImpute.html.

Algorithms↗

Variance stabilization and normalization for one-color microarray data using a data-driven multiscale approach.

MOTIVATION: Many standard statistical techniques are effective on data that are normally distributed with constant variance. Microarray data typically violate these assumptions since they come from non-Gaussian distributions with a non-trivial mean-variance relationship. Several methods have been proposed that transform microarray data to stabilize variance and draw its distribution towards the Gaussian. Some methods, such as log or generalized log, rely on an underlying model for the data. Others, such as the spread-versus-level plot, do not. We propose an alternative data-driven multiscale approach, called the Data-Driven Haar-Fisz for microarrays (DDHFm) with replicates. DDHFm has the advantage of being 'distribution-free' in the sense that no parametric model for the underlying microarray data is required to be specified or estimated; hence, DDHFm can be applied very generally, not just to microarray data. RESULTS: DDHFm achieves very good variance stabilization of microarray data with replicates and produces transformed intensities that are approximately normally distributed. Simulation studies show that it performs better than other existing methods. Application of DDHFm to real one-color cDNA data validates these results. AVAILABILITY: The R package of the Data-Driven Haar-Fisz transform (DDHFm) for microarrays is available in Bioconductor and CRAN.

Algorithms↗

GECO--linear visualization for comparative genomics.

UNLABELLED: In order to understand and interpret phylogenetic and functional relationships between multiple prokaryotic species, qualitative and quantitative data must be correlated and displayed. GECO allows linear visualization of multiple genomes using a client/server based approach by dynamically creating .png- or .pdf-formatted images. It is able to display ortholog relations calculated using BLASTCLUST by color coding ortholog representations. Irregularities on the genomic level can be identified by anomalous G/C composition. Thus, this software will enable researchers to detect horizontally transferred genes, pseudogenes and insertions/deletions in related microbial genomes. AVAILABILITY: http://bioinfo.mikrobio.med.uni-giessen.de/geco2/GecoMainServlet

Algorithms↗

The Protein Data Bank and structural genomics.

The Protein Data Bank (PDB; http://www.pdb.org/) continues to be actively involved in various aspects of the informatics of structural genomics projects--developing and maintaining the Target Registration Database (TargetDB), organizing data dictionaries that will define the specification for the exchange and deposition of data with the structural genomics centers and creating software tools to capture data from standard structure determination applications.

Animals↗

Issues in Internet survey research among cancer patients.

Considering the increasing number of cancer patients who are online, it is clear that the Internet will provide an important research medium and/or setting for oncology nurses in the near future. Despite increasing Internet usage in nursing research and practice, issues in using the Internet among cancer patients as a research tool have rarely been explored and discussed. The purpose of the article is to propose future directions for Internet research among cancer patients based on discussions of practical issues raised in an Internet survey study among 40 online cancer patients. The issues raised through the research process include (a) ethical issues, (b) recruitment issues, (c) issues in Web site development and maintenance, and (d) data entry and analysis issues. On the basis of the discussions of these issues, some future directions for Internet survey studies are proposed, including dealing with ethical issues, getting computer expertise, using motivational strategies, and using national and international approaches.

Aged↗

A computerized system for control and management of radionuclide inventory: application in nuclear medicine.

An interactive computerized system for radioisotope management and instantaneous inventory is reported. The system is capable of handling operations such as filing, nuclear imaging and disposing of various radionuclides. All radiopharmaceutical transactions are achieved with the aid of a Prime 300 mini-computer of 192K words of high speed semi-conductor memory and over 120 mega bytes of disk storage. The system automatically corrects for the appropriate decay, monitors and updates the storage file after every subsequent study. The performed study is recorded in a special file, together with the time and data retrieved from the computer's real time clock at the time of the entry. The system provides an organized and complete bookkeeping of all records concerning radionuclide transactions. It is found to be simple, efficient, highly versatile, and drastically reduces the time of operation and errors in handling the radioisotope inventory.

Computers↗

Cine MPR: interactive multiplanar reformatting of four-dimensional cardiac data using hardware-accelerated texture mapping.

Four-dimensional (4-D) imaging to capture the three-dimensional (3-D) structure and motion of the heart in real time is an emerging trend. We present here our method of interactive multiplanar reformatting (MPR), i.e., the ability to visualize any chosen anatomical cross section of 4-D cardiac images and to change its orientation smoothly while maintaining the original heart motion. Continuous animation to show the time-varying 3-D geometry of the heart and smooth dynamic manipulation of the reformatted planes, as well as large image size (100-300 MB), make MPR challenging. Our solution exploits the hardware acceleration of 3-D texture mapping capability of high-end commercial PC graphics boards. Customization of volume subdivision and caching concepts to periodic cardiac data allows us to use this hardware effectively and efficiently. We are able to visualize and smoothly interact with real-time 3-D ultrasound cardiac images at the desired frame rate (25 Hz). The developed methods are applicable to MPR of one or more 3-D and 4-D medical images, including 4-D cardiac images collected in a gated fashion.

Computing Methodologies↗

Style context with second-order statistics.

Patterns often occur as homogeneous groups or fields generated by the same source. In multisource recognition problems, such isogeny induces statistical dependencies between patterns (termed style context). We model these dependencies by second-order statistics and formulate the optimal classifier for normally distributed styles. We show that model parameters estimated only from pairs of classes suffice to train classifiers for any test field length. Although computationally expensive, the style-conscious classifier reduces the field error rate by up to 20 percent on quadruples of handwritten digits from standard NIST data sets.

Algorithms↗

An integration of online and pseudo-online information for cursive word recognition.

In this paper, we present a novel method to extract stroke order independent information from online data. This information, which we term pseudo-online, conveys relevant information on the offline representation of the word. Based on this information, a combination of classification decisions from online and pseudo-online cursive word recognizers is performed to improve the recognition of online cursive words. One of the most valuable aspects of this approach with respect to similar methods that combine online and offline classifiers for word recognition is that the pseudo-online representation is similar to the online signal and, hence, word recognition is based on a single engine. Results demonstrate that the pseudo-online representation is useful as the combination of classifiers perform better than those based solely on pure online information.

Algorithms↗

A vertical-energy-thresholding procedure for data reduction with multiple complex curves.

Due to the development of sensing and computer technology, measurements of many process variables are available in current manufacturing processes. It is very challenging, however, to process a large amount of information in a limited time in order to make decisions about the health of the processes and products. This paper develops a "preprocessing" procedure for multiple sets of complicated functional data in order to reduce the data size for supporting timely decision analyses. The data type studied has been used for fault detection, root-cause analysis, and quality improvement in such engineering applications as automobile and semiconductor manufacturing and nanomachining processes. The proposed vertical-energy-thresholding (VET) procedure balances the reconstruction error against data-reduction efficiency so that it is effective in capturing key patterns in the multiple data signals. The selected wavelet coefficients are treated as the "reduced-size" data in subsequent analyses for decision making. This enhances the ability of the existing statistical and machine-learning procedures to handle high-dimensional functional data. A few real-life examples demonstrate the effectiveness of our proposed procedure compared to several ad hoc techniques extended from single-curve-based data modeling and denoising procedures.

Algorithms↗

Data streaming in telepresence environments.

In this paper, we discuss data transmission in telepresence environments for collaborative virtual reality applications. We analyze data streams in the context of networked virtual environments and classify them according to their traffic characteristics. Special emphasis is put on geometry-enhanced (3D) video. We review architectures for real-time 3D video pipelines and derive theoretical bounds on the minimal system latency as a function of the transmission and processing delays. Furthermore, we discuss bandwidth issues of differential update coding for 3D video. In our telepresence system-the blue-c-we use a point-based 3D video technology which allows for differentially encoded 3D representations of human users. While we discuss the considerations which lead to the design of our three-stage 3D video pipeline, we also elucidate some critical implementation details regarding decoupling of acquisition, processing and rendering frame rates, and audio/video synchronization. Finally, we demonstrate the communication and networking features of the blue-c system in its full deployment. We show how the system can possibly be controlled to face processing or networking bottlenecks by adapting the multiple system components like audio, application data, and 3D video.

Computer Communication Networks↗

Population-based linkage of health records in Western Australia: development of a health services research linked database.

OBJECTIVES: To introduce the Western Australian Health Services Research Linked Database as infrastructure to support aetiologic, utilisation and outcomes research. To compare the study population, data resources, technical systems and organisational supports with international best practice in record linkage and health research. METHOD AND RESULTS: The WA Linked Database systematically links the available administrative health data within an Australian State of 1.7 million people. It brings together, initially, six core data elements (birth records, midwives' notifications, cancer registrations, in-patient hospital morbidity, in-patient and public out-patient mental health services data and death records). It will be updated regularly and is designed, in future extensions, to include data on primary, residential and domiciliary care and health surveys. Linkage uses probabilistic matching of patient names and other identifiers. Geocodes for spatial analysis are assigned using address linkage and mapping software. By June 1997, the project had taken 2 1/2 years to develop the system and link seven million core data records from 1980 to 1995. CONCLUSIONS: The system is consistent with international benchmarks, from four centres of excellence, for the study population, core datasets, matching and geocoding, and collaborative networks. There are prospects to redress deficiencies in primary medical contact and other data resources, validation studies, tracing systems and a more supportive legal framework. IMPLICATIONS: The WA Linked Database will be used in combination with medical record audits to provide a comprehensive evaluation of health system performance.

Data Collection↗

How does your searching grow? A survey of search preferences and the use of optimal search strategies in the identification of qualitative research.

The objective was to gain an overview of researchers experiences of searching the literature, with particular reference to the use of optimal search strategies (OSSs) and searching for qualitative research studies. A 13-item semi-structured questionnaire investigating search behaviour was distributed to members of the Cochrane Qualitative Methods Network. Follow-up interviews were conducted with a subset of respondents to explore issues raised and clarify points of ambiguity. Findings were analysed using data reduction, data displays and verification techniques. Eighty-six per cent of distributed questionnaires were returned. All respondents reported searching electronic databases as part of their literature search, with 80% expressing a preference for searching alone or with colleagues. Forty-one per cent indicated that they consider a database search to be only one aspect of a comprehensive literature search. The rigour and availability of OSSs was a concern for 30% of respondents. Twenty-five per cent of respondents had searched for qualitative studies, although the difficulty of locating this type of literature was considered problematic because of the varied use of the term 'qualitative'. Whilst the majority of respondents reported using OSSs in some capacity, reservations were expressed about their ability to facilitate a comprehensive search. Replies indicated a belief that OSSs can reduce the sensitivity of a search, and might limit the breadth of coverage required. A greater appreciation of the availability and purpose of OSSs-including the ability to optimize either sensitivity or recall-is needed if this enhanced approach to accurate data retrieval and potential improvement in time management is to become widespread.

Information Storage and Retrieval↗