Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Usefulness of imputation for the analysis of incomplete otoneurologic data.

The usefulness of imputation in the treatment of missing values of an otoneurologic database for the discriminant analysis was evaluated on the basis of the agreement of imputed values and the analysis results. The data consisted of six patient groups with vertigo (N=564). There were 38 variables and 11% of the data was missing. Missing values were filled in with the means, regression and Expectation-Maximisation (EM) imputation methods and a random imputation method provided the baseline results. Means, regression and EM methods agreed on 41-42% of the imputed missing values. The level of agreement between these and the random method was 20-22%. Despite the moderate agreement between the means, regression and EM methods, the discriminant functions were similar and accurate (prediction accuracy 83-99%). The discriminant functions obtained from the randomly imputed data were also accurate having prediction accuracy 88-97%. Imputation seems to be a useful method for treating the missing data in this database. However, a lot of data was missing in otoneurologic tests, which are likely to be of less importance in the diagnosis of vertiginous patients. Consequently, the disagreement of the methods did not affect clearly the discriminant analysis, and, therefore, future research requires more complete data and advanced imputation methods.

Data Collection↗

Health information on the Internet and people living with HIV/AIDS: information evaluation and coping styles.

Individuals who seek information on the Internet to cope with chronic illness may be vulnerable to misinformation and unfounded claims. This study examined the association between health-related coping and the evaluation of health information. Men (n = 347) and women (n = 72) who were living with HIV/AIDS and reported currently using the Internet completed measures assessing their Internet use. Health Web sites downloaded from the Internet were also rated for quality of information. HIV-positive adults commonly used the Internet to find health information (66%) and to learn about clinical trials (25%); they also talked to their physicians about information found online (24%). In a multivariate analysis, assigning higher credibility to unfounded Internet information was predicted by lower incomes, less education, and avoidant coping styles. People who cope by avoiding health information may be vulnerable to misinformation and unfounded claims that are commonly encountered on the Internet.

Adaptation, Psychological↗

Research in primary dental care. Part 2: Developing a research question.

The first step in planning and conducting any research is identifying the research question, that is a testable statement of the question which the research aims to answer. In this article three distinct types of research question are identified: Descriptive questions (for example Who? What? Where? When?); Questions of relationships (How are two or more things related?); Questions of comparison (often these questions will ask about cause and effect). Examples are given of each type of research question. The process of devising a research question is described, in particular searching for relevant information, and evaluating the quality of the information obtained. A list of useful resources is provided.

Data Collection↗

The Blood Stocks Management Scheme, a partnership venture between the National Blood Service of England and North Wales and participating hospitals for maximizing blood supply chain management.

BACKGROUND AND OBJECTIVES: The Blood Stocks Management Scheme (BSMS) has been established as a joint venture between the National Blood Service (NBS) in England and North Wales and participating hospitals to monitor the blood supply chain. MATERIALS AND METHODS: Stock and wastage data are submitted to a web-based data-management system, facilitating continuous and complete red cell data collection and 'real time' data extraction. RESULTS: The data-management system enables peer review of performance in respect of stock holding levels and red cell wastage. CONCLUSIONS: The BSMS has developed an innovative web-based data-management system that enables data collection and benchmarking of practice, which should drive changes in stock management practice, therefore optimizing the use of donated blood.

Blood Banks↗

Quantitative data generation for systems biology: the impact of randomisation, calibrators and normalisers.

Systems biology is an approach to the analysis and prediction of the dynamic behaviour of biological networks through mathematical modelling based on experimental data. The current lack of reliable quantitative data, especially in the field of signal transduction, means that new methodologies in data acquisition and processing are needed. Here, we present methods to advance the established techniques of immunoprecipitation and immunoblotting to more accurate and quantitative procedures. We propose randomisation of sample loading to disrupt lane correlations and the use of normalisers and calibrators for data correction. To predict the impact of each method on improving the data quality we used simulations. These studies showed that randomisation reduces the standard deviation of a smoothed signal by 55% +/- 10%, independently from most experimental settings. Normalisation with appropriate endogenous or external proteins further reduces the deviation from the true values. As the improvement strongly depends on the quality of the normaliser measurement, a criteria-based normalisation procedure was developed. Our approach was experimentally verified by application of the proposed methods to time course data obtained by the immunoblotting technique. This analysis showed that the procedure is robust and can significantly improve the quality of experimental data.

Algorithms↗

Data management in practice-based research.

OBJECTIVE: Multi-site data collection is complex and requires an effective data management system. This article explores data management issues encountered in the design, conduct, and analysis of a research project involving 74 community-based sites and a central data management system. RESULTS: Once the data arrived at the central site, data integrity was maintained at a very high level. Issues encountered in our study on low back pain reflected the practice-based nature of the study and the limitations of finances, staff, and facilities. CONCLUSION: The task of converting a research protocol to actual procedures for data collection and data management can be very challenging. The importance of early recognition of the effort and resources needed for data management and quality-control procedures cannot be overestimated.

Community Health Centers↗

Data mining and knowledge discovery in predictive toxicology.

This article describes the knowledge discovery process in predictive toxicology. This process consists of five major steps (i) feature calculation, (ii) feature selection, (iii) model induction, (iv) model validation and (v) interpretation of predictions and models. Data mining is a part of the knowledge discovery process and consists of the application of data analysis and discovery algorithms, which can be useful in all of the above steps. A brief review of suitable algorithms and their advantages and disadvantages is given for each knowledge discovery step, followed by a more detailed description of a problem-specific implementation of the lazar prediction system.

Algorithms↗

Evaluation of the usefulness of Internet searches to identify unpublished clinical trials for systematic reviews.

PRIMARY OBJECTIVE: To avoid selection and publication bias, systematic reviewers should employ a broad range of search techniques and make efforts to locate unpublished studies. We tried to establish whether searches on the World Wide Web (WWW) are useful to identify additional unpublished and ongoing clinical trials. RESEARCH DESIGN: Search strategies seven Cochrane systematic reviews were retrospectively adapted for the WWW in an attempt to find additional randomized controlled trials. METHODS AND PROCEDURES: A search strategy with the general pattern 'study methodology NEAR intervention NEAR condition' for the Internet search engine AltaVista was evaluated by measuring search time, recall of Internet searches for published studies; precision (proportion of webpages containing hints to relevant published and unpublished randomized clinical trials); number of additional unpublished or ongoing studies found on the Internet. MAIN OUTCOMES AND RESULTS: We reviewed 429 webpages in 21 hours and found hints to 14 unpublished, ongoing or recently finished trials, at least 9 were considered relevant for 4 systematic reviews. The recall of Internet searches to find references to published studies ranged between 0% and 43.6%, the precision for hints to published or unpublished studies range between 0% and 20.2%. CONCLUSIONS: Information on unpublished and particularly ongoing trials can be found on the Internet. A potential problem is the appraisal of non-peer reviewed electronic publications with questionable quality. More powerful search tools are needed. An 'Open Trial Initiative' is proposed to define a syntax for publishing trials on the web and to ensure interoperability of trial registers, so that special search engines can harvest information on ongoing and complete clinical trials.

Data Collection↗

The HIB database of annotated UniGene clusters.

SUMMARY: The HumanInfoBase (HIB) is a database of putative human gene transcripts. UniGene clusters are assembled, and the resulting consensus sequences are submitted to the PEDANT software system (Frishman,D., Albermann,K., Hani,J., Heumann,K., Metanomski,A., Zollner,A. and Mewes,H.-W., 2001, Bioinformatics, 17, 44--57) for fully automatic sequence analysis and annotation. Predicted transcripts are classified using a variety of functional and structural categories, and hyperlinks to various databases are provided for additional information. A WWW-based graphical user interface represents the assembly process as well as functionally important sites in the putative transcripts.

Data Collection↗

Ramachandran plot on the web.

A graphics package has been developed to display the main chain torsion angles phi, psi (phi, Psi); (Ramachandran angles) in a protein of known structure. In addition, the package calculates the Ramachandran angles at the central residue in the stretch of three amino acids having specified the flanking residue types. The package displays the Ramachandran angles along with a detailed analysis output. This software is incorporated with all the protein structures available in the Protein Databank.

Amino Acid Motifs↗

Fully automated ab initio protein structure prediction using I-SITES, HMMSTR and ROSETTA.

MOTIVATION: The Monte Carlo fragment insertion method for protein tertiary structure prediction (ROSETTA) of Baker and others, has been merged with the I-SITES library of sequence structure motifs and the HMMSTR model for local structure in proteins, to form a new public server for the ab initio prediction of protein structure. The server performs several tasks in addition to tertiary structure prediction, including a database search, amino acid profile generation, fragment structure prediction, and backbone angle and secondary structure prediction. Meeting reasonable service goals required improvements in the efficiency, in particular for the ROSETTA algorithm. RESULTS: The new server was used for blind predictions of 40 protein sequences as part of the CASP4 blind structure prediction experiment. The results for 31 of those predictions are presented here. 61% of the residues overall were found in topologically correct predictions, which are defined as fragments of 30 residues or more with a root-mean-square deviation in superimposed alpha carbons of less than 6A. HMMSTR 3-state secondary structure predictions were 73% correct overall. Tertiary structure predictions did not improve the accuracy of secondary structure prediction.

Algorithms↗

General framework for developing and evaluating database scoring algorithms using the TANDEM search engine.

MOTIVATION: Tandem mass spectrometry (MS/MS) identifies protein sequences using database search engines, at the core of which is a score that measures the similarity between peptide MS/MS spectra and a protein sequence database. The TANDEM application was developed as a freely available database search engine for the proteomics research community. To extend TANDEM as a platform for further research on developing improved database scoring methods, we modified the software to allow users to redefine the scoring function and replace the native TANDEM scoring function while leaving the remaining core application intact. Redefinition is performed at run time so multiple scoring functions are available to be selected and applied from a single search engine binary. We introduce the implementation of the pluggable scoring algorithm and also provide implementations of two TANDEM compatible scoring functions, one previously described scoring function compatible with PeptideProphet and one very simple scoring function that quantitative researchers may use to begin their development. This extension builds on the open-source TANDEM project and will facilitate research into and dissemination of novel algorithms for matching MS/MS spectra to peptide sequences. The pluggable scoring schema is also compatible with related search applications P3 and Hunter, which are part of the X! suite of database matching algorithms. The pluggable scores and the X! suite of applications are all written in C++. AVAILABILITY: Source code for the scoring functions is available from http://proteomics.fhcrc.org

Algorithms↗

Compilation of DNA sequences of Escherichia coli K12: description of the interactive databases ECD and ECDC.

We have compiled the DNA sequence data for Escherichia coli K12 available from the GenBank and EMBL data libraries and independently from the literature. We provide the most definitive version of the ECD Escherichia coli database now exclusively via the World Wide Web System (http://susi.bio.uni-giessen.de/ecdc.html ). Our database encloses the completed genome sequence recently published by two competing groups and an assembled set of all elder sequences. The organisation of the database allows precise physical location of each individual gene or regulatory region, even taking into consideration discrepancies in nomenclature. The WWW program allows to the user to branch into the original EMBL and SWISS-PROT datafiles. A number of links to other WWW servers dealing with E. coli is provided. A FASTA and BLAST search may be performed online. Besides the WWW format a flat file version may be obtained via ftp. A number of discrepancies between the two systematic sequence determinations and/or the literature have not yet been resolved. However, our database may serve as a reference source for resolution and/or the assignment of strain difference.

Computer Communication Networks↗