Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

SignalML: metaformat for description of biomedical time series.

This paper introduces a complete and elegant solution to the problem of inherent incompatibility of different formats used for digital storage of biomedical time series (in particular EEG) and their annotations. We define a simple XML-based language, in which information on the structure of binary data files can be simply and efficiently coded. In most cases, description of an existing format takes relatively few lines of XML code. Once written, this information can be used by any software, which, owing to this meta-description, may read the original data files, thus eliminating the need for conversions and duplication of data. This proposition is hereby submitted to an open discussion within the community involved in relevant research, clinical and commercial applications. Links to the current version of the XML Schema defining the language and pilot implementation of a compliant viewer/annotator are located at http://eeg.pl/SignalML/.

Database Management Systems↗

A low-cost digital filing system for echocardiography data with MPEG4 compression and its application to remote diagnosis.

The high cost of digital echocardiographs and the large size of data files hinder the adoption of remote diagnosis of digitized echocardiography data. We have developed a low-cost digital filing system for echocardiography data. In this system, data from a conventional analog echocardiograph are captured using a personal computer (PC) equipped with an analog-to-digital converter board. Motion picture data are promptly compressed using a moving pictures expert group (MPEG) 4 codec. The digitized data with preliminary reports obtained in a rural hospital are then sent to cardiologists at distant urban general hospitals via the internet. The cardiologists can evaluate the data using widely available movie-viewing software (Windows Media Player). The diagnostic accuracy of this double-check system was confirmed by comparison with ordinary super-VHS videotapes. We have demonstrated that digitization of echocardiography data from a conventional analog echocardiograph and MPEG 4 compression can be performed using an ordinary PC-based system, and that this system enables highly efficient digital storage and remote diagnosis at low cost.

Analog-Digital Conversion↗

Single data extraction generated more errors than double data extraction in systematic reviews.

BACKGROUND AND OBJECTIVE: To conduct a pilot study to compare the frequency of errors that accompany single vs. double data extraction, compare the estimate of treatment effect derived from these methods, and compare the time requirements for these methods. METHODS: Reviewers were randomized to the role of data extractor or data verifier, and were blind to the study hypothesis. The frequency of errors associated with each method of data extraction was compared using the McNemar test. The data set for each method was used to calculate an efficacy estimate by each method, using standard meta-analytic techniques. The time requirement for each method was compared using a paired t-test. RESULTS: Single data extraction resulted in more errors than double data extraction (relative difference: 21.7%, P = .019). There was no substantial difference between methods in effect estimates for most outcomes. The average time spent for single data extraction was less than the average time for double data extraction (relative difference: 36.1%, P = .003). CONCLUSION: In the case that single data extraction is used in systematic reviews, reviewers and readers need to be mindful of the possibility for more errors and the potential impact these errors may have on effect estimates.

Data Interpretation, Statistical↗

Patterns of memory impairment and perseverative behavior discriminate early Alzheimer's disease from subcortical vascular dementia.

Previous research suggests that the neuropsychological deficits in Alzheimer's disease (AD) are different from that of vascular dementia (VaD), especially with respect to memory, language and executive functions, but negative findings were reported. Our objective was to clarify the cognitive syndrome in AD and VaD in the early stage of these disorders. We investigated 45 patients with early AD, 23 patients with subcortical VaD and 35 normal controls. All subjects were assessed with neuropsychological battery designed to measure memory, language, praxis and executive functions. Patients with AD had significantly worse scores on Story Recall (p<0.02) and on all measures of the Free and Cued Selective Reminding Test (p<0.03 to 0.001) than did patients with VaD, as well as greater number of perseverations (p<0.02) on category fluency. Conversely, VaD patients had more perseverations (p<0.02) on the Modified Card Sorting Test. Despite the similar degree of overall cognitive deterioration, the findings show more impaired retrieval from long-term storage in AD than in VaD. Moreover, the data suggest that AD and subcortical VaD affect perseverative behavior in a different fashion. These results may be helpful in differentiating AD from VaD in the early stage of these disorders, when mental impairments are not pervasive yet.

Aged↗

Analysis of missing data in palliative care studies.

This report discusses the general problem of the analysis of data that could include missing values. In the palliative care setting, the data may not be missing at random, but instead be related to the outcome of interest, and therefore the use of standard statistical procedures may be problematic. This study summarizes differing results that were found when using three simple methods for estimating missing data in an example data set testing for differences in the use of morphine or methadone for relief of pain. Differences in the conclusions are discussed and recommendations are made to improve the reporting of studies with missing data.

Analgesics↗

Methodological issues arising from systematic reviews of the evidence of safety of vaccines.

Adaptations to the recognized methods of systematic reviewing are required when addressing questions about safety, particularly about rare and/or long-term serious adverse events. Conducting a systematic review of vaccine safety requires the implementation of novel strategies for locating studies, the use of experimental instruments to assess the quality of non-randomized studies, and the employment of pooling methods for non-randomized data, where appropriate. Standardizing both the indexing of adverse event data on electronic libraries and their reporting would improve the potential of systematic reviews of vaccine to draw accurate conclusions about the safety of a vaccine.

Data Interpretation, Statistical↗

Probability-based validation of protein identifications using a modified SEQUEST algorithm.

Database-searching algorithms compatible with shotgun proteomics match a peptide tandem mass spectrum to a predicted mass spectrum for an amino acid sequence within a database. SEQUEST is one of the most common software algorithms used for the analysis of peptide tandem mass spectra by using a cross-correlation (XCorr) scoring routine to match tandem mass spectra to model spectra derived from peptide sequences. To assess a match, SEQUEST uses the difference between the first- and second-ranked sequences (ACn). This value is dependent on the database size, search parameters, and sequence homologies. In this report, we demonstrate the use of a scoring routine (SEQUEST-NORM) that normalizes XCorr values to be independent of peptide size and the database used to perform the search. This new scoring routine is used to objectively calculate the percent confidence of protein identifications and posttranslational modifications based solely on the XCorr value.

Algorithms↗

IR-MALDI-mass analysis of electroblotted proteins directly from the membrane: comparison of different membranes, application to on-membrane digestion, and protein identification by database searching.

A systematic membrane study investigating different neutral, cationic derivatized, and hydrophilic PVDF membranes for their suitability to carry out on-membrane tryptic digestions and to obtain infrared-matrix-assisted laser desorption/ionization (IR-MALDI) mass information on the proteolytic fragments directly from the membrane was performed. Clearly, the Immobilon CD membrane (Millipore) showed the most reproducible results over a protein mass range from 12 to 66 kDa. Typical protein load to SDS-PAGE was in the 1-2 micrograms range. The protein amount used for enzymatic treatment was estimated to be in the low picomole range. Now both the intact protein mass and the masses of the specific proteolytic fragments are available directly from the membrane. Protein databases can be searched via search algorithms on the Internet using the information on the intact protein mass and the masses, e.g., of its tryptic fragments. Investigations were performed to search for neutral, enzyme-compatible IR matrixes which allow the enzymatic treatment (on-membrane digestion) while the membrane is matrix-incubated. Thiourea could be tolerated during enzymatic cleavage in solution in concentrations of 15 g/L and resulted in high-quality spectra of intact protein signals and turned, therefore, out to be the most promising candidate.

Algorithms↗

Are we stuck in the standards?

To avoid duplication of effort, slow adoption and inefficiency in development, those developing biological standards need to communicate more with each other, attract help from experts in the ontology/standards communities and keep focused on needs of users.

Biotechnology↗

A wavelet-packets based algorithm for EEG signal compression.

Transmission of biomedical signals through communication channels is being used increasingly in clinical practice. This technique requires dealing with large volumes of information, and the electroencephalographic (EEG) signal is an example of this situation. In the EEG, various channels are recorded during several hours, resulting in a great demand of storage capacity or channel bandwidth. This situation demands the use of efficient data compression systems. The objective of this work was to develop an efficient algorithm for EEG lossy compression. In this algorithm, the EEG signal is segmented and then decomposed through Wavelet Packets (WP). The WP decomposition coefficients are thresholded and those having absolute values below the threshold are deleted. The remaining coefficients are appropriately quantized and coded using a run-length coding scheme. The compressed EEG signal can be recovered by an inverse process. Extensive experimental tests were made by applying the algorithm to EEG records and measuring the compression rate (CR) and the distortion in signal segments. The WP transform showed a high robustness, allowing a reasonably low distortion after a compression-decompression process, for CR typically in the range 5-8. The algorithm has a relatively low computational cost, making it appropriate for practical applications.

Algorithms↗

A simple audio data logger for objective assessment of snoring in the home.

We have developed a portable device for patient use in logging snoring loudness in the home, for guiding treatment decisions and measuring the clinical effectiveness of treatment. The device uses a free field microphone and is positioned on a bedside table. The prototype devices contain no inherently expensive components and are simple to operate (producing only 5% patient error to date). They are portable, battery powered, rugged and produce digital data which are easily and automatically analysed, and these design parameters enable the devices to be used for first line patient assessment. Of the 75 recordings made so far from 30 patients, 85% were successful, yielding clinically useful data. Because it is sound levels which are recorded and not replayable sounds, patient privacy is maintained, resulting in excellent patient acceptance (to date no patient has refused). The device has a dynamic range of 45-90 dB sound pressure level and a frequency range of 30 Hz-5 kHz. Because snoring intensities often vary significantly throughout the night the device can measure continuously over 8 h.

Calibration↗

Data standards: a call to action.

Access to data is something that every molecular biologist takes for granted nowadays, but data alone is of little use unless it is made available in a useable form through the development and global uptake of data standards. The challenge of standards development has been taken up by grass-roots movements working within several different branches of the biomedical research community. Many of these initiatives are proving extremely successful; for example, the Gene Ontology, which provides a controlled vocabulary for describing the properties of gene products, the Microarray Gene Expression Data Society's standards for describing microarray experiments, and the emerging standards developed by the Proteomics Standards Initiative are gaining broad acceptance. Standards development now faces its greatest ever challenge--the integration of diverse data types to fulfill the goals of systems biology. Now is the time for the communities that are developing these standards, the funding bodies that have invested so heavily in high-throughput data generation, and the publishers of biomedical research papers to cooperate fully to make the goals of integrated data analysis a reality.

Animals↗

Browsing protein families via the 'Rich Family Description' format.

MOTIVATION: Multiple alignments of protein sequences are the basis of structural and functional analysis of protein families. It is however difficult even for an expert biologist to comprehend an alignment of more than 50 to 100 homologous sequences. RESULTS: This paper presents a browser for the analysis of multiple alignments of large numbers of protein sequences. Phylogenetic trees and consensus sequences are computed and used to summarise the alignments; these data are stored in a structure called Rich Family Description. Summary alignments and trees are displayed in HTML pages and can be developed or reduced by the user. This browser is used to display the ProDom domain families on the Web. Its zooming facilities allow extracting information from alignments of more than 1000 homologous sequences.

Algorithms↗

CpGProD: identifying CpG islands associated with transcription start sites in large genomic mammalian sequences.

RESULTS: CpGProD is an application for identifying mammalian promoter regions associated with CpG islands in large genomic sequences. Although it is strictly dedicated to this particular promoter class corresponding to approximately 50% of the genes, CpGProD exhibits a higher sensitivity and specificity than other tools used for promoter prediction. Notably, CpGProD uses different parameters according to species (human, mouse) studied. Moreover, CpGProD predicts the promoter orientation on the DNA strand. AVAILABILITY: http://pbil.univ-lyon1.fr/software/cpgprod.html SUPPLEMENTARY INFORMATION: http://pbil.univ-lyon1.fr/software/cpgprod.html

Animals↗

Search for structural similarity in proteins.

MOTIVATION: The expanding protein sequence and structure databases await methods allowing rapid similarity search. Geometric parameters-dihedral angle between two sequential peptide bond planes (V) and radius of curvature (R) as they appear in pentapeptide fragments in polypeptide chains-are proposed for use in evaluating structural similarity in proteins (VeaR). The parabolic (empirical) function expressing the radius of curvature's dependence on the V-angle in model polypeptides is altered in real proteins in a form characteristic for a particular protein. This can be used as a criterion for judging similarity. RESULTS: A structural comparison of proteins representing a wide spectrum of structures was assessed versus sequence similarity analysis based on the genetic semihomology algorithm. The term 'consensus structure', analogous to 'consensus sequence', was introduced for the serpine family. AVAILABILITY: Semihom-sequence comparison freely available on request from J. Leluk. VeaR-structural comparison freely available on request from I. Roterman.

Algorithms↗

Protein family annotation in a multiple alignment viewer.

SUMMARY: The Pfaat protein family alignment annotation tool is a Java-based multiple sequence alignment editor and viewer designed for protein family analysis. The application merges display features such as dendrograms, secondary and tertiary protein structure with SRS retrieval, subgroup comparison, and extensive user-annotation capabilities. AVAILABILITY: The program and source code are freely available from the authors under the GNU General Public License at http://www.pfizerdtc.com

Amino Acid Sequence↗

A computational pipeline for protein structure prediction and analysis at genome scale.

MOTIVATION: Experimental techniques alone cannot keep up with the production rate of protein sequences, while computational techniques for protein structure predictions have matured to such a level to provide reliable structural characterization of proteins at large scale. Integration of multiple computational tools for protein structure prediction can complement experimental techniques. RESULTS: We present an automated pipeline for protein structure prediction. The centerpiece of the pipeline is our threading-based protein structure prediction system PROSPECT. The pipeline consists of a dozen tools for identification of protein domains and signal peptide, protein triage to determine the protein type (membrane or globular), protein fold recognition, generation of atomic structural models, prediction result validation, etc. Different processing and prediction branches are determined automatically by a prediction pipeline manager based on identified characteristics of the protein. The pipeline has been implemented to run in a heterogeneous computational environment as a client/server system with a web interface. Genome-scale applications on Caenorhabditis elegans, Pyrococcus furiosus and three cyanobacterial genomes are presented. AVAILABILITY: The pipeline is available at http://compbio.ornl.gov/proteinpipeline/

Algorithms↗