Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

DiscoverySpace: an interactive data analysis application.

DiscoverySpace is a graphical application for bioinformatics data analysis. Users can seamlessly traverse references between biological databases and draw together annotations in an intuitive tabular interface. Datasets can be compared using a suite of novel tools to aid in the identification of significant patterns. DiscoverySpace is of broad utility and its particular strength is in the analysis of serial analysis of gene expression (SAGE) data. The application is freely available online.

Animals↗

Interpretable gene expression classifier with an accurate and compact fuzzy rule base for microarray data analysis.

An accurate classifier with linguistic interpretability using a small number of relevant genes is beneficial to microarray data analysis and development of inexpensive diagnostic tests. Several frequently used techniques for designing classifiers of microarray data, such as support vector machine, neural networks, k-nearest neighbor, and logistic regression model, suffer from low interpretabilities. This paper proposes an interpretable gene expression classifier (named iGEC) with an accurate and compact fuzzy rule base for microarray data analysis. The design of iGEC has three objectives to be simultaneously optimized: maximal classification accuracy, minimal number of rules, and minimal number of used genes. An "intelligent" genetic algorithm IGA is used to efficiently solve the design problem with a large number of tuning parameters. The performance of iGEC is evaluated using eight commonly-used data sets. It is shown that iGEC has an accurate, concise, and interpretable rule base (1.1 rules per class) on average in terms of test classification accuracy (87.9%), rule number (3.9), and used gene number (5.0). Moreover, iGEC not only has better performance than the existing fuzzy rule-based classifier in terms of the above-mentioned objectives, but also is more accurate than some existing non-rule-based classifiers.

Algorithms↗

Interpretation of analytical data on n-alkanes and polynuclear aromatic hydrocarbons in Arbacia lixula from the coasts of Tenerife (Canary Islands, Spain) by multivariate data analysis.

The hydrocarbons contents (n-alkanes, polycyclic aromatic hydrocarbons) were determined in the sea urchin Arbacia lixula. Multivariate data analysis as principal component analysis, factor analysis and, cluster analysis were applied to elucidate sources of pollution. PCA and FA were performed to establish the relationships between variables (hydrocarbons), samples (sea urchin) and sources of pollution.

Alkanes↗

[A modular method for automated evaluation of gait analysis data].

A modular methodology for automated gait data evaluation: The aim of Instrumented Gait Analysis is to measure data such as joint kinematics or kinetics during gait in a quantitative way. The data evaluation for clinical purposes is often performed by experienced physicians (diagnosis of specific motion dysfunction, planning and validation of therapy). Due to subjective evaluation and complexity of the pathologies, there exists no objective, standardized data analysis method for these tasks. This article covers the development of a modular, computer-based methodology to quantify the degree of pathological gait in comparison to normal behavior, as well as to automatically search for interpretable gait abnormalities and to visualize the results. The outcomes are demonstrated with two different patient groups.

Biomechanical Phenomena↗

Mechanisms of hydrazine toxicity in rat liver investigated by proteomics and multivariate data analysis.

A proteomics approach combined with multivariate data analysis was used to examine the hepatotoxic effect of hydrazine in 30 male Sprague Dawley rats, assigned to four treatment groups and two control groups. Liver samples from the individual animals were resolved by two-dimensional differential gel electrophoresis (2-D DIGE) and protein patterns from the 2-D gels were analyzed by principal component analysis (PCA) and partial least squares regression (PLSR). The PCA plot was able to describe the variation in the protein expression related to dose and time, by separation or clustering of different animal groups. PLSR followed by variable selection (Jack-knifing) was used to select proteins that varied significantly in relation to the dose related response of the hydrazine treatment. The 10 up-regulated and 10 down-regulated proteins with highest rank in the PLSR model were identified by mass spectrometry. Hydrazine treatment induced altered expression of proteins related to lipid metabolism, Ca(2+) homeostasis, thyroid hormone pathways and stress response. Several of the identified proteins have not previously been implicated in hydrazine toxicity and may thus be regarded as new potential biomarkers of induced liver toxicity.

Animals↗

Web-based MS/MS data analysis.

This tutorial focuses on three MS/MS data analysis programs currently available via a web interface: Mascot, Phenyx and X!Tandem. Although these programs process the same input and often produce comparable outputs, subtle differences remain. The use of parameters that are requested in the on-line forms and the subsequent interpretation of results are illustrated and explained via a single example.

Amino Acids↗

Consolidation of common parameters from multiple fits in dynamic PET data analysis.

In dynamic positron emission tomography (PET) data analysis, regions of interest (ROI's) are analyzed by fitting a parametric model to the time-activity curve acquired after a radio-labeled tracer has been introduced into the patient's bloodstream. This procedure can be carried out for multiple ROI's and/or multiple injections of the same or a different radiopharmaceutical. The approach presented here takes advantage of prior knowledge that some of the parameters of those multiple fits are the same. Reduction of the total number of parameters to be estimated results in smaller statistical uncertainty for all parameter estimates, especially those common to multiple fits.

Algorithms↗

User evaluation of an integrated medical workstation for clinical data analysis.

Results are presented of the user evaluation of an integrated medical workstation for support of clinical research. Twenty-seven users were recruited from medical and scientific staff of the University Hospital Dijkzigt, the Faculty of Medicine of the Erasmus University Rotterdam, and from other Dutch medical institutions; and all were given a written, self-contained tutorial. Subsequently, an experiment was done in which six clinical data analysis problems had to be solved and an evaluation form was filled out. The aim of this user evaluation was to obtain insight in the benefits of integration for support of clinical data analysis for clinicians and biomedical researchers. The problems were divided into two sets, with gradually more complex problems. In the first set users were guided in a stepwise fashion to solve the problems. In the second set each stepwise problem had an open counterpart. During the evaluation, the workstation continuously recorded the user's actions. From these results significant differences became apparent between clinicians and non-clinicians for the correctness (means 54% and 81%, respectively, p = 0.04), completeness (means 64% and 88%, respectively, p = 0.01), and number of problems solved (means 67% and 90%, respectively, p = 0.02). These differences were absent for the stepwise problems. Physicians tend to skip more problems than biomedical researchers. No statistically significant differences were found between users with and without clinical data analysis experience, for correctness (means 74% and 72%, respectively, p = 0.95), and completeness (means 82% and 79%, respectively, p = 0.40).(ABSTRACT TRUNCATED AT 250 WORDS)

Computer Systems↗

Treating expression levels of different genes as a sample in microarray data analysis: is it worth a risk?

One of the prevailing ideas in the literature on microarray data analysis is to pool the expression measures across genes and treat them as a sample drawn from some distribution. Several universal laws were proposed to analytically describe this distribution. This idea raises a number of concerns. The expression levels of genes are not identically distributed random variables so that treating them as a sample amounts to sampling from a mixture of equally weighted distributions, each being associated with a different gene. The expression levels of different genes are heavily dependent random variables so that the law of large numbers and statistical goodness-of-fit tests are normally inapplicable to this kind of data. This dependence represents a very serious pitfall in microarray data analysis.

Gene Expression Profiling↗

[Control survey data analysis with smoothed distribution curve].

Assay data and normal range collected in quality control surveys were analyzed with a smoothed distribution curve, in order to perform appropriate analysis of the data. In the control survey, the committee set a relative allowable limit for each test. A variable was changed continuously through the data range, and the data within the allowable limit were counted. In the computer program, the variable was changed step wise in logarithmic scale. These data were plotted on a graph with log-scale on X-axis and per cent of relative data in Y-axis. The graph thus obtained shows smoother curve than the histogram of the data, and Y-axis scale shows actual data within the allowable data at the variable scale. Normal ranges and assay data of several chemistry assays were analyzed with this method. Distribution of normal range of enzyme assays showed plural peaks in each test. Especially, in choline esterase assay, more than ten different units were used in laboratories in Japan with the same unit name, IU/1. In, amylase assay, nevertheless the normal range was separated in three peaks, the data of sample assay did not separate in peaks correspond to the normal range peaks. Thus, the smoothed distributions curve is useful for data analysis in quality control survey.

Amylases↗

Fluorescence quenching data interpretation in biological systems. The use of microscopic models for data analysis and interpretation of complex systems.

In micro-heterogeneous media (e.g. membranes, micelles and colloidal systems), the fluorescence decay in the absence of quencher is usually intrinsically complex, e.g. due to the existence of several sub-populations with different micro-environments. In this case it is impossible to analyze data in detail (accounting for transient effects) and simpler formalisms are needed. The objective of the present work is to present and discuss such simpler formalisms. The goal is to achieve simple data analysis and meaningful, clear data interpretation in complex systems using microscopic models that consider several sub-populations of chromophores. Two points are dealt with in detail. (i) It is shown that the approximation of the transient effects by the quenching sphere-of-action model is not always possible. The quenching sphere-of-action concept can be regarded as a valuable tool, although crude, only in a limited range of experimental conditions, namely time resolution. (ii) The Stern-Volmer equation usually used for data analysis is only valid for a limited range of small and moderate equilibrium association constants, Ka, although this is frequently overlooked in the literature. Self-consistency criteria are presented for the proposed methods. The well-known downward curvature due to a fraction of fluorophores which is not accessible to the quencher is only a limiting case from a set of possible situations which result in deviations to linearity. A systematic classification of the different types of quenching is presented.

Data Interpretation, Statistical↗

Planning controlled clinical trials on the basis of descriptive data analysis.

In controlled clinical trials the problem of multiplicity of desired inferential statements finds attention at an increasing rate. In this paper the previously proposed concept of Descriptive Data Analysis (DDA), situated between Confirmatory and Exploratory Data Analysis, is applied to the planning aspects of controlled trials for which the problem of multiplicity exists. The (non-Bayesian) DDA planning concept should provide the investigator with tools to draw final conclusions from data of several variables possibly observed at several time points in possibly several groups of subjects by combining his pre-trial medical experience with descriptive inferential statements (confidence intervals and test results) at nominal significance levels. DDA also provides for confirmatory statements concerning individual null hypotheses and partially global hypotheses.

Clinical Trials as Topic↗

The influence of the method of data analysis on the reported accuracy of automated blood pressure measuring devices.

OBJECTIVE: To show that different methods of data analysis affect the grading that blood pressure measuring devices achieve according to the British Hypertension Society (BHS)-protocol. METHODS: Based on the somewhat unclear description of the exact method of data analysis in the BHS-protocol four different methods can be discerned. The effect on the grading-results is calculated for these four different options. RESULTS AND CONCLUSIONS: It is shown that using these four different options the achieved grade can range for diastolic blood pressure from C (option 1) to almost A (option 4) and for systolic blood pressure from D (option 1) to B (option 4). Different researchers may well have used different methods. Option 1 is the method that should be used. Also it is stated that the systematic error and the standard deviation of differences (SDD) are measures that give more insight to describe a device's performance. Calculating the grades after correction for the systematic error shows its influence and that of the SDD on the reported accuracy of a blood pressure measuring device.

Automation↗

MALDI-MS data analysis for disease biomarker discovery.

In this chapter, we address the issue of matrix-assisted laser desorption/ionization mass spectrometry (MS) data analysis for disease biomarker discovery. We first give a general framework of MS data analysis, then focus on several key steps. After that, we show some application examples using an ovarian sera cancer dataset. Finally, we discuss the limitations of current approaches and possible future research directions.

Animals↗

Time normalization of voice signals using functional data analysis.

The harmonics-to-noise ratio (HNR) has been used to quantify the waveform irregularity of voice signals [Yumoto et al., J. Acoust. Soc. Am. 71, 1544-1550 (1982)]. This measure assumes that the signal consists of two components: a harmonic component, which is the common pattern that repeats from cycle-to-cycle, and an additive noise component, which produces the cycle-to-cycle irregularity. It has been shown [J. Qi, J. Acoust. Soc. Am. 92, 2569-2576 (1992)] that a valid computation of the HNR requires a nonlinear time normalization of the cycle wavelets to remove phase differences between them. This paper shows the application of functional data analysis to perform an optimal nonlinear normalization and compute the HNR of voice signals. Results obtained for the same signals using zero-padding, linear normalization, and dynamic programming algorithms are presented for comparison. Functional data analysis offers certain advantages over other approaches: it preserves meaningful features of signal shape, produces differentiable results, and allows flexibility in selecting the optimization criteria for the wavelet alignment. An extension of the technique for the time normalization of simultaneous voice signals (such as acoustic, EGG, and airflow signals) is also shown. The general purpose of this article is to illustrate the potential of functional data analysis as a powerful analytical tool for studying aspects of the voice production process.

Data Interpretation, Statistical↗

A method for quantification of absolute amounts of nucleic acids by (RT)-PCR and a new mathematical model for data analysis.

Accurate quantification of nucleic acids by competitive (RT)-PCR requires a valid internal standard, a reference for data normalization and an adequate mathematical model for data analysis. We report here an effective procedure for the generation of homologous RNA internal standards and a strategy for synthesizing and using a reference target RNA in quantification of absolute amounts of nucleic acids. Further, a new mathematical model describing the general kinetic features of competitive PCR was developed. The model extends the validity of quantitative competitive (RT)-PCR beyond the exponential phase. The new method eliminates the errors arising from different amplification efficiencies of the co-amplified sequences and from heteroduplex formation in the system. The high accuracy (relative error <2%) is comparable to the recently developed real time detection 5'-nuclease PCR. Also, corresponding computer software has been devised for practical data analysis.

Cell Line↗

Microarray data analysis: from hypotheses to conclusions using gene expression data.

We review several commonly used methods for the design and analysis of microarray data. To begin with, some experimental design issues are addressed. Several approaches for pre-processing the data (filtering and normalization) before the statistical analysis stage are then discussed. A common first step in this type of analysis is gene selection based on statistical testing. Two approaches, permutation and model-based methods are explained and we emphasize the need to correct for multiple testing. Moreover, powerful approaches based on gene sets are mentioned. Clustering of either genes or samples is frequently performed when analyzing microarray data. We summarize the basics of both supervised and unsupervised clustering (classification). The latter may be of use for creating diagnostic arrays, for example. Construction of biological networks, such as pathways, is a statistically challenging but complex task that is a relatively new development and hence mentioned only briefly. We finish with some remarks on literature and software. The emphasis in this paper is on the philosophy behind several statistical issues and on a critical interpretation of microarray related analysis methods.

Cluster Analysis↗