Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Using functional data analysis to summarise and interpret lactate curves.

John Tukey used the term exploratory data analysis (EDA) to describe a philosophy for analyzing data where graphical and numerical summaries are used to uncover interesting structures. The applied statistician today has a much more sophisticated set of methods to use when applying the EDA philosophy. One such collection of methods is functional data analysis (FDA), which was used to explore the structure of lactate curves. A principal components analysis and plots of the second derivatives provide new intuitive endurance markers which correlates highly with other numerical summaries of lactate curves that have been suggested in the literature.

Humans↗

A generalized additive model for microarray gene expression data analysis.

Microarray technology allows the measurement of expression levels of a large number of genes simultaneously. There are inherent biases in microarray data generated from an experiment. Various statistical methods have been proposed for data normalization and data analysis. This paper proposes a generalized additive model for the analysis of gene expression data. This model consists of two sub-models: a non-linear model and a linear model. We propose a two-step normalization algorithm to fit the two sub-models sequentially. The first step involves a non-parametric regression using lowess fits to adjust for non-linear systematic biases. The second step uses a linear ANOVA model to estimate the remaining effects including the interaction effect of genes and treatments, the effect of interest in a study. The proposed model is a generalization of the ANOVA model for microarray data analysis. We show correspondences between the lowess fit and the ANOVA model methods. The normalization procedure does not assume the majority of genes do not change their expression levels, and neither does it assume two channel intensities from the same spot are independent. The procedure can be applied to either one channel or two channel data from the experiments with multiple treatments or multiple nuisance factors. Two toxicogenomic experiment data sets and a simulated data set are used to contrast the proposed method with the commonly known lowess fit and ANOVA methods.

Algorithms↗

Uncovering psychiatric test information with graphical techniques of Exploratory Data Analysis.

This article illustrates how Exploratory Data Analysis (EDA) can complement conventional statistical methods in evaluating psychiatric tests. Using one recent EDA computer program, we evaluated the ability of repeated psychiatric screening tests (the General Health Questionnaire [GHQ]) to predict medical and psychiatric service use in a Health Maintenance Organization (HMO), the Harvard Community Health Plan (HCHP). Using a stratified random sample of 244 new HCHP enrollees and viewing three-dimensional graphs of their data from multiple perspectives, we found two subpopulations: low GHQ scorers, for whom the tests did not predict service use; and high scorers, for whom they did. Surprisingly, improving scores forecast increased use and chronically high scores predicted diminished use. Using another stratified random sample of 213 new HCHP enrollees, and with scatterplot matrices from another interactive computer program, we found that high and unchanging GHQ scores forecast HMO dropout. We examine possible interpretations--for example, that chronically distressed patients may become immobilized, diminish service use, and ultimately leave the HMO. We also explain how EDA methods may help uncover elusive results in other data (e.g., mental health outcomes).

Adult↗

Covariance models for nested repeated measures data: analysis of ovarian steroid secretion data.

We consider several covariance models for analysing repeated measures data from a study of ovarian steroid secretion in reproductive-aged women. Urinary oestradiol and serum oestrogen were repeatedly observed over three or four menstrual periods, each period separated by one year. For each menstrual period, daily first morning urine specimens were collected 8 to 18 times, and serum specimens 2 to 5 times. Thus, measurements were repeatedly observed over menstrual cycle days within menstrual periods. Owing to missing observations, the number of observations differed from subject to subject. In this study, there were two repeat factors: menstrual cycle day and menstrual period. The first repeat factor, cycle day, is nested within the second repeat factor, menstrual period. In analysing these nested repeated measures data, the correlation structure should be modelled that will account for both repeat factors. We present several covariance models for defining appropriate covariance structures for these data.

Adult↗

Regional homogeneity approach to fMRI data analysis.

Kendall's coefficient concordance (KCC) can measure the similarity of a number of time series. It has been used for purifying a given cluster in functional MRI (fMRI). In the present study, a new method was developed based on the regional homogeneity (ReHo), in which KCC was used to measure the similarity of the time series of a given voxel to those of its nearest neighbors in a voxel-wise way. Six healthy subjects performed left and right finger movement tasks in event-related design; five of them were additionally scanned in a rest condition. KCC was compared among the three conditions (left finger movement, right finger movement, and the rest). Results show that bilateral primary motor cortex (M1) had higher KCC in either left or right finger movement condition than in rest condition. Contrary to prediction and to activation pattern, KCC of ipsilateral M1 is significantly higher than contralateral M1 in unilateral finger movement conditions. These results support the previous electrophysiologic findings of increasing ipsilateral M1 excitation during unilateral movement. ReHo can consider as a complementary method to model-driven method, and it could help reveal the complexity of the human brain function. More work is needed to understand the neural mechanism underlying ReHo.

Adult↗

Automation of follow-up and data analysis of paediatric heart disease in Malta.

Widely available computer programs have been used to set up a database for patients with congenital heart disease in Malta. This database is used for clinical follow-up and research, and has been tailored to provide formatted output of specified results of follow-up as tables and graphs that are automatically updated with ongoing changes in the dataset. The system is easy to use, being menu- and icon driven and can be operated with minimal training. It has resulted in great saving of time not only in clinical practice, but also in the production of reports and analysis of data, as spreadsheets need only be created once and are then updated at will. The system also incorporates a patient summary generator and a on-screen picture library for patient explanation and teaching purposes.

Cardiac Catheterization↗

On the experimental design and data analysis of mutation accumulation experiments.

Characterizing deleterious genomic mutations is important. Most of the few current estimates come from the mutation-accumulation (M-A) approach, which has been extremely time- and labour-consuming. There is a resurgent interest in implementing this approach. However, its estimation properties under different experimental designs are poorly understood. By simulations we investigate these issues in detail. We found that many of the previous M-A experiments could have been more efficiently implemented with much less time and expense while still achieving the same estimation accuracy. If more than 100 lines are employed in M-A and if each line is replicated at least 10 times during each assay, an experiment of 10 M-A generations with two assays (at the beginning and at the end of M-A) may achieve at least the same estimation quality as a typical M-A experiment. The number of replicates per M-A line necessary for each assay largely depends on the magnitude of environmental variance. While 10 replicates are reasonable for assaying most fitness traits, many more are needed for viability, which has an exceptionally large environmental variance. The investigation is mainly carried out using Bateman-Mukai's method of moments for estimation. Estimation using Keightley's maximum likelihood is also investigated and discussed. These results should not only be useful for planning efficient M-A experiments, but also may help empiricists in deciding to adopt the M-A approach with manageable labour, time and resources.

Animals↗

Density of points clustering, application to transcriptomic data analysis.

With the increasing amount of data produced by high-throughput technologies in many fields of science, clustering has become an integral step in exploratory data analysis in order to group similar elements into classes. However, many clustering algorithms can only work properly if aided by human expertise. For example, one parameter which is crucial and often manually set is the number of clusters present in the analyzed set. We present a novel stopping rule to find the optimal number of clusters based on the comparison of the density of points inside the clusters and between them. The method is evaluated on synthetic as well as on real transcriptomic data and compared with two current methods. Finally, we illustrate its usefulness in the analysis of the expression profiles of promyelocytic cells before and after treatment with all-trans retinoic acid. Simultaneous clustering for gene regulation and absolute initial expression levels allowed the identification of numerous genes associated with signal transduction revealing the complexity of retinoic acid signaling.

Algorithms↗

ATHENA, ARTEMIS, HEPHAESTUS: data analysis for X-ray absorption spectroscopy using IFEFFIT.

A software package for the analysis of X-ray absorption spectroscopy (XAS) data is presented. This package is based on the IFEFFIT library of numerical and XAS algorithms and is written in the Perl programming language using the Perl/Tk graphics toolkit. The programs described here are: (i) ATHENA, a program for XAS data processing, (ii) ARTEMIS, a program for EXAFS data analysis using theoretical standards from FEFF and (iii) HEPHAESTUS, a collection of beamline utilities based on tables of atomic absorption data. These programs enable high-quality data analysis that is accessible to novices while still powerful enough to meet the demands of an expert practitioner. The programs run on all major computer platforms and are freely available under the terms of a free software license.

Algorithms↗

Computational cluster validation in post-genomic data analysis.

MOTIVATION: The discovery of novel biological knowledge from the ab initio analysis of post-genomic data relies upon the use of unsupervised processing methods, in particular clustering techniques. Much recent research in bioinformatics has therefore been focused on the transfer of clustering methods introduced in other scientific fields and on the development of novel algorithms specifically designed to tackle the challenges posed by post-genomic data. The partitions returned by a clustering algorithm are commonly validated using visual inspection and concordance with prior biological knowledge--whether the clusters actually correspond to the real structure in the data is somewhat less frequently considered. Suitable computational cluster validation techniques are available in the general data-mining literature, but have been given only a fraction of the same attention in bioinformatics. RESULTS: This review paper aims to familiarize the reader with the battery of techniques available for the validation of clustering results, with a particular focus on their application to post-genomic data analysis. Synthetic and real biological datasets are used to demonstrate the benefits, and also some of the perils, of analytical clustervalidation. AVAILABILITY: The software used in the experiments is available at http://dbkweb.ch.umist.ac.uk/handl/clustervalidation/. SUPPLEMENTARY INFORMATION: Enlarged colour plots are provided in the Supplementary Material, which is available at http://dbkweb.ch.umist.ac.uk/handl/clustervalidation/.

Algorithms↗

Decision support and data analysis tools for risk assessment in primary preventive study of atherosclerosis.

In the paper we show results of two programs applied for data analysis and decision support in primary preventive study of atherosclerosis. First program E.T. (Epidemiology Tools) can be used for analysis of data from retrospective (case-control) studies, prospective (cohort) studies and for standardization. Second program called CORE (COnstitution and REduction) supports the process of selection of features that are relevant for given decision making task. Program CORE is using information theory approach. Both these programs were applied to analysis of data about 1417 middle age men collected in the longitudinal study on atherosclerosis in urban population. Apart from these two new programs we have analyzed data also by the STATISTICA software.

Adult↗

Research in physical medicine and rehabilitation. VIII. Preliminary data analysis.

This paper describes important aspects of preliminary data analysis to be taken after data are checked for clerical entry errors and before the primary statistical analysis is performed. These include description and graphic display of each variable, recoding categorical data, transforming continuous data into another continuous variable and recoding continuous to categorical data. Missing values and outlying data points are identified and several techniques are recommended to minimize mistakes in variable recoding. Related variables measured with different units may be combined by using the z transformation and converted back to one of the original units for ease of interpretation. Finally, both categorical and continuous variables are checked for reliability by using kappa or the intraclass R.

Data Collection↗

SpotWhatR: a user-friendly microarray data analysis system.

SpotWhatR is a user-friendly microarray data analysis tool that runs under a widely and freely available R statistical language (http://www.r-project.org) for Windows and Linux operational systems. The aim of SpotWhatR is to help the researcher to analyze microarray data by providing basic tools for data visualization, normalization, determination of differentially expressed genes, summarization by Gene Ontology terms, and clustering analysis. SpotWhatR allows researchers who are not familiar with computational programming to choose the most suitable analysis for their microarray dataset. Along with well-known procedures used in microarray data analysis, we have introduced a stand-alone implementation of the HTself method, especially designed to find differentially expressed genes in low-replication contexts. This approach is more compatible with our local reality than the usual statistical methods. We provide several examples derived from the Blastocladiella emersonii and Xylella fastidiosa Microarray Projects. SpotWhatR is freely available at http://blasto.iq.usp.br/~tkoide/SpotWhatR, in English and Portuguese versions. In addition, the user can choose between "single experiment" and "batch processing" versions.

Blastocladiella↗

Terminal restriction fragment length polymorphism data analysis for quantitative comparison of microbial communities.

Terminal restriction fragment length polymorphism (T-RFLP) is a culture-independent method of obtaining a genetic fingerprint of the composition of a microbial community. Comparisons of the utility of different methods of (i) including peaks, (ii) computing the difference (or distance) between profiles, and (iii) performing statistical analysis were made by using replicated profiles of eubacterial communities. These samples included soil collected from three regions of the United States, soil fractions derived from three agronomic field treatments, soil samples taken from within one meter of each other in an alfalfa field, and replicate laboratory bioreactors. Cluster analysis by Ward's method and by the unweighted-pair group method using arithmetic averages (UPGMA) were compared. Ward's method was more effective at differentiating major groups within sets of profiles; UPGMA had a slightly reduced error rate in clustering of replicate profiles and was more sensitive to outliers. Most replicate profiles were clustered together when relative peak height or Hellinger-transformed peak height was used, in contrast to raw peak height. Redundancy analysis was more effective than cluster analysis at detecting differences between similar samples. Redundancy analysis using Hellinger distance was more sensitive than that using Euclidean distance between relative peak height profiles. Analysis of Jaccard distance between profiles, which considers only the presence or absence of a terminal restriction fragment, was the most sensitive in redundancy analysis, and was equally sensitive in cluster analysis, if all profiles had cumulative peak heights greater than 10,000 fluorescence units. It is concluded that T-RFLP is a sensitive method of differentiating between microbial communities when the optimal statistical method is used for the situation at hand. It is recommended that hypothesis testing be performed by redundancy analysis of Hellinger-transformed data and that exploratory data analysis be performed by cluster analysis using Ward's method to find natural groups or by UPGMA to identify potential outliers. Analyses can also be based on Jaccard distance if all profiles have cumulative peak heights greater than 10,000 fluorescence units.

Bacteria↗