Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

The pros and cons of data analysis software for qualitative research.

PURPOSE: To explore the use of computer-based qualitative data analysis software packages. SCOPE: The advantages and capabilities of qualitative data analysis software are described and concerns about their effects on methods are discussed. FINDINGS: Advantages of using qualitative data analysis software include being freed from manual and clerical tasks, saving time, being able to deal with large amounts of qualitative data, having increased flexibility, and having improved validity and auditability of qualitative research. Concerns include increasingly deterministic and rigid processes, privileging of coding, and retrieval methods; reification of data, increased pressure on researchers to focus on volume and breadth rather than on depth and meaning, time and energy spent learning to use computer packages, increased commercialism, and distraction from the real work of analysis. CONCLUSIONS: We recommend that researchers consider the capabilities of the package, their own computer literacy and knowledge of the package, or the time required to gain these skills, and the suitability of the package for their research. The intelligence and integrity that a researcher brings to the research process must also be brought to the choice and use of tools and analytical processes. Researchers should be as critical of the methodological approaches to using qualitative data analysis software as they are about the fit between research question, methods, and research design.

Humans↗

Data mining: data analysis on a grand scale?

Modern data mining has evolved largely as a result of efforts by computer scientists to address the needs of 'data owners' in extracting useful information from massive observational data sets. Because of this historical context, data mining to date has largely focused on computational and algorithmic issues rather than the more traditional statistical aspects of data analysis. This paper provides a brief review of the origins of data mining as well as discussing some of the primary themes in current research in data mining, including scalable algorithms for massive data sets, discovering novel patterns in data, and analysis of text, web, and related multimedia data sets.

Algorithms↗

[On the problem of missing data: How to identify and reduce the impact of missing data on findings of data analysis].

The impact of missing data on the analysis of empirical data is a frequently unrecognized problem. Missing data may not only result in a decrease in the actual sample size but potentially biasing effects on statistical findings have to be considered as well. Two important points are made in this article: Firstly, it is shown why the identification of potential causes of missing data should be an inherent part of any data analysis; secondly, the handling of missing data should be based on appropriate assumptions in order to avoid biased results and problems concerning the interpretation of empirical findings.

Algorithms↗

Qualitative research: data analysis techniques.

Qualitative data usually consist of the words or actions of participants. These data can be difficult to condense and organise without losing their meaning. Analysis of qualitative data requires considerable creativity on the part of the researcher.

Attitude of Health Personnel↗

An evaluation of count rate losses due to dead time in static and mobile gamma camera data analysis systems.

The count rate responses of a gamma camera and data analysis system have been shown to behave as two separate components. Five systems made up of combinations of two static gamma cameras, two data analysis systems, and a mobile gamma camera system were investigated. Camera and data analysis system dead times are quoted and percentage corrections are plotted against radioactivity of 99 Tcm in a patient phantom within the field of view. It is recommended that such a detailed study should be performed for any gamma camera system as part of a quality assurance programme.

Gamma Rays↗

Interactive spatial data analysis in medical geography.

Interactive spatial data analysis involves the use of software environments that permit the visualization, exploration and, perhaps, modelling of geographically-referenced data. Such systems are of obvious value in epidemiological research, both of an environmental and geographical nature. There is an increasing number of such software environments available on a variety of platforms and operating systems. This paper considers the use of the proprietary Geographical Information System, ARC/INFO, in a spatial analysis context, showing how the spatial analytic tools that may be added to it can be exploited by geographical epidemiologists; such tools include those for modelling possible raised incidence of disease around suspected sources of pollution. The paper also reviews the use of systems such as S-Plus and XLISP-STAT, statistical programming environments to which spatial analysis functions or libraries may be added. The use of INFO-MAP, a system designed to aid in the teaching of interactive spatial data analysis, is also highlighted. The various software environments are illustrated with reference to examples concerned with: clustering of childhood leukaemia in part of Lancashire, England; Burkitt's lymphoma in Uganda; larynx cancer in Lancashire; and childhood mortality in Auckland, New Zealand.

Adolescent↗

Comparison of different microarray data analysis programs and description of a database for microarray data management.

Data analysis and management represent a major challenge for gene expression studies using microarrays. Here, we compare different methods of analysis and demonstrate the utility of a personal microarray database. Gene expression during HIV infection of cell lines was studied using Affymetrix U-133 A and B chips. The data were analyzed using Affymetrix Microarray Suite and Data Mining Tool, Silicon Genetics GeneSpring, and dChip from Harvard School of Public Health. A small-scale database was established with FileMaker Pro Developer to manage and analyze the data. There was great variability among the programs in the lists of significantly changed genes constructed from the same data. Similarly choices of different parameters for normalization, comparison, and standardization greatly affected the outcome. As many probe sets on the U133 chip target the same Unigene clusters, the Unigene information can be used as an internal control to confirm and interpret the probe set results. Algorithms used for the determination of changes in gene expression require further refinement and standardization. The use of a personal database powered with Unigene information can enhance the analysis of gene expression data.

Database Management Systems↗

Classification using functional data analysis for temporal gene expression data.

MOTIVATION: Temporal gene expression profiles provide an important characterization of gene function, as biological systems are predominantly developmental and dynamic. We propose a method of classifying collections of temporal gene expression curves in which individual expression profiles are modeled as independent realizations of a stochastic process. The method uses a recently developed functional logistic regression tool based on functional principal components, aimed at classifying gene expression curves into known gene groups. The number of eigenfunctions in the classifier can be chosen by leave-one-out cross-validation with the aim of minimizing the classification error. RESULTS: We demonstrate that this methodology provides low-error-rate classification for both yeast cell-cycle gene expression profiles and Dictyostelium cell-type specific gene expression patterns. It also works well in simulations. We compare our functional principal components approach with a B-spline implementation of functional discriminant analysis for the yeast cell-cycle data and simulations. This indicates comparative advantages of our approach which uses fewer eigenfunctions/base functions. The proposed methodology is promising for the analysis of temporal gene expression data and beyond. AVAILABILITY: MATLAB programs are available upon request.

Algorithms↗

Health and subjective well-being: a replicated secondary data analysis.

The purposes of this article are to use replicated secondary data analysis to summarize information about the relationship between health and subjective well-being and to assess the strengths and weaknesses of replicated secondary data analysis as a mode of research synthesis. The findings from thirty-seven replications in seven surveys suggest a moderate and robust relationship between self-rated health and subjective well-being. Physician-assessed health, in contrast, exhibits weaker and less robust associations with subjective well-being. Further, the relationship between health and subjective well-being is conditioned by age and is stronger for measures of negative than positive affect. The principal advantages of replicated secondary data analysis, vis-a-vis other modes of research synthesis, are cost-effectiveness, increased ability to apply multivariate statistical techniques, and greater control and flexibility for the investigator. We suggest, nonetheless, that different modes of research synthesis can best be used for different purposes.

Health↗

Comparison of two exploratory data analysis methods for fMRI: unsupervised clustering versus independent component analysis.

Exploratory data-driven methods such as unsupervised clustering and independent component analysis (ICA) are considered to be hypothesis-generating procedures, and are complementary to the hypothesis-led statistical inferential methods in functional magnetic resonance imaging (fMRI). In this paper, we present a comparison between unsupervised clustering and ICA in a systematic fMRI study. The comparative results were evaluated by 1) task-related activation maps, 2) associated time-courses, and 3) receiver operating characteristic analysis. For the fMRI data, a comparative quantitative evaluation between the three clustering techniques, self-organizing map, "neural gas" network, and fuzzy clustering based on deterministic annealing, and the three ICA methods, FastICA, Infomax and topographic ICA was performed. The ICA methods proved to extract features relatively well for a small number of independent components but are limited to the linear mixture assumption. The unsupervised Clustering outperforms ICA in terms of classification results but requires a longer processing time than the ICA methods.

Adult↗

[The software QSR Nvivo 2.0 in qualitative data analysis: a tool for health and human sciences researches].

The QSR Nvivo 2.0 is one of the latest versions of qualitative data analysis software package. Taking on board our experience as QSR Nvivo 2.0 users we describe here the software's most important tools. The QSR Nvivo 2.0 was used to facilitate the qualitative analysis of data gathered in a Health and Education research. Considering the shortage of published materials about the issue, which suggests a lack of knowledge about the program in our context, our aim is to show how the QSR Nvivo 2.0 can assist qualitative data analysis.

Data Collection↗

Exploratory data analysis using set operations and ordinal mapping.

Exploratory data analysis requires the ability to issue ad hoc queries to filter and summarise data sets. As the sizes of health data sets grow, traditional methods of processing data have difficulty in providing acceptable response times for such queries. An alternative method is described which combines complete vertical partitioning of data with set operations on ordinal mappings (SOOM). An initial implementation of the technique provides significantly better performance than a conventional SQL database on typical exploratory data analysis queries. The use of parallel, distributed computation to further increase the performance of the technique appears to be feasible.

Algorithms↗

Modification of Hewlett-Packard Chemstation (G1034C version C.02.00 and G1701AA version C.02.00 and C.03.00) data analysis program for addition of automatic extracted ion chromatographic groups.

A modification to data analysis macros to create "ion groups" and add a menu item to the data analysis menu bar is described. These "ion groups" consist of up to 10 ions for extracted ion chromatographs. The present manuscript describes modifications to yield a drop down menu in data analysis that contains user-defined names of groups (e.g., opiates, barbiturates, benzodiazepines) of ions that, when selected, will automatically perform extracted ion chromatographs with up to 10 ions in that group for the loaded datafile.

Barbiturates↗

PC program extending the Potthoff-Roy longitudinal data analysis model to allow missing data: Kleinbaum's method.

Potthoff and Roy (Biometrika, 51 (1964) 313-326) generalized the multivariate analysis of variance model into a form that is especially useful for the study of longitudinal growth curve data. Applications of this method have, however, been limited by the requirement that each case in the sample be measured at the same set of time points, i.e. there can be no missing data. In this paper we describe, illustrate, and make available a user-friendly, interactive PC program implementing Kleinbaum's (J Mult Anal, 3 (1973) 117-124) extension of the Potthoff-Roy model to allow incomplete measurement sequences. These missing data are permitted to arise either randomly or by design as in mixed longitudinal studies.

Analysis of Variance↗

A new combined integral-light and slit-scan data analysis system (DAS) for flow cytometry.

Flow cytometry using list mode parameters such as fluorescence emission, light scatter and size on one hand and different slit-scan parameters on the other hand needs a fast, flexible, efficient and easy-to-use data analysis software. A new software package (data analysis system, DAS) has been developed that integrates data analysis for conventional (integral-light) flow cytometry and for slit-scan flow cytometry. The requirements, design and some examples are discussed and an implementation for IBM-compatible computers is presented. Special attention is directed to the handling of different data types from one-parameter histograms to multiparameter slit-scan data files. The package can be used as an interpreting programming language or as an interactive menu-driven command line interpreter with a large number of graphic, mathematical and statistical functions. DAS is not limited to use in flow cytometry only, but multidimensional data analysis, from astronomy to economics, can be done as well.

Data Display↗

A knowledge-based system for data analysis and interpretation.

Traditionally, statistical packages are employed to derive or infer facts about a Universe of Discourse through data analysis and interpretation. It is analysis that serves to transform data into information. Statistical packages provide the users with relatively easy-to-use and powerful mechanics of data analysis, but up to now they do not provide much help with the design and strategies of the analysis. As such, there is a risk of misuse of these packages by statistically inexperienced users. We propose the use of knowledge-based interfaces to support this category of users in statistical evaluations. This paper discusses our experiences from the implementation of a knowledge-based system called MAXITAB. It provides guidance in the processes of data analysis and interpretation and has been programmed as an interface to the statistical package MINITAB.

Data Interpretation, Statistical↗

[Data analysis by statistical models].

The basic idea for the realization of effective statistical data analysis is illustrated with an example. The use of statistical models is explained and the feasibility of objective comparison of the models by an information criterion AIC is demonstrated. Further, the possibility of practical use of Bayesian models for complex data analysis is explained. Finally, the necessity of cooperation between the experts of respective fields and statisticians for further development of statistical data analysis is mentioned.

Adult↗

Optical near-field data analysis through time-frequency distributions: application to the characterization and separation of the image spectral content by reassignment.

The near-field optical images have been traditionally analyzed by Fourier analysis and, recently, by wavelet analysis. Those data are nonstationary, which means that their spectral content varies with time, owing to the scanning-probe recording process; therefore time-frequency representations are, potentially, powerful tools for local characteristics extraction or shape separation, since they distribute the energy of the analyzed signal over the time and frequency variables and faithfully depict the signal local behavior. In this study we show that Cohen's class time-frequency distributions and their modified version by the reassignment method are appropriate tools for the analysis of near-field optical data. We demonstrate this by using these tools first on simulated data and second on experimental near-field optical images. Within this context we observe that time-frequency analysis allows one to easily characterize local frequencies, which involves a possible separation of relevant optical signal from artifacts.

Journal Article↗