Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

[Classification, a suitable method for nursing research--practical advice for data collection and data analysis].

Arranging subjective data in rank order is a simple method for making comparisons of judgements and evaluating them. Since it is difficult to find practical instructions for this which can be easily understood, this paper uses data obtained in a study of patients' opinions about their privacy. Results are analysed, presented and explained. The method is appropriate for a wide range of research in which individuals are asked for opinions.

Data Collection↗

Are standardized mortality ratios valid for public health data analysis?

Standardized mortality ratios (SMRs) have been criticized as lacking validity, and it has been recommended to use standardized rate ratios (SRRs) instead. A review of the epidemiology literature and standard epidemiology textbooks showed disagreement concerning the validity of SMRs and a lack of data to support claims concerning their validity. Therefore, we sought to determine the validity of SMRs in public health data analysis. Simulations were carried out using widely disparate study population age distributions and disease rates encountered in public health data analysis. We compared SMRs and SRRs as absolute measures of increased mortality in a population, and for ranking mortality in different populations. The simulations showed that SMRs changed by 6 per cent to 8 per cent when the age distribution was changed from that of a 'young' age distribution to that of an 'old' age distribution. In comparison, SRRs changed by 4 per cent to 5 per cent when the age-adjustment standard was changed from the 1940 U.S. Census population to the 1990 U.S. Census population. County rankings by SRR were somewhat more similar among themselves than when compared with rankings by SMR, but the differences were not large. Based on our findings, SMRs are of similar usefulness to SRRs in public health data analysis, will lead to similar conclusions, and may be used to compare different geographic areas.

Adolescent↗

Computer-assisted diagnosis by a model-free system of direct data analysis.

The basis of the method of data analysis presented is, in the case of any diagnostic test, the automatic compilation of separate frequency distributions for each diagnostic classification. The distinction of different test results for different diseases (the correlation for which the tests are used) can thus be quantitatively monitored. This offers opportunities for more specific control of the accuracy of the data base. Measurements of relative frequencies obtained from the frequency distributions of individuals with and without a given disease can serve as a quantitative handle for the selection of the combination of tests, and for adjustments of individual parameters, which will maximize the discrimination. The usual cutoffs are not used. A data-processing system can serve for the direct incorporation of patient chart data (including test results), and for the automation of the analysis described, with pattern recognition or cluster-seeking techniques. The ability of this system of analysis to minimize some of the problems associated with methods utilizing mathematical models is discussed.

Diagnosis, Computer-Assisted↗

Longitudinal data analysis (repeated measures) in clinical trials.

Longitudinal data is often collected in clinical trials to examine the effect of treatment on the disease process over time. This paper reviews and summarizes much of the methodological research on longitudinal data analysis from the perspective of clinical trials. We discuss methodology for analysing Gaussian and discrete longitudinal data and show how these methods can be applied to clinical trials data. We illustrate these methods with five examples of clinical trials with longitudinal outcomes. We also discuss issues of particular concern in clinical trials including sequential monitoring and adjustments for missing data. A review of current software for analysing longitudinal data is also provided. Published in 1999 by John Wiley & Sons, Ltd. This article is a US Government work and is the public domain in the United States.

Clinical Trials as Topic↗

Microcomputer-assisted univariate survival data analysis using Kaplan-Meier life table estimators.

We describe a microcomputer program (KMSURV) for exploratory univariate statistical analysis of survival data which is directly applicable to the evaluation of clinical trials and to retrospective epidemiological studies of hospital registry-based data. The program calculates life-table-like information based on Kaplan-Meier's product-limit estimators of the survivorship function S(t) and provides summary measures of average survival times. In addition, two non-parametric tests for the comparison of survival distributions are performed. A report-quality, high resolution plot of the S(t) estimates for all groups being compared complements each set of analyses. KMSURV is not a simple adaptation of a mainframe statistical analysis package and, thus, it utilizes efficiently the interactive environment which is inherent in microcomputing.

Actuarial Analysis↗

Reduction of interlaboratory variability in flow cytometric immunophenotyping by standardization of instrument set-up and calibration, and standard list mode data analysis.

Two workshops addressed the question to which degree standardization of instrument set-up and calibration, and standard list mode data analysis would reduce interlaboratory variability of flow cytometric results on prestained peripheral blood mononuclear cells (PBMC). Standard instrument set-up included uniform positioning of the "windows of analysis" for the forward and sideward light scatter and fluorescence (FL) 1 (i.e., fluorescein isothiocyanate [FITC]) and 2 (i.e., phycoerythrin [PE]) parameters. Reference standards and PBMC, double-stained with FITC- and PE-conjugated monoclonal antibodies covering a wide range of FL intensities and coexpression patterns, were sent out to 25 laboratories in Workshop 1 and to 35 laboratories in Workshop 2 with the following requests: a) to set up instruments according to local and standard protocols, b) to acquire list mode data on the PBMC with both instrument settings, and c) to analyze both datasets according to local protocols. Standard analysis of the list mode data acquired with uniform instrument settings was performed centrally using so-called "latent class model" software (Van Putten et al., Cytometry 14:86-96, 1993). This software provides an automated, "no-gating" analytical method of lymphocyte immunophenotypes and employs fixed FL marker settings as defined prior to each analytical run. In Workshop 1, these markers were set in identical histogram channels for all instruments based on results obtained with a reference instrument. Standard analysis of list mode data acquired after uniform instrument set-up led only to a 13% reduction of interlaboratory variability of results as compared to data analysis using local protocols. The standard protocol for instrument set-up led to uniform positioning of relatively strong FL signals but variable positioning of unstained cells on the FL histogram scales. Hence, standard FL marker settings were inappropriate for some instruments. Therefore, instrument responses to FITC and PE signals in Workshop 2 were calibrated using microbeads labeled with FITC or PE in a range of predefined FL intensities expressed in MESF units (molecules of equivalent soluble fluorochrome). That approach allowed the positioning of the FL markers for the standard analysis on the basis of identical FL1 and FL2 intensities, expressed in MESF units, for all instruments. Standard analysis of list mode data acquired after uniform instrument set-up and calibrated FL marker settings led to a 43% reduction of interlaboratory variability as compared to data analysis to local protocols. We conclude that standard list mode data analysis using fixed FL marker settings reduces the interlaboratory variability of flow cytometric results on prestained PBMC, provided that the instruments have been set up in a uniform way and that FL markers have been standardized on the basis of calibration of each instrument's response to the corresponding FL signals.

Antigens, CD↗

Toward a computer assisted analysis of NOESY spectra: a multivariate data analysis of an RNA NOESY spectrum.

A multivariate data-representation of a portion of the H-NOESY spectrum of an RNA octamer duplex was used to explore the possibility of using Principal Component Analysis and Partial Least Squares Discrimination for pattern recognition. In this case, it is found that the methods can: (i) distinguish slices containing signal from those containing only noise, (ii) locate slices containing overlapping signals, and (iii) in some cases to segregate slices with unique aspects such as those from terminal nucleotides, overlapping signals, purine-H8, pyrimidine-H6 and adenine-H2 containing slices. These properties can easily be included in a scheme to automate spectral analysis. The formulation described here does not distinguish patterns needed to automate sequential assignment of resonances in NOESY spectra of RNA.

Base Sequence↗

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software↗

Interpretation of analytical data on n-alkanes and polynuclear aromatic hydrocarbons in Arbacia lixula from the coasts of Tenerife (Canary Islands, Spain) by multivariate data analysis.

The hydrocarbons contents (n-alkanes, polycyclic aromatic hydrocarbons) were determined in the sea urchin Arbacia lixula. Multivariate data analysis as principal component analysis, factor analysis and, cluster analysis were applied to elucidate sources of pollution. PCA and FA were performed to establish the relationships between variables (hydrocarbons), samples (sea urchin) and sources of pollution.

Alkanes↗

Consolidation of common parameters from multiple fits in dynamic PET data analysis.

In dynamic positron emission tomography (PET) data analysis, regions of interest (ROI's) are analyzed by fitting a parametric model to the time-activity curve acquired after a radio-labeled tracer has been introduced into the patient's bloodstream. This procedure can be carried out for multiple ROI's and/or multiple injections of the same or a different radiopharmaceutical. The approach presented here takes advantage of prior knowledge that some of the parameters of those multiple fits are the same. Reduction of the total number of parameters to be estimated results in smaller statistical uncertainty for all parameter estimates, especially those common to multiple fits.

Algorithms↗

User evaluation of an integrated medical workstation for clinical data analysis.

Results are presented of the user evaluation of an integrated medical workstation for support of clinical research. Twenty-seven users were recruited from medical and scientific staff of the University Hospital Dijkzigt, the Faculty of Medicine of the Erasmus University Rotterdam, and from other Dutch medical institutions; and all were given a written, self-contained tutorial. Subsequently, an experiment was done in which six clinical data analysis problems had to be solved and an evaluation form was filled out. The aim of this user evaluation was to obtain insight in the benefits of integration for support of clinical data analysis for clinicians and biomedical researchers. The problems were divided into two sets, with gradually more complex problems. In the first set users were guided in a stepwise fashion to solve the problems. In the second set each stepwise problem had an open counterpart. During the evaluation, the workstation continuously recorded the user's actions. From these results significant differences became apparent between clinicians and non-clinicians for the correctness (means 54% and 81%, respectively, p = 0.04), completeness (means 64% and 88%, respectively, p = 0.01), and number of problems solved (means 67% and 90%, respectively, p = 0.02). These differences were absent for the stepwise problems. Physicians tend to skip more problems than biomedical researchers. No statistically significant differences were found between users with and without clinical data analysis experience, for correctness (means 74% and 72%, respectively, p = 0.95), and completeness (means 82% and 79%, respectively, p = 0.40).(ABSTRACT TRUNCATED AT 250 WORDS)

Computer Systems↗

[Control survey data analysis with smoothed distribution curve].

Assay data and normal range collected in quality control surveys were analyzed with a smoothed distribution curve, in order to perform appropriate analysis of the data. In the control survey, the committee set a relative allowable limit for each test. A variable was changed continuously through the data range, and the data within the allowable limit were counted. In the computer program, the variable was changed step wise in logarithmic scale. These data were plotted on a graph with log-scale on X-axis and per cent of relative data in Y-axis. The graph thus obtained shows smoother curve than the histogram of the data, and Y-axis scale shows actual data within the allowable data at the variable scale. Normal ranges and assay data of several chemistry assays were analyzed with this method. Distribution of normal range of enzyme assays showed plural peaks in each test. Especially, in choline esterase assay, more than ten different units were used in laboratories in Japan with the same unit name, IU/1. In, amylase assay, nevertheless the normal range was separated in three peaks, the data of sample assay did not separate in peaks correspond to the normal range peaks. Thus, the smoothed distributions curve is useful for data analysis in quality control survey.

Amylases↗

Fluorescence quenching data interpretation in biological systems. The use of microscopic models for data analysis and interpretation of complex systems.

In micro-heterogeneous media (e.g. membranes, micelles and colloidal systems), the fluorescence decay in the absence of quencher is usually intrinsically complex, e.g. due to the existence of several sub-populations with different micro-environments. In this case it is impossible to analyze data in detail (accounting for transient effects) and simpler formalisms are needed. The objective of the present work is to present and discuss such simpler formalisms. The goal is to achieve simple data analysis and meaningful, clear data interpretation in complex systems using microscopic models that consider several sub-populations of chromophores. Two points are dealt with in detail. (i) It is shown that the approximation of the transient effects by the quenching sphere-of-action model is not always possible. The quenching sphere-of-action concept can be regarded as a valuable tool, although crude, only in a limited range of experimental conditions, namely time resolution. (ii) The Stern-Volmer equation usually used for data analysis is only valid for a limited range of small and moderate equilibrium association constants, Ka, although this is frequently overlooked in the literature. Self-consistency criteria are presented for the proposed methods. The well-known downward curvature due to a fraction of fluorophores which is not accessible to the quencher is only a limiting case from a set of possible situations which result in deviations to linearity. A systematic classification of the different types of quenching is presented.

Data Interpretation, Statistical↗

Planning controlled clinical trials on the basis of descriptive data analysis.

In controlled clinical trials the problem of multiplicity of desired inferential statements finds attention at an increasing rate. In this paper the previously proposed concept of Descriptive Data Analysis (DDA), situated between Confirmatory and Exploratory Data Analysis, is applied to the planning aspects of controlled trials for which the problem of multiplicity exists. The (non-Bayesian) DDA planning concept should provide the investigator with tools to draw final conclusions from data of several variables possibly observed at several time points in possibly several groups of subjects by combining his pre-trial medical experience with descriptive inferential statements (confidence intervals and test results) at nominal significance levels. DDA also provides for confirmatory statements concerning individual null hypotheses and partially global hypotheses.

Clinical Trials as Topic↗

Time normalization of voice signals using functional data analysis.

The harmonics-to-noise ratio (HNR) has been used to quantify the waveform irregularity of voice signals [Yumoto et al., J. Acoust. Soc. Am. 71, 1544-1550 (1982)]. This measure assumes that the signal consists of two components: a harmonic component, which is the common pattern that repeats from cycle-to-cycle, and an additive noise component, which produces the cycle-to-cycle irregularity. It has been shown [J. Qi, J. Acoust. Soc. Am. 92, 2569-2576 (1992)] that a valid computation of the HNR requires a nonlinear time normalization of the cycle wavelets to remove phase differences between them. This paper shows the application of functional data analysis to perform an optimal nonlinear normalization and compute the HNR of voice signals. Results obtained for the same signals using zero-padding, linear normalization, and dynamic programming algorithms are presented for comparison. Functional data analysis offers certain advantages over other approaches: it preserves meaningful features of signal shape, produces differentiable results, and allows flexibility in selecting the optimization criteria for the wavelet alignment. An extension of the technique for the time normalization of simultaneous voice signals (such as acoustic, EGG, and airflow signals) is also shown. The general purpose of this article is to illustrate the potential of functional data analysis as a powerful analytical tool for studying aspects of the voice production process.

Data Interpretation, Statistical↗

A method for quantification of absolute amounts of nucleic acids by (RT)-PCR and a new mathematical model for data analysis.

Accurate quantification of nucleic acids by competitive (RT)-PCR requires a valid internal standard, a reference for data normalization and an adequate mathematical model for data analysis. We report here an effective procedure for the generation of homologous RNA internal standards and a strategy for synthesizing and using a reference target RNA in quantification of absolute amounts of nucleic acids. Further, a new mathematical model describing the general kinetic features of competitive PCR was developed. The model extends the validity of quantitative competitive (RT)-PCR beyond the exponential phase. The new method eliminates the errors arising from different amplification efficiencies of the co-amplified sequences and from heteroduplex formation in the system. The high accuracy (relative error <2%) is comparable to the recently developed real time detection 5'-nuclease PCR. Also, corresponding computer software has been devised for practical data analysis.

Cell Line↗