Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Empirical considerations in orthopaedic research design and data analysis. Part II: The application of data analytic techniques.

To assure that a hypothesis is tested as rigorously as possible, the proper statistical method must be used to analyze the data. But without a strong background in statistics, it may be difficult to determine the efficacy of the data analytic technique used in the study. This paper describes several widely used data analytic techniques and offers examples of their proper application in orthopaedic research design.

Data Interpretation, Statistical↗

Application of the exploratory data analysis for evaluating the toxicity of chlorinated phenol derivatives by various cell models.

Exploratory data analysis based on multivariate statistical analysis techniques was introduced as a new approach to expressing the toxicity of chemical substances at the simultaneous acceptance of various cell models. Using principal component analysis and cluster analysis methods the toxicity of chlorinated phenol derivatives on employing some of the cell models (chlorococcal algae, cyanobacteria, bacteria, micromycetes, plant and animal cells) was characterized. The previous empirical experience that the toxicity of chlorinated phenol derivatives will increase with a growing degree of chlorination and that the presence of the methoxy group will cause a lowering of the toxic effect was demonstrated. The relationship between groups of tests used was presented.

Allium↗

A categorical data analysis of contacts with the Family Health Clinic, Calabar, Nigeria.

The relationships of population, environmental and accessibility variables to registration and attendance by mothers of children under 6 at the Family Health Clinic in Calabar, Nigeria are investigated. The technique used to analyze the data collected is categorical data analysis which proceeds in two stages, variable selection to reduce the variable set and fitting a log-linear model to the reduced set. Details of the statistical procedures used are provided to indicate how categorical data analysis can be used as a valuable tool of analysis in medical geographical studies that employ count or frequency data. It was found that younger mothers and Ibibio women registered more often at the clinic than did their counterparts. However, if the relatively sparse data on fathers is accepted, the association between age and registration is found to be spurious and a model can be substituted which shows younger fathers and fathers who spoke a non-Efik/Ibibio language to be associated with higher clinic registration of mothers. It was further found that for registered mothers the probability of a clinic visit was decreased by mother's age, increased by distance given no travel cost, unaffected by distance given some travel cost, increased by travel cost given a short distance to the clinic and decreased by travel cost given a longer distance from the clinic. These results are discussed in relation to population characteristics such as socio-economic status, clinic procedures such as health worker activities, transportation availability in Calabar, the spatial ecology of the city and local environmental conditions.

Adult↗

Is mixed effects modeling or naïve pooled data analysis preferred for the interpretation of single sample per subject toxicokinetic data?

The purpose of this study was to evaluate whether mixed effects modeling (MEM) performs better than either noncompartmental or compartmental naïve pooled data (NPD) analysis for the interpretation of single sample per subject pharmacokinetic (PK) data. Using PK parameters determined during a toxicokinetic study in rats, we simulated data sets that might emerge from similar experiments. Data sets were simulated with varying numbers of animals at each sampling time (4-48) and the number of samples taken (1-3) from each individual. Each data set was replicated 50 times and analyzed using several variations of MEM that differed in the assumptions made regarding intraindividual error, NPD, and a graphical noncompartmental method. These analyses attempted to retrieve the underlying parameter and covariate effect values. We compared these analysis methods with respect to how well the underlying values were retrieved. All analysis methods performed poorly with single sample per subject data but MEM gave less biased estimates under the simulated conditions used here. MEM performance increased when covariate effects were sought in the analysis compared with analyses seeking only PK parameters. Decreasing the number of animals used per sampling time from 48 to 16 did not influence the quality of parameter estimates but further reductions (< 16 animals per sampling time) resulted in a reduced proportion of acceptable estimates. Parameter estimate quality improved and worsened with MEM and NPD, respectively, when additional samples were obtained from each individual. Assumptions made regarding the magnitude of intraindividual error were unimportant with single sample per subject data but influenced parameter estimates if more samples were obtained from each individual. MEM is preferable to both NPD and noncompartmental approaches for the analysis of single sample per subject data but even with MEM estimates of clearance are often biased.

Animals↗

Data analysis in behavioral cerebral blood flow activation studies using xenon-133 clearance.

BACKGROUND AND PURPOSE: Three mainstream strategies exist to detect the responses of regional cerebral blood flow to functional activation. We tested the significance of changes in raw regional cerebral blood flow data, regional cerebral blood flow data normalized by division by global cerebral blood flow (dependent model of the regional-to-global cerebral blood flow relation), and regional cerebral blood flow data treating global cerebral blood flow as a covariate (independent model). Both latter models attempt to enhance regional sensitivity by removing global effects. We examined the sensitivity and pitfalls of these three strategies in behavioral activation studies. METHODS: These three strategies of data analysis were applied to changes in regional cerebral blood flow induced by a visuospatial problem-solving task in 38 healthy subjects as measured by the intravenous xenon-133 method with 32 stationary detectors. RESULTS: Mental activation increased blood flow in all regions of interest. Raw data were most sensitive and reliable to detect responses to mental stimulation. Both the independent and dependent models to remove global effects were less sensitive and falsely indicated deactivation in regions that were clearly stimulated. CONCLUSIONS: In behavioral activation paradigms, safe data analysis should be restricted to using raw regional cerebral blood flow increases without normalization or separation of global from regional effects. Studies using complex stimulation tasks should be scrutinized for global cerebral blood flow effects confounding regional responses.

Behavior↗

A visual data analysis system for the medical image processing.

We developed a visual data analysis system that can easily manage a large volume of medical imaging data. This system can analyze sets of imaging data using general image processing methods, so that various kinds of medical imaging data such as ECG charts, X ray image films, and MRI images, can be processed. The system has a graphical user interface (GUI). A physician who is novice at the system can manipulate the imaging data intuitively by pull down menus, pop up menus and buttons within the window system. The system can run on a standard UNIX workstation which is faster and more powerful than most personal computers. The system needs an X window system/Motif and C compiler. These are standard system programs already available on most UNIX workstations. The source code of the system can be retrieved from our anonymous ftp site via Internet.

Computer Graphics↗

Evaluation of replication studies, combined data analysis, and analytical methods in complex diseases.

Due to genetic heterogeneity, phenocopies, incomplete penetrance, misdiagnosis, and unknown mode of inheritance, linkage studies of most complex diseases are unlikely to provide conclusive findings with unambiguously high lod scores. Typically, several marginally significant lod scores or elevated lod scores are observed in a genome-wide screen. However, it is usually difficult to differentiate these findings from false positives (type I errors). Two approaches are commonly used to guard against false positives: replication studies in independent samples and combined data analysis. In the current paper, we evaluated these two common approaches using simulated data where data from multiple groups were available and locations of disease genes were known. We found replication studies and combined data analysis performed similarly in terms of their ability to identify true and false positive linkages. Both approaches confirmed two true linkages and did not confirm any false positive linkages. The results also indicated that it is not appropriate to apply the criteria proposed for confirming significant evidence for linkage to confirm regions with only suggestive evidence for linkage. The current results support previous findings that parametric analysis using an incorrect genetic model can still identify a true linkage.

Environment↗

[Comparison of software programs for data analysis of complex surveys].

OBJECTIVE: To compare specific software programs for data analysis of complex surveys regarding the following characteristics: ease of application, computer efficiency and accuracy of the results. METHODS: Secondary data from the Pesquisa Nacional sobre Demografia e Saúde (National survey on demography and health) (1996) with a target population of women aged 15 to 49 years old were used. This was a probabilistic subsampling drawn in two stages, then stratified, with the probability proportional to size in the first stage. The northern and mid-western regions of the country were selected for the study. The parameters of interest were mean for the age variable, and the proportion for five other qualitative variables. The software programs used were Epi Info, Stata and WesVarPC. RESULTS: The programs have two common options for the files import: the dBASE and text type files. The number of steps previous to the execution of the analyses were twenty- one for Epi Info, eleven for Stata and nine for WesVarPC. Efficiency was high for all them, that is, less that three seconds. The standard errors estimated using Epi Info and Stata were the same, with approximation up to the third decimal; those for WesVarPC were generally higher. CONCLUSIONS: Epi Info is the most limited software program regarding the analyses currently performed; however it is easy to use and free. Stata and WesVarPC are far more complete, however the disadvantage is their cost. The choice of the software program will depend mainly on the user's specific needs.

Adolescent↗

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i)&#xa0;perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii)&#xa0;introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans↗

Estimating fertility potential via semen analysis data.

The aim of this study was to evaluate diagnostic profiles for the assessment of semen analysis data with respect to male fertility potential. Semen samples taken from 208 patients of known fertility and suspected infertility were studied by conventional semen analysis methods. The data throw doubt upon the validity of an approach based on the number of deviations from the normal standard values defined by the World Health Organization. The alternative approach of a specific semen characteristic (particularly morphology) as the major predictor of fertility produced no beneficial results. However, the semen analysis index based on semen volume, sperm count, percentage motility and normal forms resulted in a high accuracy of classification but for only 44% of the cases, with 3% false negatives and 10% false positives using cut-off indices of > or = 0.6 and < or = -1.0 for defining 'fertile' and 'infertile' zones, respectively. In conclusion, it is emphasized that there are a number of specific semen analysis variables, each expressing a different aspect of male fertility potential which, when combined in correct proportion, do provide the optimal evaluation of the male fertility status. However, in order to increase the prognostic potential of the semen sample, new and meaningful parameters must be discovered.

Adult↗

Novel data analysis for synchronised spontaneous neuromagnetic activity.

A novel approach to neuromagnetic data analysis is presented. This technique is aimed at studying synchronised spontaneous activity (SSA) and has been used to resolve two different signals from one single evoked response, providing evidence for two possibly distinct sources. The data presented are consistent with a model that permits the generators of spontaneous activity to be synchronised by sensory stimuli.

Brain↗

Distributed intelligent data analysis in diabetic patient management.

This paper outlines the methodologies that can be used to perform an intelligent analysis of diabetic patients' data, realized in a distributed management context. We present a decision-support system architecture based on two modules, a Patient Unit and a Medical Unit, connected by telecommunication services. We stress the necessity to resort to temporal abstraction techniques, combined with time series analysis, in order to provide useful advice to patients; finally, we outline how data analysis and interpretation can be cooperatively performed by the two modules.

Computer Communication Networks↗

A hierarchical, count-based model highlights challenges in scATAC-seq data analysis and points to opportunities to extract finer-resolution information.

BACKGROUND: Data from Single-cell Assay for Transposase Accessible Chromatin with Sequencing (scATAC-seq) is highly sparse. While current computational methods feature a range of transformation procedures to extract meaningful information, major challenges remain. RESULTS: Here, we discuss the major scATAC-seq data analysis challenges such as sequencing depth normalization and region-specific biases. We present a hierarchical count model that is motivated by the data generating process of scATAC-seq data. Our simulations show that current scATAC-seq data, while clearly containing physical single-cell resolution, are too sparse to infer true informational-level single-cell, single-region of chromatin accessibility states. CONCLUSIONS: While the broad utility of scATAC-seq at a cell type level is undeniable, describing it as fully resolving chromatin accessibility at single-cell resolution, particularly at individual locus level, may overstate the level of detail currently achievable. We conclude that chromatin accessibility profiling at true single-cell, single-region resolution is challenging with current data sensitivity, but that it may be achieved with promising developments in optimizing the efficiency of scATAC-seq assays.

Single-Cell Analysis↗

[Computer-assisted data analysis in a pediatric intensive care unit].

Computer assisted real time data analysis introduces a reasonable method of judgment into patient monitoring systems. From fast changing vital parameters discrete heart and respiration rate samples are immediately evaluated and presented as graphs near the bedside. Thus, statistical routines can increase the better understanding of instable clinical conditions and lend support to the decision making process. The early detection of a pathological trend in a patient whose ability to compensate is still present provides necessary time for diagnostic or preventive countermeasures in case of emergency.

Computers↗

Optical illusions from visual data analysis: example of the New Zealand asthma mortality epidemic.

The abundance of health-related statistics routinely collected worldwide invites their misuse from haphazard associations between secular trends of these data. This misuse is often compounded by assessing these associations simply on the basis of a visual inspection of the data. The visual approach to data analysis, known to have several pitfalls, is particularly tempting in the context of asthma where it has often been used. For example, the epidemic of asthma deaths that occurred in New Zealand during the last two decades has been imputed to fenoterol, a medication for asthma, on the basis of a visual assessment of ecological data. The simultaneity of time trends in the asthma death rate and fenoterol market share in that country formed an important part of the statistical basis of the evidence. We verified whether the results of such visual analyses are corroborated by more objective quantitative statistical methods of analysis. We reanalyzed these same data, namely the time trend data of New Zealand asthma death rates, fenoterol market share, sales of beta-agonists and inhaled corticosteroids, measured yearly for the 16-year span 1976-1991, using Poisson weighted loglinear regression. We found that the protective effect of inhaled corticosteroids (rate ratio 0.5 per canister per month; 95% confidence interval 0.4 to 0.7; p = 0.0001) was more closely associated with changes in asthma mortality than either fenoterol (RR 2.7 per canister per month; 95% CI: 0.9 to 7.5; p = 0.06) or all beta-agonists combined (RR 1.6; 95% CI: 0.8 to 3.0; p = .19). We conclude from this quantitative analysis that these ecological asthma mortality data provide evidence of a stronger association with inhaled corticosteroids, little used in New Zealand at the onset of the epidemic but used abundantly at its termination, than with fenoterol. This conclusion is diametrically opposite to that found by the visual approach. The quantitative analysis demonstrates that the visual approach to the analysis of ecological data, although seemingly convincing, can be misleading by creating an optical illusion. This purely visual approach to data analysis may thus have serious implications when the resulting scientific information is used to make vital public health and policy decisions.

Administration, Inhalation↗

Who 'controls' quality control data analysis?

A common quality control tool is peer group comparison of data from commercial controls. While its real-time effectiveness is limited, inappropriate statistical management of the data can cause an individual lab's performance to be misrepresented. Here are two examples where vendor-directed data analysis contained flagrant errors. The finding that vendors use inappropriate algorithms to compare accuracy and precision of peer performance suggests a need, the author believes, to set rigorous standards of reporting for the protection of participating laboratories.

California↗

An evaluation of five commercial immunoassay data analysis software systems.

An evaluation of five commercial software systems used for immunoassay data analysis revealed numerous deficiencies. Often, the utility of statistical output was compromised by poor documentation. Several data sets were run through each system using a four-parameter calibration function, and the results were compared to those from an independent method. Comparable results between systems were obtained, but often several attempts at analysis were necessary. The evaluation process revealed that it is difficult to monitor the numerous options available on these types of programs, and that incorrect results could easily be obtained if comparison analyses were not used. Recommendations for improved software functionality and for using the four-parameter calibration model are presented.

Data Interpretation, Statistical↗