Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Discrimination of Francisella tularensis subspecies using surface enhanced laser desorption ionization mass spectrometry and multivariate data analysis.

Francisella tularensis causes the zoonotic disease tularemia, and is considered a potential bioterrorist agent due to its extremely low infection dose and potential for airborne transmission. Presently, F. tularensis is divided into four subspecies; tularensis, holarctica, mediasiatica and novicida. Phenotypic discrimination of the closely related subspecies with traditional methods is difficult and tedious. Furthermore, the results may be vague and they often need to be complemented with virulence tests in animals. Here, we have used surface enhanced laser desorption ionization time-of-flight mass spectrometry (SELDI-TOF-MS) to discriminate between the four subspecies of F. tularensis. The method is based on the differential binding of protein subsets to chemically modified surfaces. Bacterial thermolysates were added to anionic, cationic, and copper ion-loaded immobilized metal affinity SELDI chip surfaces. After binding, washing, and SELDI-TOF-MS different protein profiles were obtained. The spectra generated from the different surfaces were then used to characterize each bacterial strain. The results showed that the method was reproducible, with an average intensity variation of 21%, and that the mass precision was good (300-450 ppm). Moreover, in subsequent cluster analysis and principal component analysis (PCA) data for the analyzed Francisella strains grouped according to the recognized subspecies. Partial least squares-discriminant analysis (PLS-DA) of the protein profiles also identified proteins that differed between the strains. Thus, the protein profiling approach based on SELDI-TOF-MS holds great promise for rapid high-resolution phenotypic identification of bacteria.

Animals↗

Technical note: the initial stages of statistical data analysis.

OBJECTIVE: To provide an overview of several important data-related considerations in the design stage of a research project and to review the levels of measurement and their relationship to the statistical technique chosen for the data analysis. BACKGROUND: When planning a study, the researcher must clearly define the research problem and narrow it down to specific, testable questions. The next steps are to identify the variables in the study, decide how to group and treat subjects, and determine how to measure, and the underlying level of measurement of, the dependent variables. Then the appropriate statistical technique can be selected for data analysis. DESCRIPTION: The four levels of measurement in increasing complexity are nominal, ordinal, interval, and ratio. Nominal data are categorical or "count" data, and the numbers are treated as labels. Ordinal data can be ranked in a meaningful order by magnitude. Interval data possess the characteristics of ordinal data and also have equal distances between levels. Ratio data have a natural zero point. Nominal and ordinal data are analyzed with nonparametric statistical techniques and interval and ratio data with parametric statistical techniques. ADVANTAGES: Understanding the four levels of measurement and when it is appropriate to use each is important in determining which statistical technique to use when analyzing data.

Journal Article↗

[Which dosage concept for adrenaline is correct in cardiopulmonary resuscitation? A data analysis of preclinical resuscitations].

AIM: Epinephrine is the drug of choice in cardiopulmonary resuscitation. Its dosage, however, is controversially discussed. The American Heart Association recommends for standard use in adults 1 mg epinephrine every 3-5 minutes, but classifies a medium dose, a high dose and a step-by-step escalating dosage concept as potentially useful alternatives. Aim of this study was to develop a rationale for the escalating dosage concept using an analysis of preclinical resuscitation data. METHODS: Bonn city (141 km2, 310,000 residents, 52% female, 13.9% > 65 years) was served by a double-response system of two ALS-units (staffed by physicians) and four BLS-units (staffed by paramedics). All patients were included in this data analysis, which were cardiopulmonary resuscitated according to the AHA guidelines by the ALS-unit Bonn-North (66% of area and 240,000 residents) from 1989 to 1994. All relevant data were documented by the emergency physicians using preformed treatment sheets. Discharge rates were determined by reviewing the hospital records, and one-year survival data were collected by mail contact with the primary physicians. The correlations between duration of cardiac arrest, dosage of epinephrine and outcome were determined by regression analysis in patients older than 17 years suffering from unwitnessed and bystander-witnessed cardiac arrest. Statistical significance was assumed for p < 0.05. RESULTS: Within 1989 to 1994 the ALS-team of Bonn-North resuscitated 685 cardiac arrest patients with presumed cardiac aetiology in 383 (56%) return of spontaneous circulation (ROSC) could be achieved. The epinephrine dosage required to achieve ROSC increased with prolongation of cardiac arrest interval (1989-1994; patients with ROSC selected, n = 263; r = 0.2276; p = 0.0002). Resuscitation success, however, decreased with increasing, dosage of epinephrine (1989-1992; n = 345; ROSC: r = -0.2643; p < 0.001; survival > 24 h: r = -0.3393; p < 0.001; discharge: r = -0.1677; p = 0.0018; survival > 1 year: r = -0.2685; p < 0.001). CONCLUSION: Based on these data, we recommend an escalating epinephrine dosage concept, which facilitates titration of the drug to an effective level and meets the needs of the individual patient. This concept avoids overdosage in patients who had just collapsed shortly before initiation of CPR, attains higher levels of epinephrine in patients suffering from prolonged cardiac arrest, and takes into consideration that the effective epinephrine dose varies individually and increases with prolongation of the cardiac arrest interval.

Adolescent↗

Evaluation of Bayesian estimation in comparison to NONMEM for population pharmacokinetic data analysis: application to pefloxacin in intensive care unit patients.

The pharmacokinetics of pefloxacin (PF) were investigated in a population of 74 intensive care unit patients receiving 400 mg bid as 1-hr infusion using (i) Bayesian estimation (BE) of individual patient parameters followed by multiple linear regression (MLR) analysis and (ii) NONMEM analysis. The data consisted of 3 to 9 PF plasma levels per patient measured over 1 to 3 dosage intervals (total 113) according to four different limited (suboptimal) sampling 3-point protocols. Twenty-nine covariates (including 15 comedications) were considered to explain the interpatient variability. Predicted PF CL for a patient with median covariates values was similar in both BE/MLR and NONMEM analysis (4.02 and 3.92 L/hr, respectively). Bilirubin level and age were identified as the major determinants of PF CL by both approaches with similar predicted magnitude of effects (about 40 and 30% decrease of median CL, respectively). Confounding effects were observed between creatinine clearance (26% decrease of PF CL in the BE/MLR model), simplified acute physiology score (a global score based on 14 biological and clinical variables) (18% decrease of median CL in the NONMEM model) and age (entered in both models) which were highly correlated in our data base. However, both models predicted similar PF CL for actual subpopulations by using actual covariate values. Finally, the NONMEM analysis allowed identification of an effect of weight on CL (decrease of CL for weight < 65 kg) whereas the BE/MLR analysis predicted an increase of CL in patients treated with phenobarbital. In conclusion, both approaches allowed identification of the major risk factors of PF pharmacokinetics in ICU patients. Their potential use at different stages of drug development is discussed.

Adolescent↗

Application of neural networks to flow cytometry data analysis and real-time cell classification.

Conventional analysis of flow cytometric data requires that population identification be performed graphically after a sample has been run using two-parameter scatter plots. As more parameters are measured, the number of possible two-parameter plots increases geometrically, making data analysis increasingly cumbersome. Artificial Neural Systems (ANS), also known as neural networks, are a powerful and convenient method for overcoming this data bottleneck. ANS "learn" to make classifications using all of the measured parameters simultaneously. Mathematical models and programming expertise are not required. ANS are inherently parallel so that high processing speed can be achieved. Because ANS are nonlinear, curved class boundaries and other nonlinearities can emerge naturally. Here, we present biomedical and oceanographic data to demonstrate the useful properties of neural networks for processing and analyzing flow cytometry data. We show that ANS are equally useful for human leukocytes and marine plankton data. They can easily accommodate nonlinear variations in data, detect subtle changes in measurements, interpolate and classify cells they were not trained on, and analyze multiparameter cell data in real time. Real-time classification of a mixture of six cyanobacteria strains was achieved with an average accuracy of 98%.

Animals↗

An aspect of discrete data analysis: fitting a beta-binomial distribution to the hospitals' data.

Statistical analysis for discrete data, particularly for probability models such as the binomial, Poisson and multinomial, is by now very well understood, with a wealth of suitable software. It can happen that the standard generalized linear modelling (glm) software is not completely appropriate, since over-dispersion is present, relative to the standard distributions such as the Poisson or the binomial. Failure to take account of this over-dispersion, for example in fitting a model such as log(p/(1 - p)) = alpha + beta x (where the covariate x is the dose) will mean that our estimates of beta will be less precise than the binomial-based formula gives us. Thus for example we will be quoting confidence intervals for beta that are too narrow. One way of coping with this problem is to use a probability model which is more general than the binomial, and one such model is the beta-binomial. This paper discusses beta-binomial modelling (in S-Plus) in relation to the interesting data set given in the 1998 BMJ paper by Spiegelhalter and Marshall on success rates of 52 in vitro fertilisation clinics in the UK.

Data Interpretation, Statistical↗

Distributed data analysis in a multicenter study: the CARDIA Study.

Unlike distributed data entry, which is used in many large epidemiologic studies and multicenter clinical trials, distributed data analysis is a relatively new concept. This paper reports on the usefulness of such a system in the Coronary Artery Risk Development in Young Adults (CARDIA) Study. CARDIA distributes the entire examination dataset to participating centers soon after completion of each round of data collection. The process was designed to encourage more numerous, diverse, and rapid publications, and to allow for more efficient use of the manpower and expertise in centers. Responsibilities of the coordinating center have changed from a conventional coordinating center but remain substantial due to the need for collating, monitoring, verifying, and documenting the distributed data analysis (DDA) system. DDA is successful from the standpoint of implementation and operation--21 manuscripts representing work analyzed at six participating centers had been submitted for publication within 3.5 years of the completion of the baseline examination.

Adolescent↗

Intrasubject repeatability of gait analysis data in normal and spastic children.

OBJECTIVE: To evaluate intrasubject repeatability of data obtained from computer-aided motion analysis in normal and spastic children. DESIGN: Prospective controlled study. BACKGROUND: Information from gait analysis is used in selecting therapeutic interventions for gait improvement in cerebral palsy. While there are several studies regarding repeatability of normal gait, there are no studies evaluating the repeatability of spastic gait. METHODS: Forty children (20 normal, 20 with diplegic type of cerebral palsy) were subjected to gait analysis. Kinematic, kinetic and time distance parameters obtained from gait analysis were studied for intrasubject variability within-day and between-day using statistical measures. RESULTS: Normal children had lower variability in time distance parameters than spastic children both within and between days. The repeatability of kinetics was better than those of kinematics, and values for normal children were better than those for spastic children. Within-day repeatability of kinematics and kinetics was better in normal children. Between-day repeatability of kinematics was better in normal children, while spastic children showed better repeatability for kinetics. CONCLUSIONS: We found lower repeatability of gait analysis data in spastic children compared to normal children. Restricted joint range of motion due to spasticity in the group of cerebral palsy patients may be responsible for the lower repeatability of data. Some errors due to marker placement are inadvertent and contribute to the lower between-day repeatability. RELEVANCE: The results of this study should be of interest to clinicians who make therapeutic decisions in patients with cerebral palsy using gait analysis data, and for scientists studying normal and pathological gait.

Adolescent↗

Methods of data analysis in the emergency medicine literature.

The authors hypothesized that data analysis in the current emergency medicine literature uses relatively few methods and sought to determine the frequency distributions of each method of analysis. The authors defined their population as original contributions in three refereed emergency medicine journals from September, 1985 through July, 1989. Letters to the editor, brief reports, reviews, and case reports were excluded. The authors reviewed 250 randomly selected articles and identified the method(s) of data analysis in each. The absolute frequency distribution of statistics were as follows: descriptive statistics only, 31%; contingency tables, 35% (chi 2, 28.4%; Fisher's exact test, 13.2%; McNemar's test, 0.4%); Student's t-test, 34%; ANOVA/ANCOVA, 12%; regression techniques, 8% (simple linear regression, 4.0%; multiple regression, 3.6%; logistic regression, 1.6%); nonparametric tests, 7% (Mann-Whitney, 2.8%; Wilcoxon, 2.4%; Dunnett, 0.8%; Kolmogorv-Smirnov, 0.4%; Kruskal-Wallis, 0.4%); multiple comparisons, 6% (Scheffé, 4.4%; Newman-Keuls, 2.0%); correlation techniques, 4% (Pearson product-moment correlation coefficient, 2.8%; Kendall's tau, 0.8%; Spearman's rho, 0.4%); confidence intervals, 2%. Correction techniques were used in 9% (Dunn-Bonnferoni, 4.8%; Yates correction, 4.4%). No statistics were found in 2% of the articles reviewed. Five statistical methods account for the vast majority (97% cumulative) of statistical uses in emergency medicine literature. This information should prove useful in deciding which tests should be emphasized in educating emergency physicians.

Curriculum↗

Examples for the improvements in AES depth profiling of multilayer thin film systems by application of factor analysis data evaluation.

Factor analysis has proved to be a powerful tool for the full exploitation of the chemical information included in the peak shapes and peak positions of spectra measured by AES depth profiling. Due to its ability to extract the number of independent chemical components, their spectra and their depth distributions, its information content exceeds the one of the usual peak-to-peak height evaluation of AES depth profile data. Using modern software with a graphically interactive user interface the analyst is put into a position, where he can work with Factor Analysis on a physically intuitive level despite of all the matrix algebra mathematics which it is based upon. The progress brought about by Factor Analysis to AES depth profiles of thin films is demonstrated by the analysis of two thin film systems. The first one is a Pt/Ti metallisation used as bottom electrode for ferroelectric thin films, the second one is a multilayer system where a Ti silicide formation of buried Ti/Si bilayers has been induced. Both examples show that Factor Analysis evaluation of AES depth profile data is capable to give access to stoichiometry information and to reveal interfacial layer phases, information which is hardly obtained from the conventional peak-to-peak height data evaluation.

Journal Article↗

UALCAN Mobile, an app for cancer proteogenomic data analysis.

Cancer is a complex disease affecting various organs and is a major cause of death worldwide. During cancer initiation, disease progression, and tumor metastasis, various genomic and proteomic alterations are observed. Recent technological advances have led to the generation of large amounts of molecular data, including genomics and transcriptomics. These large-scale datasets can be utilized to analyze and identify sub-class-specific cancer biomarkers and targets. However, there is a need for the development of user-friendly tools for large-scale data analysis, disseminating the analyzed data in a visualizable format to cancer researchers with no programming skills. We developed UALCAN, a comprehensive platform that allows users to integrate disparate data to better understand the genes, proteins, and pathways perturbed in cancer and make discoveries of potential biomarkers and targets. In the current study, we describe the development of the UALCAN Mobile application (app) that will provide cancer transcriptomic data obtained from The Cancer Genome Atlas (TCGA) project to evaluate protein-coding gene expression based on various stratifications, including stage, grade, race, gender, and molecular-subtypes across over 30 types of cancers. In addition, the UALCAN mobile provides data analysis options for epigenetic changes due to DNA promoter methylation and Clinical Proteomic Tumor Analysis Consortium (CPTAC) cancer proteomic data. The app provides access to large cancer molecular datasets on the go. To find changes in the expression of causative genes and proteins and to identify biomarkers and therapeutic targets, UALCAN mobile app will be extremely valuable. The "UALCAN Mobile" app is free to use and can be downloaded from both the iOS/Apple and the Android Play Store and has been downloaded over 100 times in each of iOS and android app stores.

app↗

Data analysis for continuous EEG monitoring in the ICU: seeing the forest and the trees.

Continuous EEG monitoring (CEEG) is a powerful tool for evaluating cerebral function in obtunded and comatose critically ill patients. The ongoing analysis of CEEG data is a major task because of the volume of data generated during monitoring and the need for near real-time interpretation of a patient's EEG patterns. Advances in digital EEG data acquisition, computer processing, data transmission, and data display have made CEEG monitoring in the intensive care unit technically feasible. A variety of quantitative EEG tools such as Fourier analysis and amplitude-integrated EEG, and other methods of data analysis such as computerized seizure detection, increasingly allow for focused review of EEG epochs of potential interest. These tools reduce the tremendous time burdens that accompany analysis of the complete CEEG data stream, and allow bedside personnel and nonexpert staff to potentially recognize significant EEG changes in a timely fashion. This article uses literature review and clinical case examples to illustrate techniques for the display and analysis of intensive care unit CEEG recordings. Areas requiring further research and development are discussed.

Algorithms↗

A synthesis technique for grounded theory data analysis.

AIMS: The purposes of this paper are to examine the issues surrounding current changes in grounded theory (GT) research methods and to explicate an innovative synthesis technique to GT data analysis. BACKGROUND: In recent years there has been a steady rise in the number of published research reports that use the GT method. However, this growing body of GT literature has been criticized for its lack of adherence to the method as explicated by its originators, Glaser and Strauss. METHODS: Recent and past literature that explicates, describes, and discusses GT methods is reviewed. A synthesis technique for grounded theory data analysis was developed to analyse qualitative data collected for a grounded theory study on caregiving. This synthesis technique was derived from the works of four grounded theorists (Kathy Charmaz, Mark Chesler, Juliet Corbin and Anselm Strauss). RESULTS: The lack of clarity and the inconsistencies surrounding GT analysis, as reported in the literature, resulted in the development of a synthesis technique based on the works of the aforementioned-grounded theorists. The product was a synthesis approach that included analytical steps from each of these authors. CONCLUSIONS: This synthesis approach increased understanding and enhanced clarity of GT data analysis techniques. This paper illustrates how integration of works by noted qualitative scholars is an appropriate and effective means to advance the discourse on data analysis for GT research studies.

Black or African American↗

Quality control for functional magnetic resonance imaging using automated data analysis and Shewhart charting.

A data acquisition and analysis protocol for quality control (QC) of functional magnetic resonance imaging studies is presented. Two sets of data are acquired, single-timepoint data for measurement of signal-to-ghost and signal-to-noise ratios, and multiple-timepoint data for measurement of short-term drift. Since manual data analysis can be time consuming and an impediment to regular QC, an automated data processing scheme is presented. The use of automated Shewhart charting is proposed to identify significant changes in each parameter over the long term. The protocol has successfully identified system faults and deteriorations undetected by conventional QC.

Artifacts↗

Quantifying errors in tableting data analysis using the Ryshkewitch equation due to inaccurate true density.

Although inaccurate true density affects analysis of powder compaction data, such effects have not been systematically evaluated in the literature. This work is aimed at quantitatively evaluating effects of inaccurate true density on tableting data analysis using the Ryshkewitch equation, sigma = sigma0 e - bepsilon, where epsilon is tablet porosity, sigma is tensile strength, sigma(0) and b are constants that are used to characterize tableting properties of a powder. Mathematical expressions are derived to enable reliable prediction of the influence of inaccurate true density on sigma(0) and b. The validity of the expressions is suggested by modeling the effects of inaccurate true density based on a set of accurate literature tableting data and confirmed using multiple sets of tableting data of water-containing powders that exhibit inaccurate helium pycnometry densities. Percentage errors in fitted sigma(0) and b as functions of errors in true density follow the derived mathematical expressions. With increasing percentage error in true density, percentage errors increase exponentially in fitted sigma(0) and increase linearly in fitted b, while R(2) is not affected. According to the mathematical expressions, true density with <0.28% error is required to achieve 4% accuracy in fitted sigma(0) for typical pharmaceutical powders.

Chemistry, Pharmaceutical↗

Research in physical medicine and rehabilitation. V. Data entry and early exploratory data analysis.

The process of data entry and initial analysis to locate data errors is described. Basic terms are defined and a simple method of entering data by using word processing software is illustrated. Data checking is done by using visual check of the raw data. Statistical programs are then used to locate possible data errors by finding data points (outliers) that are very different from the average. Special graphic output of statistical programs, scatterplots and box and whisker plots can be used to further locate questionable data. Examples of data entry forms and annotated step by step data cleaning with the use of inexpensive programs for personal computers are presented.

Computers↗

The RUMBA software: tools for neuroimaging data analysis.

The enormous scale and complexity of data sets in functional neuroimaging makes it crucial to have well-designed and flexible software for image processing, modeling, and statistical analysis. At present, researchers must choose between general purpose scientific computing environments (e.g., Splus and Matlab), and specialized human brain mapping packages that implement particular analysis strategies (e.g., AFNI, SPM, VoxBo, FSL or FIASCO). For the vast majority of users in Human Brain Mapping and Cognitive Neuroscience, general purpose computing environments provide an insufficient framework for a complex data-analysis regime. On the other hand, the operational particulars of more specialized neuroimaging analysis packages are difficult or impossible to modify and provide little transparency or flexibility to the user for approaches other than massively multiple comparisons based on inferential statistics derived from linear models. In order to address these problems, we have developed open-source software that allows a wide array of data analysis procedures. The RUMBA software includes programming tools that simplify the development of novel methods, and accommodates data in several standard image formats. A scripting interface, along with programming libraries, defines a number of useful analytic procedures, and provides an interface to data analysis procedures. The software also supports a graphical functional programming environment for implementing data analysis streams based on modular functional components. With these features, the RUMBA software provides researchers programmability, reusability, modular analysis tools, novel data analysis streams, and an analysis environment in which multiple approaches can be contrasted and compared. The RUMBA software retains the flexibility of general scientific computing environments while adding a framework in which both experts and novices can develop and adapt neuroimaging-specific analyses.

Algorithms↗

Coronary plaque classification with intravascular ultrasound radiofrequency data analysis.

BACKGROUND: Atherosclerotic plaque stability is related to histological composition. However, current diagnostic tools do not allow adequate in vivo identification and characterization of plaques. Spectral analysis of backscattered intravascular ultrasound (IVUS) data has potential for real-time in vivo plaque classification. METHODS AND RESULTS: Eighty-eight plaques from 51 left anterior descending coronary arteries were imaged ex vivo at physiological pressure with the use of 30-MHz IVUS transducers. After IVUS imaging, the arteries were pressure-fixed and corresponding histology was collected in matched images. Regions of interest, selected from histology, were 101 fibrous, 56 fibrolipidic, 50 calcified, and 70 calcified-necrotic regions. Classification schemes for model building were computed for autoregressive and classic Fourier spectra by using 75% of the data. The remaining data were used for validation. Autoregressive classification schemes performed better than those from classic Fourier spectra with accuracies of 90.4% for fibrous, 92.8% for fibrolipidic, 90.9% for calcified, and 89.5% for calcified-necrotic regions in the training data set and 79.7%, 81.2%, 92.8%, and 85.5% in the test data, respectively. Tissue maps were reconstructed with the use of accurate predictions of plaque composition from the autoregressive classification scheme. CONCLUSIONS: Coronary plaque composition can be predicted through the use of IVUS radiofrequency data analysis. Autoregressive classification schemes performed better than classic Fourier methods. These techniques allow real-time analysis of IVUS data, enabling in vivo plaque characterization.

Algorithms↗