Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Correlation Of Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Staphylococcus aureus phage typing, antimicrobial susceptibility patterns and patient data correlated using a personal computer: advantages for monitoring the epidemiology of MRSA.

Staphylococcus aureus, especially methicillin-resistant S. aureus (MRSA) is an important nosocomial pathogen. A problem encountered in the control of staphylococcal nosocomial infection is the difficulty correlating data available from antimicrobial susceptibility patterns (usually available locally) with phage typing patterns which are generally determined in a reference laboratory. Data problems are exaggerated because many patients have numerous samples taken over long periods of time. This paper describes the use of a database, 'DataEase', on a personal computer to correlate available data with patient demographics and to present the resulting information in whatever format is required. The information might be presented as the phage pattern of each isolate from patients in a particular unit per month or as a combination of the phage types and antimicrobial susceptibility patterns of isolates from a particular unit. The system can be used to generate a list of all isolates for phage typing in a format that allows the phage typing pattern to be recorded directly onto the list worksheet. In addition to providing clinically relevant reports, the system simplifies the collation of epidemiological information, data from which showed that since 1987 the incidence of MRSA has been rising in the group of hospitals served by this laboratory; MRSA of phage type 83A became a problem during 1990 and 1991; the incidence of gentamicin-sensitive phage type non-typable MRSA increased fourfold between 1988 and 1991; and the incidence of gentamicin-resistant MRSA rose sharply in 1992.

Bacteriophage Typing↗

Development of interpretive breakpoints for antifungal susceptibility testing: conceptual framework and analysis of in vitro-in vivo correlation data for fluconazole, itraconazole, and candida infections. Subcommittee on Antifungal Susceptibility Testing of the National Committee for Clinical Laboratory Standards.

The availability of reproducible antifungal susceptibility testing methods now permits analysis of data correlating susceptibility in vitro with outcome in vivo in order to define interpretive breakpoints. In this paper, we have examined the conceptual framework underlying interpretation of antimicrobial susceptibility testing results and then used these ideas to drive analysis of data packages developed by the respective manufacturers that correlate fluconazole and itraconazole MICs with outcome of candidal infections. Tentative fluconazole interpretive breakpoints for MICs determined by the National Committee for Clinical Laboratory Standards' M27-T broth macrodilution methodology are proposed: isolates for which MICs are < or = 8 microg/mL are susceptible to fluconazole, whereas those for which MICs are > or = 64 microg/mL appear resistant. Isolates for which the MIC of fluconazole is 16-32 microg/mL are considered susceptible dependent upon dose (S-DD), on the basis of data indicating clinical response when > 100 mg of fluconazole per day is given. These breakpoints do not, however, apply to Candida krusei, as it is considered inherently resistant to fluconazole. Tentative interpretive MIC breakpoints for itraconazole apply only to mucosal candidal infections and are as follows: susceptible, < or = 0.125 microg/mL; S-DD, 0.25-0.5 microg/mL; and resistant, > or = 1.0 microg/mL. These tentative breakpoints are now open for public commentary.

Animals↗

[Analysis of correlated data in occupational medicine: examples with binary data].

In a previous paper in this Journal we presented and discussed examples of analysis of correlated data when the response variable was continuous and normally distributed (the measurement of exposure to a toxic substance was the case in point). In this paper we extend the analysis and the discussion to take into account categorical (binary) variables (described in terms of proportions or odds); to favour the comprehension of the analogies (and discrepancies) between the two contexts we have fully developed an example that mimics the situation presented in the previous paper. Marginal, conditional, random effects and transitional models for correlated data are introduced in practical terms; the meaning of the different estimates obtained are interpreted for epidemiological purposes; the disadvantages of not considering correlation in the analysis are explained and the complexities connected to this type of analysis are fully appreciated. It is concluded that correlated data are very frequently encountered in occupational settings and that an appropriate analysis is necessary. This analysis requires sophisticated computer programs and statistical expertise, particularly in the case of categorical data.

Cluster Analysis↗

A note on robust variance estimation for cluster-correlated data.

There is a simple robust variance estimator for cluster-correlated data. While this estimator is well known, it is poorly documented, and its wide range of applicability is often not understood. The estimator is widely used in sample survey research, but the results in the sample survey literature are not easily applied because of complications due to unequal probability sampling. This brief note presents a general proof that the estimator is unbiased for cluster-correlated data regardless of the setting. The result is not new, but a simple and general reference is not readily available. The use of the method will benefit from a general explanation of its wide applicability.

Analysis of Variance↗

[Analysis of correlated data: problems and examples in industrial medicine].

In occupational health we are frequently faced with data that are not independent, but to recognize the lack of such an independence is not yet a common practice. To help researchers in the field to treat correlated data in a proper way this paper has two aims: to highlight practical situations in which the data are not independent and to show the main differences between a statistical analysis which does consider the correlation appropriately and one which doesn't. As to the first aim, four typical examples are discussed: repeated measurements in the same subject (e.g., cells in which the number of sister chromatid exchanges is counted); health effects observed in multiple organs (e.g., visual impairment in both eyes); evaluation of prevention programs (e.g., exposure assessment to styrene before and after environment remediation); longitudinal studies of health effects (e.g., changes over time of pulmonary function parameters). With respect to the second aim a practical exercise is described completely. Measurements of exposure to a toxic substance in two different departments before and after environment remediation are evaluated with statistical tools which both do and do not consider the correlation between such measurements. Differences in the results obtained, with particular reference to indices of variability (e.g., standard errors), point toward the need of analysing correlated data with appropriate statistical tools that take correlation into account.

Analysis of Variance↗

Graphical model checking with correlated response data.

Correlated response data arise often in biomedical studies. The generalized estimation equation (GEE) approach is widely used in regression analysis for such data. However, there are few methods available to check the adequacy of regression models in GEE. In this paper, a graphical method is proposed based on Cook and Weisberg's marginal model plot. A bootstrap method is applied to obtain the reference band to assess statistical uncertainties in comparing two marginal mean functions. We also propose using the generalized additive model (GAM) in a similar fashion. The proposed two methods are easy to implement by taking advantage of existing smoothing and GAM softwares for independent data. The usefulness of the methodology is demonstrated through application to a correlated binary data set drawn from a clinical trial, the Lung Health Study.

Adult↗

Sample size and power calculations with correlated binary data.

Correlated binary data are common in biomedical studies. Such data can be analyzed using Liang and Zeger's generalized estimating equations (GEE) approach. An attractive point of the GEE approach is that one can use a misspecified working correlation matrix, such as the working independence model (i.e., the identity matrix), and draw (asymptotically) valid statistical inference by using the so-called robust or sandwich variance estimator. In this article we derive some explicit formulas for sample size and power calculations under various common situations. The given formulas are based on using the robust variance estimator in GEE. We believe that these formulas will facilitate the practice in planning two-arm clinical trials with correlated binary outcome data.

Clinical Trials as Topic↗

A comparison of 3-D data correlation methods for fractionated stereotactic radiotherapy.

PURPOSE: Stereotactic radiosurgery is currently used to treat patients who are not good candidates for conventional neurosurgical procedures. For treatments of nonvascular tumor cells, it appears that fractionation offers a radiobiological advantage between tumor and normal tissues. Therefore, fractionated stereotactic radiotherapy (FSR) is preferred because it minimizes normal tissue complications and maximizes local tumor control probability. We have implemented a methodology clinically to perform the noninvasive patient repositioning technique. The 3-D data correlation method for high-precision and multiple fraction stereotactic treatments has been presented. METHODS AND MATERIALS: Three different optimization algorithms (Hooke and Jeeves optimization, simplex optimization, and simulated annealing optimization) are evaluated to calculate the transformation parameters necessary for FSR. A least-square object function is created to perform the 3-D data matching process. By minimizing the unconstrained object function value the best fit can be approached for the reference 3-D data sets. Simulation shows that these algorithms deliver results that are comparable to the previously published correlation algorithm (1,2) (singular value decomposition [SVD] method). The advantage for optimization algorithms is easily understood and can be readily implemented by using a personal computer (PC). The mathematical framework provides a tool to calculate the transformation matrix which can be used to adjust patient position for fractionated treatments. Therefore, using these algorithms for a high-precision fractionated treatment is possible without an invasive repeat fixation device and has been implemented clinically. A bite plate system was incorporated to acquire 3-D patient data. With a 3-D digital camera localization device, the patient motion can be followed in real time with the system calibrated to the isocenter. RESULTS: Two types of data sets are utilized to study the correlation results. One is using the digitized patient data which were retrieved clinically. The other is using the randomly generated data sets. Simulation errors for the optimization algorithms are all less than 1 mm in translation and less than 1 degree in rotation. Currently, FSR is performed using special designed repeat fixation devices which assure reproducible patient position for multiple fractions of radiation treatment. Clinical results indicated that this technique provided excellent treatment results. CONCLUSION: Three optimization algorithms have been applied and evaluated in calculating the transformation parameters between two 3-D contours or digitized data points. The mathematical functions behind these optimization algorithms are straightforward and can be easily implemented. When incorporated with the proper CT/MR image data with an electronic portal imaging (EPI) system, this process can possibly verify the patient's treatment position whenever there is doubt about the movement during the treatment procedure.

Algorithms↗

Analysis of correlated data in human biomonitoring studies. The case of high sister chromatid exchange frequency cells.

Sister chromatid exchange (SCE) analysis in peripheral blood lymphocytes is a well established technique that aims to evaluate human exposure to toxic agents. The individual mean value of SCE per cell had been the only recommended index to measure the extent of this cytogenetic damage until the early 1980's, when the concept of high frequency cells (HFC) was introduced to increase the sensitivity of the assay. All statistical analyses proposed thus far to handle these data are based on measures which refer to the individual mean values and not to the single cell. Although this approach allows the use of simple statistical methods, part of the information provided by the distribution of SCE per single cell within the individual is lost. Using the appropriate methods developed for the analysis of correlated data, it is possible to exploit all the available information. In particular, the use of random-effects models seems to be very promising for the analysis of clustered binary data such as HFC. Logistic normal random-effects models, which allow modelling of the correlation among cells within individuals, have been applied to data from a large study population to highlight the advantages of using this methodology in human biomonitoring studies. The inclusion of random-effects terms in a regression model could explain a significant amount of variability, and accordingly change point and/or interval estimates of the corresponding coefficients. Examples of coefficients that change across different regression models and their interpretation are discussed in detail. One model that seems particularly appropriate is the random intercepts and random slopes model.

Adult↗

Ordinal regression methodology for ROC curves derived from correlated data.

We present an approach for the analysis of correlated ROC data, using ordinal regression models in conjunction with generalized estimating equations. The approach applies to the analysis of degree-of-suspicion data derived from multiple interpretations of the same diagnostic study and from the examination of the same patients with multiple diagnostic modalities. The regression models make it possible to incorporate patient and reader characteristics into the analysis, without having to resort to stratification. We illustrate the potential of the approach with analysis of data from two studies in diagnostic oncology.

Diagnosis↗

[Use of GEE for modeling censored correlated data: application to the study of risk factors for withdrawal of totally implantable vascular access devices in cystic fibrosis].

BACKGROUND: The proportional hazards model proposed by Cox for modeling censored data is not suited for correlated delays, for instance when several events can be observed on each subject. METHODS: To analyze correlated delays, we propose to use a log-linear marginal model equivalent to Cox model. Correlations are taken into account through the use of Liang and Zeger's Generalized Estimating Equations (GEE) and of their robust variance estimator. An advantage of this method is that it can be implemented through the SAS GENMOD procedure. When ties are observed, we propose to use multiple imputations, creating M data sets without ties from the original one. RESULTS: This method is applied to a retrospective survey on the risk of withdrawing totally implantable vascular access devices (TIVAD) because of complication in cystic fibrosis patients: 265 TIVAD implanted in 200 patients were observed. Risk factors were characteristics of the device or of the patient. Results obtained with the robust variance estimator and ten imputations show that the use of the device for taking blood (vs exclusive perfusion of antibiotics), polyurethane catheter (vs. silicon), use of counterpressure for upkeeping and pulmonary colonization by Pseudomonas Aeruginosa are significantly associated to withdrawal. Under the Cox model which does not account for the correlations, some conclusions differ because the robust variance of the estimators is smaller than the variance obtained under the working assumption of independent delays. CONCLUSION: This approach allows the modeling of correlated survival data with SAS software. Our results illustrate the necessity of accounting for existing correlations.

Catheters, Indwelling↗

Teacher discipline and child misbehavior in day care: untangling causality with correlational data.

Day-care centers provide an ideal, underused setting for studying the developmental processes of child psychopathology. The influence of day-care teachers' lax and overreactive discipline on children's behavior problems was examined, as was the influence of children's behavior problems on teachers' discipline. Participants were 145 children and 16 day-care teachers from 8 classrooms in a day-care center for children from low-income families. Two techniques are presented for estimating causal relations based on correlational data gathered from day-care centers: 2-stage least squares and simultaneous structural equation modeling. Across techniques, teachers' laxness strongly influenced child misbehavior, and child misbehavior influenced both teachers' overreactivity and laxness. Teachers' overreactivity did not influence child misbehavior

Adult↗

A Monte Carlo procedure for two-stage tests with correlated data.

One strategy for mapping disease loci using marker-disease associations is to test for association with case-control samples and follow up a positive result with a family-based test. Using a family-based test in the second stage can help provide protection against false-positive results that can result from use of inappropriate controls and provides assurance that association identified in the first stage is occurring between linked loci. It is crucial for this two-stage strategy that the first stage be as powerful as possible to detect association since only positive results are tested in the second stage. In certain situations, the power of the first-stage test can be increased by combining the case-control and family data. However, this introduces correlation between the first- and second-stage tests, and treating them as independent tests causes a bias. Here we propose a Monte Carlo method that accounts for the correlation and provides the correct significance level for the second-stage test. We also discuss the use of a two-stage procedure when doing a genome scan for the data presented in the Genetic Analysis Workshop 9 study.

Genetic Markers↗

What can go wrong when you assume that correlated data are independent: an illustration from the evaluation of a childhood health intervention in Brazil.

The key analytical challenge presented by longitudinal data is that observations from one individual tend to be correlated. Although longitudinal data commonly occur in medicine and public health, the issue of correlation is sometimes ignored or avoided in the analysis. If longitudinal data are modelled using regression techniques that ignore correlation, biased estimates of regression parameter variances can occur. This bias can lead to invalid inferences regarding measures of effect such as odds ratios (OR) or risk ratios (RR). Using the example of a childhood health intervention in Brazil, we illustrate how ignoring correlation leads to incorrect conclusions about the effectiveness of the intervention.

Age Factors↗

Cervical biopsy/cytology correlation data can be collected prospectively and shared clinically.

Cervical cytology (Cy) and biopsy (Bx) correlation is used by institutions for the evaluation of their cytodiagnostic capabilities as a part of overall laboratory quality improvement (QI). However, the data obtained from correlation are not routinely included in most surgical pathology (SP) reports. Our laboratory's procedure is to include the correlation of the patient's previous (most recent) cytology smear in the surgical pathology report of all/any gynecologic surgical pathology specimens. We reviewed this process for the time period between July 1998-June 1999. Any noncorrelating cases were assigned a correlation review code by the reviewing cytopathologist: major Cy diagnostic error (DE1), minor Cy diagnostic error (DE2), Cy sampling error (Cy SE), or biopsy sampling error (Bx SE). Of 3,486 cases reviewed, 3,229 cases were satisfactory for correlation studies. Concordant results were found in 86.9%. Cy DE1 due to either Cy screening or interpretation errors or both were found in 0.2% (n = 7) of all cases, while Cy DE2 due to the same were found in 1% (n = 32). Bx SE accounted for discrepancies in 6.8% (n = 220) of all cases, while 5.1% (n = 164) of the total cases were discrepancies due to Cy SE. Follow-up Bx was available in 97.2% (n = 214) of the Bx SE, and showed 16.4% (n = 35) to be major discrepancies and 83.6% (n = 179) to be minor discrepancies. Cervical Cy/Bx correlation is useful for the evaluation of a laboratory's QI. It is also useful for the identification of either Cy or Bx SE. While QI data exist as "internal use only" documents, SE data (as part of the CC (correlation comment) included in SP reports) are vital to a specific/given patient. Bx SE was identified in 6.3% of our patients, indicating a possible need for rebiopsy. This type of QI data may be shared clinically, and may direct the management for maximum diagnostic and patient benefit.

Biopsy↗

[Methodology for analyzing censored correlated data: application of marginal and frailty approaches in human genetics. The European Community Alport Syndrome Concerted Action Group (ECASCA)].

BACKGROUND: Statistical analysis for correlated censored data allows to study censored events in clustered structure designs. Considering a possible correlation among failure times of the same group, standard methodology is no longer applicable. We investigated proposed models in this context to study familial data about a genetic disease, Alport syndrome. Alport syndrome is a severe hereditary disease due to abnormal collagenous chains. Renal failure is the main symptom of the disease. It progresses toward end-stage renal failure (IRT) according to a high time variability. As shown by genetic studies, mutations of COL4A5 gene are involved in the X-linked Alport Syndrome. Due to the large range of the mutation types, the aim of this study was to search for a possible genetic origin of the heterogeneity of the disease severity. METHODS: Marginal survival models and mixed effects survival models (so-called frailty models) were proposed to take into account the possible non independence of the observations. In this study, time until end-stage renal failure is a rightly censored end point. Possible intra-familial correlations due to shared environmental and/or genetic factors could induce dependence among familial failure times. In this paper, we fit marginal and frailty proportional hazards models to evaluate the effect of mutation type on the risk of IRT and an interfamilial heterogeneity of failure times. RESULTS: In this study, the use of these models allows to show the presence of an interfamilial heterogeneity of the failure times to IRT. Moreover, the results suggest that some mutation types are linked to a higher risk of fast evolution to IRT, which explains partially the interfamilial heterogeneity of the failure times. CONCLUSIONS: This paper shows the interest of marginal and frailty models to evaluate the heterogeneity of censored responses and to study relationships between a censored criterion and covariables. This study puts forward the importance of characterizing the mutation at a molecular level to understand the relationship between genotype and phenotype.

Data Interpretation, Statistical↗

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation↗