Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Correlation Of Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Atmospheric metal deposition in a moss data correlation study with mortality and disease in the Netherlands.

The present paper addresses the correlations between moss metal concentrations and epidemiological data on health and mortality rates in The Netherlands. Attention was given to both total and fractionated metal concentrations in the moss tissues, the latter by factor-analytical (mathematical) approaches, and to both grouped and specific diseases. Better than 95% probability correlations were found both for total moss elements and mortality due to specific diseases and for fractionated moss elements and mortality rates summed for grouped diseases. Overall, the presented data suggest that correlation studies between biomonitoring data on metal air pollution and (epidemiological) health data may prove valuable in turning attention to specific metal-health issues and in directing further study into possible dose-response mechanisms in air-associated metal epidemiology.

Air Pollutants↗

Cervical biopsy/cytology correlation data can be collected prospectively and shared clinically.

Cervical cytology (Cy) and biopsy (Bx) correlation is used by institutions for the evaluation of their cytodiagnostic capabilities as a part of overall laboratory quality improvement (QI). However, the data obtained from correlation are not routinely included in most surgical pathology (SP) reports. Our laboratory's procedure is to include the correlation of the patient's previous (most recent) cytology smear in the surgical pathology report of all/any gynecologic surgical pathology specimens. We reviewed this process for the time period between July 1998-June 1999. Any noncorrelating cases were assigned a correlation review code by the reviewing cytopathologist: major Cy diagnostic error (DE1), minor Cy diagnostic error (DE2), Cy sampling error (Cy SE), or biopsy sampling error (Bx SE). Of 3,486 cases reviewed, 3,229 cases were satisfactory for correlation studies. Concordant results were found in 86.9%. Cy DE1 due to either Cy screening or interpretation errors or both were found in 0.2% (n = 7) of all cases, while Cy DE2 due to the same were found in 1% (n = 32). Bx SE accounted for discrepancies in 6.8% (n = 220) of all cases, while 5.1% (n = 164) of the total cases were discrepancies due to Cy SE. Follow-up Bx was available in 97.2% (n = 214) of the Bx SE, and showed 16.4% (n = 35) to be major discrepancies and 83.6% (n = 179) to be minor discrepancies. Cervical Cy/Bx correlation is useful for the evaluation of a laboratory's QI. It is also useful for the identification of either Cy or Bx SE. While QI data exist as "internal use only" documents, SE data (as part of the CC (correlation comment) included in SP reports) are vital to a specific/given patient. Bx SE was identified in 6.3% of our patients, indicating a possible need for rebiopsy. This type of QI data may be shared clinically, and may direct the management for maximum diagnostic and patient benefit.

Biopsy↗

[Methodology for analyzing censored correlated data: application of marginal and frailty approaches in human genetics. The European Community Alport Syndrome Concerted Action Group (ECASCA)].

BACKGROUND: Statistical analysis for correlated censored data allows to study censored events in clustered structure designs. Considering a possible correlation among failure times of the same group, standard methodology is no longer applicable. We investigated proposed models in this context to study familial data about a genetic disease, Alport syndrome. Alport syndrome is a severe hereditary disease due to abnormal collagenous chains. Renal failure is the main symptom of the disease. It progresses toward end-stage renal failure (IRT) according to a high time variability. As shown by genetic studies, mutations of COL4A5 gene are involved in the X-linked Alport Syndrome. Due to the large range of the mutation types, the aim of this study was to search for a possible genetic origin of the heterogeneity of the disease severity. METHODS: Marginal survival models and mixed effects survival models (so-called frailty models) were proposed to take into account the possible non independence of the observations. In this study, time until end-stage renal failure is a rightly censored end point. Possible intra-familial correlations due to shared environmental and/or genetic factors could induce dependence among familial failure times. In this paper, we fit marginal and frailty proportional hazards models to evaluate the effect of mutation type on the risk of IRT and an interfamilial heterogeneity of failure times. RESULTS: In this study, the use of these models allows to show the presence of an interfamilial heterogeneity of the failure times to IRT. Moreover, the results suggest that some mutation types are linked to a higher risk of fast evolution to IRT, which explains partially the interfamilial heterogeneity of the failure times. CONCLUSIONS: This paper shows the interest of marginal and frailty models to evaluate the heterogeneity of censored responses and to study relationships between a censored criterion and covariables. This study puts forward the importance of characterizing the mutation at a molecular level to understand the relationship between genotype and phenotype.

Data Interpretation, Statistical↗

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation↗

[Modeling correlated data in epidemiology: mixed or marginal model?].

Correlated observations (within centers, families, subjects,.) are common in epidemiology. Even when one is only interested in the modeling of means according to risk factors, it is also necessary to model the variance-covariance matrix of the observations in order to make correct inferences on the parameters of interest. All the more so when the aim of the survey is the measurement of these correlations or of the variance of the random effects from which they are assumed to originate. We discuss, within the framework of the linear and of the logistic models, the implications of two choices for the modeling of covariances. The mixed model shows the unobserved elements responsible for the similarity between certain observations. In a longitudinal survey, for instance, one can use a random effect, specific to each subject, expressing how much a subject's trajectory is translated as compared to what is expected according to its characteristics (age, sex,.). The marginal approach leads to modeling separately the means and the covariance matrix of the observations. The distinction between these two approaches is important for non linear models, in particular the logistic one. We insist on the interconnection between a mixed model formulation and a marginal one, as well as on the implication of the choice in terms of the parameters' interpretation.

Epidemiologic Research Design↗

Statistical analysis of correlated data using generalized estimating equations: an orientation.

The method of generalized estimating equations (GEE) is often used to analyze longitudinal and other correlated response data, particularly if responses are binary. However, few descriptions of the method are accessible to epidemiologists. In this paper, the authors use small worked examples and one real data set, involving both binary and quantitative response data, to help end-users appreciate the essence of the method. The examples are simple enough to see the behind-the-scenes calculations and the essential role of weighted observations, and they allow nonstatisticians to imagine the calculations involved when the GEE method is applied to more complex multivariate data.

Body Height↗

Separation of individual-level and cluster-level covariate effects in regression analysis of correlated data.

The focus of this paper is regression analysis of clustered data. Although the presence of intracluster correlation (the tendency for items within a cluster to respond alike) is typically viewed as an obstacle to good inference, the complex structure of clustered data offers significant analytic advantages over independent data. One key advantage is the ability to separate effects at the individual (or item-specific) level and the group (or cluster-specific) level. We review different approaches for the separation of individual-level and cluster-level effects on response, their appropriate interpretation and give recommendations for model fitting based on the intent of the data analyst. Unlike many earlier papers on this topic, we place particular emphasis on the interpretation of the cluster-level covariate effect. The main ideas of the paper are highlighted in an analysis of the relationship between birth weight and IQ using sibling data from a large birth cohort study.

Birth Weight↗

Rerandomization tests for analyzing correlated data from dental studies.

Dental research studies often produce relatively small data sets in which observations are serially or spatially correlated. Rerandomization tests are presented as alternatives to analysis of variance and multivariate analysis for assessing group differences using such data. Rerandomization tests are particularly useful when the investigator is unwilling to make strong assumptions about the nature of the serial correlation or the distribution of the data. Two examples are discussed that demonstrate these techniques.

Data Interpretation, Statistical↗

Probabilistic structure calculations: a three-dimensional tRNA structure from sequence correlation data.

Algorithms based on probability theory can address issues of uncertainty directly through their representational framework and their theory for data combination. In this paper, we discuss the advantages of probabilistic formulations for molecular-structure calculations, describe one implementation of such a formulation, and show its performance on a data set derived from analysis of the statistical correlations within a set of aligned transfer RNA sequences. By assigning reasonable physical interpretations to certain statistical correlations, we are able to calculate three-dimensional structures for tRNA from a random starting structure. The constraints that we use are associated with different variances, and so their effects are not uniform, and must be reconciled by a probabilistic algorithm to yield the most likely structure. As might be predicted, the uncertainty in the position for each base is a function of both the number and strength of the constraints, and is reflected in the variances in atomic position calculated by the algorithm. For example, the hinge region in the tRNA is shown to be the most uncertain. In addition, the algorithm retains information about positional covariation that is useful for understanding the relationships between different parts of the structure. These experiments also demonstrate that we can define a single-sphere representation for each base that is useful for nucleic acid structural calculations in the same way that alpha-carbon representations are useful for protein structural calculations.

Algorithms↗

Statistical inference for correlated data in ophthalmologic studies.

In ophthalmologic studies, each subject usually contributes important information for each of two eyes and the values from the two eyes are generally highly correlated. Previous studies showed that test procedures for binary paired data that ignore the presence of intraclass correlation could lead to inflated significance levels. Furthermore, it is possible that asymptotic versions of these procedures that take the intraclass correlation into account could also produce unacceptably high type I error rates when the sample size is small or the data structure is sparse. We propose two alternatives for these situations, namely the exact unconditional and approximate unconditional procedures. According to our simulation results, the exact procedures usually produce extremely conservative empirical type I error rates. That is, the corresponding type I error rates could greatly underestimate the pre-assigned nominal level (e.g. (empirical type I error rate/nominal type I error rate) 0.8). On the other hand, the approximate unconditional procedures usually yield empirical type I error rates close to the pre-chosen nominal level. We illustrate our methodologies with a data set from a retinal detachment study.

Biometry↗

A semiparametric bootstrap approach to correlated data analysis problems.

In this note, we outline a simple to use yet powerful bootstrap algorithm for handling correlated outcome variables in terms of either hypothesis testing or confidence intervals using only the marginal models. This new method can handle combinations of continuous and discrete data and can be used in conjunction with other covariates in a model. The procedure is based upon estimating the family-wise error (FWE) rate and then making a Bonferroni-type correction. A simulation study illustrates the accuracy of the algorithm over a variety of correlation structures.

Algorithms↗

Signal recovery from autocorrelation and cross-correlation data.

A signal recovery technique is motivated and derived for the recovery of several nonnegative signals from measurements of their autocorrelation and cross-correlation functions. The iterative technique is shown to preserve nonnegativity of the signal estimates and to produce a sequence of estimates whose correlations better approximate the measured correlations as the iterations proceed. The method is demonstrated on simulated data for active imaging with dual-frequency or dual-polarization illumination.

Algorithms↗

Implications of regression analysis and correlational data between subtests of the WISC-R and the PPVT-R for a delinquent population.

Computed correlations between the subscales of the Wechsler Intelligence Scales for Children-Revised (WISC-R) and the Peabody Picture Vocabulary Test-Revised (PPVT-R) using as a sample 72 adjudicated male delinquents aged 13-10 to 16-11. Significant relationships at the .0001 level were obtained for 10 subtests with only one, Object Assembly, computed at the .001 level. A forward selection multiple regression analysis resulted in six subtests of the WISC-R correlating to the PPVT-R with a R2 value of .78. The significance and the implications of this relationship for the juvenile delinquent population were discussed.

Adolescent↗

Database tools for integrating and searching membrane property data correlated with neuronal morphology.

A critical problem in neuroscience is the lack of database tools for integrating neuronal property data. We report here the development of a combined object oriented-relational database (NeuronDB, http://senselab.med.yale.edu/neurondb) that meets these needs by providing tools for integrating data within neurons and comparing data across neurons. It focuses on three types of neuronal properties voltage-gated channels, neurotransmitter receptors, and neurotransmitters. The data are organized in relation to different regions of neurons as represented in canonical forms; using simple canonical models of complex cells as a vehicle for indexing information permits the database to be searchable across different neurons. Using these multidimensional search tools, users can locate specific properties in specific regions of a neuron; obtain integrated summaries of all properties within a region; and carry out searches to compare properties across equivalent compartments in different neurons. These tools thus permit searches of the multidimensional neuron property space equivalent to homology searches of sequence databases. NeuronDB is accessible over the Internet; it provides immediate links to citation indexes and abstracts supporting the deposited data, and annotations that indicate the state of acceptance of the data. Users are encouraged to contribute data. The ability to input the data from NeuronDB directly to NEURON and GENESIS is being developed. As a shared Web resource, NeuronDB should enhance the efforts of neuroscientists and neuronal modellers to analyze and compare the functional operations of different types of neurons.

Automation↗

The value of monitoring frozen section-permanent section correlation data over time.

CONTEXT: The effectiveness of the long-term monitoring of errors detected by frozen section-permanent section correlation is unknown. OBJECTIVE: To determine factors important in laboratory improvement in frozen section-permanent section discordant and deferral rates by participation in a multi-institutional continuous quality improvement program. DESIGN: Participants in the College of American Pathologists Q-Tracks program self-reported the number of anatomic pathology frozen-permanent section discordant and deferred cases in their laboratories by prospectively performing secondary review of intraoperative consultations. Laboratories participated in the program for 1 to 5 years and reported their data every quarter. We calculated mean and median discordant and deferred case frequencies and used mixed linear modeling to determine if length of participation in the program was associated with improved performance. PARTICIPANTS: One hundred seventy-four laboratories self-reported data. MAIN OUTCOME MEASURES: Mean frozen-permanent section discordant and deferred diagnostic frequencies and changes in these frequencies over time were measured. RESULTS: The mean and median frozen-permanent section discordant frequencies were 1.36% and 0.70%, respectively. The mean and median deferred diagnostic frequencies were 2.35% and 1.20%, respectively. Longer participation in the Q-Tracks program was significantly associated (P = .04) with lower discordant frequencies; 4- or 5-year participation showed a decrease in discordant frequency of 0.99%, whereas 1-year participation showed a decrease in discordant frequency of 0.84%. Longer participation in the Q-Tracks monitor was associated with lower microscopic sampling frequencies for discordant diagnoses (P = .04). Increased length of participation in the Q-Tracks program was significantly associated (P = .04) with lower deferred diagnostic frequencies. CONCLUSIONS: Long-term monitoring of frozen-permanent section correlation is associated with sustained improvement in performance.

Diagnostic Errors↗

Analog processing of vestibular nystagmus for on-line cross- correlation data analysis.

An analog processing circuit is described which allow accurate measurement of the phase relationships between input angular acceleration and resulting eye velocity. Vestibular nystagmic data are processed via analog technics to yield slowphase eye velocity. The turntable velocity input is cross-correlated with the eye velocity output, using a Nicolet MED-80 minicomputer system. The resulting correlograms are further processed to obtain precise phase information. Test data analysis shows a system resolution within 1 degree. Data from human and animal subjects are portrayed.

Acceleration↗

New hydrophilicity scale derived from high-performance liquid chromatography peptide retention data: correlation of predicted surface residues with antigenicity and X-ray-derived accessible sites.

A new set of hydrophilicity high-performance liquid chromatography (HPLC) parameters is presented. These parameters were derived from the retention times of 20 model synthetic peptides, Ac-Gly-X-X-(Leu)3-(Lys)2-amide, where X was substituted with the 20 amino acids found in proteins. Since hydrophilicity parameters have been used extensively in algorithms to predict which amino acid residues are antigenic, we have compared the profiles generated by our new set of hydrophilic HPLC parameters on the same scale as nine other sets of parameters. Generally, it was found that the HPLC parameters obtained in this study correlated best with antigenicity. In addition, it was shown that a combination of the three best parameters for predicting antigenicity further improved the predictions. These predicted surface sites or, in other words, the hydrophilic, accessible, or mobile regions were then correlated to the known antigenic sites from immunological studies and accessible sites determined by X-ray crystallographic data for several proteins.

Amino Acids↗