Search PubMed⌕ Search

PubMed · 11927463

Advanced statistics: statistical methods for analyzing cluster and cluster-randomized data.

Abstract

Sometimes interventions in randomized clinical trials are not allocated to individual patients, but rather to patients in groups. This is called cluster allocation, or cluster randomization, and is particularly common in health services research. Similarly, in some types of observational studies, patients (or observations) are found in naturally occurring groups, such as neighborhoods. In either situation, observations within a cluster tend to be more alike than observations selected entirely at random. This violates the assumption of independence that is at the heart of common methods of statistical estimation and hypothesis testing. Failure to account for the dependence between individual observations and the cluster to which they belong can have profound implications on the design and analysis of such studies. Their p-values will be too small, confidence intervals too narrow, and sample size estimates too small, sometimes to a dramatic degree. This problem is similar to that caused by the more familiar "unit of analysis error" seen when observations are repeated on the same subjects, but are treated as independent. The purpose of this paper is to provide an introduction to the problem of clustered data in clinical research. It provides guidance and examples of methods for analyzing clustered data and calculating sample sizes when planning studies. The article concludes with some general comments on statistical software for cluster data and principles for planning, analyzing, and presenting such studies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Robert L Wears. 2002. Advanced statistics: statistical methods for analyzing cluster and cluster-randomized data.. https://doi.org/10.1111/j.1553-2712.2002.tb01332.x

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

Fuzzy species among recombinogenic bacteria.

BACKGROUND: It is a matter of ongoing debate whether a universal species concept is possible for bacteria. Indeed, it is not clear whether closely related isolates of bacteria typically form discrete genotypic clusters that can be assigned as species. The most challenging test of whether species can be clearly delineated is provided by analysis of large populations of closely-related, highly recombinogenic, bacteria that colonise the same body site. We have used concatenated sequences of seven house-keeping loci from 770 strains of 11 named Neisseria species, and phylogenetic trees, to investigate whether genotypic clusters can be resolved among these recombinogenic bacteria and, if so, the extent to which they correspond to named species. RESULTS: Alleles at individual loci were widely distributed among the named species but this distorting effect of recombination was largely buffered by using concatenated sequences, which resolved clusters corresponding to the three species most numerous in the sample, N. meningitidis, N. lactamica and N. gonorrhoeae. A few isolates arose from the branch that separated N. meningitidis from N. lactamica leading us to describe these species as 'fuzzy'. CONCLUSION: A multilocus approach using large samples of closely related isolates delineates species even in the highly recombinogenic human Neisseria where individual loci are inadequate for the task. This approach should be applied by taxonomists to large samples of other groups of closely-related bacteria, and especially to those where species delineation has historically been difficult, to determine whether genotypic clusters can be delineated, and to guide the definition of species.

Cluster Analysis↗

Scaling and clustering in the study of semantic disruptions in patients with schizophrenia: a re-evaluation.

Some recent studies of semantics in schizophrenia have employed multidimensional scaling and clustering techniques to analyse verbal fluency and triadic comparison data. The conclusions have been: (i) patients generate fewer words in fluency tasks and display more variable similarity groupings of words in triadic tasks, and (ii) this is due to deficits in semantics. We analysed data from both tasks. On the verbal fluency task, patients produced significantly fewer responses than controls. The results also showed little patient-specific inter-individual consistency. Similarly, for triadic comparison data, we did not find much patient-specific inter-individual consistency. When correlating patients' results at different measurement times with means of controls, the data of individual patients (at either of the two measurement times) were not predicted better from their data at the other measurement time than from controls. This latter finding suggests little patient-specific intra-individual consistency and, thus, pleads against idiosyncratic semantic deficits. Our findings do not refute the hypothesis that schizophrenia is associated with semantic disruptions. However, our results demonstrate that because of severe statistical restrictions and requirements associated with some scaling and clustering techniques, these methods may not be as useful in this enterprise as previously thought.

Cluster Analysis↗