Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

scATAnno: Automated Cell Type Annotation for Single-cell ATAC-seq Data.

Recent advances in single-cell epigenomic techniques have increased the demand for single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) analysis. One key analytical task is to determine cell type identity based on epigenetic data. Here, we introduce scATAnno, a Python package designed to automatically annotate scATAC-seq data using large-scale scATAC-seq reference atlases. This workflow generates reference atlases from publicly available datasets, enabling accurate cell type annotation by integrating query data with reference atlases without the use of single-cell RNA sequencing (scRNA-seq) data. To enhance annotation accuracy, we incorporated k-nearest neighbors (KNN)-based and weighted distance-based uncertainty scores to effectively detect cell populations within the query data that are distinct from all cell types in the reference data. We compared and benchmarked scATAnno against five other published cell annotation approaches, demonstrating its superior performance across multiple datasets and metrics. We further showcased the utility of scATAnno across multiple datasets, including peripheral blood mononuclear cells (PBMCs), triple-negative breast cancer (TNBC), and basal cell carcinoma (BCC), and demonstrated that scATAnno accurately annotates cell types across diverse biological conditions. Overall, scATAnno is a useful tool for scATAC-seq reference atlas construction and cell type annotation and can facilitate the interpretation of new scATAC-seq datasets in complex biological systems. scATAnno is publicly available at https://scatanno-main.readthedocs.io/.

Single-Cell Analysis↗

Comprehensive Evaluation and Explainable Interpretation of Peptide-HLA Binding Prediction Tools.

Accurate prediction of peptide binding to human leukocyte antigen class I (HLA-I) molecules is critical for advancing immunological research, particularly in vaccine design and immunotherapy. However, limitations in model performance, interpretability, and dataset quality impede the widespread adoption of existing predictive tools. Here, we present a comprehensive evaluation of 17 HLA-I peptide binding prediction models, utilizing a meticulously curated dataset comprising over 290,000 peptides spanning 44 HLA-I alleles. We assessed model accuracy, robustness, and interpretability, employing explainability techniques such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) to elucidate underlying prediction mechanisms. Our results reveal substantial performance disparities, with self-attention-based models, including STMHCpan and BigMHC, exhibiting superior accuracy. Notably, the capsule network model CapsNet-MHC_AN demonstrated robust performance. Models trained on eluted ligand datasets outperformed those relying on binding affinity data, underscoring the critical role of high-quality training data. Ensemble and multi-algorithm approaches further improved prediction reliability. These findings highlight the need for ongoing innovation in model architecture, integration of diverse and high-quality datasets, and incorporation of structural predictors to develop more accurate, interpretable, and clinically applicable HLA-I peptide binding prediction tools.

HLA-I binding↗

Suppression of LKB1-mutant lung adenocarcinoma by natural killer cells from females.

BACKGROUND: This study addressed the enigma of sex differences in smoking-related lung cancer, particularly focusing on the low LKB1 mutation frequency in female patients with lung adenocarcinoma. METHODS: Sex bias was studied with a genetically engineered mouse model and various tail-vein injection models. Immune cells were analyzed by antibody-depletion study, flow cytometry, and immunofluorescence. The relevance of our findings to human disease was validated by evaluating various lung adenocarcinoma datasets. All statistical tests are 2-sided. RESULTS: A statistically significant percentage of females are resistant to LKB1-mutant tumor formation in our models, reflecting this sex difference in humans. Natural killer (NK) cells were identified as a critical factor in this sex-biased response. This sex difference was observed primarily in LKB1-mutant lung adenocarcinoma, probably due to their low major histocompatibility complex class I level, making them the ideal target for NK cells through the missing-self recognition. Although females resistant to LKB1-mutant lung adenocarcinoma formation did not have enhancement of any specific NK subpopulation, our immunofluorescence analysis revealed high numbers of NKs in female lungs even with the presence of LKB1-mutant lung adenocarcinoma. Our gene set enrichment analysis of The Cancer Genome Atlas-lung adenocarcinoma dataset also showed that female LKB1-mutant lung adenocarcinoma patients have a stronger NK-mediated response after adjusting for other male-female differences using the LKB1 wild-type lung adenocarcinoma dataset. CONCLUSION: Females have a stronger NK-mediated response against LKB1-mutant lung adenocarcinoma, which was present in our mouse model and the human lung adenocarcinoma dataset. This study revealed a novel role of NK cells in suppressing LKB1-mutant lung adenocarcinoma in females, which should be assessed in the clinical setting in the future.

Killer Cells, Natural↗

ERCnet: Phylogenomic Prediction of Interaction Networks in the Presence of Gene Duplication.

Assigning gene function from genome sequences is a rate-limiting step in molecular biology research. A protein's position within an interaction network can potentially provide insights into its molecular mechanisms. Phylogenetic analysis of evolutionary rate covariation (ERC) in protein sequence has been shown to be effective for large-scale prediction of functional relationships and interactions. However, gene duplication, gene loss, and other sources of phylogenetic incongruence are barriers for analyzing ERC on a genome-wide basis. Here, we developed ERCnet, a bioinformatic program designed to overcome these challenges, facilitating efficient all-versus-all ERC analyses for large protein sequence datasets. We simulated proteome datasets and found that ERCnet achieves combined false positive and negative error rates well below 10% and that our novel "branch-by-branch" length measurements outperforms "root-to-tip" approaches in most cases, offering a valuable new strategy for performing ERC. We also compiled a sample set of 35 angiosperm genomes to test the performance of ERCnet on empirical data, including its sensitivity to user-defined analysis parameters such as input dataset size and branch-length measurement strategy. We investigated the overlap between ERCnet runs with different species samples to understand how species number and composition affect predicted interactions and to identify the protein sets that consistently exhibit ERC across angiosperms. Our systematic exploration of the performance of ERCnet provides a roadmap for design of future ERC analyses to predict functional interactions in a wide array of genomic datasets. ERCnet code is freely available at https://github.com/EvanForsythe/ERCnet.

Gene Duplication↗

SOURCE: a unified genomic resource of functional annotations, ontologies, and gene expression data.

The explosion in the number of functional genomic datasets generated with tools such as DNA microarrays has created a critical need for resources that facilitate the interpretation of large-scale biological data. SOURCE is a web-based database that brings together information from a broad range of resources, and provides it in manner particularly useful for genome-scale analyses. SOURCE's GeneReports include aliases, chromosomal location, functional descriptions, GeneOntology annotations, gene expression data, and links to external databases. We curate published microarray gene expression datasets and allow users to rapidly identify sets of co-regulated genes across a variety of tissues and a large number of conditions using a simple and intuitive interface. SOURCE provides content both in gene and cDNA clone-centric pages, and thus simplifies analysis of datasets generated using cDNA microarrays. SOURCE is continuously updated and contains the most recent and accurate information available for human, mouse, and rat genes. By allowing dynamic linking to individual gene or clone reports, SOURCE facilitates browsing of large genomic datasets. Finally, SOURCEs batch interface allows rapid extraction of data for thousands of genes or clones at once and thus facilitates statistical analyses such as assessing the enrichment of functional attributes within clusters of genes. SOURCE is available at http://source.stanford.edu.

Animals↗

GANA--a genetic algorithm for NMR backbone resonance assignment.

NMR data from different experiments often contain errors; thus, automated backbone resonance assignment is a very challenging issue. In this paper, we present a method called GANA that uses a genetic algorithm to automatically perform backbone resonance assignment with a high degree of precision and recall. Precision is the number of correctly assigned residues divided by the number of assigned residues, and recall is the number of correctly assigned residues divided by the number of residues with known human curated answers. GANA takes spin systems as input data and uses two data structures, candidate lists and adjacency lists, to assign the spin systems to each amino acid of a target protein. Using GANA, almost all spin systems can be mapped correctly onto a target protein, even if the data are noisy. We use the BioMagResBank (BMRB) dataset (901 proteins) to test the performance of GANA. To evaluate the robustness of GANA, we generate four additional datasets from the BMRB dataset to simulate data errors of false positives, false negatives and linking errors. We also use a combination of these three error types to examine the fault tolerance of our method. The average precision rates of GANA on BMRB and the four simulated test cases are 99.61, 99.55, 99.34, 99.35 and 98.60%, respectively. The average recall rates of GANA on BMRB and the four simulated test cases are 99.26, 99.19, 98.85, 98.87 and 97.78%, respectively. We also test GANA on two real wet-lab datasets, hbSBD and hbLBD. The precision and recall rates of GANA on hbSBD are 95.12 and 92.86%, respectively, and those of hbLBD are 100 and 97.40%, respectively.

Algorithms↗

Information-theoretic identification of predictive SNPs and supervised visualization of genome-wide association studies.

The size, dimensionality and the limited range of the data values makes visualization of single nucleotide polymorphism (SNP) datasets challenging. The purpose of this study is to evaluate the usefulness of 3D VizStruct, a novel multi-dimensional data visualization technique for SNP datasets capable of identifying informative SNPs in genome-wide association studies. VizStruct is an interactive visualization technique that reduces multi-dimensional data to three dimensions using a combination of the discrete Fourier transform and the Kullback-Leibler divergence. The performance of 3D VizStruct was challenged with several diverse, biologically relevant published datasets including the human lipoprotein lipase (LPL) gene locus, the human Y-chromosome in several populations and a multi-locus genotype dataset of coral samples from four populations. In every case, the SNPs and or polymorphic markers identified by the 3D VizStruct mapping were predictive of the underlying biology.

Animals↗

The HIV positive selection mutation database.

The HIV positive selection mutation database is a large-scale database available at http://www.bioinformatics.ucla.edu/HIV/ that provides detailed selection pressure maps of HIV protease and reverse transcriptase, both of which are molecular targets of antiretroviral therapy. This database makes available for the first time a very large HIV sequence dataset (sequences from approximately 50 000 clinical AIDS samples, generously contributed by Specialty Laboratories, Inc.), which makes possible high-resolution selection pressure mapping. It provides information about not only the selection pressure on individual sites but also how selection pressure at one site is affected by mutations on other sites. It also includes datasets from other public databases, namely the Stanford HIV database [S. Y. Rhee, M. J. Gonzales, R. Kantor, B. J. Betts, J. Ravela and R. W. Shafer (2003) Nucleic Acids Res., 31, 298-303]. Comparison between these datasets in the database enables cross-validation with independent datasets and also specific evaluation of the effect of drug treatment.

Acquired Immunodeficiency Syndrome↗

Statistical power analysis to estimate how many months of data are required to identify operating room staffing solutions to reduce labor costs and increase productivity.

UNLABELLED: We performed a statistical power analysis to determine how many historical data are needed for optimal operating room (OR) management decision making. The work applies to hospitals that provide service for all of its surgeons' elective cases on whatever workday the surgeons and patients choose. The hospital and anesthesia group adjust OR staffing and patient scheduling to care for the patients while minimizing OR staffing costs and maximizing labor productivity. Two years of data were obtained from a seven-OR surgical suite. The data were repeatedly split into training and testing datasets. The optimal staffing solution was calculated for each training dataset to maximize the efficiency of OR time usage and was then applied to the corresponding testing dataset. Training datasets ranged in size from 30 to 270 consecutive workdays. With 30 workdays of data, the statistical method identified staffing solutions that had an average of 35% decreased costs and 27% increased productivity as compared to the existing staffing plan. There was no significant improvement in performance with more than 210 workdays (10 mo) of data. With 30 workdays of OR or anesthesia group data, the optimization method can significantly reduce staffing costs and increase productivity compared with existing staffing. When applied routinely for adjusting staffing (e.g., on a quarterly basis), 9 to 12 mo of data should be used. IMPLICATIONS: With 30 workdays of operating room or anesthesia group data, the optimization method can propose staffing solutions that significantly decrease costs and increase productivity compared with existing staffing solutions. We recommend that, when the statistical method is applied routinely for adjusting staffing (e.g., on a quarterly basis), 9 to 12 mo of data be used.

Efficiency↗

Improved radiologic staging of lung cancer with 2-[18F]-fluoro-2-deoxy-D-glucose-positron emission tomography and computed tomography registration.

PURPOSE: To determine if volumetric nonlinear registration or registration of thoracic computed tomography (CT) and 2-[18F]-fluoro-2-deoxy-D-glucose-positron emission tomography (FDG-PET) datasets changes the detection of mediastinal and hilar nodal disease in patients undergoing staging for lung cancer and if it has any impact on radiologic lung cancer staging. METHOD: Computer-based image registration was performed on 45 clinical thoracic helical CT and FDG-PET scans of patients with lung cancer who were staged by mediastinoscopy and/or thoracotomy. Thoracic CT, FDG-PET, and registration datasets were each interpreted by 2 readers for the presence of metastatic nodal disease and were staged independently of each other. Results were compared with surgical pathologic findings. RESULTS: One hundred and thirty lymph node stations in the mediastinum and hila were evaluated each on CT, PET, and registration datasets. Sensitivity, specificity, positive predictive value, and negative predictive value, respectively, for detecting metastatic nodal disease for CT were 74%, 78%, 55%, 88%; for PET with CT side by side, 59% to 76%, 77% to 89%, 48% to 68%, and 84% to 91%; and for CT-PET registration, 71% to 76%, 89% to 96%, 70% to 86%, and 90% to 91%. Registration images were significantly more sensitive in detecting nodal disease over PET for 1 reader (P = 0.0156) and were more specific than PET (P = 0.0107 and 0.0017) in identifying the absence of mediastinal disease for both readers. Registration was significantly more accurate for staging when compared with PET for both readers (P = 0.002 and 0.035). CONCLUSION: Registration of CT and FDG-PET datasets significantly improved the specificity of detecting metastatic disease. In addition, registration improved the radiologic staging of lung cancer patients when compared with CT or FDG-PET alone.

Adult↗

Compression and reconstruction of sorted PET listmode data.

BACKGROUND: In nuclear medicine data can be stored in histogram or listmode format. The most popular histogram format is the planar projection format. Due to the increase in detector blocks, the improved energy resolution and the trends towards time of flight, dynamic and gated imaging, it can be more appropriate to store the data in listmode format. The size of the storage in this format increases linearly with the number of properties (positions, energy, time info) while the histogram format increases exponentially. However, the datasize of listmode data also increases linearly with the number of coincidences. Due to the high number of counts in 3D PET this will lead to very large datasets. Therefore a good compression algorithm for listmode data is very important. METHODS: A sorting and compression method is proposed to reduce the amount of space needed to store the listmode dataset. One event is represented by one number without any information loss compared to the original listmode file. The next step is to sort all events into an array of increasing numbers. These data are compressed by the gzip routine. One of the advantages of 3D PET listmode reconstructions is that they result in a more uniform resolution across the field of view (FOV), which is not always true for other reconstruction algorithms. This improved resolution is shown for the listmode data of a gamma camera operating in PET mode. RESULTS: First the effect of positional accuracy in the listmode dataset is evaluated by comparing resolution in the reconstructions. It is shown that the highest accuracy is not necessary and a significant reduction in the size of the dataset can be obtained prior to lossless compression. A further reduction can be obtained by using the proposed sorting and compression techniques. It is shown that the storage space decreases linearly with the logarithm of the number of coincidences. The compression obtained by different acquisition matrices was compared. Finally it is shown that the 3D listmode reconstruction of sorted listmode data is faster because of improved cache behaviour. The method can be applied to any kind of listmode data. The compression factors will improve when the ratio of measured events to possible events increases.

Algorithms↗

Impact of differences in diagnostic criteria when determining the incidence of contact lens-associated keratitis.

PURPOSE: The purpose of this study is to examine the effect of differences in within-study and between-study diagnostic criteria in determining the incidence of contact lens-associated keratitis. METHODS: We applied the sets of criteria for "microbial keratitis" as described in five previous studies to the dataset of Morgan et al., which documents 118 cases of contact lens-associated keratitis across a wide range of clinical severities. For each set of criteria, the incidence of contact lens-associated keratitis was calculated for the following five lens type/modality combinations: daily-wear rigid, daily-wear daily disposable hydrogel, daily-wear hydrogel, extended-wear hydrogel, and extended-wear silicone hydrogel. The effect of varying the clinical severity score for the differentiation of nonsevere versus severe keratitis was also examined with respect to the dataset of Morgan et al. RESULTS: The size and location of the corneal infiltrative events identified as representing "microbial keratitis" for each of the different sets of criteria are illustrated in a series of cartograms. A key between-study difference in the incidence values calculated for the various sets of criteria relates to the categories of extended-wear hydrogel and extended-wear silicone hydrogel lenses. Specifically, the incidence of "microbial keratitis" was found to be statistically significantly greater for extended-wear hydrogel compared with extended-wear silicone hydrogel lenses when the set of criteria of Morgan et al. was applied, but not when the other sets of criteria were applied, to the dataset of Morgan et al. Increasing the threshold clinical severity criterion for differentiating between nonsevere and severe keratitis within this dataset resulted in lower incidence values; however, such changes in threshold had a minimal impact on relative risk values. CONCLUSIONS: The choice of criteria for diagnosing contact lens-associated "microbial keratitis" has a significant impact on calculations of the incidence of this condition.

Contact Lenses↗

Dissection of phylogenetic relationships among 19 rapidly growing Mycobacterium species by 16S rRNA, hsp65, sodA, recA and rpoB gene sequencing.

The current classification of non-pigmented and late-pigmenting rapidly growing mycobacteria (RGM) capable of producing disease in humans and animals consists primarily of three groups, the Mycobacterium fortuitum group, the Mycobacterium chelonae-abscessus group and the Mycobacterium smegmatis group. Since 1995, eight emerging species have been tentatively assigned to these groups on the basis of their phenotypic characters and 16S rRNA gene sequence, resulting in confusing taxonomy. In order to assess further taxonomic relationships among RGM, complete sequences of the 16S rRNA gene (1483-1489 bp), rpoB (3486-3495 bp) and recA (1041-1056 bp) and partial sequences of hsp65 (420 bp) and sodA (441 bp) were determined in 19 species of RGM. Phylogenetic trees based upon each gene sequence, those based on the combined dataset of the five gene sequences and one based on the combined dataset of the rpoB and recA gene sequences were then compared using the neighbour-joining, maximum-parsimony and maximum-likelihood methods after using the incongruence length difference test. Combined datasets of the five gene sequences comprising nearly 7000 bp and of the rpoB+recA gene sequences comprising nearly 4600 bp distinguished six phylogenetic groups, the M. chelonae-abscessus group, the Mycobacterium mucogenicum group, the M. fortuitum group, the Mycobacterium mageritense group, the Mycobacterium wolinskyi group and the M. smegmatis group, respectively comprising four, three, eight, one, one and two species. The two protein-encoding genes rpoB and recA improved meaningfully the bootstrap values at the nodes of the different groups. The species M. mucogenicum, M. mageritense and M. wolinskyi formed new groups separated from the M. chelonae-abscessus, M. fortuitum and M. smegmatis groups, respectively. The M. mucogenicum group was well delineated, in contrast to the M. mageritense and M. wolinskyi groups. For phylogenetic organizations derived from the hsp65 and sodA gene sequences, the bootstrap values at the nodes of a few clusters were <70 %. In contrast, phylogenetic organizations obtained from the 16S rRNA, rpoB and recA genes were globally similar to that inferred from combined datasets, indicating that the rpoB and recA genes appeared to be useful tools in addition to the 16S rRNA gene for the investigation of evolutionary relationships among RGM species. Moreover, rpoB gene sequence analysis yielded bootstrap values higher than those observed with recA and 16S rRNA genes. Also, molecular signatures in the rpoB and 16S rRNA genes of the M. mucogenicum group showed that it was a sister group of the M. chelonae-abscessus group. In this group, M. mucogenicum ATCC 49650(T) was clearly distinguished from M. mucogenicum ATCC 49649 with regard to analysis of the five gene sequences. This was in agreement with phenotypic and biochemical characteristics and suggested that these strains are representatives of two closely related, albeit distinct species.

Bacterial Proteins↗

Exploring endothelial cell environments across organs in spatially resolved omics data.

Endothelial cells are ubiquitously present in the human body and line the luminal surface of blood and lymphatic vessels. The oxygen-dependence of cells impacts their proximity to blood vessels, and consequently, to endothelial cells depending on their functional properties and priorities. This paper presents cell-to-nearest-endothelial-cell distance distributions for various cell types using 399 spatially resolved omics datasets from 14 studies comprising 12 tissue types with a total of 47,349,496 cells. Additionally, we developed an open-source web-based interactive tool, Cell Distance Explorer, that allows researchers to interactively visualize cell graphs and linkages in 2D and 3D datasets. Finally, we present a hierarchical neighborhood analysis focused on the endothelial cell neighborhoods in small and large intestine datasets. This paper provides an open-access resource (datasets, tools, and analyses) to characterize and compare cell distances and cell neighborhoods in spatially resolved omics data.

Journal Article↗

A robust meta-classification strategy for cancer diagnosis from gene expression data.

One of the major challenges in cancer diagnosis from microarray data is to develop robust classification models which are independent of the analysis techniques used and can combine data from different laboratories. We propose a meta-classification scheme which uses a robust multivariate gene selection procedure and integrates the results of several machine learning tools trained on raw and pattern data. We validate our method by applying it to distinguish diffuse large B-cell lymphoma (DLBCL) from follicular lymphoma (FL) on two independent datasets: the HuGeneFL Affmetrixy dataset of Shipp et al. (www. genome.wi.mit.du/MPR /lymphoma) and the Hu95Av2 Affymetrix dataset (DallaFavera's laboratory, Columbia University). Our meta-classification technique achieves higher predictive accuracies than each of the individual classifiers trained on the same dataset and is robust against various data perturbations. We also find that combinations of p53 responsive genes (e.g., p53, PLK1 and CDK2) are highly predictive of the phenotype.

Algorithms↗

Automatic classification of heartbeats using ECG morphology and heartbeat interval features.

A method for the automatic processing of the electrocardiogram (ECG) for the classification of heartbeats is presented. The method allocates manually detected heartbeats to one of the five beat classes recommended by ANSI/AAMI EC57:1998 standard, i.e., normal beat, ventricular ectopic beat (VEB), supraventricular ectopic beat (SVEB), fusion of a normal and a VEB, or unknown beat type. Data was obtained from the 44 nonpacemaker recordings of the MIT-BIH arrhythmia database. The data was split into two datasets with each dataset containing approximately 50,000 beats from 22 recordings. The first dataset was used to select a classifier configuration from candidate configurations. Twelve configurations processing feature sets derived from two ECG leads were compared. Feature sets were based on ECG morphology, heartbeat intervals, and RR-intervals. All configurations adopted a statistical classifier model utilizing supervised learning. The second dataset was used to provide an independent performance assessment of the selected configuration. This assessment resulted in a sensitivity of 75.9%, a positive predictivity of 38.5%, and a false positive rate of 4.7% for the SVEB class. For the VEB class, the sensitivity was 77.7%, the positive predictivity was 81.9%, and the false positive rate was 1.2%. These results are an improvement on previously reported results for automated heartbeat classification systems.

Algorithms↗

Composite rectilinear deformation for stretch and squish navigation.

We present the first scalable algorithm that supports the composition of successive rectilinear deformations. Earlier systems that provided stretch and squish navigation could only handle small datasets. More recent work featuring rubber sheet navigation for large datasets has focused on rendering and on application-specific issues. However, no algorithm has yet been presented for carrying out such navigation methods; our paper addresses this problem. For maximum flexibility with large datasets, a stretch and squish navigation algorithm should allow for millions of potentially deformable regions. However, typical usage only changes the extents of a small subset k of these n regions at a time. The challenge is to avoid computations that are linear in n, because a single deformation can affect the absolute screen-space location of every deformable region. We provide an O(klogn) algorithm that supports any application that can lay out a dataset on a generic grid, and show an implementation that allows navigation of trees and gene sequences with millions of items in sub-millisecond time.

Journal Article↗

A likelihood-based method for testing for nonstochastic variation of diversification rates in phylogenies.

Observed variations in rates of taxonomic diversification have been attributed to a range of factors including biological innovations, ecosystem restructuring, and environmental changes. Before inferring causality of any particular factor, however, it is critical to demonstrate that the observed variation in diversity is significantly greater than that expected from natural stochastic processes. Relative tests that assess whether observed asymmetry in species richness between sister taxa in monophyletic pairs is greater than would be expected under a symmetric model have been used widely in studies of rate heterogeneity and are particularly useful for groups in which paleontological data are problematic. Although one such test introduced by Slowinski and Guyer a decade ago has been applied to a wide range of clades and evolutionary questions, the statistical behavior of the test has not been examined extensively, particularly when used with Fisher's procedure for combining probabilities to analyze data from multiple independent taxon pairs. Here, certain pragmatic difficulties with the Slowinski-Guyer test are described, further details of the development of a recently introduced likelihood-based relative rates test are presented, and standard simulation procedures are used to assess the behavior of the two tests in a range of situations to determine: (1) the accuracy of the tests' nominal Type I error rate; (2) the statistical power of the tests; (3) the sensitivity of the tests to inclusion of taxon pairs with few species; (4) the behavior of the tests with datasets comprised of few taxon pairs; and (5) the sensitivity of the tests to certain violations of the null model assumptions. Our results indicate that in most biologically plausible scenarios, the likelihood-based test has superior statistical properties in terms of both Type I error rate and power, and we found no scenario in which the Slowinski-Guyer test was distinctly superior, although the degree of the discrepancy varies among the different scenarios. The Slowinski-Guyer test tends to be much more conservative (i.e., very disinclined to reject the null hypothesis) in datasets with many small pairs. In most situations, the performance of both the likelihood-based test and particularly the Slowinski-Guyer test improve when pairs with few species are excluded from the computation, although this is balanced against a decline in the tests' power and accuracy as fewer pairs are included in the dataset. The performance of both tests is quite poor when they are applied to datasets in which the taxon sizes do not conform to the distribution implied by the usual null model. Thus, results of analyses of taxonomic rate heterogeneity using the Slowinski-Guyer test can be misleading because the test's ability to reject the null hypothesis (equal rates) when true is often inaccurate and its ability to reject the null hypothesis when the alternative (unequal rates) is true is poor, particularly when small taxon pairs are included. Although not always perfect, the likelihood-based test provides a more accurate and powerful alternative as a relative rates test.

Biodiversity↗