Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Virus-mediated fate of antimicrobial resistance genes in livestock manure anaerobic digestion.

Antimicrobial resistance (AMR) poses a critical global health challenge, with livestock manure acting as a significant environmental reservoir for antimicrobial resistance genes (ARGs). Anaerobic digestion (AD) is a pivotal process for mitigating ARG dissemination at the livestock-environment-human interface. This study aims to elucidate the global dynamics of ARGs in AD systems, focusing on virus-host interactions and arms race, to identify actionable strategies for AMR control. We analyzed 205 metagenomic (4.5 Tb) and 36 meta-transcriptomic (640 Gb) datasets, including 15 newly generated datasets, revealing that pig manure AD harbors the highest ARG abundance (0.668 ARGs/16S rRNA), while AD systems generally exhibit limited transcriptional activation of ARGs. We constructed a viral dataset for livestock manure AD (GVD_LMAD), comprising 59,316 DNA and 727 RNA viral operational taxonomic units (vOTUs). Virus-host interactions established by CRISPR-Cas spacer, tRNA and homology matches revealed 889 lytic infections of antimicrobial-resistant bacteria (ARB) compared to only 18 ARG transduction events. Further analysis showed that the relative abundance of vOTUs assigned to the reduction role (4.11% ± 3.19%) was substantially higher than that of reproduction (0.72% ± 0.64%) and transduction (0.19% ± 0.30%), demonstrating that, among viral processes, lysis outweighs transduction in contributing to ARG abundance reduction in AD. Furthermore, an antiviral defense system (ADS) catalogue (GADSC_LMAD), derived from 2760 high-quality metagenome-assembled genomes (MAGs) containing 39,307 ADS, with ADS prevalence in ARB (7.8 ± 6.0 per MAG), indicating an intensified virus-host arms race in AD that may shield ARB from phage lysis. The resulting CRISPR-Cas immune network with expressed spacers targets foreign ARG-carrying sequences (primarily plasmids and ICEs), suggesting a mechanism that restricts horizontal gene transfer (HGT) via conjugation and transformation, despite shielding ARB from phage lysis. Collectively, these findings highlight that viral communities significantly contribute to ARG reduction through phage lysis relative to transduction, while the ADS-mediated arms race, despite protecting ARB, constructs a biological firewall that potentially limits HGT of ARGs. This study provides novel insights into virus-host dynamics as a key mechanism for controlling ARG dissemination in AD systems.

Animals↗

SIMS: A deep-learning label transfer tool for single-cell RNA sequencing analysis.

Cell atlases serve as vital references for automating cell labeling in new samples, yet existing classification algorithms struggle with accuracy. Here we introduce SIMS (scalable, interpretable machine learning for single cell), a low-code data-efficient pipeline for single-cell RNA classification. We benchmark SIMS against datasets from different tissues and species. We demonstrate SIMS's efficacy in classifying cells in the brain, achieving high accuracy even with small training sets (<3,500 cells) and across different samples. SIMS accurately predicts neuronal subtypes in the developing brain, shedding light on genetic changes during neuronal differentiation and postmitotic fate refinement. Finally, we apply SIMS to single-cell RNA datasets of cortical organoids to predict cell identities and uncover genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Single-Cell Analysis↗

Mudskipper detects combinatorial RNA binding protein interactions in multiplexed CLIP data.

The uncovering of protein-RNA interactions enables a deeper understanding of RNA processing. Recent multiplexed crosslinking and immunoprecipitation (CLIP) technologies such as antibody-barcoded eCLIP (ABC) dramatically increase the throughput of mapping RNA binding protein (RBP) binding sites. However, multiplex CLIP datasets are multivariate, and each RBP suffers non-uniform signal-to-noise ratio. To address this, we developed Mudskipper, a versatile computational suite comprising two components: a Dirichlet multinomial mixture model to account for the multivariate nature of ABC datasets and a softmasking approach that identifies and removes non-specific protein-RNA interactions in RBPs with low signal-to-noise ratio. Mudskipper demonstrates superior precision and recall over existing tools on multiplex datasets and supports analysis of repetitive elements and small non-coding RNAs. Our findings unravel splicing outcomes and variant-associated disruptions, enabling higher-throughput investigations into diseases and regulation mediated by RBPs.

RNA-Binding Proteins↗

Quantification of escape from X chromosome inactivation with single-cell omics data reveals heterogeneity across cell types and tissues.

Several X-linked genes escape from X chromosome inactivation (XCI), while differences in escape across cell types and tissues are still poorly characterized. Here, we developed scLinaX for directly quantifying relative gene expression from the inactivated X chromosome with droplet-based single-cell RNA sequencing (scRNA-seq) data. The scLinaX and differentially expressed gene analyses with large-scale blood scRNA-seq datasets consistently identified the stronger escape in lymphocytes than in myeloid cells. An extension of scLinaX to a 10x multiome dataset (scLinaX-multi) suggested a stronger escape in lymphocytes than in myeloid cells at the chromatin-accessibility level. The scLinaX analysis of human multiple-organ scRNA-seq datasets also identified the relatively strong degree of escape from XCI in lymphoid tissues and lymphocytes. Finally, effect size comparisons of genome-wide association studies between sexes suggested the underlying impact of escape on the genotype-phenotype association. Overall, scLinaX and the quantified escape catalog identified the heterogeneity of escape across cell types and tissues.

X Chromosome Inactivation↗

Assessment and integration of publicly available SAGE, cDNA microarray, and oligonucleotide microarray expression data for global coexpression analyses.

Large amounts of gene expression data from several different technologies are becoming available to the scientific community. A common practice is to use these data to calculate global gene coexpression for validation or integration of other "omic" data. To assess the utility of publicly available datasets for this purpose we have analyzed Homo sapiens data from 1202 cDNA microarray experiments, 242 SAGE libraries, and 667 Affymetrix oligonucleotide microarray experiments. The three datasets compared demonstrate significant but low levels of global concordance (rc<0.11). Assessment against Gene Ontology (GO) revealed that all three platforms identify more coexpressed gene pairs with common biological processes than expected by chance. As the Pearson correlation for a gene pair increased it was more likely to be confirmed by GO. The Affymetrix dataset performed best individually with gene pairs of correlation 0.9-1.0 confirmed by GO in 74% of cases. However, in all cases, gene pairs confirmed by multiple platforms were more likely to be confirmed by GO. We show that combining results from different expression platforms increases reliability of coexpression. A comparison with other recently published coexpression studies found similar results in terms of performance against GO but with each method producing distinctly different gene pair lists.

Gene Expression Profiling↗

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits.

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

Animals↗

Exploring differences across pangenome-graph representations using Escherichia coli O157:H7 as a model.

Pangenome graphs are increasingly used to represent population-scale bacterial diversity, yet construction methods span fundamentally different representation paradigms whose outputs and sensitivities to assembly quality remain poorly quantified. We systematically reviewed microbial pangenome graph tools and benchmarked seven representative methods spanning gene-cluster, compacted coloured de Bruijn graph, one hybrid approach and one multiple sequence alignment method. Using a repeat-rich Escherichia coli O157:H7 dataset with complete genomes and matched short-read data, we constructed graphs from identical inputs and observed orders-of-magnitude differences in graph size and fragmentation, indicating that global topology is driven by representation strategy. Varying completeness composition revealed that assembly fragmentation is a first-order determinant of graph structure: gene-cluster graphs contracted as draft assemblies replaced complete genomes, whereas compacted coloured de Bruijn graphs expanded, with distinct degree-prevalence fingerprints across tools. In contrast, the multiple sequence alignment method could not be evaluated across fragmented inputs because it did not run reliably on draft-assembly datasets. Computational cost mirrored these shifts and depended strongly on completeness composition, including a pronounced runtime penalty for one compacted coloured de Bruijn graph implementation on all-draft inputs. Finally, analysis of Shiga toxin loci showed that pangenome-level reconciliation by gene-cluster-based tools does not reliably correct assembly artefacts at challenging multi-copy genes and that performance varies by locus. Together, these findings show that pangenome graphs are representation-dependent models of bacterial diversity, and that, in this repeat-rich O157:H7 benchmark dataset, assembly completeness is a primary determinant of their topology, scalability, and locus-level accuracy.

Escherichia coli O157↗

Convergent functional genomics: a Bayesian candidate gene identification approach for complex disorders.

Identifying genes involved in complex neuropsychiatric disorders through classic human genetic approaches has proven difficult. To overcome that barrier, we have developed a translational approach called Convergent Functional Genomics (CFG), which cross-matches animal model microarray gene expression data with human genetic linkage data as well as human postmortem brain data and biological role data, as a Bayesian way of cross-validating findings and reducing uncertainty. Our approach produces a short list of high probability candidate genes out of the hundreds of genes changed in microarray datasets and the hundreds of genes present in a linkage peak chromosomal area. These genes can then be prioritized, pursued, and validated in an individual fashion using: (1) human candidate gene association studies and (2) cell culture and mouse transgenic models. Further bioinformatics analysis of groups of genes identified through CFG leads to insights into pathways and mechanisms that may be involved in the pathophysiology of the illness studied. This simple but powerful approach is likely generalizable to other complex, non-neuropsychiatric disorders, for which good animal models, as well as good human genetic linkage datasets and human target tissue gene expression datasets exist.

Animals↗

True and false discovery in DNA microarray experiments: transcriptome changes in the hippocampus of presenilin 1 mutant mice.

In transcriptome profiling experiments using DNA microarrays, it is critical to maximize putatively true data discovery while keeping the false discovery rate at acceptable levels. Using previously published and verified transcriptome datasets of mice with genetically altered PS1 physiology, we present a simple, robust, and system-specific assessment of type I and type II errors in two independent microarray experimental series. We provide evidence to suggest that for maximizing true discovery and minimizing false discovery, statistical criteria alone are inferior to statistical significance plus magnitude of change criteria. Furthermore, we found that, regardless of the exact criteria used for determining differential expression, different data extraction protocols give rise to different discovery and false discovery rates. In addition, a large proportion of expression differences were both dataset and analytical approach dependent. The data assessment methods presented and discussed in this manuscript can be easily carried out on any microarray dataset using basic spreadsheet functions as the only tool needed. Finally, we provide an in-depth analysis of the hippocampal transcriptome of DeltaE9 hPS1 transgenic mice and mice with a conditional ablation of the PS1 gene.

Animals↗

Evidence for two independent lineages of Griffithsia (Ceramiaceae, Rhodophyta) based on plastid protein-coding psaA, psbA, and rbcL gene sequences.

The ceramiaceous red algal genus Griffithsia has characteristic large vegetative cells visible to the unaided eye and thousands of nuclei in a single cell at maturity. Its members often occur intertidally along temperate to tropical coasts. Although previous morphological studies indicated that Griffithsia is subdivided into four groups, there is no molecular phylogeny for the genus. We present the multigene phylogeny of the genus based on plastid protein-coding psaA, psbA, and rbcL genes from ten samples of eight Griffithsia species, eight samples of five putative relatives, such as Anotrichium and Halurus, and three outgroup taxa. Saturation plots for each of the three datasets showed no evidence of saturation at any codon position. The partition homogeneity test indicated that none of the individual datasets resulted in significantly incongruent trees. All the analyses of individual and concatenated datasets separated Griffithsia into two well-defined lineages: Lineage 1 was composed of Griffithsia corallinoides, Griffithsia pacifica, and Griffithsia tomo-yamadae, while lineage 2 encompassed Griffithsia antarctica, Griffithsia japonica, Griffithsia teges, Griffithsia traversii, and Griffithsia sp. Our results support the monophyly of the four Anotrichium species and cast a question on the autonomy of Halurus. The monophyly of the tribe Griffithsieae is well resolved, although interrelationships among Griffithsia, Anotrichium, and Halurus were unclear. Our study indicates that the psaA and psbA genes are powerful new tools for the genus-level phylogeny of red algal groups, such as Griffithsia. This is the first report on the multigene phylogeny of the Ceramiales algae based on three protein-coding plastid genes.

Algal Proteins↗

Investigation of molluscan phylogeny using large-subunit and small-subunit nuclear rRNA sequences.

The Mollusca represent one of the most morphologically diverse animal phyla, prompting a variety of hypotheses on relationships between the major lineages within the phylum based upon morphological, developmental, and paleontological data. Analyses of small-ribosomal RNA (SSU rRNA) gene sequence have provided limited resolution of higher-level relationships within the Mollusca. Recent analyses suggest large-subunit (LSU) rRNA gene sequences are useful in resolving deep-level metazoan relationships, particularly when combined with SSU sequence. To this end, LSU (approximately 3.5 kb in length) and SSU (approximately 2 kb) sequences were collected for 33 taxa representing the major lineages within the Mollusca to improve resolution of intraphyletic relationships. Although the LSU and combined LSU+SSU datasets appear to hold potential for resolving branching order within the recognized molluscan classes, low bootstrap support was found for relationships between the major lineages within the Mollusca. LSU+SSU sequences also showed significant levels of rate heterogeneity between molluscan lineages. The Polyplacophora, Gastropoda, and Cephalopoda were each recovered as monophyletic clades with the LSU+SSU dataset. While the Bivalvia were not recovered as monophyletic clade in analyses of the SSU, LSU, or LSU+SSU, the Shimodaira-Hasegawa test showed that likelihood scores for these results did not differ significantly from topologies where the Bivalvia were monophyletic. Analyses of LSU sequences strongly contradict the widely accepted Diasoma hypotheses that bivalves and scaphopods are closely related to one another. The data are consistent with recent morphological and SSU analyses suggesting scaphopods are more closely related to gastropods and cephalopods than to bivalves. The dataset also presents the first published DNA sequences from a neomeniomorph aplacophoran, a group considered critical to our understanding of the origin and early radiation of the Mollusca.

Animals↗

Phylogenetic investigations of Antarctic notothenioid fishes (Perciformes: Notothenioidei) using complete gene sequences of the mitochondrial encoded 16S rRNA.

The Notothenioidei dominates the fish fauna of the Antarctic in both biomass and diversity. This clade exhibits adaptations related to metabolic function and freezing avoidance in the subzero Antarctic waters, and is characterized by a high degree of morphological and ecological diversity. Investigating the macroevolutionary processes that may have contributed to the radiation of notothenioid fishes requires a well-resolved phylogenetic hypothesis. To date published molecular and morphological hypotheses of notothenioids are largely congruent, however, there are some areas of significant disagreement regarding higher-level relationships. Also, there are critical areas of the notothenioid phylogeny that are unresolved in both molecular and morphological phylogenetic analyses. Previous molecular phylogenetic analyses of notothenioids using partial mtDNA 12S and 16S rRNA sequence data have resulted in limited phylogenetic resolution and relatively low node support. One particularly controversial result from these analyses is the paraphyly of the Nototheniidae, the most diverse family in the Notothenioidei. It is unclear if the phylogenetic results from the 12S and 16S partial gene sequence dataset are due to limited character sampling, or if they reflect patterns of evolutionary diversification in notothenioids. We sequenced the complete mtDNA 16S rRNA gene for 43 notothenioid species, the largest sampling to-date from all eight taxonomically recognized families. Phylogenetic analyses using both maximum parsimony and maximum likelihood resulted in well-resolved trees with most nodes supported with high bootstrap pseudoreplicate scores and significant Bayesian posterior probabilities. In all analyses the Nototheniidae was monophyletic. Shimodaira-Hasegawa tests were able to reject two hypotheses that resulted from prior morphological analyses. However, despite substantial resolution and node support in the 16S rRNA trees, several phylogenetic hypotheses among closely related species and clades were not rejected. The inability to reject particular hypotheses among species in apical clades is likely due to the lower rate of nucleotide substitution in mtDNA rRNA genes relative to protein coding regions. Nevertheless, with the most extensive notothenioid taxon sampling to date, and the much greater phylogenetic resolution offered by the complete 16S rRNA sequences over the commonly used partial 12S and 16S gene dataset, it would be advantageous for future molecular investigations of notothenioid phylogenetics to utilize at the minimum the complete gene 16S rRNA dataset.

Animals↗

Phylogeny of the bears (Ursidae) based on nuclear and mitochondrial genes.

The taxomic classification and phylogenetic relationships within the bear family remain argumentative subjects in recent years. Prior investigation has been concentrated on the application of different mitochondrial (mt) sequence data, herein we employ two nuclear single-copy gene segments, the partial exon 1 from gene encoding interphotoreceptor retinoid binding protein (IRBP) and the complete intron 1 from transthyretin (TTR) gene, in conjunction with previously published mt data, to clarify these enigmatic problems. The combined analyses of nuclear IRBP and TTR datasets not only corroborated prior hypotheses, positioning the spectacled bear most basally and grouping the brown and polar bear together but also provided new insights into the bear phylogeny, suggesting the sister-taxa association of sloth bear and sun bear with strong support. Analyses based on combination of nuclear and mt genes differed from nuclear analysis in recognizing the sloth bears as the earliest diverging species among the subfamily ursine representatives while the exact placement of the sun bear did not resolved. Asiatic and American black bears clustered as sister group in all analyses with moderate levels of bootstrap support and high posterior probabilities. Comparisons between the nuclear and mtDNA findings suggested that our combined nuclear dataset have the resolving power comparable to mtDNA dataset for the phylogenetic interpretation of the bear family. As can be seen from present study, the unanimous phylogeny for this recently derived family was still not produced and additional independent genetic markers were in need.

Animals↗

Taxon sampling effects in molecular clock dating: an example from the African Restionaceae.

Three commonly used molecular dating methods for correction of variable rates (non-parametric rate smoothing, penalized likelihood, and Bayesian rate correction) as well as the assumption of a global molecular clock were tested for sensitivity to taxon sampling. The test dataset of 6854 basepairs for 300 terminals includes a nearly complete sample of the Restio-clade of the African Restionaceae (272 of the 288 species), as well as 26 outgroup species. Of this, nested subsets of 35, 51, 80, 120, 150, and the full 300 species were used. Molecular dating experiments with these datasets showed that all methods are sensitive to undersampling, but that this effect is more severe in analyses that use more extreme rate smoothing. Additionally, the undersampling effect is positively related to distance from the calibration node. The combined effect of undersampling and distance from the calibration node resulted in up to threefold differences in the age estimation of nodes from the same dataset with the same calibration point. We suggest that the most suitable methods are penalized likelihood and Bayesian when a global clock assumption has been rejected, as these methods are more successful at finding optimal levels of smoothing to correct for rate heterogeneity, and are less sensitive to undersampling.

Africa↗

Modeling nucleotide evolution at the mesoscale: the phylogeny of the neotropical pitvipers of the Porthidium group (viperidae: crotalinae).

We analyzed the phylogeny of the Neotropical pitvipers within the Porthidium group (including intra-specific through inter-generic relationships) using 1.4 kb of DNA sequences from two mitochondrial protein-coding genes (ND4 and cyt-b). We investigated how Bayesian Markov chain Monte-Carlo (MCMC) phylogenetic hypotheses based on this 'mesoscale' dataset were affected by analysis under various complex models of nucleotide evolution that partition models across the dataset. We develop an approach, employing three statistics (Akaike weights, Bayes factors, and relative Bayes factors), for examining the performance of complex models in order to identify the best-fit model for data analysis. Our results suggest that: (1) model choice may have important practical effects on phylogenetic conclusions even for mesoscale datasets, (2) the use of a complex partitioned model did not produce widespread increases or decreases in nodal posterior probability support, and (3) most differences in resolution resulting from model choice were concentrated at deeper nodes. Our phylogenetic estimates of relationships among members of the Porthidium group (genera: Atropoides, Cerrophidion, and Porthidium) resolve the monophyly of the three genera. Bayesian MCMC results suggest that Cerrophidion and Porthidium form a clade that is the sister taxon to Atropoides. In addition to resolving the intra-specific relationships among a majority of Porthidium group taxa, our results highlight phylogeographic patterns across Middle and South America and suggest that each of the three genera may harbor undescribed species diversity.

Animals↗

Phylogenetic relationships among Syndermata inferred from nuclear and mitochondrial gene sequences.

Phylogenetic relationships among Syndermata have been extensively debated, mainly because the sister-group of the Acanthocephala has not yet been clearly identified from analyses of morphological and molecular data. Here we conduct phylogenetic analyses on samples from the 4 classes of Acanthocephala (Archiacanthocephala, Eoacanthocephala, Polyacanthocephala, and Palaeacanthocephala) and the 3 Rotifera classes (Bdelloidea, Monogononta, and Seisonidea). We do so using small-subunit (SSU) and large-subunit (LSU) ribosomal DNA and cytochrome c oxidase subunit 1 (cox 1) sequences. These nuclear and mitochondrial DNA sequences were obtained for 27 acanthocephalans, 9 rotifers, and representatives of 6 phyla that were used as outgroups. Maximum parsimony (MP), maximum likelihood (ML), and Bayesian analyses were conducted on the nuclear rDNA(SSU+LSU) and the combined sequence dataset(SSU+LSU+cox 1 genes). Phylogenetic analyses of the combined rDNA and cox 1 data uniformly provided strong support for a clade including rotifers plus acanthocephalans (Syndermata). Strong support was also found for monophyly of Acanthocephala in analyses of the combined dataset or rDNA sequences alone. Within the Acanthocephala the monophyletic grouping of the representatives of each class was strongly supported. Our results depicted Archiacanthocephala as the sister-group to the remaining acanthocephalans. Analyses of the combined dataset recovered a sister-group relationship between Acanthocephala and Bdelloidea by parsimony, likelihood, and Bayesian methods. Support for this clade was generally strong. Alternative topologies that depicted a different rotifer sister-group of Acanthocephala (or monophyly of Rotifera) were significantly worse. In this paraphyletic assemblage of rotifers, the relative positions of Seisonidea and Monogononta to the clade Bdelloidea+Acanthocephala were inconsistent among trees based on different inference methods. These results indicate that Bdelloidea is the free-living sister-group to acanthocephalans, which should prove key for comparative investigations of the morphological, molecular, and ecological changes accompanying the evolution of parasitism.

Acanthocephala↗

Phylogenetic utility of rapidly evolving DNA at high taxonomical levels: contrasting matK, trnT-F, and rbcL in basal angiosperms.

The prevailing view in molecular systematics is that relationships among distantly related taxa should be inferred using DNA segments with low rates of evolution. However, recent analyses of sequences from the rapidly evolving matK and trnT-trnF regions yielded well resolved and highly supported trees for early diverging angiosperms. We compare here the phylogenetic structure in matK, trnT-F, and rbcL datasets for the same 42, primarily basal angiosperm taxa. Phylogenetic trees based on matK or trnT-F are far more robust than those based on rbcL. Combined analysis of the rapidly evolving regions provides support for higher-level relationships stronger than that derived from analyses of multi-gene datasets of up to several fold the number of characters analyzed here. In addition to displaying a higher percentage of parsimony-informative characters, the average phylogenetic signal per informative character is significantly higher in the datasets from rapidly evolving DNA than in the more slowly evolving rbcL, as detected using resampling of identical numbers of parsimony-informative characters from the data matrices and subjecting different statistics for overall tree robustness and phylogenetic signal to significance tests. Automated via a set of scripts, the method used here should be easily extendable to comparisons of a broader range of genomic regions for varying taxon samplings. The relative performance of markers correlates not only with a lower mean homoplasy in matK and trnT-trnF compared to rbcL, but in particular correlates negatively with the percentage of sites exhibiting maximum or close to maximum homoplasy. A likelihood ratio test confirms that the rapidly evolving gene matK evolves significantly closer to neutrality, which may be one of the underlying factors for lower levels of overall homoplasy. Our results are in line with evidence from simulation studies suggesting that the deleterious effect of multiple hits in using rapidly evolving DNA at rather deep phylogenetic levels may have been overestimated, and thus promote extending the use of rapidly evolving DNA to deeper phylogenetic levels.

Codon↗

Comparison of paralog identification methods and their impact on species tree topologies in target capture phylogenomics within the Sindora clade (Detarioideae: Leguminosae).

Target capture is a common method of generating high throughput DNA sequencing data for phylogenetic reconstruction of species relationships, for which single copy genes are usually most informative. However, a pervasive problem with target capture is that putatively single copy genes may in fact be paralogs resulting from gene duplication, which are problematic for phylogenetic inference because their evolutionary history may differ from the divergence history of species. Here, we use as a case study a target enrichment dataset of 88 species of Detarioideae (Leguminosae) with a focus on the Sindora clade to examine approaches for handling paralogs, including the built-in paralog handling functions in HybPiper and CAPTUS, plus subsequent steps using Putative Paralog Detection and the tree-based Yang & Smith orthology inference approach. We compare the paralogs flagged using these methods and verify their performance with BLAST mapping against a reference genome sequence of Sindora glabra, and then subsequently compare the species tree topologies produced across these methods. Our comparisons of paralogs flagged across the Sindora clade show that the Putative Paralog Detection pipeline was the most accurate in identifying paralogs in terms of its similarity to the BLAST mapping, followed by the built-in paralog identification function of CAPTUS. However, the results we recovered for the Detarioideae subfamily suggest that the largest differences in species tree topology resulted from the use of paralog-filtered alignments (such as with the Putative Paralog Detection pipeline and the Yang & Smith orthology inference approaches) rather than just by removing the sequences of identified paralogous genes. This was the true for HybPiper-assembled datasets but was not seen in CAPTUS-assembled datasets. In all comparisons, the topological differences caused by different paralog handling methods tended to be confined to clades where processes such as hybridisation and introgression are prevalent. Our study provides a roadmap to establish the best approach to identify, eliminate or separate paralogs in the absence of a chromosomally contiguous reference genome for a study group, and highlights the importance of careful data inspection and processing in addition to understanding the extent of paralogy and paralog characteristics (e.g. sequence divergence between copies) for their study group.

Phylogeny↗