Search PubMedSearch

SEARCH · Search PubMed

Results for “Artifacts”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Multi-omic characterization of the Hispanic/Latino blood lipidome reveals an additional locus and attenuated genetic prediction.

While lipids have been extensively investigated, genetic regulation of the circulating lipidome in diverse populations remains poorly understood. We conducted a lipidome-wide genome-wide association study (GWAS) of 830 lipid species in 2,287 Hispanic/Latino participants and performed predictive modeling across omics layers. We identified 7,593 genome-wide significant SNPs mapping to 208 genes. Conditional analysis disentangled the long-range linkage disequilibrium artifacts from the pleiotropic FADS1/2/3 cluster. Separately, we discovered an association at the GPLD1 locus for a circulating ceramide. Colocalization revealed shared genetic architecture with conventional lipids alongside distinct, species-specific pathways. Incorporating Native/Indigenous American expression quantitative trait loci (eQTLs) within a multi-omic framework uncovered 62 likely regulatory genes missed by European-centric gene expression models. Finally, genetically regulated predictive models demonstrated performance declining from transcriptomics to proteomics to lipidomics, reflecting increased distance from gene action along the molecular cascade. Our study provides a genetic landscape of lipid metabolism in a highly burdened population and highlights the challenges in predicting lipid abundance.

Hispanic/Latino population

Paralogous evolution of the ITS2 region in Xiphophorus.

Ribosomal ITS2 is widely used in phylogenetic studies, yet its multigene organization and potential paralogy can obscure true species relationships. This proof-of-concept study investigates whether ITS2 sequences derived from long-read genomic data in multiple Xiphophorus species primarily reflect orthologous history or are shaped by ancient and local duplications. Phylogenetic analyses reveal two major, reciprocally mirroring ITS2 clades that represent long-standing paralogous rDNA lineages rather than simple allelic variants. The two paralogons show strong asymmetry in copy retention and loss for the majority of the species analyzed in this study. Exceptionally some other species are confined to one paralogon group and exhibit alternating ITS2 variants consistent with persistent ancestral polymorphism. A striking copy number imbalance in X. variatus, combined with its phylogenetic incongruence relative to the established species tree, is best explained by historical rDNA introgression followed by biased concerted evolution that nearly erased one paralogous copy. Despite incomplete homogenization, heterogeneous evolutionary rates, and occasional long-branch artifacts, the recovered paralog-specific topologies largely recapitulate the accepted Xiphophorus species phylogeny, indicating that ITS2 retains a robust organismal signal while also recording episodes of introgression and differential paralog evolution. These results demonstrate that explicit recognition of ITS2 paralogs can both improve phylogenetic interpretation and open avenues for future sequence-structure-based analyses of rDNA evolution and genus-level systematics in Xiphophorus.

Gene duplication

Protein Language Model Decoys for Target Decoy Competition in Proteomics: Quality Assessment and Benchmarks.

Large-scale proteomics relies heavily on target-decoy competition for false discovery rate estimation in peptide identification, and the performance of this strategy depends strongly on the design of the decoy database. Classical generators such as reversal and shuffling remain widely used. Here, we introduce the first protein language model-based (PLM) decoy generation for peptide identification and benchmark it against classical strategies. We evaluate these approaches using three complementary quality-control layers: sequence-based separability, search-engine-agnostic spectral-space diagnostics, and end-to-end mass spectrometry benchmarks, including pipelines with rescoring. Across these analyses, PLM-based decoys are harder for sequence-only neural networks to distinguish than most classical generators, suggesting fewer obvious sequence-level artifacts. However, this signal is only weakly informative for search performance. Spectral diagnostics further show that short peptides occupy a particularly crowded target-decoy space and are therefore especially prone to local collisions across all generators. In full search pipelines, reverse decoys remain a strong baseline, and current PLM-based generators do not yet provide a clear overall advantage. We therefore view PLM-based decoys not as universal replacements for reverse decoys but as tunable tools for benchmarking, diagnostics, stress testing, and future adaptive decoy optimization, with increasing value as search models become more expressive.

Proteomics

Improved spike-in normalization clarifies the relationship between active histone modifications and transcription.

Spike-in normalization enables quantitative analysis of chromatin immunoprecipitation sequencing (ChIP-seq) signal. Here we introduce a robust dual spike-in normalization approach for ChIP-seq (ChIP-wrangler), optimize parameters and verify its accuracy in quantifying changes in ChIP-seq signal and detecting technical artifacts. We use ChIP-wrangler to revisit recent claims that active histone marks depend on transcription. We show that acute depletion of RNA polymerase II (RNAPII) has a modest impact on H3K27ac levels, with only 6% of peaks significantly changing after RNAPII depletion, indicating that histone acetylation maintenance is not entirely dependent on ongoing transcription. Promoters and enhancers are differentially affected, with 82% of decreasing acetylation peaks located at promoter-distal elements with enhancer-related motifs. ChIP-wrangler provides increased rigor and 'guardrails' for successful spike-in normalization and, as applied here, refines the understanding of crosstalk between RNAPII activity and transcription-associated histone marks.

Histones

Genetic architectures of brain-related traits are shaped by strong selective constraints.

Genome-wide association studies (GWAS) have identified hundreds of significant loci for psychiatric disorders, yet the strength of these associations remains modest compared to other human complex traits with similar numbers of hits. Whether this pattern reflects statistical artifacts or real biological differences-and, if the latter, what underlies it-remains unclear. In addition to psychiatric disorders, we find that other traits with functional enrichment in the central nervous system (CNS), whether binary or quantitative, also share similar genetic architectures, characterized by GWAS hits of limited statistical significance and generally higher allele frequencies. In comparing the architecture of binary and quantitative traits, we adjust for statistical power in their respective studies. After this adjustment, we fit an evolutionary model of architecture and show that CNS-enriched traits have large mutational target sizes, with contributing variants and genes experiencing stronger selection than those for other traits. Our findings reveal heterogeneity among complex traits and provide insights into traits that more effectively capture fitness-relevant processes. More broadly, our results suggest that the genetic architectures of complex traits are shaped by the tissues through which these traits are mediated.

Humans

OmicsQ: a user-friendly platform for interactive quantitative omics data analysis.

MOTIVATION: High-throughput omics technologies generate complex datasets with thousands of features that are quantified across multiple experimental conditions, but often suffer from incomplete measurements, missing values, and individually fluctuating variances. This requires analytical tools for accurate, deep and insightful biological interpretation, capable of dealing with a large variety of data properties and different amounts of completeness. Software capable of handling such data complexity and integrating with external applications for downstream analysis remains rare and mostly relies on programming-based environments, limiting accessibility for researchers without computational expertise. RESULTS: We present OmicsQ, an interactive, web-based platform designed to streamline quantitative omics data analysis. OmicsQ provides an intuitive, browser-based visualization interface that integrates established statistical processing tools. Those include robust batch correction, automated experimental design annotation, and handling of missing data without imputation, which maintains data integrity and avoids artifacts from a priori assumptions. OmicsQ seamlessly interacts with external applications (e.g. PolySTest, VSClust, ComplexBrowser) for statistical testing, clustering, analysis of protein complex behavior, and pathway enrichment, offering a comprehensive and flexible workflow from data import to biological interpretation that is broadly applicable across domains. AVAILABILITY AND IMPLEMENTATION: OmicsQ is implemented in R and Shiny and is available at https://computproteomics.bmb.sdu.dk/app_direct/OmicsQ. Source code and installation instructions: https://github.com/computproteomics/OmicsQ, DOI: 10.5281/zenodo.17778420.

Software

MethylModes: computationally efficient detection of multimodal distributions in DNA methylation data.

SUMMARY: MethylModes is an R package and Shiny application to identify multimodal distributions in human DNA methylation at individual CpG sites. Multimodal distributions, which can be the result of nearby genetic variation, environmental exposures, or assay artifacts, are susceptible to confounding and important to identify for methylation analysis. MethylModes is easily incorporated into existing quality control pipelines of array-based DNA methylation data. The underlying algorithm uses kernel smoothing of probe-level data to locate the number and location of peaks. The algorithm can be parallelized across probes for efficient implementation at genome-scale. We provide a case study implementation of MethylModes in the Health and Retirement Study as well as the Airwave Health Monitoring Study. AVAILABILITY AND IMPLEMENTATION: MethylModes is available on GitHub at https://github.com/lutiffan/methylModes as an R package wrapping an R Shiny application. We include a toy dataset to validate installation. The codebase is also published on Zenodo at https://doi.org/10.5281/zenodo.17448517.

DNA Methylation

QCatch: a framework for quality control assessment and analysis of single-cell sequencing data.

MOTIVATION: Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstream results. While tools like alevin-fry and simpleaf (and augmented execution context for the alevin-fry), offer flexibility and computational efficiency to process single-cell data, this ecosystem will further benefit from a standardized QC reporting tailored for its outputs. RESULTS: We introduce QCatch, a Python-based command-line tool that generates comprehensive and interactive HTML QC reports designed specifically for single-cell quantification results. Taking the output directory of alevin-fry or simpleaf as the input, QCatch is able to perform essential processing steps, like cell calling, and generate detailed QC reports that contain informative visualizations and statistics, including unique molecular identifier (UMI) count distributions, sequencing saturation estimates, and splicing status information, for QC assurance. Built for seamless integration into downstream analysis workflows, QCatch exports the processed results in a richly-annotated H5AD format file, a widely used data format common among many downstream single-cell data analysis tools. AVAILABILITY AND IMPLEMENTATION: The source code and documentation of QCatch are available on GitHub at https://github.com/COMBINE-lab/QCatch. QCatch can be installed via both Bioconda and PyPI.

Single-Cell Analysis

Duplex-Indel: a Snakemake pipeline for somatic Indel calling in Tn5 transposase-based duplex sequencing data.

SUMMARY: Duplex-Indel is a novel Snakemake workflow for detecting somatic small insertions and deletions (Indels) from Tn5 transposase-based duplex sequencing data. Duplex-Indel enhances the accuracy of mutation calling at the single-molecule level by requiring consensus support from both DNA strands for each somatic Indel, minimizing confounding from technical artifacts. Duplex-Indel extends somatic mutation calling in Tn5 transposase-based duplex sequencing data to include Indels. We have demonstrated the accuracy and robustness of Duplex-Indel using cancer cell lines. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are available under the MIT license on GitHub at https://github.com/ealee-lab/duplex-indel and archived on Zenodo at https://doi.org/10.5281/zenodo.19228799.

Transposases

Agentomics: an agentic system that autonomously develops novel state-of-the-art solutions for biomedical machine learning tasks.

MOTIVATION: Extracting knowledge from biomedical data is crucial for advancing our understanding of biological systems and developing novel therapeutics. The quantity, quality, and resolution of biomedical data constantly evolves, requiring the automation of biomedical machine learning (ML). Existing Automated ML tools lack flexibility, while large language models (LLMs) struggle to consistently deliver reproducible machine learning codebases, and existing LLM Agent-powered solutions lag behind human-engineered ML models. RESULTS: Here, we introduce Agentomics, an autonomous LLM-powered agentic system for end-to-end ML experimentation. Given a biomedical dataset, Agentomics implements various ML modeling strategies, and produces a ready-to-use ML model. Agentomics introduces strict validation checkpoints for standard ML development steps, allowing gradual development on top of working code with defined interfaces and validated artifacts. Further, it offers native support for biomedical foundation models that can be leveraged during experimentation. The generic nature of Agentomics allows the user to create ML solutions for a large variety of datasets and use various LLMs. We evaluate Agentomics across 20 datasets from the domains of Protein Engineering, Drug Discovery, and Regulatory Genomics. When benchmarked against other agentic systems, Agentomics outperformed them in all tested domains. When benchmarked against human expert solutions, Agentomics generated novel state-of-the-art models for 11/20 established benchmark datasets. AVAILABILITY AND IMPLEMENTATION: Agentomics is implemented in Python. Source code and documentation are freely available at: https://github.com/BioGeMT/Agentomics-ML.

Machine Learning

Computational tool choice impacts CRISPR spacer-protospacer detection.

MOTIVATION: CRISPR spacer-protospacer matching is widely used to infer host-virus interactions in microbial and viromics studies, but the choice of sequence search or alignment tool and its reporting behavior is often under-evaluated for this specific task. RESULTS: Using synthetic, semi-synthetic, and real datasets, we benchmarked commonly used tools and observed substantial differences in recall, runtime, and resource usage across distance metrics and thresholds. Our analyses support practical defaults for large-scale spacer-target matching and clarify trade-offs between exhaustive and heuristic approaches. AVAILABILITY: Source code and benchmark workflows are available at https://github.com/UriNeri/spacer_matching_bench. Data and run artifacts are archived on Zenodo (https://doi.org/10.5281/zenodo.15171878).

Software

Genomic Footprints of Historical Introgression Between Ancient Lineages of Wild Oryza AA-Genome Species With Widely Separated Contemporary Distributions.

Phylogenetic incongruence is increasingly recognized as pervasive, yet the extent to which reticulate evolution occurs between groups separated by substantial geographical distances and deep phylogenetic divergence remains poorly characterized. In the Oryza AA-genome group-a model for plant speciation and domestication-the traditional bifurcation model posits that Australian Oryza meridionalis and African Oryza longistaminata occupy basal branches, distinct from the more recently diversified monophyletic clade comprising Asian and other African lineages, including major cultivars. However, recent evidence from endogenous viral sequences has hinted at unexpected genetic relatedness between African O. longistaminata and Asian Oryza sativa, which are geographically and phylogenetically distant. Here, we conducted a genome-wide survey across 11 Oryza species to systematically identify genomic regions exhibiting phylogenetic incongruence. Widespread phylogenetic discordance was observed, notably involving genomic segments in which O. longistaminata showed phylogenetic proximity to Asian species, contradicting their established deep divergence. To distinguish between introgression and incomplete lineage sorting, we performed four-taxon ABBA-BABA tests, which provided statistical support for introgression. Furthermore, divergence time estimates for these incongruent regions were younger than the species divergence times, suggesting historical introgression between the ancestors of lineages that are currently separated by vast geographical distances. Systematic assessments indicated that potential analytical artifacts, such as compositional bias and substitution saturation, were unlikely to explain the observations. These convergent lines of evidence suggest that ancient introgression had occurred between currently geographically separated and evolutionarily divergent Oryza lineages, leaving detectable footprints across their modern genomes.

Oryza

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95 % or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Phylogenomic Analyses Reveal that Panguiarchaeum Is a Clade of Genome-Reduced Asgard Archaea Within the Njordarchaeia.

The Asgard archaea are a diverse archaeal phylum important for our understanding of cellular evolution because they include the lineage that gave rise to eukaryotes. Recent phylogenomic work has focused on characterizing the diversity of Asgard archaea in an effort to identify the closest extant relatives of eukaryotes. However, resolving archaeal phylogeny is challenging, and the positions of 2 recently described lineages-Njordarchaeales and Panguiarchaeales-are uncertain, in ways that directly bear on hypotheses of early evolution. In initial phylogenetic analyses, these lineages branched either with Asgards or with the distantly related Korarchaeota, and it has been suggested that their genomes may be affected by metagenomic contamination. Resolving this debate is important because these clades include genome-reduced lineages that may help inform our understanding of the evolution of symbiosis within Asgard archaea. Here, we performed phylogenetic analyses revealing that the Njordarchaeales and Panguiarchaeales constitute the new class Njordarchaeia within Asgard archaea. We found no evidence of metagenomic contamination affecting phylogenetic analyses. Njordarchaeia exhibit hallmarks of adaptations to (hyper-)thermophilic lifestyles, including biased sequence compositions that can induce phylogenetic artifacts unless adequately modeled. Panguiarchaeum is metabolically distinct from its relatives, with reduced metabolic potential and various auxotrophies. Phylogenetic reconciliation recovers a complex common ancestor of Asgard archaea that encoded the Wood-Ljungdahl pathway. The subsequent loss of this pathway during the reductive evolution of Panguiarchaeum may have been associated with the switch to a symbiotic lifestyle, potentially based on H2-syntrophy. Thus, Panguiarchaeum may contain the first obligate symbionts within Asgard archaea besides the lineage leading to eukaryotes.

Phylogeny

Validation and Optimization of Breeding Strategy for miR-141/200c Knockout Mice to Eliminate Off-Target Gene Silencing using FLPo Deleter.

MicroRNAs (miRNAs) of the miR-200 family specifically miR-141 and miR-200c regulate neurogenesis, differentiation, and epithelial-mesenchymal transitions in development and several diseases including cancer and stroke. The STOCK Mirc13tm1Mtm /Mmjax mouse line, which targets the miR-141/200c cluster, was originally generated and described by Park et al. 2012 as a conditional "knockout-first" allele requiring a two-step breeding strategy: FLP recombination to excise lacZ/neo cassettes followed by Cre recombination to delete the floxed miRNA cluster (1). However, subsequent studies either bypassed this step and reported knockouts based on direct crosses with Cre mouse lines, leaving residual lacZ/neo sequences that may silence upstream elements or introduce transcriptional artifacts or rare studies used less efficient FLPe Deleter mice. Here, we present a detailed and refined strategy to conditional miR-141/200c knockouts mice using FLPo Deleter mice to efficiently eliminate lacZ/neo cassettes. Our approach not only confirmed complete deletion of miR-141 and miR-200c in various organs such olfactory bulbs and lungs where these miRNAs are robustly expressed using various approach such as genotyping qPCR validation and in situ hybridization but showed that without the use of FLPo deleter mice deletion of miR-141/200c cluster amy also lead to loss of several close proximity physiologically important genes such as ptpn6, phb2, atn1 and eno1. By restoring a clean floxed allele using FLPo deleter mice prior to Cre deletion, we establish a reliable and interpretable mouse model for dissecting the roles of the miR-141/200c cluster miRNA in various disease models.

Journal Article

Comprehensive benchmarking of somatic structural variant detection at ultra-low allele fractions.

Postzygotic mosaicism gives rise to somatic structural variants (SVs) at ultra-low variant allele fractions (VAFs), which pose challenges for detection due to the high-coverage sequencing required and noise introduced by sequencing artifacts. Although somatic SV detection has been extensively studied in cancer, these studies are not directly applicable to the study of tissue mosaicism, as they rely on matched normals, target higher VAF ranges, and are enriched for different types of SVs. We present comprehensive benchmark data and best practices for non-cancer somatic SV detection. We created a synthetic mosaic sample by combining six HapMap individuals at varying proportions, generating allele fractions as low as 0.25%. This sample was sequenced to ~2,300x total coverage using Illumina, PacBio, and Nanopore technologies across multiple sequencing centers. A high-confidence benchmark SV set containing over 21,000 pseudo-somatic insertions and deletions ≥50bp was derived from haplotype-resolved assemblies. We evaluated 12 SV discovery pipelines and identified caller-specific strengths and sequencing platform-specific shortcomings. We find that short read-based approaches show reduced recall for insertions and repeat-associated SVs, whereas long-read sequencing achieves high accuracy throughout the genome, increasing linearly with coverage. The best algorithm's sensitivity exceeded 80% for VAFs ≥4% and 15% for VAFs of 0.5-1% with 60x coverage. The publicly available benchmarking data and comparative analysis of current methods provide a foundation for robust discovery of SV mosaicism in non-cancer tissues..

Journal Article

Genomic regionality in rates of evolution is not explained by clustering of genes of comparable expression profile.

In mammalian genomes, linked genes show similar rates of evolution, both at fourfold degenerate synonymous sites (K4) and at nonsynonymous sites (KA). Although it has been suggested that the local similarity in the synonymous substitution rate is an artifact caused by the inclusion of disparately evolving gene pairs, we demonstrate here that this is not the case: after removal of disparately evolving genes, both (1) linked genes and (2) introns from the same gene have more similar silent substitution rates than expected by chance. What causes the local similarity in both synonymous and nonsynonymous substitution rates? One class of hypotheses argues that both may be related to the observed clustering of genes of comparable expression profile. We investigate these hypotheses using substitution rates from both human-mouse and mouse-rat comparisons, and employing three different methods to assay expression parameters. Although we confirm a negative correlation of expression breadth with both K4 and KA, we find no evidence that clustering of similarly expressed genes explains the clustering of genes of comparable substitution rates. If gene expression is not responsible, what about other causes? At least in the human-mouse comparison, the local similarity in KA can be explained by the covariation of KA and K4. As regards K4, our results appear consistent with the notion that local similarity is due to processes associated with meiotic recombination.

Animals

Estimation of chloroplast macromolecular complex copy numbers and subunit stoichiometries during the Chlamydomonas reinhardtii cell cycle.

An unbiased, quantitative view of biomolecules in a living cell is a prerequisite for accurate modeling approaches and informs our understanding of cellular metabolism at scale. In this work, we used the total protein approach (TPA), in which the total protein mass of a given proteomics sample is used as a calibrator for absolute protein quantification, to determine protein abundances during the Chlamydomonas reinhardtii diurnal cycle. We use external, independently measured quantitative markers (metals, pigments) to assess the absolute protein abundances in unlabeled whole cell extracts. We calculate protein abundances in fg cell-1 of 7322 Chlamydomonas proteins, 2266 of which were captured in every time point, including the major proteins involved in the light reactions, photoprotection, proteostasis, and fatty acid metabolism during a cell cycle. As expected, Rubisco large and small subunits are present in a 1:1 stoichiometry, with the large subunit being the most abundant protein in our data set, averaging 5.05 × 106 molecules per cell, reflecting 2.7% of the total protein mass. We noticed that PSII is the most abundant complex involved in the light reactions with 2.08 × 106 complexes per cell. PSI averages 1.75 × 106 complexes per cell and cytochrome b6f averages 0.77 × 106 complexes per cell. The TPA is a robust tool to study proteome dynamics quantitatively, while avoiding artifacts due to biochemical fractionation. Our proteome data set with an unprecedented temporal resolution is a valuable resource to assess protein abundances during the cell cycle in the reference alga Chlamydomonas.

Chlamydomonas reinhardtii