Search PubMedSearch

SEARCH · Search PubMed

Results for “sequencing artifacts”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

32 records · Page 2Linked to original sources

RP-REP Ribosomal Profiling Reports: an open-source cloud-enabled framework for reproducible ribosomal profiling data processing, analysis, and result reporting.

Ribosomal profiling is an emerging experimental technology to measure protein synthesis by sequencing short mRNA fragments undergoing translation in ribosomes. Applied on the genome wide scale, this is a powerful tool to profile global protein synthesis within cell populations of interest. Such information can be utilized for biomarker discovery and detection of treatment-responsive genes. However, analysis of ribosomal profiling data requires careful preprocessing to reduce the impact of artifacts and dedicated statistical methods for visualizing and modeling the high-dimensional discrete read count data. Here we present Ribosomal Profiling Reports (RP-REP), a new open-source cloud-enabled software that allows users to execute start-to-end gene-level ribosomal profiling and RNA-Seq analysis on a pre-configured Amazon Virtual Machine Image (AMI) hosted on AWS or on the user's own Ubuntu Linux server. The software works with FASTQ files stored locally, on AWS S3, or at the Sequence Read Archive (SRA). RP-REP automatically executes a series of customizable steps including filtering of contaminant RNA, enrichment of true ribosomal footprints, reference alignment and gene translation quantification, gene body coverage, CRAM compression, reference alignment QC, data normalization, multivariate data visualization, identification of differentially translated genes, and generation of heatmaps, co-translated gene clusters, enriched pathways, and other custom visualizations. RP-REP provides functionality to contrast RNA-SEQ and ribosomal profiling results, and calculates translational efficiency per gene. The software outputs a PDF report and publication-ready table and figure files. As a use case, we provide RP-REP results for a dengue virus study that tested cytosol and endoplasmic reticulum cellular fractions of human Huh7 cells pre-infection and at 6 h, 12 h, 24 h, and 40 h post-infection. Case study results, Ubuntu installation scripts, and the most recent RP-REP source code are accessible at GitHub. The cloud-ready AMI is available at AWS (AMI ID: RPREP RSEQREP (Ribosome Profiling and RNA-Seq Reports) v2.1 (ami-00b92f52d763145d3)).

AMI

ZILA-SRM: a probabilistic framework with zero-inflated latent models for robust strain reconstruction from metagenomes.

UNLABELLED: Resolving bacterial strain diversity from shotgun metagenomic data is fundamental to understanding intra-host evolution, transmission dynamics, and phenotypic heterogeneity. However, current probabilistic approaches face a severe "identifiability limit" when disentangling highly similar genomes. Under high-noise conditions, sequencing errors, coverage overdispersion, and collinearity confound standard expectation-maximization algorithms, resulting in overfitting and spurious "ghost" strains. Here, we introduce zero-inflated latent allocation for strain reconstruction from metagenomes with adaptive sparsity regularization (ZILA-SRM) to overcome this barrier through three innovations. First, we integrate a zero-inflated Poisson mixture model to decouple "structural zeros" (true strain absence) from "sampling zeros" (stochastic dropout), addressing overdispersion in standard Poisson-based tools. Second, we impose a convex adaptive sparsity regularization penalty that leverages biological sparsity priors to shrink noise artifacts dynamically. Third, we implement a graph-theoretic refinement step using maximal clique enumeration to resolve haplotype collinearity. Benchmarking against StrainFinder and MixtureS on 702 synthetic data sets shows that ZILA-SRM achieves a 20% improvement in precision in high-complexity scenarios while maintaining over 80% recall for minor variants at 0.5% abundance. Re-analysis of deep-sequencing data from 195 Mycobacterium tuberculosis clinical samples reveals cryptic low-abundance drug-resistant variants in 12% of patients, including a minor clone carrying the rpoB S450L mutation. Furthermore, application to skin microbiome data sets further reveals a strong negative correlation between dominant Staphylococcus aureus and Staphylococcus epidermidis strains, providing genomic evidence for competitive exclusion. These findings establish ZILA-SRM as a robust tool for resolving strain-level diversity in complex metagenomes. IMPORTANCE: Understanding microbial communities at the strain level is critical because closely related strains can differ dramatically in traits such as drug resistance, virulence, and ecological interactions. However, resolving individual strains from metagenomic sequencing data remains difficult, especially when strains are highly similar or present at low abundance. As a result, biologically meaningful diversity is often obscured or misinterpreted as noise. In this study, we introduce a new framework that improves the reliability of strain reconstruction from complex metagenomic data. By reducing false-positive strain detection while preserving sensitivity to rare variants, our approach enables more accurate characterization of microbial populations. This improved resolution reveals previously hidden subpopulations in clinical and microbiome datasets, providing clearer insights into microbial evolution, competition, and the emergence of clinically relevant traits such as antibiotic resistance.

Metagenomics

An enhanced multisegment RT-PCR method for influenza A virus sequencing: Improved performance and reduced preparation time over traditional methods.

Influenza A viruses (IAVs) remain a major global health threat, affecting both human and animal populations. Whole-genome sequencing is essential for monitoring viral evolution, zoonotic transmission, and emerging variants. However, conventional RT-PCR methods often result in incomplete gene coverage, amplification biases, and reduced sequencing accuracy, particularly in clinical samples. We developed a robust In-house method for IAV full-genome sequencing using the Oxford Nanopore Technologies (ONT) long-read sequencing platform. This method integrates an in-house multisegment Reverse Transcription PCR (RT-PCR) method with a streamlined 2-pool primer design targeting all eight IAV gene segments. RNA extracted from clinical and stock virus samples was reverse-transcribed and amplified using Superscript IV-based chemistry, followed by magnetic bead purification to ensure high-quality amplicons. Sequencing libraries were prepared with the Native Barcoding Kit 24 (SQK-NBD114.24) and sequenced on R10.4.1 flow cells on the MinION MK1C device. Data analysis using the Iterative Refinement Meta-Assembler (IRMA) confirmed improved read depth, uniform coverage, and complete genome recovery. Compared to conventional methods, our In-House Multisegment 2-Pool (IH-MS2P) RT-PCR method generated higher numbers of matched read counts, minimized chimeric artifacts, and delivered superior genome coverage across human, swine, and avian isolates. This optimized RT-PCR method provides a high-performance, time-efficient, and portable solution for influenza genomics, demonstrating robust applicability even with clinical samples of low RNA yield.

Influenza A virus

Paralogous evolution of the ITS2 region in Xiphophorus.

Ribosomal ITS2 is widely used in phylogenetic studies, yet its multigene organization and potential paralogy can obscure true species relationships. This proof-of-concept study investigates whether ITS2 sequences derived from long-read genomic data in multiple Xiphophorus species primarily reflect orthologous history or are shaped by ancient and local duplications. Phylogenetic analyses reveal two major, reciprocally mirroring ITS2 clades that represent long-standing paralogous rDNA lineages rather than simple allelic variants. The two paralogons show strong asymmetry in copy retention and loss for the majority of the species analyzed in this study. Exceptionally some other species are confined to one paralogon group and exhibit alternating ITS2 variants consistent with persistent ancestral polymorphism. A striking copy number imbalance in X. variatus, combined with its phylogenetic incongruence relative to the established species tree, is best explained by historical rDNA introgression followed by biased concerted evolution that nearly erased one paralogous copy. Despite incomplete homogenization, heterogeneous evolutionary rates, and occasional long-branch artifacts, the recovered paralog-specific topologies largely recapitulate the accepted Xiphophorus species phylogeny, indicating that ITS2 retains a robust organismal signal while also recording episodes of introgression and differential paralog evolution. These results demonstrate that explicit recognition of ITS2 paralogs can both improve phylogenetic interpretation and open avenues for future sequence-structure-based analyses of rDNA evolution and genus-level systematics in Xiphophorus.

Gene duplication

pSTRminer: integrated bioinformatic software for genome-wide identification and population-scale evaluation of polymorphic short tandem repeats.

Animal forensic genetics plays a critical role in criminal investigations by providing crucial evidence through domestic animal individualization and wildlife species identification. While human forensic genetics benefits from standardized short tandem repeats (STR) genotyping systems, animal forensic applications encounter significant challenges, including the limited availability of validated STR markers, the prevalence of error-prone dinucleotide STRs (di-STRs), and insufficient integration of population data. To address these challenges, we developed pSTRminer, an integrated bioinformatic tool that automates genome-wide STR mining and polymorphism evaluation. By applying pSTRminer to domestic cattle (Bos taurus), we identified 775,444 STRs de novo from the reference genome and genotyped them using whole-genome sequencing data from 60 Chinese and 111 African cattle to evaluate polymorphism across diverse genetic backgrounds. This led to the development of the cattle STR database (CSDB), comprising loci with a genotyping success rate&#x2009;&#x2265;&#x2009;40% and polymorphism information content (PIC)&#x2009;&#x2265;&#x2009;0.5. Experimental validation of 30 randomly selected tetranucleotide STRs (tetra-STRs) and 33 di-STRs via next-generation sequencing in a local Chinese cattle population (n&#x2009;=&#x2009;145) confirmed marker reliability. Although tetra-STRs had lower average polymorphism levels, they exhibited significantly lower stutter ratios (p&#x2009;<&#x2009;0.05), providing a viable path for identifying discriminative markers with fewer artifacts. Systematic screening revealed that certain tetra-STRs could surpass di-STRs in polymorphism. In conclusion, pSTRminer provides a scalable framework for developing standardized STR panels, facilitating the identification of robust and informative markers in forensic applications.

Bioinformatic software

Molecular Landscape and Advanced Diagnostic Technologies for BRAF Mutations in Cancer: From Quantitative PCR and ddPCR to CRISPR-Based Platforms.

BRAF mutations are key oncogenic alterations across multiple malignancies, including melanoma, thyroid carcinoma, colorectal cancer, non-small cell lung cancer, glioma, and hairy cell leukemia. The most prevalent variant, BRAF-V600E, induces constitutive activation of the MAPK signaling pathway, promoting tumor progression and influencing therapeutic responsiveness. Accurate detection of BRAF alterations is therefore essential for molecular classification, prognostic assessment, treatment selection, and resistance surveillance. This review summarizes the molecular heterogeneity of BRAF mutations and critically evaluates current diagnostic methodologies. Conventional approaches such as allele-specific PCR and Sanger sequencing are compared with advanced quantitative platforms, including high-resolution melting analysis, droplet digital PCR, and next-generation sequencing, with emphasis on analytical sensitivity, mutation coverage, and clinical applicability. Emerging technologies such as CRISPR-based assays, rolling circle amplification systems, and nanoparticle-based biosensors and point-of-care diagnostic platforms are also discussed for their potential to enhance ultra-sensitive detection, particularly in liquid biopsy settings. These emerging tools are highlighted for their potential to enable ultra-sensitive, rapid, and decentralized mutation detection, particularly in liquid biopsy settings. Key challenges, including intratumoral heterogeneity, low allele-frequency variants, FFPE-associated artifacts, and clonal evolution under therapeutic pressure, are examined within a translational framework. In addition, we examine critical barriers to clinical implementation, including standardization, cost, and global accessibility of molecular diagnostics, and outline potential solutions through scalable technologies and decentralized testing strategies. We propose that optimal BRAF testing requires a mutation subclass-informed and clinically integrated strategy combining comprehensive baseline profiling with longitudinal molecular monitoring. Future diagnostic paradigms will likely integrate multi-omics data and artificial intelligence (AI)-assisted interpretation to refine precision oncology implementation. Looking forward, we propose that optimal BRAF testing will require integration of multi-omics profiling with AI-assisted interpretation, enabling automated variant classification, real-time clinical decision support, and improved prediction of therapeutic response and resistance.

Humans

Computational tool choice impacts CRISPR spacer-protospacer detection.

MOTIVATION: CRISPR spacer-protospacer matching is widely used to infer host-virus interactions in microbial and viromics studies, but the choice of sequence search or alignment tool and its reporting behavior is often under-evaluated for this specific task. RESULTS: Using synthetic, semi-synthetic, and real datasets, we benchmarked commonly used tools and observed substantial differences in recall, runtime, and resource usage across distance metrics and thresholds. Our analyses support practical defaults for large-scale spacer-target matching and clarify trade-offs between exhaustive and heuristic approaches. AVAILABILITY: Source code and benchmark workflows are available at https://github.com/UriNeri/spacer_matching_bench. Data and run artifacts are archived on Zenodo (https://doi.org/10.5281/zenodo.15171878).

Software

Metax enables accurate cross-domain taxonomic profiling of metagenomes.

Taxonomic profiling is fundamental to microbiome research, yet achieving high species-level accuracy remains challenging for complex communities that span bacteria, viruses, eukaryotes, and archaea, and these limitations are exacerbated in low-biomass, host-dominated samples. We introduce Metax, a cross-domain taxonomic profiler that integrates coverage-based probabilistic modeling with an expectation-maximization framework to distinguish true microbial signals from artifacts. Across >600 samples from host-associated, environmental, wastewater, and low-biomass clinical settings, including benchmarks with limited reference representation, Metax improved profiling accuracy, achieving on average 55% higher F1 scores and 45% lower Bray-Curtis dissimilarity than other methods. Moreover, this broad evaluation demonstrated that Metax resolved bacterial and viral signatures of peri-implantitis in oral microbiomes and revealed signals suggestive of reagent-borne contaminants and reference misassemblies in plasma-cell-free DNA. By leveraging genome-wide coverage evidence, Metax enables robust cross-domain profiling across diverse sample types and sequencing depths, including settings where reference databases are highly incomplete.

abundance estimation

DNA-FISH Metaphase Spreads to Distinguish Extrachromosomal DNA from Homogeneously Staining Regions in Human Cancer Cell Lines.

Whole-genome sequencing identifies focal DNA amplifications with base-pair resolution but cannot determine whether amplified sequences reside on extrachromosomal DNA (ecDNA, also known as double minutes) or within chromosomally integrated homogeneously staining regions (HSRs). DNA fluorescence in situ hybridization (DNA-FISH) metaphase spreads remain the gold standard for distinguishing these amplification states at single-cell resolution. Here, we present a detailed protocol for DNA-FISH metaphase spreads using human cancer cell lines, encompassing cell culture, metaphase arrest, hypotonic treatment, fixation, chromosome spreading, fluorescent probe hybridization, and fluorescence imaging. The protocol incorporates intermediate quality-control steps to verify successful chromosome dispersion and optimize metaphase spread quality, making the workflow accessible to laboratories without specialized cytogenetics expertise. Results demonstrate clear visualization of ecDNA and HSR amplification states using locus-specific probes and illustrate common technical artifacts that can affect interpretation. This protocol provides a robust and reproducible approach for studying the structural organization of oncogene amplification in cancer cells.

Humans

DNA-FISH Metaphase Spreads to Distinguish Extrachromosomal DNA from Homogeneously Staining Regions in Human Cancer Cell Lines.

UNLABELLED: Whole-genome sequencing identifies focal DNA amplifications with base-pair resolution but cannot determine whether amplified sequences reside on extrachromosomal DNA (ecDNA, also known as double minutes) or within chromosomally integrated homogeneously staining regions (HSRs). DNA fluorescence in situ hybridization (DNA-FISH) metaphase spreads remain the gold standard for distinguishing these amplification states at single-cell resolution. Here, we present a detailed protocol for DNA-FISH metaphase spreads using human cancer cell lines, encompassing cell culture, metaphase arrest, hypotonic treatment, fixation, chromosome spreading, fluorescent probe hybridization, and fluorescence imaging. The protocol incorporates intermediate quality-control steps to verify successful chromosome dispersion and optimize metaphase spread quality, making the workflow accessible to laboratories without specialized cytogenetics expertise. Results demonstrate clear visualization of ecDNA and HSR amplification states using locus-specific probes and illustrate common technical artifacts that can affect interpretation. This protocol provides a robust and reproducible approach for studying the structural organization of oncogene amplification in cancer cells. SUMMARY: We report a DNA-FISH metaphase spread protocol that visually detects locus copy number and location within the genome. This approach enables single-cell resolution of amplification states, specifically in cancer cell lines containing extrachromosomal DNA and homogeneously staining regions.

Journal Article

Genomic Footprints of Historical Introgression Between Ancient Lineages of Wild Oryza AA-Genome Species With Widely Separated Contemporary Distributions.

Phylogenetic incongruence is increasingly recognized as pervasive, yet the extent to which reticulate evolution occurs between groups separated by substantial geographical distances and deep phylogenetic divergence remains poorly characterized. In the Oryza AA-genome group-a model for plant speciation and domestication-the traditional bifurcation model posits that Australian Oryza meridionalis and African Oryza longistaminata occupy basal branches, distinct from the more recently diversified monophyletic clade comprising Asian and other African lineages, including major cultivars. However, recent evidence from endogenous viral sequences has hinted at unexpected genetic relatedness between African O. longistaminata and Asian Oryza sativa, which are geographically and phylogenetically distant. Here, we conducted a genome-wide survey across 11 Oryza species to systematically identify genomic regions exhibiting phylogenetic incongruence. Widespread phylogenetic discordance was observed, notably involving genomic segments in which O. longistaminata showed phylogenetic proximity to Asian species, contradicting their established deep divergence. To distinguish between introgression and incomplete lineage sorting, we performed four-taxon ABBA-BABA tests, which provided statistical support for introgression. Furthermore, divergence time estimates for these incongruent regions were younger than the species divergence times, suggesting historical introgression between the ancestors of lineages that are currently separated by vast geographical distances. Systematic assessments indicated that potential analytical artifacts, such as compositional bias and substitution saturation, were unlikely to explain the observations. These convergent lines of evidence suggest that ancient introgression had occurred between currently geographically separated and evolutionarily divergent Oryza lineages, leaving detectable footprints across their modern genomes.

Oryza

Innovative CRISPR/Cas9-Based Strategy for Allele-Specific HLA Peptidome Analysis Using a Pan-HLA Antibody.

Human leukocyte antigen (HLA) immunopeptidomics is restricted by the limited availability of allele-specific antibodies and by potential artifacts introduced by HLA overexpression systems. To address these challenges, we developed a CRISPR/Cas9-based strategy that selectively deletes undesired classical class I alleles while preserving a single endogenous allele, thereby enabling allele-resolved peptidome profiling with a pan-HLA class I antibody. As a proof of concept, we edited JY cells to eliminate HLA-B&#x2217;07:02 and HLA-C&#x2217;07:02 while retaining HLA-A&#x2217;02:01 (&#x394;BC clones). Peptide-HLA complexes were immunoprecipitated from WT and &#x394;BC clones using either the pan-HLA class I antibody W6/32 or the A&#x2217;02:01-specific antibody PA2.1, followed by nanoLC-MS/MS and computational HLA assignment. Deletion of HLA-B and HLA-C alleles caused an expected &#x223c;55% reduction in total class I surface expression. Despite this, W6/32 immunoprecipitation from &#x394;BC clones recovered a comparable peptide yield to PA2.1 in WT cells. Binding predictions showed that most peptides identified in &#x394;BC clones using W6/32 were assigned to HLA-A&#x2217;02:01, with near-complete loss of HLA-B&#x2217;07:02- and HLA-C&#x2217;07:02-derived peptides. Sequence logo analysis confirmed the canonical A&#x2217;02:01 motif across conditions. The &#x394;BC W6/32 immunopeptidome exhibited a high degree of overlap (&#x223c;88%) with the WT PA2.1 repertoire, supporting the specificity and fidelity of the approach. These findings establish CRISPR-based editing of HLA alleles as a viable strategy for allele-specific immunopeptidome analysis using pan-HLA antibodies, supporting its potential application beyond this proof-of-concept system, reducing reliance on allele-specific reagents and facilitating the study of underrepresented HLA alleles.

Humans

Genomic regionality in rates of evolution is not explained by clustering of genes of comparable expression profile.

In mammalian genomes, linked genes show similar rates of evolution, both at fourfold degenerate synonymous sites (K4) and at nonsynonymous sites (KA). Although it has been suggested that the local similarity in the synonymous substitution rate is an artifact caused by the inclusion of disparately evolving gene pairs, we demonstrate here that this is not the case: after removal of disparately evolving genes, both (1) linked genes and (2) introns from the same gene have more similar silent substitution rates than expected by chance. What causes the local similarity in both synonymous and nonsynonymous substitution rates? One class of hypotheses argues that both may be related to the observed clustering of genes of comparable expression profile. We investigate these hypotheses using substitution rates from both human-mouse and mouse-rat comparisons, and employing three different methods to assay expression parameters. Although we confirm a negative correlation of expression breadth with both K4 and KA, we find no evidence that clustering of similarly expressed genes explains the clustering of genes of comparable substitution rates. If gene expression is not responsible, what about other causes? At least in the human-mouse comparison, the local similarity in KA can be explained by the covariation of KA and K4. As regards K4, our results appear consistent with the notion that local similarity is due to processes associated with meiotic recombination.

Animals

Functional unknomics of the SAR11 clade reveal hidden genetic potential underlying adaptation to bottom-up and top-down pressures.

UNLABELLED: A substantial fraction of the genes in bacteria lack detectable sequence similarity to genes with known functions. These functionally uncharacterized genes-collectively referred to as the "unknome"-represent a largely unexplored genetic repertoire harboring insights into marine bacterial ecology. In this study, we explored the function of the unknome of the SAR11 clade, the most abundant bacterial lineage in the ocean, with a particular focus on genes that provide insight into its ecology. Based on the Clusters of Orthologous Genes and Kyoto Encyclopedia of Genes and Genomes classifications, approximately 56% of SAR11 ortholog groups were classified as members of the unknome. Among the SAR11 unknome, we successfully inferred the functions of 57 ortholog groups that are conserved in the SAR11 clade by protein structure similarity searches and genomic context analyses. These ortholog groups include putative transporter components, supporting the current ecological understanding that the SAR11 clade is specialized in substrate uptake to adapt to oligotrophic marine environments. Furthermore, structural analysis indicated that the DUF2237-containing protein, enriched in marine environments, may interact with purine nucleotide-containing compounds. This may suggest the existence of unique nucleotide utilization mechanisms in marine bacteria. In addition, we identified candidate viral defense systems within the unknome, indicating that diverse defense systems are present in at least one-third of cultured SAR11 strains. The presence of these defense systems, even within streamlined SAR11 genomes, suggests that they confer significant ecological advantages. Our analyses provide insights into the genetic basis of bottom-up processes (adaptation to oligotrophic environments) and top-down processes (antiviral defense strategy) contributing to ecological success. IMPORTANCE: Many microbial genes have no experimentally established function, limiting our ability to explain how microorganisms adapt to their environments. We examined this uncharacterized gene space, or "unknome" in SAR11, the most abundant bacterial clade in the ocean, by integrating evolutionary conservation, genomic context, predicted protein structure, and environmental distribution. This approach enabled us to prioritize components of the SAR11 unknome, including a core unknome conserved across the clade and genes enriched in specific lineages, and to identify several candidates with possible ecological roles in nutrient acquisition and defense against viruses. Our results suggest that the SAR11 unknome contains important clues to the ecological success of SAR11 rather than merely reflecting incomplete annotation or gene-prediction artifacts. Our study highlights the potential value of unknome analysis for identifying ecologically relevant genes in environmental microorganisms.

Pelagibacterales