Search PubMedSearch

SEARCH · Search PubMed

Results for “Somatic variant calling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

ONCOLINER: A new solution for monitoring, improving, and harmonizing somatic variant calling across genomic oncology centers.

The characterization of somatic genomic variation associated with the biology of tumors is fundamental for cancer research and personalized medicine, as it guides the reliability and impact of cancer studies and genomic-based decisions in clinical oncology. However, the quality and scope of tumor genome analysis across cancer research centers and hospitals are currently highly heterogeneous, limiting the consistency of tumor diagnoses across hospitals and the possibilities of data sharing and data integration across studies. With the aim of providing users with actionable and personalized recommendations for the overall enhancement and harmonization of somatic variant identification across research and clinical environments, we have developed ONCOLINER. Using specifically designed mosaic and tumorized genomes for the analysis of recall and precision across somatic SNVs, insertions or deletions (indels), and structural variants (SVs), we demonstrate that ONCOLINER is capable of improving and harmonizing genome analysis across three state-of-the-art variant discovery pipelines in genomic oncology.

Humans

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article

A Comprehensive Bioinformatics Approach to Analysis of Variants: Variant Calling, Annotation, and Prioritization.

Next-Generation Sequencing (NGS), also known as high-throughput sequencing technologies, has enabled rapid and efficient sequencing of large amounts of DNA and RNA. These technologies have revolutionized the field of genomics, transcriptomics, and proteomics and have been widely used in cancer research, leading to advances in clinical diagnosis and treatment. Improvements in the NGS technologies enabled millions of fragments to be sequenced simultaneously in a time- and cost-effective manner and resulted in large amount of genomic data which require efficient analysis methods. Analysis of the genomic data requires both efficient computer resources and bioinformatics approaches. This chapter details a comprehensive computational approach and analysis steps for genomic data analysis.

Computational Biology

Somatic likelihood tiering: an interpretable post-calling triage protocol for tumor-only whole-exome variant review.

Tumor-only whole-exome sequencing (WES) is used when matched normal tissue is unavailable, but one sample can produce thousands of variants. Somatic likelihood tiering (SLT) is an interpretable post-calling protocol that ranks Mutect2 calls into four review-priority tiers using population-frequency, germline-quality, cancer-knowledge, PureCN posterior, and clonal-hematopoiesis evidence. Layer 2 distinguishes common, rare-callable, and unevaluable gnomAD states; missing or unmatchable gnomAD evidence is not positive rarity evidence. On the SEQC2 HCC1395 benchmark, the callability-aware SLT-A row contained 101 calls, 78 truth variants, 77.2% PPV (95% Wilson confidence interval 68.1%-84.3%), and a Number Needed to Review (NNR) of 1.29 (1.19-1.47). The conservative SLT-C catchment retained 352 of 455 truth variants (77.4%, 73.3%-81.0%) and all tiers together retained 430 of 455 truth variants. SNV performance is the primary calibration frame: SLT-C retained 341 of 439 SNV truth variants, whereas indel results were exploratory because only 16 truth indels were available. Clinical cohorts are reported as recall and concordance versus partially dependent matched-normal Mutect2 references, not independent clinical sensitivity. Patient-level bootstrap intervals were principal: HdM-BLCA-1 SLT-A recall was 18.2% (14.0%-23.5%), and LUAD-TW SLT-A recall was 49.1% (26.6%-63.3%) among 32 evaluable patients. The HdM-BLCA-1 median SLT-A queue remained 1277 variants per patient, so SLT reduces first-pass candidate counts but does not measure review time or eliminate FFPE candidate-count burden. SLT provides an auditable tumor-only WES review queue, not a substitute for matched-normal sequencing, independent orthogonal validation, or definitive somatic classification.

Humans

An alignment-free strategy for circulating tumor DNA detection and tumor fraction estimation from whole-genome sequencing data.

Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed alignment-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth ~ 28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmer's tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.

Circulating Tumor DNA

Backtracking Cell Phylogenies in the Human Brain with Somatic Mosaic Variants.

Somatic mosaic variants, and especially somatic single nucleotide variants (sSNVs), occur in progenitor cells in the developing human brain frequently enough to provide permanent, unique, and cumulative markers of cell divisions and clones. Here, we describe an experimental workflow to perform lineage studies in the human brain using somatic variants. The workflow consists in two major steps: (1) sSNV calling through whole-genome sequencing (WGS) of bulk (non-single-cell) DNA extracted from human fresh-frozen tissue biopsies, and (2) sSNV validation and cell phylogeny deciphering through single nuclei whole-genome amplification (WGA) followed by targeted sequencing of sSNV loci.

Humans

OctopuSV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis.

MOTIVATION: Structural variants (SVs) influence gene regulation, disease progression, and diagnostics, yet integrating SV calls across platforms remains difficult due to inconsistent annotations, limited merging flexibility, and fragmented workflows. Ambiguous breakend (BND) annotations, which comprise many variant calls, are often discarded or misclassified, hindering variant characterization. Existing tools lack advanced merging operations essential for precise identification of disease-specific or somatic variants across samples or patient groups. Additionally, current SV analysis pipelines require extensive manual intervention and complex parameter tuning, compromising reproducibility and scalability. Addressing these gaps is crucial for improving the accuracy, interpretability, and clinical utility of SV analyses. RESULTS: We developed OctopuSV and TentacleSV to address these long-standing challenges in SV analysis. OctopuSV features a specialized BND correction module that converts ambiguous BND annotations into canonical SV types, recovering important variants that are often overlooked by existing tools. Additionally, it provides advanced set operations (difference, complement, custom-defined) that enable sophisticated variant filtering without programming expertise, critical for identifying tumor-specific SVs or variants unique to specific sample groups. TentacleSV completes our solution by automating the entire SV analysis process from raw sequencing data to high-confidence callsets, ensuring consistency and reproducibility across projects. Benchmarking across short-read and long-read platforms showed superior F1 score, complete SV type consistency compared to existing tools. Our framework enables experimental biologists and clinical researchers to perform sophisticated analyses ranging from cancer subtype-specific SV identification to multi-sample comparative studies without requiring specialized programming skills. AVAILABILITY AND IMPLEMENTATION: All codes are available at https://github.com/ylab-hi/OctopuSV; https://github.com/ylab-hi/TentacleSV.

Software

SeqQC-former: A sequence-quality fusion framework for QC-aware review prioritization of candidate somatic SNVs in cancer genomics.

The accurate prioritization of candidate somatic single-nucleotide variants (SNVs) remains a challenge due to the substantial variability in sequencing quality across genomic loci. SeqQC-Former is a sequence-quality fusion framework that integrates the local nucleotide context with read-level quality-control (QC) covariates derived from matched tumor-normal sequencing data. This integration generates QC-aware prioritization scores for the downstream review of candidate variants. Unlike conventional variant callers, SeqQC-Former is designed not to infer biological truth but to support post-calling review and prioritization under heterogeneous sequencing conditions. The framework was trained and evaluated on a SEQC2-derived dataset comprising 89,447 candidate loci, including 1378 positive and 88,069 negative loci. In chromosome-held-out validation, which aims to reduce potential genomic-position leakage, SeqQC-Former demonstrated strong discrimination (AUROC = 0.9479; AUPRC = 0.9448), indicating good generalization to previously unseen chromosomes. Given that the SEQC2-derived labels contain QC-associated information; these results should be interpreted as an evaluation of QC-aware prioritization capability rather than an independent validation of biological variant correctness. Ablation analyses revealed that structured QC covariates provided the dominant predictive signal under the current SEQC2-derived labeling regime. SeqQC-Former achieved a significantly higher AUROC than classical machine-learning baselines, as determined by DeLong's test (p&#x202f;<&#x202f;0.01). Application to 53,164 glioblastoma variants demonstrated that external predictions were sensitive to QC scaling and threshold selection, underscoring that model outputs should be interpreted as QC-dependent prioritization scores rather than calibrated probabilities or definitive biological classifications. Overall, SeqQC-Former offers a reproducible post-calling QC-aware prioritization framework for large-scale somatic SNV review and underscores the importance of explicitly modeling sequencing-quality information when interpreting structured cancer genomics datasets.

Humans

Genetic control of local mutation rates.

Mutations are the source of evolutionary novelty but also the cause of genetic diseases and cancer. Mutation rates are known to be heterogeneous along the genome, however the extent to which local mutation rates vary among individuals in a population and are genetically determined is unknown. To test this, we analyzed the chromosomal distribution of somatic mutations in cell lines from 1,662 individuals, controlling for the confounding effects of DNA replication timing on local mutation rates and of trans-acting modulators on global mutation rates. We describe substantial interindividual variation in mutation rates across the human genome. By comparing mutation-rate variation to individuals' genotypes, we identified 35 instances in which polymorphic alleles in the population associate with somatic mutation rates in their vicinity. We call these mutation quantitative trait loci (mutQTLs). mutQTLs associated with somatic mutations in lymphoblastoid cell lines and in chronic lymphocytic leukemia, and with germline genetic variants. Two of the four mutQTLs inferred to be associated with germline mutation-rate variation were located within large clusters of zinc-finger genes and transposable elements, where they functioned as cis-mutators conferring an increased rate of mutation in their vicinity. mutQTLs provide a portal into the evolution of mutation rate heterogeneity across the genome and across individuals.

Humans

Severus detects somatic structural variation and complex rearrangements in cancer genomes using long-read sequencing.

For the detection of somatic structural variation (SV) in cancer genomes, long-read sequencing is advantageous over short-read sequencing with respect to mappability and variant phasing. However, most current long-read SV detection methods are not developed for the analysis of tumor genomes characterized by complex rearrangements and heterogeneity. Here, we present Severus, a breakpoint graph-based algorithm for somatic SV calling from long-read cancer sequencing. Severus works with matching normal samples, supports unbalanced cancer karyotypes, can characterize complex multibreak SV patterns and produces haplotype-specific calls. On a comprehensive multitechnology cell line panel, Severus consistently outperforms other long-read and short-read methods in terms of SV detection F1 score (harmonic mean of the precision and recall). We also illustrate that compared to long-read methods, short-read sequencing systematically misses certain classes of somatic SVs, such as insertions or clustered rearrangements. We apply Severus to several clinical cases of pediatric leukemia/lymphoma, revealing clinically relevant cryptic rearrangements missed by standard genomic panels.

Humans

High early death rates, treatment resistance, and short survival of&#xa0;Black adolescents and young adults with AML.

Survival of patients with acute myeloid leukemia (AML) is inversely associated with age, but the impact of race on outcomes of adolescent and young adult (AYA; range, 18-39 years) patients is unknown. We compared survival of 89 non-Hispanic Black and 566 non-Hispanic White AYA patients with AML treated on frontline Cancer and Leukemia Group B/Alliance for Clinical Trials in Oncology protocols. Samples of 327 patients (50 Black and 277 White) were analyzed via targeted sequencing. Integrated genomic profiling was performed on select longitudinal samples. Black patients had worse outcomes, especially those aged 18 to 29 years, who had a higher early death rate (16% vs 3%; P=.002), lower complete remission rate (66% vs 83%; P=.01), and decreased overall survival (OS; 5-year rates: 22% vs 51%; P<.001) compared with White patients. Survival disparities persisted across cytogenetic groups: Black patients aged 18 to 29 years with non-core-binding factor (CBF)-AML had worse OS than White patients (5-year rates: 12% vs 44%; P<.001), including patients with cytogenetically normal AML (13% vs 50%; P<.003). Genetic features differed, including lower frequencies of normal karyotypes and NPM1 and biallelic CEBPA mutations, and higher frequencies of CBF rearrangements and ASXL1, BCOR, and KRAS mutations in Black patients. Integrated genomic analysis identified both known and novel somatic variants, and relative clonal stability at relapse. Reduced response rates to induction chemotherapy and leukemic clone persistence suggest a need for different treatment intensities and/or modalities in Black AYA patients with AML. Higher early death rates suggest a delay in diagnosis and treatment, calling for systematic changes to patient care.

Adolescent

Genomic and epigenomic diversity of breast cancer across Western and MENA populations: implications for precision oncology.

Breast cancer is the most common malignancy in women worldwide and is increasingly recognized as a biologically diverse disease shaped by both molecular and ancestral context. Women from the Middle East and North Africa (MENA) populations, including Saudi Arabia, often present at a younger age and with more aggressive subtypes such as HER2-positive and triple-negative breast cancer (TNBC) compared with Western cohorts. These clinical patterns reflect a distinctive genomic background marked by high consanguinity, founder mutations in key susceptibility genes, and population-specific somatic alterations that are not fully captured in global reference datasets. This review brings together current evidence on somatic, germline, transcriptomic, and epigenomic diversity in breast cancer across Western and MENA populations, with a focus on Saudi cohorts. Drawing on a previously published systematic review of more than 2,500 MENA breast cancer cases, TP53 accounted for approximately 24% and PIK3CA for roughly 10% of curated somatic mutation records pooled across 44 studies (proportions of mutation calls, not per-patient prevalence); in a separate single-center Saudi cohort, only 3.7% of patients underwent BRCA testing, and 37.5% of this clinically selected, testing-referred subgroup carried a pathogenic variant, a figure that should not be read as general-population BRCA prevalence. Variants of uncertain significance exceeded 20% across several regional genomic studies. We summarize conserved driver events, such as recurrent TP53 and PIK3CA mutations, while highlighting regional features, including unique stop-gain and loss-of-function variants, a high copy-number burden, and early-onset disease linked to ancestral architecture. We also discuss emerging data on MENA-specific regulatory signatures, including immune-enriched and basal-myo transcriptomic clusters, CIMP-like methylation patterns, and non-coding RNA networks; these associations are numerically suggestive in available cohorts but have not reached statistical significance in existing studies and warrant validation in larger, dedicated MENA/Saudi cohorts before being considered established determinants of treatment response and resistance. Finally, we examine the clinical implications of this diversity for biomarker development, pharmacogenomics, and access to targeted therapies, and outline practical steps toward ancestry-aware precision oncology in the region.

BRCA

Clinical Variant Interpretation with the Integrative Genomics Viewer (IGV) for Molecular Pathologists.

The integrative genomics viewer (IGV) is a pivotal tool in clinical genomics, enabling the visualization and interpretation of complex sequencing data. Bringing clinical knowledge to bear with visual evaluation of sequencing results is the primary means by which molecular pathologists and other professionals assess and finalize cases. A variety of software tools can assist, but their relationship to the underlying data must be understood and applied systematically. This study includes essential background on next-generation sequencing (NGS) data file types (e.g., FASTQ, BAM, VCF) with a discussion of their format and purpose. We then describe features of IGV that derive nuances from these files. We utilize a series of curated practical cases based on clinical vignettes through which the reader will interact with clinical NGS sequencing data using the IGV software to review various types of clinically relevant variants relative to the human reference genome. These clinical vignettes have been curated to describe examples of some of the complexities of interpretation of genomic data, and how utilizing IGV as part of a routine workflow can provide additional interpretive information for variants beyond routine bioinformatic software algorithm variant calls. The visual inspection of genomic variants utilizing the tools within IGV can unmask subtle contextual cues (i.e., variant allele frequency, strand bias, tissue-specific context) that can influence the interpretation of genomic variants. Although this study focuses on using IGV for the detection and interpretation of somatic variants, the provided applications can be extrapolated for use in the germline setting, including analysis of complex variants and detection of mosaicism.

Humans

Genomic Characterization of ETV6::RUNX1-Positive Childhood B-ALL in a Chinese Cohort: Novel Fusion Partners, Co-Occurring Mutations, and Risk-Stratifying Biomarkers.

BACKGROUND: ETV6::RUNX1 is the most common genetic abnormality in pediatric B-cell acute lymphoblastic leukemia (ALL; &#x223c;25%), yet the comprehensive genetic architecture and molecular predictors of intermediate-risk (IR) stratification remain incompletely characterized. METHODS: We performed whole-transcriptome sequencing (Illumina NovaSeq 6000, rRNA depletion, 41.70 Gb/sample) on bone marrow samples from 93 pediatric ETV6::RUNX1-positive B-ALL patients. Bioinformatics analysis included STAR alignment, MuTect2 variant calling, FusionCatcher fusion detection, and VEP annotation. The Jaccard index with permutation testing assessed mutation co-occurrence; logistic regression identified independent predictors of IR classification. RESULTS: Beyond ETV6::RUNX1, we identified 51 distinct fusion genes across the cohort, including the reciprocal RUNX1-ETV6 (73.1%), chr8::KLF1210 (38.7%), and KLF12-chr8 (34.4%). Somatic mutations in 249 genes were detected; the most frequent were KIAA1715 (17.2%), KRAS (11.8%), and NSD2 (10.8%). Network analysis revealed significant chromatin modifier co-occurrence (KIAA1715-KMT2C: J = 0.136, p = 0.015) and KRAS-NRAS mutual exclusivity (J = 0.000, p = 0.042). PTCH1 (OR = 3.50, 95% CI 0.21-58.49, p = 0.41) and GNB1 (OR = 6.5, 95% CI 1.2-34.8, p = 0.029) mutations independently predicted IR classification. chr8::KLF1210 fusion correlated with higher Day-19 MRD levels (p = 0.038). CONCLUSIONS: GNB1 mutation represents a novel independent predictor of IR stratification in ETV6::RUNX1-positive B-ALL. The chromatin modifier co-occurrence module and extensive fusion architecture reveal biological heterogeneity within this favorable-risk subtype, with potential implications for risk-adapted therapeutic strategies.

B&#x2010;ALL

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software

Genome-wide detection of human 5' UTR variants that impact protein translation.

The 5' untranslated region (5' UTR) of messenger RNAs (mRNAs) plays a central role in regulating protein synthesis initiation, particularly through the Kozak sequence and upstream open reading frames (uORFs). Genetic variants within these regulatory elements could affect translation, altering gene expression and contributing to clinical phenotypes in humans. We developed a computational method called 5ULTRA (5' Untranslated Region Annotation) for analysis of whole-exome sequencing and whole-genome sequencing data to detect, annotate, and prioritize 5' UTR variants with potential translation impact. 5ULTRA identifies single-nucleotide variants, indels, and splicing variants that affect uORFs by creating or disrupting start/stop codons and that alter Kozak sequence strength of either the uORFs or the main coding sequence. 5ULTRA incorporates recent uORF databases and provides comprehensive annotations. 5ULTRA implements a machine-learning score to prioritize candidate variants with predicted effects on translation and also provides specific mechanistic predictions. The score correlates strongly with experimentally measured protein-level effects of 5' UTR variants. We applied 5ULTRA to multiple genetics datasets across diverse disease contexts, identifying candidate variants including potential cancer-driving somatic mutations predicted to decrease ABI1 level or increase NRAS abundance; common variants associated with traits such as multiple sclerosis, lung function, and cardiovascular function, by altering protein levels of TAGAP, VRTN, and SPAAR, respectively; and rare germline variants in our cohort, including a splicing variant of RPSA leading to 5' UTR sequence alteration that causes congenital asplenia and a variant of TNF that could predispose to tuberculosis.

Humans