Search PubMedSearch

Biomedical subjects

Muhammad Zubair

Publications and source records attributed to Muhammad Zubair.

2 recordsLinked to original sources

SeqQC-former: A sequence-quality fusion framework for QC-aware review prioritization of candidate somatic SNVs in cancer genomics.

The accurate prioritization of candidate somatic single-nucleotide variants (SNVs) remains a challenge due to the substantial variability in sequencing quality across genomic loci. SeqQC-Former is a sequence-quality fusion framework that integrates the local nucleotide context with read-level quality-control (QC) covariates derived from matched tumor-normal sequencing data. This integration generates QC-aware prioritization scores for the downstream review of candidate variants. Unlike conventional variant callers, SeqQC-Former is designed not to infer biological truth but to support post-calling review and prioritization under heterogeneous sequencing conditions. The framework was trained and evaluated on a SEQC2-derived dataset comprising 89,447 candidate loci, including 1378 positive and 88,069 negative loci. In chromosome-held-out validation, which aims to reduce potential genomic-position leakage, SeqQC-Former demonstrated strong discrimination (AUROC = 0.9479; AUPRC = 0.9448), indicating good generalization to previously unseen chromosomes. Given that the SEQC2-derived labels contain QC-associated information; these results should be interpreted as an evaluation of QC-aware prioritization capability rather than an independent validation of biological variant correctness. Ablation analyses revealed that structured QC covariates provided the dominant predictive signal under the current SEQC2-derived labeling regime. SeqQC-Former achieved a significantly higher AUROC than classical machine-learning baselines, as determined by DeLong's test (p&#x202f;<&#x202f;0.01). Application to 53,164 glioblastoma variants demonstrated that external predictions were sensitive to QC scaling and threshold selection, underscoring that model outputs should be interpreted as QC-dependent prioritization scores rather than calibrated probabilities or definitive biological classifications. Overall, SeqQC-Former offers a reproducible post-calling QC-aware prioritization framework for large-scale somatic SNV review and underscores the importance of explicitly modeling sequencing-quality information when interpreting structured cancer genomics datasets.

Humans

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning