Search PubMedSearch

Biomedical subjects

Yasminka A Jakubek

Publications and source records attributed to Yasminka A Jakubek.

3 recordsLinked to original sources

Systematic contextual biases in SegmentNT potentially relevant to other nucleotide transformer models.

Recent advances in large language models have extended to genomic applications, yet model robustness relative to context is unclear. Here, we demonstrate two intrinsic biases (input sequence length and nucleotide position) affecting SegmentNT results, a model included with the Nucleotide Transformer that provides nucleotide-level predictions of biological features. We demonstrate that nucleotide position within the input sequence (beginning, middle, or end) alters the nature of SegmentNT's raw prediction probabilities, which can be standardized to improve prediction consistency. While longer input sequence length improves model performance, diminishing returns suggest a surprisingly small input length of ∼3072 nucleotides might be sufficient for many applications. We further identify a 24-nucleotide periodic oscillation in SegmentNT's prediction probabilities, revealing an intrinsic bias potentially linked to the model's training tokenization (6-mers) and architecture. We identify potential approaches to account for these biases and provide generalizable insights for utilizing nucleotide-resolution functional prediction models.

Nucleotides

qcCHIP: an R package to identify clonal hematopoiesis variants using cohort-specific data characteristics.

SUMMARY: Clonal hematopoiesis (CH) is a molecular biomarker associated with various adverse outcomes in both healthy individuals and those with underlying conditions, including cancer. Detecting CH usually involves genomic sequencing of individual blood samples followed by robust bioinformatics data filtering. We report an R package, qcCHIP, a bioinformatics pipeline that implements permutation-based parameter optimization to guide quality control filtering and cohort-specific CH identification. We benchmark qcCHIP under various data settings, including different sequencing depths, ranges of cohort sizes, with and without normal-tumor paired samples, and across different cancer types. We show that qcCHIP allows users to customize analysis needs to generate CH calls based on cohort-specific data characteristics. AVAILABILITY AND IMPLEMENTATION: qcCHIP R package is freely accessible at GitHub https://github.com/tenglab/qcCHIP and DOI: 10.5281/zenodo.16421861.

Humans

The Genetic Determinants and Genomic Consequences of Non-Leukemogenic Somatic Point Mutations.

Clonal hematopoiesis (CH) is defined by the expansion of a lineage of genetically identical cells in blood. Genetic lesions that confer a fitness advantage, such as point mutations or mosaic chromosomal alterations (mCAs) in genes associated with hematologic malignancy, are frequent mediators of CH. However, recent analyses of both single cell-derived colonies of hematopoietic cells and population sequencing cohorts have revealed CH frequently occurs in the absence of known driver genetic lesions. To characterize CH without known driver genetic lesions, we used 51,399 deeply sequenced whole genomes from the NHLBI TOPMed sequencing initiative to perform simultaneous germline and somatic mutation analyses among individuals without leukemogenic point mutations (LPM), which we term CH-LPMneg. We quantified CH by estimating the total mutation burden. Because estimating somatic mutation burden without a paired-tissue sample is challenging, we developed a novel statistical method, the Genomic and Epigenomic informed Mutation (GEM) rate, that uses external genomic and epigenomic data sources to distinguish artifactual signals from true somatic mutations. We performed a genome-wide association study of GEM to discover the germline determinants of CH-LPMneg. After fine-mapping and variant-to-gene analyses, we identified seven genes associated with CH-LPMneg (TCL1A, TERT, SMC4, NRIP1, PRDM16, MSRA, SCARB1), and one locus associated with a sex-associated mutation pathway (SRGAP2C). We performed a secondary analysis excluding individuals with mCAs, finding that the genetic architecture was largely unaffected by their inclusion. Functional analyses of SMC4 and NRIP1 implicated altered HSC self-renewal and proliferation as the primary mediator of mutation burden in blood. We then performed comprehensive multi-tissue transcriptomic analyses, finding that the expression levels of 404 genes are associated with GEM. Finally, we performed phenotypic association meta-analyses across four cohorts, finding that GEM is associated with increased white blood cell count and increased risk for incident peripheral artery disease, but is not significantly associated with incident stroke or coronary disease events. Overall, we develop GEM for quantifying mutation burden from WGS without a paired-tissue sample and use GEM to discover the genetic, genomic, and phenotypic correlates of CH-LPMneg.

Journal Article