Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

National genomic projects in Asia and Africa: a review.

National genome projects (NGPs) are increasingly shaping precision medicine by improving representation of population-specific genetic diversity. This review compiles findings from NGPs across Asia and Africa, regions that remain underrepresented in global genomic databases despite their extensive demographic and genetic diversity. A total of 53 studies from 24 countries were identified to understand (1) the genomic approach utilized, (2) novel findings that have emerged, and (3) strategies for improving research in these regions. The NGPs implement population-based variome databases (20 NGPs), linear reference genome assemblies (8 NGPs), and graph-based pangenome assemblies (1 NGP). Novel variants ranged between 0.28% (China) and 19.6% (Iran), whereas rare variants accounted for up to 88.9% of the detected variants in the Chinese population. Each NGP documents its country's evolutionary and migration history, which impacts disease frequency and pharmacogenomic variants. Clinically, NGPs revealed strong population stratification in disease-associated and pharmacogenomic variants. For example, the GJB2 rs72474224 hearing-loss variant ranged from 13% in Vietnam and 12% in Hong Kong to 0.0894% in Turkey, while the VKORC1 rs9923231 pharmacogenomic variant reached 89.2% in Taiwan but was 20%-25% in European-related Russian subpopulations. These findings demonstrate that clinically relevant allele frequencies, pathogenicity assessments, and drug-response markers differ substantially across ancestries. This review highlights ongoing efforts and strategies to enhance the representativeness of genomic data through NGPs in Asia and Africa. We also suggest future directions for national projects, including integrating family-based studies, multi-omic data, and standardized pipelines to accelerate discovery and support the equitable implementation of precision medicine.

Humans

Privacy-preserving framework for genomic computations via multi-key homomorphic encryption.

MOTIVATION: The affordability of genome sequencing and the widespread availability of genomic data have opened up new medical possibilities. Nevertheless, they also raise significant concerns regarding privacy due to the sensitive information they encompass. These privacy implications act as barriers to medical research and data availability. Researchers have proposed privacy-preserving techniques to address this, with cryptography-based methods showing the most promise. However, existing cryptography-based designs lack (i) interoperability, (ii) scalability, (iii) a high degree of privacy (i.e. compromise one to have the other), or (iv) multiparty analyses support (as most existing schemes process genomic information of each party individually). Overcoming these limitations is essential to unlocking the full potential of genomic data while ensuring privacy and data utility. Further research and development are needed to advance privacy-preserving techniques in genomics, focusing on achieving interoperability and scalability, preserving data utility, and enabling secure multiparty computation. RESULTS: This study aims to overcome the limitations of current cryptography-based techniques by employing a multi-key homomorphic encryption scheme. By utilizing this scheme, we have developed a comprehensive protocol capable of conducting diverse genomic analyses. Our protocol facilitates interoperability among individual genome processing and enables multiparty tests, analyses of genomic databases, and operations involving multiple databases. Consequently, our approach represents an innovative advancement in secure genomic data processing, offering enhanced protection and privacy measures. AVAILABILITY AND IMPLEMENTATION: All associated code and documentation are available at https://github.com/farahpoor/smkhe.

Computer Security

Real-world estrogen receptor alpha 1 (ESR1) testing patterns and results for ER+/HER2- metastatic breast cancer in the United States, 2018-2024.

PURPOSE: To understand historical and recent ESR1 testing rates, and when ESR1 mutations emerge during first-line (1 L) treatment. METHODS: This retrospective, observational cohort study used the Flatiron Health Research Database (FHRD) and the Flatiron Health-Foundation Medicine metastatic breast cancer (mBC) Clinico-Genomic Database (CGDB). Adult patients with a confirmed diagnosis of hormone receptor-positive/human epidermal growth factor receptor 2-negative mBC from 1/1/2018 to 6/30/2024 were included. ESR1 testing patterns and test results were descriptively analyzed. RESULTS: Among 7772 patients with mBC in the FHRD who initiated 1 L therapy, tumor ESR1 mutation status was evaluated for 222 (3%) patients at baseline (≤ 90 days before 1 L) and 1355 (17%) during 1 L. The percentage of patients who had an ESR1 test result reported during 1 L increased over time (11% in 2018-19, 19% in 2020-21, 22% in 2022-24). Median time from 1 L start to first ESR1 test was 7.4 months (mos) among tested patients. A positive test result was reported for 29/222 (13%) patients tested at baseline and 240/1355 (18%) tested during 1 L. Most (60%) tests during 1 L used tissue specimens, while the remaining 40% were liquid biopsies, and the median time from specimen collection to result reporting in 1 L was 28 (IQR:10-84) days. Focusing on time periods wherein specimens were provided, ESR1 test positivity was 6.7% (76/1,127) for specimens provided at baseline, 23% (15/65) for specimens provided 9 to 12 months into 1L therapy, 38% (26/69) for those provided 15 to 18 months into 1L, and 40% (38/94) for those provided from 18 to 24 months into 1L. CONCLUSIONS: ESR1 mutations can be detected at any time interval during 1L. CLINICAL TRIAL NUMBER: Not applicable.

Adult

Pharmacogenomic and drug interactions risk in cardio-oncology: A precision medicine perspective for India.

Cardio-oncology patients may face complex treatment regimens due to the concurrent existence of cancer and cardiovascular disease, leading to a considerable polypharmacy burden. This significantly increases the prospect of drug-drug interactions (DDIs) and gene-drug interactions. The majority of these interactions arise from comparable pharmacokinetic and pharmacological pathways associated with drug transporters and cytochrome P450 enzymes. The significance of pharmacogenomics in tailored treatment strategies are emphasised by the fact that genetic variability enhances individual differences in drug response, safety, and efficacy. This narrative review focus on the effects of key genetic polymorphisms (e.g., DPYD, CYP2C19, and CYP2C9) on the metabolism and efficacy of commonly prescribed anticancer and cardiovascular medications such as fluoropyrimidines, clopidogrel, and warfarin. In addition it explore the role of pharmacogenomic variants on drug-drug interactions within the field of cardio-oncology. The study ultimately emphasizes the necessity of precision medicine in India to address the genetic diversity and underrepresentation in global genomic databases. The absence of pharmacogenomic testing, infrastructural deficiencies, financial constraints, and insufficient clinical integration hinder the widespread use of this technology in India. The Genome India Project and other national initiatives establish the foundation for pharmacogenomic-guided therapy. Utilizing genetic data, together with artificial intelligence-based predictive tools, for clinical decision-making may enhance medication safety and yield optimal outcomes in Indian cardio-oncology patients.

Humans

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface

ZBTB16-associated NK cell alterations reveal shared immunometabolic signatures linking primary Sjögren's syndrome and type 1 diabetes mellitus.

BACKGROUND: Primary Sjögren's syndrome (pSS) and type 1 diabetes mellitus (T1DM) share immune-inflammatory features, yet conserved pathogenic signatures linking these autoimmune disorders remain incompletely understood. The present research sought to uncover common molecular markers and dissect the underlying immune-metabolic cross-talk underlying pSS and T1DM. METHODS: Gene expression profiles of patients with pSS and T1DM were retrieved from the Gene Expression Omnibus database, normalized, and corrected for batch effects prior to downstream analyses. Overlapping potential biomarkers were screened by integrating differential expression analysis, weighted gene co-expression network analysis and least absolute shrinkage and selection operator regression. Functional enrichment based on Gene Ontology and Kyoto Encyclopedia of Genes and Genomes databases was implemented to interpret gene biological properties, and a protein-protein interaction network was further established afterwards. Diagnostic performance was evaluated using receiver operating characteristic analysis. Experimental validation was conducted in non-obese diabetic (NOD) mice using quantitative PCR, immunohistochemistry, and flow cytometry. The CIBERSORT algorithm was adopted to quantify immune cell infiltration levels. RESULTS: ZBTB16 was identified as a shared hub biomarker in both pSS and T1DM and exhibited favorable diagnostic performance. Experimental validation confirmed significantly reduced ZBTB16 expression in peripheral blood mononuclear cells, salivary gland tissues, and pancreatic tissues of NOD mice. Gene Set Enrichment Analysis indicated that ZBTB16-associated signatures were enriched in mitochondrial-related processes, neuroactive ligand-receptor interactions, and ribosome-related pathways. Immune infiltration analysis revealed that resting natural killer (NK) cells were positively correlated with ZBTB16 expression in both diseases. Flow cytometric analysis further confirmed a reduced proportion of resting NK cells in peripheral blood of NOD mice, consistent with the CIBERSORT-based prediction. CONCLUSION: This study identifies ZBTB16 as a shared biomarker linking pSS and T1DM. Reduced resting NK-cell abundance was consistently observed in both computational and experimental analyses, and bioinformatic correlation analysis suggested a positive association with ZBTB16 expression. These findings provide evidence for shared molecular and immunological signatures underlying the two autoimmune disorders and support further investigation of the biological role and diagnostic value of ZBTB16 in pSS and T1DM.

Sjogren's Syndrome

Semiparametric efficient estimation of small genetic effects in large-scale population cohorts.

Population genetics seeks to quantify DNA variant associations with traits or diseases, as well as interactions among variants and with environmental factors. Computing millions of estimates in large cohorts in which small effect sizes and tight confidence intervals are expected, necessitates minimizing model-misspecification bias to increase power and control false discoveries. We present TarGene, a unified statistical workflow for the semi-parametric efficient and double robust estimation of genetic effects including $ k $-point interactions among categorical variables in the presence of confounding and weak population dependence. $ k $-point interactions, or Average Interaction Effects (AIEs), are a direct generalization of the usual average treatment effect (ATE). We estimate genetic effects with cross-validated and/or weighted versions of Targeted Minimum Loss-based Estimators (TMLE) and One-Step Estimators (OSE). The effect of dependence among data units on variance estimates is corrected by using sieve plateau variance estimators based on genetic relatedness across the units. We present extensive realistic simulations to demonstrate power, coverage, and control of type I error. Our motivating application is the targeted estimation of genetic effects on trait, including two-point and higher-order gene-gene and gene-environment interactions, in large-scale genomic databases such as UK Biobank and All of Us. All cross-validated and/or weighted TMLE and OSE for the AIE $ k $-point interaction, as well as ATEs, conditional ATEs and functions thereof, are implemented in the general purpose Julia package TMLE.jl. For high-throughput applications in population genomics, we provide the open-source Nextflow pipeline and software TarGene which integrates seamlessly with modern high-performance and cloud computing platforms.

Humans

SPEN inactivation drives resistance to androgen receptor pathway inhibitors in metastatic prostate cancer.

PURPOSE: Treatment intensification with androgen receptor pathway inhibitors (ARPIs) has become the standard of care for patients with metastatic prostate cancer. However, there remains an unmet need to identify biomarkers for treatment resistance. Here, we identify SPEN inactivation as a driver of ARPI resistance. EXPERIMENTAL DESIGN: Pre-clinical studies were performed in LNCaP and VCaP cell lines. Data from a nationwide prostate cancer clinico-genomic database were extracted. Log-rank test and Cox proportional hazards models were used to compare time to next treatment (TTNT) on ARPI with/without SPEN mutations. SPEN immunohistochemistry was performed on a rapid autopsy metastatic tissue microarray. RESULTS: SPEN was identified as a top enzalutamide resistance hit in an unbiased genome-wide loss-of-function screen. SPEN inactivation results in upregulation of cell cycle proliferation and basal/stem cell activity as well as increased translation of pro-oncogenic genes. In a large patient cohort (N=6828), SPEN mutations are enriched following treatment with ARPIs (2.1% to 3.6%, p=0.001) and correlate with shorter TTNT on ARPI in patients with metastatic hormone-sensitive prostate cancer (6.4 vs 29.7 months, HR 2.67, p=0.02). In a metastatic rapid autopsy cohort (N=181), low SPEN H-score is associated with shorter time on abiraterone (5.0 vs 7.9 months, p=0.023) in metastatic castration-resistant prostate cancer. CONCLUSIONS: In real-world cohorts, loss of SPEN function across genomic, transcriptomic, and protein levels is associated with reduced benefit from ARPI therapy in metastatic prostate cancer. These findings identify SPEN inactivation as a clinically relevant biomarker of ARPI resistance that warrants prospective evaluation to guide treatment selection.

Journal Article

An open-source clinical bioinformatics pipeline for real-world NGS implementation: translating genomic variants into actionable treatment strategies in oncology.

BACKGROUND: Next-Generation Sequencing (NGS) has become a cornerstone technology in clinical practice, yet its adoption presents significant challenges. Physicians and oncologists must manage vast amounts of genome-scale data and transform it into actionable insights for complex decision-making. While commercial systems exist to synthesize data from NGS experiments into clinical reports, many are hindered by limitations such as closed-source designs that restrict transparency and customization. Additionally, some fail to leverage publicly available genomic databases, missing opportunities to integrate valuable external data. Furthermore, the rigidity of many tools in accommodating diverse NGS panels limits their applicability across varied clinical scenarios. METHODS: To address these limitations, we developed OncoReport, an open-source tool that generates comprehensive reports from NGS analyses. By integrating publicly accessible databases, OncoReport provides a robust, user-friendly environment equipped with essential tools for NGS analysis. This design aims to enhance data interpretation and support informed clinical decision-making. RESULTS: Rigorous testing has demonstrated OncoReport’s effectiveness in producing detailed, actionable reports that are clear and easy to use. By automating key aspects of the workflow, the tool significantly reduces manual effort and expedites the synthesis and interpretation of NGS results, making genomic insights more accessible to clinicians. CONCLUSION: OncoReport offers a transparent, flexible, and efficient framework for clinicians to analyze and apply genomic data in patient care. By streamlining workflows and leveraging open-source principles, it empowers healthcare professionals to make informed, data-driven decisions. OncoReport is freely available at https://oncoreport.atlas.dmi.unict.it, with source code and issue tracking on GitHub: https://github.com/knowmics-lab/oncoreport .

Humans

Lifestyle Differentiation Among Marine Denitrifying Microorganisms.

Microorganisms carrying out denitrification in marine anoxic zones drive bioavailable nitrogen loss. Sequencing datasets have demonstrated the modularity of denitrification, with most populations having the genetic capability for only a subset of the pathway (NO3-➔NO2-➔NO➔N2O➔N2). Although previous work provided ecological explanations for this diversity among the functional modules, large trait variations exist within each functional module, and this within-module diversity and its biogeochemical implications remain unexplored. Here, we combine genomic data and modeling to explore how metabolic "lifestyle" strategies influence denitrifier community structure. We build a comprehensive genomic database of marine denitrifiers, and identify lifestyle differentiation among denitrifier functional groups. We then extend a mathematical ecosystem model by resolving two microbial functional types for each module representing a metabolic trade-off: a copiotroph, optimized for fast growth, and an oligotroph, optimized for high nutrient affinity. In the model, as the supply of organic matter relative to nitrate increases, the degree of copiotrophy among the community increases and then decreases. This suggests that oligotrophs are associated with either organic-matter- or nitrate-limiting conditions, whereas copiotrophic lifestyles are associated with an intermediate regime. Our model further associates NO2- reducers with oligotrophy and NO3- reducers with copiotrophy, particularly those producing greenhouse gas nitrous oxide (N2O), linking N2O production to substrate-replete conditions, which is consistent with our genome-based lifestyle estimates. Results provide insight into denitrifier ecological niches and thus the biogeochemical conditions that are associated with the production of intermediates, such as N2O, improving our understanding of how nitrogen cycling will change in a warming ocean.

Marine denitrifiers

Maternal contact and age-dependent succession influence the assembly of the calf rumen microbiome and virome.

Early-life colonization of the rumen is particularly important; however, the processes by which microbial and viral communities are transmitted and developed remain poorly understood. Here, we present a genome-resolved investigation of the effects of maternal contact and age-dependent succession on the calf rumen microbiome and DNA virome by comparing calves raised with or without maternal contact across early life using the metagenome-assembled genomes (MAGs) and viral operational taxonomic units (vOTUs) reconstructed from whole- and virus-like particle metagenomes. Across longitudinal samples from calves and their mothers, we identified 694 MAGs and 30,479 vOTUs, substantially expanding current genome databases and revealing extensive microbial and viral novelty. Our analyses demonstrated that both prokaryotes and DNA viruses are shared between dams and calves, with greater sharing observed in calves raised with maternal contact than in calves raised without maternal contact. Notably, viral sharing between cow-calf pairs was markedly lower compared to prokaryotes, suggesting high turnover and rapid viral diversification. Age-associated analyses further revealed coordinated shifts in prokaryotes and their viruses, with dominant genera such as Prevotella, Ruminococcus, and Fibrobacter, and their corresponding viruses increasing after day 40. These findings indicate that the early-life rumen microbiome and DNA virome undergo substantial age-dependent succession and are associated with maternal contact, providing new insights into host-microbe-virus interactions during rumen development.IMPORTANCEThis study provides one of the first genome-resolved views of DNA viral community development during early rumen colonization in calves (from 1 week to 70 days of age) and reveals how maternal contact and age influence the establishment of the calf rumen microbiome and virome. By analyzing longitudinal samples from calves raised with or without their mothers, we show that prokaryotes and their viruses undergo coordinated, age-dependent succession. Our results demonstrate that maternal separation alters the assembly of the calf rumen microbiome, highlighting the influence of maternal contact during early-life rumen development. These findings underscore the high plasticity of the early-life rumen ecosystem and suggest that early management practices, such as maternal separation, can have lasting effects on rumen development. This work provides fundamental insights into the establishment and succession of the calf rumen microbiome and DNA virome during early life and may contribute to future microbiome manipulation studies.

Animals

Exploiting the weak link: Ataxia-Telangiectasia Mutated dysfunction in oesophagogastric tumours.

ATM (ataxia-telangiectasia mutated) is a central regulator of the DNA damage response, coordinating double-strand break repair, checkpoint control, and cell fate decisions. Its disruption drives genomic instability and has been implicated across multiple tumour types. In oesophagogastric cancers, ATM alterations occur in a clinically relevant subset of cases, encompassing both somatic and germline events, and are associated with distinct molecular features including reduced co-occurrence with TP53 mutations and elevated homologous recombination deficiency scores. This narrative review synthesises published literature and publicly available genomic databases to examine ATM biology, the spectrum of ATM alterations across oesophageal adenocarcinoma, oesophageal squamous cell carcinoma, and gastric cancer subtypes, and the challenges of defining true ATM deficiency. The therapeutic implications of ATM dysfunction are evaluated across radiotherapy, platinum-based chemotherapy, ATR inhibition, and PARP inhibition. ATM alterations are detected in approximately 6% of tumours pan-cancer and in up to 10% of oesophagogastric cases. Defining ATM deficiency remains challenging, as immunohistochemistry, next-generation sequencing, and functional assays each carry distinct limitations. ATR inhibition emerges as the most consistently supported therapeutic strategy, with converging preclinical and early clinical evidence across oesophagogastric models. By contrast, available data do not support treating ATM deficiency as equivalent to BRCA-like homologous recombination deficiency, and PARP inhibitor monotherapy has not demonstrated consistent benefit. Prospective validation of functional ATM assays, histology-stratified trial design, and integration of genomic, protein-level, and functional evidence represent key priorities for translating ATM-guided strategies into oesophagogastric cancer practice.

Humans

Identification of CD55 as a downstream factor of EP4 receptor signaling in colorectal cancer cells.

Prostaglandin E2 (PGE2) signaling through the E-type prostanoid 4 (EP4) receptor has been implicated in the pathophysiology of colorectal cancer (CRC). We herein identified decay-accelerating factor, also known as CD55, as a novel CRC-associated downstream factor of the EP4 receptor. The integration of transcriptomic profiling of PGE2-stimulated HCA-7 human colon cancer cells with analyses of cancer genomic databases predicted CD55 as a potential EP4 receptor-regulated target. Inhibitor-based experiments showed the induction of CD55 after a PGE2 stimulation required the EP4 receptor and Gi protein in HCA-7 cells, whereas protein kinase A signaling was dispensable. In combination with a toxicogenomic database analysis, p38 mitogen-activated protein kinase (MAPK) was identified as the predominant effector connecting the EP4 receptor to CD55 upregulation. A single-cell RNA-seq re-analysis of human CRC tissues revealed CD55 upregulation and p38 MAPK-related gene set enrichment in epithelial cells expressing the EP4 receptor, suggesting that this induction mechanism may operate in a subset of epithelial cells in clinical specimens. Collectively, these results delineate a PGE2/EP4 receptor/Gi protein/p38 MAPK signaling axis that induces CD55 expression in HCA-7 cells and epithelial tumor cells, provide new mechanistic clues for understanding the regulation of complement regulatory molecule CD55 expression by prostaglandin signaling.

Humans

Identification of a Novel Splice-Site variant in TACR3 (c.888 + 1G > A) Associated with Asthenozoospermia and Hypogonadotropic Hypogonadism in an Iranian Family.

BACKGROUND: TACR3 encodes the receptor for neurokinin B, a key regulator of the hypothalamic-pituitary-gonadal axis. Disruption of this pathway can impair gonadotropin release and male reproductive function. Given the genetic heterogeneity of male infertility, this study aimed to identify novel variants in TACR3 that may underlie asthenozoospermia and related hormonal abnormalities. METHODS: Fifteen infertile men with confirmed asthenozoospermia were enrolled. Whole-exome sequencing (WES) was performed on genomic DNA from peripheral blood, and the candidate variant was validated by Sanger sequencing. Functional predictions were made using PolyPhen-2, SIFT, MutationTaster, and REVEL. TACR3 mRNA expression levels were assessed by real-time PCR in available samples. RESULTS: A novel splice-site variant, TACR3 (NM_001059.3:c.888 + 1G > A), was detected and found to segregate with infertility in one family, appearing homozygously in two infertile brothers and heterozygously in the proband with severe asthenozoospermia. The variant was absent in public and local genomic databases, suggesting its extremely rare frequency. Furthermore, RT-PCR showed a dramatic reduction or complete loss of TACR3 expression in affected individuals, confirming its deleterious effect on splicing and mRNA stability. CONCLUSION: We identified a previously unreported splice-site mutation in TACR3 (c.888 + 1G > A) that likely causes familial infertility by disrupting the neurokinin B/NK3R signaling pathway. While the heterozygous proband exhibited severe asthenozoospermia, the homozygous brothers displayed hormonal profiles typical of hypogonadotropic hypogonadism. These findings extend the mutational landscape of TACR3 and highlight its essential contribution to male reproductive endocrinology.

Humans

The impact of EGFR subtype combined with TP53 co-mutation status on survival outcomes with front-line osimertinib in non-small cell lung cancer (NSCLC).

BACKGROUND: Osimertinib is a standard therapy for EGFR-mutant NSCLC. However, markers to better identify those at risk for poor outcomes are needed. This is the largest study to date evaluating the impact of EGFR subtype combined with TP53 co-mutation status on survival endpoints with front-line osimertinib. METHODS: Patients from a U.S. clinical-genomic database with advanced EGFR-mutant NSCLC receiving front-line osimertinib were studied. Real-world progression-free survival (rwPFS) and overall survival (OS) were determined using Kaplan-Meier methods, and multivariable Cox regression compared outcomes after accounting for relevant clinical covariates. RESULTS: Of 606 patients, 277 (46%) had EGFR L858R, 384 (63%) had TP53 co-mutations, and 186 (30.7%) had both. Bearing L858R vs. exon 19 deletions (rwPFS: hazard ratio [HR] 1.4, P&#xa0;=&#xa0;0.001; OS: HR 1.3, P&#xa0;=&#xa0;0.01) or a TP53 co-mutation vs. wildtype (rwPFS: HR 1.5, P&#xa0;<&#xa0;0.001; OS: HR 1.6, P&#xa0;<&#xa0;0.001) predicted inferior outcomes. Especially short median rwPFS (10.1 vs. 21.4&#xa0;months, HR 2.2, P&#xa0;<&#xa0;0.001) and OS (21.3 vs. 53.4&#xa0;months, HR 2.3, P&#xa0;<&#xa0;0.001) were observed in patients with both markers (L858R/TP53-mutant) as compared to neither (exon 19 deletion/TP53-wildtype). CONCLUSIONS: Having EGFR L858R or a TP53 co-mutation were independent predictors of inferior rwPFS and OS with front-line osimertinib. Patients with both unfavorable alterations had the shortest survival. Risk stratifying using a combination of these markers can assist in identifying patients for novel trials or approved intensified therapies.

Humans

Association between smoking, tumor genetics, and outcomes in men with metastatic prostate cancer.

PURPOSE: Smoking has been associated with increased metastatic prostate cancer mortality, but the mechanisms behind this are largely unknown. We hypothesized that smoking increases the risk of genetic alterations associated with aggressive disease and/or the transformation to neuroendocrine prostate cancer (NEPC). PATIENTS AND METHODS: We utilized the Prostate Cancer Precision Medicine Multi-institutional Collaborative Effort (PROMISE) clinical genomic database for this retrospective analysis. We associated patient characteristics and tumor genetic data with smoking exposure at diagnosis (current, former, never and pack years) and with clinical outcomes, including overall survival (OS) from diagnosis or time to developing metastatic disease and NEPC status. RESULTS: We identified 2353 men with prostate cancer and next generation somatic tumor sequencing evaluable for analysis in PROMISE, including 8% current, 39% former, and 52% never smokers. Current smokers were more likely to be younger and to have metastatic (M1 or N1) disease at diagnosis, and less likely to have prior local therapy (all p&#x2009;<&#x2009;0.001). Current smoking was associated with worse OS from diagnosis (99.9 mo vs 137.6 mo, HR 1.42, 95% CI 1.14-1.77), which remained significant after adjusting for disease characteristics. We found no difference in the percentage of NEPC at initial diagnosis or at any time between current, former, and never smokers (p&#x2009;=&#x2009;0.8). We found positive associations between smoking status and genetic alterations in SPOP (current: 15%, former 6.7%, never 3.8%; p&#x2009;=&#x2009;0.018), FGFR1 (current 10%, former 0.4%, never 1.1% p&#x2009;=&#x2009;0.001), and ARID1A (current 5.1%, former 2.2%, never 0.4%; p&#x2009;=&#x2009;0.035) in patients with&#xa0;metastatic androgen pathway modulator sensitive prostate cancer (APMS). CONCLUSION: Active smoking is associated with worse overall and prostate cancer specific survival as compared to never/former smoking and was associated with specific tumor genetic alterations but not small cell/NEPC transformation.

Journal Article

Digital Kennison: A bioinformatics pipeline for rapid mapping of sequences to the Drosophila melanogaster Y chromosome.

The Drosophila melanogaster Y chromosome is currently known to contain 13 single-copy protein-coding genes, six of which are essential for male fertility, as well as several non-coding genes and abundant repetitive DNA. Localization of Y-linked sequences has traditionally relied on labor-intensive crosses using Kennison's translocation strains, which map Y-linked loci by generating flies deficient for each of the six Y-chromosome fertility regions (ks-1, ks-2, kl-1, kl-2, kl-3, and kl-5). Here we present Digital Kennison, a computational pipeline that recasts this classical mapping strategy as a sequence-based analysis. The pipeline queries eight genomic databases derived from Kennison's strains using BLAST and read coverage, assigning sequences to fertility regions or the centromeric region with a calibrated confidence score. We benchmarked the method on 60 Y-linked sequences spanning all seven regions, including single-copy protein-coding genes, Mst77Y family members, non-coding RNAs, and the centromere. Digital Kennison achieved 97% precision while resolving challenging cases, including boundary-spanning genes (PRY and Ppr-Y), fragmented Mst77Y copies, and FDY, which has a closely related autosomal paralog. Beyond validating known localizations, the pipeline localized the unmapped gene CG41561 to the kl-1region and reassigned the transcript CR40629-RC from the kl-2 region to kl-5. It also localized 7 of 16 recently transferred Y-linked sequences, including 4 with high confidence. Applied to 904 R6 scaffolds, Digital Kennison assigned 75% to fertility regions, including five currently annotated as autosomal-pericentromeric. Digital Kennison reduces sequence localization from weeks of genetic crosses to minutes of computation while preserving the power of classical translocation mapping.

Drosophila melanogaster