Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic Library”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

polars-bio-fast, scalable, and out-of-core operations on large genomic interval datasets.

MOTIVATION: Genomic studies very often rely on computationally intensive analyses of relationships between features, which are typically represented as intervals along a 1D coordinate system (such as positions on a chromosome). In this context, the Python programming language is extensively used for manipulating and analyzing data stored in a tabular form of rows and columns, called a DataFrame. Pandas is the most widely used Python DataFrame package and has been criticized for inefficiencies and scalability issues, which its modern alternative-Polars-aims to address with a native backend written in the Rust programming language. RESULTS: polars-bio is a Python library that enables fast, parallel and out-of-core operations on large genomic interval datasets. Its main components are implemented in Rust, using the Apache DataFusion query engine and Apache Arrow for efficient data representation. It is compatible with Polars and Pandas DataFrame formats. In a real-world comparison (107 versus 1.2×106 intervals), our library runs overlap queries 6.5×, nearest queries 15.5×, count_overlaps queries 38×, and coverage queries 15× faster than Bioframe. On equally sized synthetic sets (107 versus 107), the corresponding speedups are 1.6×, 5.5×, 6×, and 6×. In streaming mode, on real and synthetic interval pairs, our implementation uses 90× and 15× less memory for overlap, 4.5× and 6.5× less for nearest, 60× and 12× less for count_overlaps, and 34× and 7× less for coverage than Bioframe. Multi-threaded benchmarks show good scalability characteristics. To the best of our knowledge, polars-bio is the most efficient single-node library for genomic interval DataFrames in Python. AVAILABILITY AND IMPLEMENTATION: polars-bio is an open-source Python package distributed under the Apache License available for major platforms, including Linux, macOS, and Windows in the PyPI registry. The online documentation is https://biodatageeks.org/polars-bio/ and the source code is available on GitHub: https://github.com/biodatageeks/polars-bio and Zenodo: https://doi.org/10.5281/zenodo.16374290. are available at Bioinformatics online.

Software

Protocol to decode the role of transcriptionally active microbes in SARS-CoV-2-positive patients using an RNA-seq-based approach.

The elucidation of the role of microorganisms in human infections has been hindered by difficulties using conventional culture-based techniques. Here, we present a protocol for the investigation of transcriptionally active microbes (TAMs) using an RNA sequencing (RNA-seq)-based approach. We describe the steps for RNA isolation, viral genome sequencing, RNA-seq library preparation, and metatranscriptomic and transcriptomic analysis. This protocol permits a comprehensive evaluation of TAMs' contributions to the differential severity of infectious diseases, with a particular focus on diseases such as COVID-19. For complete details on the use and execution of this protocol, please refer to Devi et al.1.

Humans

Using Chromosome Conformation Capture Combined with Deep Sequencing (Hi-C) to Study Genome Organization in Bacteria.

Genome organization is fundamental to all living organisms. Long DNA molecules are organized in hierarchical orders to be accommodated into eukaryotic nuclei or bacterial cells, which are thousands of folds shorter. Over the past two decades, chromosome conformation capture (3C) techniques substantially advanced our understanding of genome folding inside cells. 3C involves crosslinking and proximity ligation, and quantifies the physical contacts between two DNA regions within the genome. Coupled with high-throughput sequencing, 3C-seq and Hi-C techniques detect genome-wide DNA interactions, providing a comprehensive view of global genome organization. Here, we describe a detailed method to prepare Hi-C libraries using Bacillus subtilis, which includes procedures of crosslinking chromatin, digesting the crosslinked genome, labeling DNA ends with biotin, ligating DNA, and preparing the DNA library for sequencing using an Illumina platform.

High-Throughput Nucleotide Sequencing

Effectiveness of mass spectrometry and genomic analysis in the surveillance of nontuberculous Mycobacterium in Taiwan.

Nontuberculous mycobacteria (NTM) are diverse, and species-level identification remains challenging in routine diagnostics. We analyzed NTM isolates collected at three regional centers of the National Taiwan University Hospital (NTUH) from 2019 to 2024 to assess geographic variation and identification performance after implementation of matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS). Among 3,188 cases meeting the microbiological criteria for probable pulmonary NTM disease, the species distribution differed by region: Mycobacterium avium complex predominated in central Taiwan (Yunlin, 47.3%), whereas M. abscessus complex (Taipei, 26.5%) and M. kansasii (Hsinchu, 12.4%) were more common in northern Taiwan. In 2019, 14.5% of isolates were reported to be unidentified by MALDI-TOF MS; with workflow optimization and database updates, this percentage decreased but plateaued at 4.5-4.8%. Whole-genome sequencing (WGS) of 61 randomly selected persistently unidentified isolates revealed eight average nucleotide identity (ANI)-defined clusters; 55 isolates (90.2%) could not be assigned to known species using current reference databases. Two clusters detected only in Hsinchu were phylogenetically closest to M. kyorinense, with ANI values below the species demarcation threshold. Overall, we observed marked regional heterogeneity of NTM in Taiwan and a persistent identification gap that remained after MALDI-TOF MS optimization and follow-up WGS.IMPORTANCEThis study characterized regional differences in the NTM species distribution across Taiwan, and the results highlight the limitations of current identification approaches. MALDI-TOF MS identifies most isolates, but locally circulating lineages represent a persistent gap in global reference libraries. Even with whole-genome sequencing (WGS), 90.2% (55/61) of persistently unresolved isolates could not be assigned to known species in the current reference databases despite the formation of clear ANI- and phylogeny-defined clusters. These findings show that both proteomic and genomic reference resources for clinical NTM remain incomplete. Expanding regionally representative databases and performing WGS for isolates that remain unresolved by MALDI-TOF MS will be necessary to improve species-level resolution for surveillance and clinical interpretation.

Taiwan

Systematic performance evaluation and application validation of an end-to-end NGS workstation.

Next-generation sequencing (NGS) library preparation is a core component of precision genomics, but it is commonly constrained by inefficiency, variability, and low throughput of manual protocols. To address these limitations, we developed and systematically evaluated a fully automated NGS workstations and further validated its performance across representative application scenarios. The automated system reduced total processing time from 8 to 10 to 4–6 h. At the same time, it maintained similar performance in pre-library metric, including DNA yield and fragment size, as well as post-capture sequencing metrics (Q30 > 90%, mapping rates > 95%, on-target rates 85–90%). The duplication rate was reduced to 5–8%, compared with 10–15% for manual methods, indicating increased library complexity. Bioinformatic evaluation of inter-species read mapping showed minimal cross-contamination, with a maximum contamination ratio of 0.0003%, indicating effective sample isolation in the automated workflow. High concordance in variant detection was observed between automated and manual workflows. Overall, this automated workstation provides a standardized and reproducible workflow that supports scalable precision genomics applications.

High-Throughput Nucleotide Sequencing

Effect of repeated mass drug administration on the transmission of yaws: a retrospective genomic epidemiology study.

BACKGROUND: Yaws, a neglected tropical disease caused by Treponema pallidum subspecies pertenue (T p pertenue), has evaded eradication, in part due to a high proportion of asymptomatic cases. Repeated mass drug administration (MDA), whereby an entire population is repeatedly treated irrespective of disease, could provide a solution. Here, we aimed to investigate the effect of MDA on the genomic epidemiology of T p pertenue. METHODS: We conducted a retrospective genomic epidemiology study on samples collected during a cluster-randomised trial of mass administration of azithromycin for yaws eradication in the Namatanai District of Papua New Guinea. Participants were in 38 wards (administrative units encompassing several villages) in three local-level government areas (LLGs). The experimental group received an initial round of MDA followed by two further rounds 6 months and 12 months after the first round. The control group received one round of MDA followed by two rounds of treatment targeting clinical cases and contacts only, on the same schedule as the MDA in the experimental group. A follow-up survey on both groups was done 18 months after the first MDA round. Swab samples were collected at each round from ulcerative and nodular skin lesions, and blood was collected by finger-prick for serological testing at 18 months. Metadata on ulcer size (cm) and duration (days) were recorded at each round, and treponemal and non-treponemal antibodies were recorded at 18 months. Samples from swabs positive for T p pertenue underwent library preparation and whole-genome sequencing. We examined the phylogenetic relationships between genomes, linking them with geospatial and patient metadata to understand the impact of MDA on T p pertenue diversity and transmission. FINDINGS: Swabs collected from 297 individuals with active yaws from April 30, 2018, to Nov 2, 2019, yielded 222 good-quality Tp pertenue genomes. We identified 20 sublineages of T p pertenue in the control group and 21 in the experimental group at the beginning of the study. At the end of the study, there were 13 sublineages in the control group and three in the experimental group, of which two persisted in both groups. Three sublineages not detected at baseline were observed in the control group after commencing MDA. The two sublineages that persisted in both groups had non-synonymous mutations in penicillin-binding proteins. One of these sublineages evolved macrolide resistance in three individuals and was associated with lowered treponemal antibody (p=0&#xb7;0036) and longer ulcer duration (p=0&#xb7;015). Despite the study taking place within a small island, sublineages were geographically clustered, with pairs of samples from the same ward (odds ratio 7&#xb7;1, 95% CI 5&#xb7;7-8&#xb7;8; p<0&#xb7;0001) or neighbouring wards (4&#xb7;3, 3&#xb7;3-5&#xb7;4; p<0&#xb7;0001) more likely to share the same sublineages compared with pairs from different LLGs. Additionally, older individuals were more likely to share sublineages than were younger individuals (1&#xb7;5, 1&#xb7;2-1&#xb7;9; p<0&#xb7;0001). INTERPRETATION: Repeated MDA was successful in reducing and maintaining the genetic diversity of T p pertenue at a low level but was associated with the development of macrolide resistance. Yaws re-emergence after MDA was attributed to multiple sublineages, of which the majority were detected in the population before MDA. Participants within the same ward were more likely to share sublineages than those that were more widely geographically separated, suggesting that re-emergence was driven by local transmission. These findings could inform future yaws elimination strategies. FUNDING: European Research Council, EU, Provincial Deputation of Barcelona, Barber&#xe0; Solid&#xe0;ria Foundation, Wellcome, and Fundaci&#xf3; "la Caixa".

Adolescent

Genome-Wide In Vivo RNAi Screening Identifies HOXD4 as a Tumor Metastasis Suppressor in Colorectal Cancer.

Metastasis remains a major therapeutic challenge in colorectal cancer, highlighting an urgent need to elucidate its underlying molecular mechanisms. In this study, an in vivo screening system integrating genome-wide short hairpin RNA library and next-generation sequencing identifies six candidate metastasis suppressors, among which Homeobox D4 (HOXD4) shows the most pronounced effects. Clinicopathological analyses reveal significant HOXD4 downregulation in tumor tissues relative to adjacent normal tissues, with reduced expression strongly correlating with aggressive tumor features. Functional assays demonstrate that HOXD4 depletion enhances migration, invasion, and tumorsphere formation in HCT116 cells, while ectopic HOXD4 overexpression reverses these malignant phenotypes in SW620 cells. Mechanistically, HOXD4 suppresses epithelial-mesenchymal transition (EMT) by directly binding to the promoter of Forkhead box Q1 (FOXQ1), a key driver of EMT and stemness, and thereby transcriptionally repressing its expression. Immunohistochemistry confirms an inverse correlation between HOXD4 and FOXQ1 expression in clinical specimens. Rescue experiments substantiate that HOXD4 exerts its metastasis-suppressing functions via FOXQ1 regulation. Collectively, these findings not only establish an efficient platform for screening tumor metastasis suppressors, but also identify HOXD4 as a master transcriptional regulator of the FOXQ1-EMT axis, providing a promising target for metastasis interception.

Humans

Gene polymorphisms associated with progression of primary open-angle glaucoma: A systematic review.

Glaucoma is the leading global cause of irreversible blindness, with primary open-angle glaucoma (POAG) its most prevalent subtype. While elevated intraocular pressure is a major risk factor, glaucoma progression is multifactorial, influenced by genetic, environmental, vascular and mechanical factors. Genetic polymorphisms have been linked to both POAG susceptibility and progression, yet most studies focus on risk factors for disease onset rather than progression. We provide an overview of the current literature on gene polymorphisms associated with POAG progression. We conducted a systematic search following PRISMA guidelines in MEDLINE, EMBASE, Web of Science, Cochrane Library, Scopus and Public Health Genomics and Precision Health Knowledge Base. Eligible studies investigated associations between genetic variants and structural or functional markers of glaucoma progression in adult-onset POAG patients. Eighteen articles were included. HLA class I haplotypes (A1-B8 and A2-B40) and MYOC.mt1+ carriers showed faster progression of optic nerve head damage. The APOE &#x3b5;4 allele was linked to faster macular thinning in normal tension glaucoma patients. BDNF rs6265 Val/Val homozygotes exhibited accelerated retinal nerve fiber layer loss, particularly in females. TGFBR3-CDC7 (rs1192415: G) and MYOC.mt1+ carriers experienced accelerated visual field deterioration. Carriers of GAS7 (rs9913911: AA), IL1B (rs1143627: CT and rs16944: CT) and OPTN (rs2234968) had a higher likelihood of requiring surgery. Variants in ABCA1, CDKN2B-AS, eNOS and Piezo1 showed inconclusive results. These findings support a role for genetic polymorphisms in POAG progression and highlight the potential of genetic screening to identify patients at increased risk for rapid disease progression.

Disease progression

Amplification-Free Nanopore Sequencing for Herpesvirus DNA Detection in Intraocular Fluids.

PURPOSE: To evaluate the feasibility of amplification-free nanopore sequencing for detecting herpesvirus DNA in intraocular fluid using multiplex polymerase chain reaction (mPCR)-characterized herpesvirus-positive and herpesvirus-negative samples. DESIGN: Retrospective, single-center, cross-sectional study. PARTICIPANTS: This study included 42 patients with uveitis whose intraocular fluid samples were examined by mPCR, including 20 mPCR-positive samples (all positive for herpesviruses) and 22 mPCR-negative samples. METHODS INTERVENTION OR TESTING: DNA extracted from intraocular fluid samples underwent ligation-based library preparation without whole-genome amplification and was sequenced on the MinION platform with Flongle flow cells for untargeted analysis. Nanopore sequencing results were compared with mPCR findings, and associations between nanopore-derived virus-specific read counts and corresponding herpesvirus DNA copy numbers measured by mPCR were assessed. MAIN OUTCOME MEASURES: Primary outcome measure was concordance between nanopore sequencing and mPCR in herpesvirus species identification. Secondary outcome measures included nanopore sequencing detection rates stratified according to mPCR-measured herpesvirus DNA copy numbers and correlations between nanopore sequencing-derived virus-specific read counts and mPCR-measured herpesvirus DNA copy numbers. RESULTS: Among 20 mPCR-positive intraocular fluid samples, nanopore sequencing identified viral DNA from the same herpesvirus species detected by mPCR in 15 (75.0%), indicating species-level concordance. None of the 22 mPCR-negative samples contained virus-specific reads. Among the 22 herpesvirus targets identified in the 20 mPCR-positive samples, herpesvirus DNA copy numbers measured by mPCR were significantly higher in nanopore-positive than in nanopore-negative targets (P = 0.015). Nanopore detection rates increased with increasing herpesvirus DNA copy numbers measured by mPCR: 3 of 6 targets (50.0%) with <105 copies/mL, 2 of 4 (50.0%) with 105-106 copies/mL, and 12 of 12 (100%) with >106 copies/mL (P = 0.021). Nanopore sequencing-derived virus-specific read counts correlated positively with herpesvirus DNA copy numbers measured by mPCR (r = 0.76, P = 0.0004). CONCLUSIONS: Amplification-free nanopore sequencing demonstrated the feasibility of detecting herpesvirus DNA in intraocular fluid samples, with detection performance dependent on herpesvirus DNA load. This simplified workflow may provide complementary information regarding viral DNA burden in minute ocular samples. FINANCIAL DISCLOSURES: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Herpesvirus

The rat serum albumin gene: analysis of cloned sequences.

The rat serum albumin gene has been isolated from a recombinant library containing the entire rat genome cloned in the lambda phage Charon 4A. Preliminary R-loop and restriction analysis has revealed that this gene is split into at least 14 fragments (exons) by 13 intervening sequences (introns), and that it occupies a minimum of 14.5 kilobases of genomic DNA.

Animals

IMAGE cDNA clones, UniGene clustering, and ACeDB: an integrated resource for expressed sequence information.

In this study we describe a new information resource that provides integrated access to information on IMAGE (integrated molecular analysis of genomes and their expression) cDNA library clones and derived expressed sequence tags (ESTs). We have developed an automated procedure that collates data from various public sources into a single ACeDB database. This database is a valuable tool for electronic cloning experiments and gene expression studies. It allows researchers to find information about cDNA libraries, plate addresses, insert sizes, and sequence data for IMAGE clones, the assignment of ESTs to UniGene clusters, and the chromosomal location of those genes in an efficient, graphically oriented manner.

Cloning, Molecular

OligoSeq: Rapid nanopore-sequencing of single-stranded oligonucleotides.

Nanopore-based DNA sequencing technology has achieved remarkable success in sequencing increasingly long DNA strands (e.g., over a million nucleotides long) for genomics research and biotechnology applications. However, the same level of progress has not been achieved for DNA oligonucleotides (usually &#x2264; 300 nucleotides long). Oligonucleotides play a crucial role in genome engineering efforts through oligo library generation and in DNA data storage, where they are used to encode computer information, such as binary (digital) data in DNA libraries. To enable these applications, accurate sequencing of oligonucleotides in a way that allows to assess for sequence variability, quality and length is essential. But sequencing solutions for oligonucleotides - particularly DNA primers for PCR, oligo DNA libraries used for mutagenesis or cDNA libraries used in gene expression analysis - remain inadequate. To address this gap, OligoSeq is presented as an innovative approach that integrates two complementary techniques: AmpliSeq (based on PCR) and RevSeq (based on reverse complementation with sequence-specific or random primers) to facilitate sequencing of single-stranded oligonucleotides using reference sequence anchor matches of more than &#x2265; 90% identity spanning from about 70% to 10% with AmpliSeq or RevSeq with random nonamers, respectively, and resolving the final reference sequence based on the most likely candidate from basecall frequencies, regardless of length and double-stranding method. OligoSeq can be integrated with nanopore sequencing technology pipelines and can be used as a reference for other sequencing platforms requiring double-stranded adapters, offering a practical and scalable alternative for standard quality control in single-stranded oligonucleotide synthesis. The use of nanopore technology, compatible with the double-stranding methods showcased, is shown to be the most cost-effective method for resolving original DNA sequences of different length and quality, and to assess its sequence variability, compared to other methods such as Illumina, PacBio or HPLC/MS.

Sequence Analysis, DNA

Fast and flexible minimizer digestion with digest.

SUMMARY: Minimizer digestion is an increasingly common component of bioinformatics tools, including tools for de Bruijn graph assembly and sequence classification. We describe a new open source tool and library to facilitate efficient digestion of genomic sequences. It can produce digests based on the related ideas of minimizers, modimizers or syncmers. Digest uses efficient data structures, scales well to many threads, and produces digests with expected spacings between digested elements. AVAILABILITY AND IMPLEMENTATION: Digest is implemented in C++17 with a Python API, and is available open-source at https://github.com/VeryAmazed/digest. The python library is available on Bioconda. Rust bindings are available as a public crate at https://crates.io/crates/digest-rs.

Software

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study

Phage Immunoprecipitation and Sequencing-a Versatile Technique for Mapping the Antibody Reactome.

Characterizing the antibody reactome for circulating antibodies provide insight into pathogen exposure, allergies, and autoimmune diseases. This is important for biomarker discovery, clinical diagnosis, and prognosis of disease progression, as well as population-level insights into the immune system. The emerging technology phage display immunoprecipitation and sequencing (PhIP-seq) is a high-throughput method for identifying antigens/epitopes of the antibody reactome. In PhIP-seq, libraries with sequences of defined lengths and overlapping segments are bioinformatically designed using naturally occurring proteins and cloned into phage genomes to be displayed on the surface. These libraries are used in immunoprecipitation experiments of circulating antibodies. This can be done with parallel samples from multiple sources, and the DNA inserts from the bound phages are barcoded and subjected to next-generation sequencing for hit determination. PhIP-seq is a powerful technique for characterizing the antibody reactome that has undergone rapid advances in recent years. In this review, we comprehensively describe the history of PhIP-seq and discuss recent advances in library design and applications.

Humans

Establishment of a CRISPR-Cas9 Library for Indica Rice and Identification of OsOPR5 (LOC_Os06g11210) as a Regulator of Root Architecture.

Functional characterization of a large number of rice genes remains a major challenge despite the availability of genome sequences and large-scale transcriptomic datasets. CRISPR-Cas9 library is a powerful approach for high-throughput targeted mutagenesis; however, its application in indica rice cultivars remains limited due to low transformation and regeneration efficiencies. In this study, we developed a CRISPR-Cas9 library targeting 12,000 rice genes and evaluated its utility for functional genomics in the indica cultivar MTU-1010. Sanger sequencing and NGS analysis of the plasmid library revealed high sgRNA coverage and more than 80% accuracy. Transformation of the developed library into the indica cultivar MTU-1010 resulted in a high target editing efficiency, with 90% of analyzed transgenic plants carrying mutations at the intended target site. Functional analysis of one homozygous mutant identified a previously uncharacterized role for OsOPR5 (LOC_Os06g11210), a member of the 12-oxophytodienoate reductase family in root architecture. The opr5 mutants exhibited significant reductions in lateral root number, seminal and crown root number, and root length, demonstrating that OsOPR5 positively regulates root system architecture in rice. Notably, endogenous jasmonic acid (JA) and JA-isoleucine levels were not significantly altered in the mutant, suggesting potential functional specialization or redundancy among rice OPR family members for JA accumulation. The root system architecture is a key determinant of water and nutrient acquisition; our results suggest that OsOPR5 may play an important role in adaptation under adverse environmental conditions. Collectively, this study establishes an efficient genome-editing platform for indica rice and identifies OsOPR5 as a novel regulator of root development.

Oryza

Enzymatic depletion of transposable elements in sequencing libraries and its application for genotyping multiplexed CRISPR-edited plants.

Whole-genome sequencing has become a common strategy to genotype individual plants of interest. Although a limited number of genomic regions usually need to be surveyed with this strategy, excess sequencing information is almost always generated at an appreciable financial cost. Repetitive sequences (e.g., transposons), which can account for more than 80% of the genome of some plants, are often not required in these genotyping projects. Therefore, strategies that enrich DNA coding for the protein-coding genes prior to sequencing can lower the cost to obtain sufficient sequence information. Here, we present the development and application of methylation-sensitive reduced representation sequencing (MsRR-Seq), which relies on the cytosine methylation-sensitive restriction enzyme MspJI to deplete constitutive heterochromatic DNA before library construction. By applying MsRR-Seq to citrus and maize, we show that protein-coding genes can be enriched in sequencing datasets. We then describe the application of MsRR-Seq to facilitate the identification of complex mutants from populations of citrus plants resulting from multiplex CRISPR/Cas9 editing of four genes. Overall, this work demonstrates an easy and low-cost method to enrich non-repetitive DNA in high-throughput sequencing libraries, an approach that is especially useful for large plant genomes with an excessively high proportion of methylated repetitive sequences.

DNA Transposable Elements

Genome-scale CRISPR screening uncovers SRSF6 as a target to sensitize hepatocellular carcinoma to radiotherapy.

BACKGROUND & AIMS: Radiotherapy confers clinical benefits to patients with hepatocellular carcinoma (HCC) across all stages, yet its clinical efficacy is limited by radioresistance. This study aimed to identify key regulators of HCC radiosensitivity through genome-wide functional screening. METHODS: A genome-wide CRISPR-Cas9 screen in Huh7 cells identified radiosensitivity regulators, with SRSF6 validated by siRNA knockdown and &#x3b3;-H2AX assessment. Stable shRNA-mediated SRSF6 knockdown was established in Huh7 and HepG2 cells, followed by clonogenic, EdU incorporation, apoptosis, micronucleus, and comet assays. Mechanistically, RNA-seq, Western blotting, mRNA stability assays, RIP-qPCR, and RAD51 overexpression rescue assays were performed. The therapeutic potential of the SRSF6 inhibitor indacaterol was evaluated using MTS assays, HCC xenograft mouse models (BALB/c-nu/nu, n = 28), and HCC patient-derived organoids (PDOs) (n = 3). In addition, SRSF6 expression and its correlation with patient survival were analyzed using data from The Cancer Genome Atlas and a tissue microarray (n = 14 HCC and 14 paired adjacent non-tumorous liver samples). RESULTS: We identified the RNA-binding protein SRSF6 as a driver of HCC radioresistance. SRSF6 depletion enhanced the radiosensitivity of HCC cells (p <0.05-0.0001) by post-transcriptionally destabilizing the mRNAs of critical DNA repair genes (p <0.05-0.0001), thereby impairing radiation-induced DNA damage repair. The radiosensitizing effect of SRSF6 depletion was partially abrogated by ectopic overexpression of the core DNA repair protein RAD51 (p <0.05-0.001). Indacaterol exhibited cytotoxic effects on HCC cells (p <0.05-0.0001) and enhanced the antitumor efficacy of radiation in vivo (p <0.05-0.0001), as further validated across multiple HCC patient-derived organoids (p <0.05-0.0001). CONCLUSIONS: SRSF6 is a key regulator of HCC radioresistance through its post-transcriptional control of DNA repair capacity, and represents a novel therapeutic target to sensitize HCC to radiotherapy. IMPACT AND IMPLICATIONS: In this study, we performed a genome-wide CRISPR-Cas9 knockout library screen to dissect the molecular determinants governing HCC radiosensitivity, and identified RNA-binding protein SRSF6 as a driver of HCC radioresistance. We demonstrate that SRSF6 depletion disrupts the post-transcriptional stability of key DNA repair gene mRNAs and enhances HCC radiosensitivity. These findings are important for radiation oncologists and translational researchers, as they identify SRSF6-dependent RNA regulation as a critical determinant of radiotherapy response in HCC. Practically, we show that the clinically approved bronchodilator indacaterol suppresses SRSF6 function and enhances the antitumor efficacy of radiotherapy, offering a readily repurposable pharmacological strategy to overcome radioresistance. These implications are based on preclinical evidence across multiple models; however, future clinical trials are needed to validate the safety and efficacy of indacaterol-based radiosensitization in patients with HCC.

DNA repair