Search PubMedSearch

SEARCH · Search PubMed

Results for “indel”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Discovering human transcription factor physical interactions with genetic variants, novel DNA motifs, and repetitive elements using enhanced yeast one-hybrid assays.

Identifying transcription factor (TF) binding to noncoding variants, uncharacterized DNA motifs, and repetitive genomic elements has been technically and computationally challenging. Current experimental methods, such as chromatin immunoprecipitation, generally test one TF at a time, and computational motif algorithms often lead to false-positive and -negative predictions. To address these limitations, we developed an experimental approach based on enhanced yeast one-hybrid assays. The first variation of this approach interrogates the binding of >1000 human TFs to repetitive DNA elements, while the second evaluates TF binding to single nucleotide variants, short insertions and deletions (indels), and novel DNA motifs. Using this approach, we detected the binding of 75 TFs, including several nuclear hormone receptors and ETS factors, to the highly repetitive Alu elements. Further, we identified cancer-associated changes in TF binding, including gain of interactions involving ETS TFs and loss of interactions involving KLF TFs to different mutations in the TERT promoter, and gain of a MYB interaction with an 18-bp indel in the TAL1 superenhancer. Additionally, we identified TFs that bind to three uncharacterized DNA motifs identified in DNase footprinting assays. We anticipate that these enhanced yeast one-hybrid approaches will expand our capabilities to study genetic variation and undercharacterized genomic regions.

Algorithms

Whole-Exome and Whole-Genome Sequencing of Candidate Pharmacogenomic and Schizophrenia-Related Genes in Sudanese Families with Schizophrenia.

BACKGROUND: Schizophrenia is considered a neuro-developmental disorder leading to disastrous lifelong disability of the patients and their families. There is a lack of data regarding pharmacogenomics of schizophrenia in Sudan. This study aimed to identify different genes affecting the treatment outcomes in Sudanese patients with schizophrenia. METHODS: A case-control study was conducted on seven families having more than one member diagnosed with schizophrenia. This was a small exploratory family-based sequencing study involving 18 affected individuals and 8 controls from seven families. Ethical clearance and informed consent were obtained. Demographic data were collected using a standardized data collection sheet. DNA was extracted from blood samples collected from patients and control groups. Then, whole-exome and genome sequencing were performed. Sixty-six genes associated with schizophrenia, treatment, and treatment resistance were selected from the variant calling file. Variants showing single-nucleotide polymorphisms (SNPs) were identified. These variants were then classified based on their impact on the protein-coding sequence into high- and moderate-impact. Moreover, indel mutations were also identified. RESULTS: Twelve variants of seven genes (COMT, FMO1, LPL, CYP2E1, ABCC1, GRM3, CYP2C9) were identified as genes with impact and potential association with schizophrenia (p-value=0.006632). Forty-three genes had a moderate impact, and they showed a potential association with schizophrenia (p-value=0.0004436). Two variants were indel mutations (CYP2D6, DTNBP1) and showed association with schizophrenia (p-value=0.004741). The p-values were generated from different databases. CONCLUSION: This exploratory family-based sequencing study identified several potentially relevant pharmacogenomic and schizophrenia-associated variants in Sudanese families, warranting validation in larger and ethnically diverse cohorts.

antipsychotics

Targeted multiplex gene knockouts in Lemna minor using CRISPR/Cas9.

Lemna minor (commonly known as duckweed) is a fast-growing aquatic plant recognized as a promising green bioreactor for recombinant protein production. Its rapid proliferation, high protein yield, environmental adaptability, and edibility make it highly attractive for biotechnological applications. It is essential to develop and expand genetic tools tailored to this species to maximize these advantages and further unlock its biotechnological potential. A key strategy for achieving this goal is the implementation of advanced genome editing technologies, such as the CRISPR/Cas9 system. Although multiplex CRISPR/Cas9 gene editing has previously been successfully applied in Lemna aequinoctialis, the capability of the endogenous plant tRNA processing system for multiplex editing in L. minor using the polycistronic tRNA-sgRNA (PTG)/Cas9 system has not yet been explored. In this study, a PTG construct was engineered to include four sgRNAs designed to simultaneously target two plant-specific glycosyltransferase genes: α-1,3-fucosyltransferase (FucT) and β-1,2-xylosyltransferase (XylT). As anticipated, the PTG-Cas9 system successfully induced frameshift mutations, characterized by insertions and deletions (indels), in regenerated L. minor plants derived from transformed calli. Validation via PCR and RT-PCR analysis, followed by sequencing of the target loci, confirmed the presence of indels at the target sites. Furthermore, western blot analyses utilizing antibodies specific to XylT and FucT in two homozygous lines (lines 44 and 217) revealed truncated XylT proteins in both lines. Moreover, an in-frame FucT protein was detected in line 217, whereas FucT expression was absent in line 44. This study marked the first successful demonstration of PTG-Cas9 system for multiplex genome editing in L. minor, paving the way for advanced genetic engineering in this species.

CRISPR-Cas Systems

STX1B variant-specific synaptic dysfunction is associated with network hyperexcitability in human iPSC-derived neurons.

BACKGROUND: Variants in STX1B/syntaxin-1B are linked to a spectrum of fever-associated epilepsy syndromes. While studies in murine models have provided mechanistic insights, their relevance to human disease in a heterozygous context may be limited. METHODS: We investigated two pathogenic STX1B variants using isolated single neurons and neuronal network cultures derived from patient-specific induced pluripotent stem cells. These carried either a de novo p.G226R variant, associated with severe developmental epilepsy, or an InDel variant (p.K45delinsRCMIE/p.L46M) linked to a transient familial seizure syndrome. Synaptic function and network excitability were assessed using patch-clamp and multi-electrode array recordings, alongside morphological and transcriptomic profiling. FINDINGS: G226R exhibited both gain- and loss-of-function characteristics, with increased miniature excitatory postsynaptic current frequency in networks but not in autapses, and synaptic failure during sustained high-frequency stimulation. For the InDel variant, the predicted loss-of-function phenotype based on reduced syntaxin-1B levels was not detectable at the single-cell level, likely masked by compensatory synaptic upregulation. At the network level, however, both variants were associated with neuronal hyperexcitability, characterised by more frequent and prolonged bursting activity, with a much stronger phenotype in G226R-containing networks. Transcriptomic profiling revealed a differential dysregulation of synaptic and other neuronal genes. INTERPRETATION: The divergence between morphological, electrophysiological and transcriptomic findings suggests that compensatory mechanisms may contribute to network hyperexcitability. Initially engaged to maintain homoeostasis, they may ultimately contribute to a pathological network state. The graded severity of network alterations across STX1B variants correlates with the clinical phenotypes. FUNDING: BMBF (Treat ION-01GM2210A, SNAREopathies-01EW1809A), 2023 FEBS Summer Fellowship, Fortüne programme (2610-0-0), EKFS college precise.net, Open Access Publishing Fund of University of Tübingen.

Humans

Genome-wide variation analysis of two Salvia hispanica L. genotypes and implication for associations with metabolic and adaptive traits.

BACKGROUND: Advances in next-generation sequencing have accelerated genome-wide exploration of genetic diversity in underutilized oilseed crops. Salvia hispanica L. (chia), a high-nutrient pseudocereal rich in omega-3 fatty acids, is increasingly valued for its health benefits and commercial potential, yet it remains poorly characterized at the genomic level. Understanding the scale and nature of genomic variation is essential for improving complex traits such as oil yield, stress tolerance, and seed quality. METHODS: Two contrasting chia genotypes, Black-chia (CACH-B) and White- chia (CACH-W), were resequenced using the Bio-Resequencing Toolkit (BRT) pipeline. High-coverage sequencing, with a mapping rate exceeding 99% and an average depth of approximately 28×, facilitated the detection and annotation of single-nucleotide polymorphisms (SNPs), insertions and deletions (InDels), copy-number variations (CNVs), and structural variants (SVs). The functional classification of variant impacts enabled the identification of genes potentially linked to metabolic and adaptive traits. RESULTS: A total of 1.97 million SNPs, 401,493 InDels, 836 CNVs, and 15,288 SVs were identified across the chia genome. Notably, approximately 53% of exonic SNPs were non-synonymous (dN/dS ≈ 1.28), predominantly affecting lipid metabolism, transcriptional regulation, and stress response pathways, potentially altering key agronomic traits. In addition, CNV hotspots were concentrated in chromosomes 3 and 6, overlapping MYB, WRKY, and bZIP transcription factor loci, may potentially be involved in stress tolerance and yield. Furthermore, structural rearrangements, including inversions and duplications within the FAD2, FAD3, and CYP450 gene clusters, were potentially associated with seed pigmentation and omega-3 biosynthesis, pointing to their potential breeding relevance. Observed heterozygosity (Hₒ ≈ 0.71) and nucleotide diversity (π ≈ 7 × 10-3) indicated moderate to high allelic richness. In addition, the low FST value (0.038) indicates substantial genomic similarity between the two genotypes. CONCLUSION: This study presents the first comprehensive map integrating SNPs, CNVs, and SVs in S. hispanica L. The results reveal a structurally dynamic genome characterized by substantial sequence and structural variation, providing valuable insights into genomic diversity and potential adaptive mechanisms in chia. The coexistence of high SNP diversity and abundant structural variation underpins chia's nutritional specialization and environmental resilience. These results deliver a foundational genomic resource for marker-assisted breeding, genome-wide association studies, and the development of climate-resilient chia cultivars.

Copy-number variation, structural variation

Somatic likelihood tiering: an interpretable post-calling triage protocol for tumor-only whole-exome variant review.

Tumor-only whole-exome sequencing (WES) is used when matched normal tissue is unavailable, but one sample can produce thousands of variants. Somatic likelihood tiering (SLT) is an interpretable post-calling protocol that ranks Mutect2 calls into four review-priority tiers using population-frequency, germline-quality, cancer-knowledge, PureCN posterior, and clonal-hematopoiesis evidence. Layer 2 distinguishes common, rare-callable, and unevaluable gnomAD states; missing or unmatchable gnomAD evidence is not positive rarity evidence. On the SEQC2 HCC1395 benchmark, the callability-aware SLT-A row contained 101 calls, 78 truth variants, 77.2% PPV (95% Wilson confidence interval 68.1%-84.3%), and a Number Needed to Review (NNR) of 1.29 (1.19-1.47). The conservative SLT-C catchment retained 352 of 455 truth variants (77.4%, 73.3%-81.0%) and all tiers together retained 430 of 455 truth variants. SNV performance is the primary calibration frame: SLT-C retained 341 of 439 SNV truth variants, whereas indel results were exploratory because only 16 truth indels were available. Clinical cohorts are reported as recall and concordance versus partially dependent matched-normal Mutect2 references, not independent clinical sensitivity. Patient-level bootstrap intervals were principal: HdM-BLCA-1 SLT-A recall was 18.2% (14.0%-23.5%), and LUAD-TW SLT-A recall was 49.1% (26.6%-63.3%) among 32 evaluable patients. The HdM-BLCA-1 median SLT-A queue remained 1277 variants per patient, so SLT reduces first-pass candidate counts but does not measure review time or eliminate FFPE candidate-count burden. SLT provides an auditable tumor-only WES review queue, not a substitute for matched-normal sequencing, independent orthogonal validation, or definitive somatic classification.

Humans

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals

Mutation accumulation in a hybrid parthenogenetic vertebrate.

Asexual lineages are thought to experience elevated extinction rates compared with sexual species, yet direct evidence for the underlying genetic causes remains scarce. Muller's ratchet predicts that the absence of recombination in asexual organisms facilitates the accumulation of deleterious mutations, thereby reducing long-term fitness. Here, we test this hypothesis in the hybrid-origin, parthenogenetic whiptail lizard Aspidoscelis tesselatus by integrating short-read RNAseq and long-read IsoSeq data from both the asexual lineage and its parental sexual species. We reconstructed phased transcripts for A. tesselatus to quantify mutation accumulation relative to the parental sexual species. Comparative analyses revealed elevated ω ratios in both parental genomic complements (subgenomes) of the parthenogenetic lineage, consistent with accelerated accumulation of nonsynonymous mutations. Structural variant analyses identified multiple indels in expressed transcripts predicted to disrupt protein domains. Functional annotation indicated that genes affected by both single-nucleotide variants and indels were enriched for roles in chromatin organization, apoptosis regulation, and transcriptional control. While both parental subgenomes showed similar evolutionary patterns, the maternal complement exhibited more structural and missense mutations than the paternal complement. Together, these results provide evidence that mutations accumulate in asexual A. tesselatus in genes involved in core cellular functions, supporting theoretical predictions that Muller's ratchet contributes to mutation accumulation in asexual lineages.

Animals

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

Ribosomal DNA and RNA (rDNA and rRNA) sequences are usually discarded from sequencing analyses. But with hundreds of copies of rDNA genes it is unknown whether they possess sequence variations that form different types of ribosomes that affect human physiology and disease. Here, we developed an algorithm for variant-calling between paralog genes (termed RGA) and compared rDNA variations found in short- and long-read sequencing data from the 1,000 Genomes Project (1KGP) and Genome In A Bottle (GIAB). We additionally developed a novel protocol for long-read sequencing full-length rRNA (RIBO-RT) from actively translating ribosomes. Our analyses identified hundreds of rDNA variants, most of which, surprisingly, are short insertion-deletions (indels) and dozens of highly abundant rRNA variants that are incorporated into translationally active ribosomes. To visualize variant ribosomes at the single cell level, we developed an in-situ rRNA sequencing method (SWITCH-seq) which revealed that variants are co-expressed within individual cells. Strikingly, by analyzing rDNA, we found that variants assemble into distinct ribosome subtypes. We discovered that these subtypes acquire different rRNA structures by successfully employing dimethyl sulfate (DMS) probing of full length rRNA. With this atlas we investigated rRNA variation changes across human tissues and cancer types. This revealed tissue-specific rRNA subtype expression in endoderm/ectoderm-derived tissues. In cancer, low abundant rRNA variants can become highly expressed, which suggests the presence of cancer-specific ribosomes. Together, this study identifies and comprehensively characterizes the diversity of ribosomes at the level of rRNA variants which is dominated by indel variants, their chromosomal location and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Journal Article

DeepSomatic: Accurate somatic small variant discovery for multiple sequencing technologies.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies now offer potential advantages in terms of repeat mapping and variant phasing. We present DeepSomatic, a deep learning method for detecting somatic SNVs and insertions and deletions (indels) from both short-read and long-read data, with modes for whole-genome and exome sequencing, and able to run on tumor-normal, tumor-only, and with FFPE-prepared samples. To help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available a dataset of five matched tumor-normal cell line pairs sequenced with Illumina, PacBio HiFi, and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples and technologies (short-read and long-read), DeepSomatic consistently outperforms existing callers, particularly for indels.

Journal Article

MCALIGN: stochastic alignment of noncoding DNA sequences based on an evolutionary model of sequence evolution.

A method is described for performing global alignment of noncoding DNA sequences based on an evolutionary model parameterized by the frequency distribution of lengths of insertion/deletion events (indels) and their rate relative to nucleotide substitutions. A stochastic hill-climbing algorithm is used to search for the most probable alignment between a pair of sequences or three sequences of known phylogenetic relationship. The performance of the procedure, parameterized according to the empirical distribution of indel lengths in noncoding DNA of Drosophila species, is investigated by simulation. We show that there is excellent agreement between true and estimated alignments over a wide range of sequence divergences, and that the method outperforms other available alignment methods.

Algorithms

Genome-wide association studies of plant traits and functional analysis of leaf development-related genes in citrus.

Labor-saving and high-light-efficiency tree architecture is a key breeding objective for woody fruit trees like citrus. However, population genetics information on these traits remains limited. In this study, tree architecture, thorn, and leaf traits were evaluated in 353 F2 progeny derived from a cross between Clementine mandarin and precocious trifoliate orange-an early-flowering variety. A random subset of 300 offspring was sequenced for a genome-wide association study (GWAS), which detected 10 216 significantly associated SNPs and defined several major quantitative trait loci (QTLs) for the target traits. Subsequent bulked segregant analysis (BSA) and GWAS on individuals with extreme compound leaf phenotypes mapped the causal gene(s) to a 0.8 Mb region (22.15-22.95 Mb) on chromosome 4. Genetic analysis across multiple hybrid combinations confirmed that the compound leaf trait in trifoliate orange is dominantly inherited and follows Mendelian segregation. Transcriptome profiling of parental leaves at different developmental stages identified a KNOX gene, CiKNAT6, as a candidate. Further validation using CAPS markers and Hi-Tom sequencing demonstrated tight linkage between an InDel polymorphism in CiKNAT6 and leaf shape across diverse citrus species and the F2 population, with co-segregation observed for the compound leaf trait. Due to alternative splicing producing seven splice variants, the CiKNAT6 DNA sequence was selected for genetic transformation experiments. Functional analysis revealed that the Clementine mandarin allele of CiKNAT6 is non-functional owing to an InDel, whereas ectopic expression of the trifoliate orange allele in tobacco and lemon induced leaf curling and reduced leaf size. CRISPR-Cas9 knockout of CiKNAT6 in trifoliate orange resulted in increased leaf area. These findings provide valuable genetic resources and insights for future studies on tree architecture and leaf morphology.

Plant Leaves

Mutagenic Impact and Evolutionary Influence of Chemoradiotherapy in Hematologic Malignancies.

UNLABELLED: Ionizing radiotherapy (RT) is a widely used treatment strategy for malignancies. In solid tumors, RT-induced double-strand breaks lead to the accumulation of insertion-deletions (indels; ID), and their repair by nonhomologous end joining has been linked to the ID8 mutational signature in surviving cells. However, the extent of RT-induced mutagenesis in hematologic malignancies and its impact on their mutational profiles and interplay with commonly used chemotherapies has not yet been explored. In this study, we interrogated 580 whole-genome sequence (WGS) samples from patients with large B-cell lymphoma, multiple myeloma, and myeloid neoplasms and identified ID8 only in relapsed disease. Yet ID8 was detected after exposure to both RT and mutagenic chemotherapy (i.e., platinum and melphalan). Using WGS of single-cell colonies derived from treated lymphoma cells, we revealed a dose-response relationship between RT and platinum and ID8. Finally, using ID8 as a genomic barcode, we demonstrate that a single RT-surviving cell may seed distant relapse. SIGNIFICANCE: RT and the ID8 indel signature are related, but their genomic impact on hematologic malignancies is unclear. Leveraging WGS, we linked ID8 to both RT and mutagenic chemotherapy and validated that platinum can induce ID8. We used ID8 as a genomic barcode to reveal that RT-resistant cells may seed systemic relapse.

Humans

Genetic diversity, phylogenetic relationships, and marker development between Hydrangea serrata and H. macrophylla based on plastome and 45S nrDNA.

Ornamental hydrangeas (genus Hydrangea) are cultivated worldwide for their diverse flower colors and attractive morphology. Here, we assembled the complete plastid genome (plastome) and 45S nuclear ribosomal DNA (45S nrDNA) sequences of 22 individuals representing H. serrata, H. macrophylla, and related species (H. arborescens, H. paniculata, H. petiolaris, and H. hydrangeoides). The plastomes contained up to 2,344 single-nucleotide polymorphisms (SNPs) and 367 insertions/deletions (InDels) within the genus, whereas the assembled 45S nrDNA sequences showed 119 SNPs and 10 InDels. Phylogenetic analyses based on plastome and 45S nrDNA sequences clearly separated H. serrata and H. macrophylla from the other Hydrangea species. In the plastome-based tree, H. petiolaris was placed in the same clade as H. arborescens, whereas in the 45S nrDNA-based tree it showed a close relationship to H. hydrangeoides. The H. serrata and H. macrophylla samples were not always separated according to their species boundaries, as observed in samples Hse8-Hse12. Notably, one H. serrata sample (Hse8), collected from a wild mountainous region of Japan, exhibited a closer genetic relationship to H. macrophylla samples, indicating that cultivated hydrangeas may have originated from a specific wild lineage of H. serrata adapted to mountainous habitats. Using plastome-derived molecular markers, 66 Hydrangea samples were further classified into five groups, with Group II comprising both cultivated H. macrophylla and a subset of wild H. serrata samples, suggesting a close genetic affinity between this group and the ancestral gene pool of cultivated H. macrophylla. Based on these genomic resources, eight plastome-derived molecular markers were developed to differentiate cultivated hydrangeas from wild genotypes and to assess genetic diversity within H. serrata and H. macrophylla, providing practical tools for germplasm identification, breeding, and genetic resource management of Hydrangea species.

hydrangea

A quick guide to evaluating prime editing efficiency in mammalian cells.

According to the Clinvar database, modeling the diseases associated with pathogenic mutations requires the installation of base substitutions, small insertions or deletions. Prime editor (PE) was recently developed to precisely install any base substitutions and/or small insertions/deletions (indels) in mammalian cells and animals without requiring DSBs or donor DNA templates. PE also offers greater editing and targeting flexibility compared to other precision CRISPR editing methods because the versatile editing information is encoded in the reverse-transcription template of its prime editing guide RNA. However, optimal PE system selection and experimental design can be complex, and there are various factors that can affect PE efficiency. This chapter serves as a rapid entry-level guideline for the application of PE, providing an experimental framework for using PE at a specific genomic locus. RUNX1 was selected as a representative target site to illustrate the detailed methodology for constructing PE plasmids and the process of transfecting these plasmids into 293FT cells. We further examined the efficiency of PE-mediated genome editing in mammalian cells by using next-generation sequencing.

Gene Editing

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

ONCOLINER: A new solution for monitoring, improving, and harmonizing somatic variant calling across genomic oncology centers.

The characterization of somatic genomic variation associated with the biology of tumors is fundamental for cancer research and personalized medicine, as it guides the reliability and impact of cancer studies and genomic-based decisions in clinical oncology. However, the quality and scope of tumor genome analysis across cancer research centers and hospitals are currently highly heterogeneous, limiting the consistency of tumor diagnoses across hospitals and the possibilities of data sharing and data integration across studies. With the aim of providing users with actionable and personalized recommendations for the overall enhancement and harmonization of somatic variant identification across research and clinical environments, we have developed ONCOLINER. Using specifically designed mosaic and tumorized genomes for the analysis of recall and precision across somatic SNVs, insertions or deletions (indels), and structural variants (SVs), we demonstrate that ONCOLINER is capable of improving and harmonizing genome analysis across three state-of-the-art variant discovery pipelines in genomic oncology.

Humans

Efficient and precise programmable DNA knock-in without double-strand breaks.

Programmable gene knock-in holds substantial promise for treating genetic diseases and advancing cell therapies. However, achieving precise and efficient kilobase-scale DNA fragment integration remains challenging1,2. Here we report CRISPR kilobase-scale nickase-targeting (KNIT) editing for efficient, precise and programmable kilobase-scale DNA insertion without double-strand DNA cleavage, which is enabled through the coupling of a Cas9 nickase with a DNA donor recruiting system. KNIT editing facilitates programmable integration of DNA fragments from 0.7 kb to more than 10 kb and is effective across genomic loci and cell types. It achieves up to 89% efficiency and markedly reduces unintended insertion-deletion mutation (indels) rates, translocations and off-target editing. The system supports repeated insertion editing and multiloci gene knock-in with minimal translocations. Its enhanced version, KNIT editor 2, further improves efficiency via a single transfection. Moreover, in mutant cells with a pathological mutation, KNIT editing restores normal gene expression by inserting a therapeutic gene into a safe harbour locus or its native locus. Notably, KNIT editing enables non-viral and programmable chimeric antigen receptor T cell (CAR-T cell) engineering without double-strand breaks and with clinically relevant efficiencies. Moreover, the engineered CAR-T cells exhibit effective antitumour activity in vitro and in mouse models. Therefore, by achieving programmable and site-specific kilobase-scale DNA insertions without double-strand breaks while reducing unintended outcomes, KNIT editing provides a versatile platform for advancing personalized medicine.

Animals