Search PubMedSearch

SEARCH · Search PubMed

Results for “indel”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Duplex-Indel: a Snakemake pipeline for somatic Indel calling in Tn5 transposase-based duplex sequencing data.

SUMMARY: Duplex-Indel is a novel Snakemake workflow for detecting somatic small insertions and deletions (Indels) from Tn5 transposase-based duplex sequencing data. Duplex-Indel enhances the accuracy of mutation calling at the single-molecule level by requiring consensus support from both DNA strands for each somatic Indel, minimizing confounding from technical artifacts. Duplex-Indel extends somatic mutation calling in Tn5 transposase-based duplex sequencing data to include Indels. We have demonstrated the accuracy and robustness of Duplex-Indel using cancer cell lines. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are available under the MIT license on GitHub at https://github.com/ealee-lab/duplex-indel and archived on Zenodo at https://doi.org/10.5281/zenodo.19228799.

Transposases

NanoFilter: enhancing phasing performance by utilizing highly consistent INDELs and SNVs in nanopore sequencing.

MOTIVATION: Nanopore sequencing data offer longer reads compared to other technologies, which is beneficial for phasing and genome assembly. INDELs provide valuable haplotype information and have significant potential to improve phasing performance. However, accurately identifying INDELs with variant callers is challenging, and incorporating INDELs into phasing remains a complex task. To address these issues, we developed NanoFilter, a novel filtering strategy designed to filter out INDELs that contain wrong phasing information based on their consistency. RESULTS: Our assessment using Nanopore R10 simplex data shows that filtering out low-consistency INDELs increases their precision from 88.3% to 98.8%, nearly matching the precision of SNVs. In the phasing results of Margin, incorporating these filtered INDELs leads to a 12.77% increase in N50 length and fewer switch errors. Furthermore, we found that SNVs filtered by NanoFilter will enhance assembly performance. When NanoFilter is integrated into the HapDup assembly pipeline, NanoFilter reduces the Hamming error rate and increases N50 length by 7.8%. AVAILABILITY AND IMPLEMENTATION: NanoFilter is available at https://github.com/Chenshanming-repo/NanoFilter (DOI: 10.5281/zenodo.16777826) and HapDup-NanoFilter is available at https://github.com/Chenshanming-repo/HapDup-NanoFilter (DOI: 10.5281/zenodo.16777890).

Nanopore Sequencing

Development of Genome-Derived InDel Markers and Genetic Diversity Analysis of Caragana acanthophylla in Xinjiang, China.

Caragana acanthophylla Kom. is an ecologically important drought-tolerant shrub in Xinjiang, China, but species-specific molecular markers for germplasm characterization remain limited. We sampled 93 individuals from 11 localities representing the currently known distribution of C. acanthophylla in Xinjiang. Three individuals per locality (33 in total) were whole-genome resequenced, yielding 2,873,410 high-quality SNPs and 5,679,915 InDels. Genome-wide SNP-based PCA and genetic relationship analysis provided an independent high-resolution assessment of the 33 resequenced individuals. From 34 candidate primer pairs, eight polymorphic InDel markers with stable amplification and clear genotyping profiles were retained and applied to all 93 individuals. The SNP dataset revealed clear regional differentiation and finer locality-associated relationships. Analysis of the same 33 individuals with the eight InDel loci recovered part of this broad pattern, particularly the differentiation of the western YL materials, but showed lower fine-scale resolution. Across all 93 individuals, the InDel panel revealed moderate to low marker-level genetic diversity and detectable regional differentiation. AMOVA attributed 67.00% of the variation to differences among the 11 original sampling localities, while the five exploratory analytical groups showed a similar among-group component (68.37%). The Mantel correlation detected across all 93 individuals (r = 0.801, p < 0.001) disappeared after YL was excluded (r = -0.032, p = 0.724), indicating that the overall spatial signal was largely driven by the geographic separation of YL. These results support the eight-marker panel as a practical, low-cost tool for preliminary germplasm characterization and broader sample screening, while genome-wide SNP data provide substantially greater resolution for population-level inference.

Caragana acanthophylla

Pilot study of allele-specific multi-InDel markers for the detection of extremely unbalanced DNA mixtures.

Mixtures are common in forensic casework, and they represent one of the most challenging types of biological evidence. Traditional short tandem repeat analyses are often associated with limitations when dealing with extremely unbalanced mixtures because alleles from minor contributors can easily be masked by those of major contributors. Consequently, researchers have developed new technologies and methods for improving the analysis of mixtures, spanning upstream DNA extraction and downstream software analysis. Among these, strategies combining allele-specific amplification with compound markers have drawn particular interest because of their ability to selectively detect minor contributors in complex mixtures. In this study, we screened multi-InDels across the entire genome, designed allele-specific primers compatible with the capillary electrophoresis platform, and further explored their potential in unbalanced DNA mixtures and cell-free fetal DNA (cffDNA). Ultimately, a set comprising 10 multi-InDels was developed, and this included two groups of primers that separately amplified the long alleles (L primer set) and short alleles (S primer set). The results demonstrated that each primer pair could detect the minor component at a 1:1000 mixture ratio, whereas the L and S primer sets successfully detected the minor contributors at mixture ratios of 1:200 and 1:500, respectively. Furthermore, in the cffDNA analysis, 60 of 78 informative markers were successfully detected, with the complete detection of all informative markers achieved in 18 mother-child reference pairs. Overall, allele-specific amplification-based multi-InDel markers enabled the sensitive detection of minor contributors, providing a potential strategy for the analysis of unbalanced two-person mixtures.

Allelic-specific amplification

Genome-Wide Identification of SSR and InDel Markers and Experimental Validation of SSR Markers for Distinguishing Cold-Tolerant and Cold-Sensitive Lily Cultivars.

In this study, whole-genome resequencing was performed on the cold-tolerant variety ND-6 and the cold-sensitive variety 'Sorbonne'. After evaluation, the Lilium davidii var. unicolor reference genome was selected to analyze SSR distribution characteristics. Whole-genome InDel identification and comparative analysis were conducted for the two varieties, yielding 34,812,909 and 24,497,857 InDels, respectively. Short InDels were predominant, with deletions slightly outnumbering insertions, mostly located in intergenic regions. Twenty pairs of SSR primers were screened and synthesized. Among them, 10 pairs amplified clearly, with a polymorphism rate of 82.6%, effectively distinguishing the two cultivars examined in this study. This study provides systematic data and a reliable marker resource for the analysis of lily genomic variation, laying a foundation for the identification of cold-tolerant germplasm; validation across additional cultivars and individuals will be required to extend their utility to broader germplasm.

cold resistant lilies

Incorporating indel channels into average-case analysis of seed-chain-extend.

MOTIVATION: Given a sequence s1 of n letters drawn independently and identically (i.i.d.) from an alphabet of size &#x3c3; and a mutated substring s2 of length m<n, we want to recover the mutation history that generated s2 from s1. Many modern sequence aligners for this task use seed-chain-extend with k-mer seeds. Previously, Shaw and Yu showed linear-gap cost chaining can produce a chain with 1-O(1m) recoverability, the proportion of the mutation history that is recovered, in O(mn2.43&#x3b8;&#x2009;log&#x2009;n) expected time for seed-chain-extend (assuming pre-seeded reference), where &#x3b8;<0.206 is the mutation rate under a substitution-only channel and s1 is uniformly random. A gap remains between theory and practice, as real genomes include insertions and deletions (indels). RESULTS: We introduce mathematical machinery to deal with the two new obstacles introduced by indel channels: the dependence of neighbouring anchors and the presence of anchors that are only partially correct. We prove that expected recoverability of an optimal chain is &#x2265;1-O(1m) and expected runtime is O(mn3.15&#xb7;&#x3b8;T&#x2009;log&#x2009;n), given the total mutation rate &#x3b8;T=&#x3b8;i+&#x3b8;d+&#x3b8;s (sum of substitution, insertion, and deletion rates) is &#x3b8;T&#x2264;0.159. We thus narrow (but not close) the gap between theory and practice. AVAILABILITY AND IMPLEMENTATION: https://github.com/Lazarus42/seed_chainer_indels.

INDEL Mutation

Indel mutation in transcription factor PabHLH2 regulates amygdalin accumulation and kernel bitterness in apricot.

Amygdalin, the phytochemical responsible for the characteristic bitterness of apricot (Prunus armeniaca L.) kernels, also exhibits significant bioactive properties and therapeutic potential. Genetic regulation of amygdalin content is therefore a key objective in apricot breeding programs aimed at quality improvement. In this study, we conducted quantitative trait loci (QTL) mapping to uncover the genetic basis of sweet-bitter differentiation in apricot kernels. We identified a 15-bp insertion/deletion (indel) polymorphism strongly related to kernel bitterness, with marker validation achieving 100% concordance across 601 apricot germplasm accessions. Notably, this polymorphic site is located within the helix-loop-helix (HLH) domain of the basic HLH (bHLH) transcription factor PabHLH2. Protein interaction analyses revealed that the 15-bp deletion variant impaired dimerization capacity, reducing transcriptional activation of downstream targets. Using yeast one-hybrid screening and dual-luciferase reporter assays, we identified PaCYP71AN24 and PaCYP79D16 as direct transcriptional targets of PabHLH2. Functional characterization further indicated that the PabHLH2a variant (harboring the 15-bp insertion) significantly enhanced the promoter activity of these cytochrome P450 genes compared with the deletion variant. Transient overexpression and silencing experiments in apricot kernels further confirmed that the 15-bp insertion positively regulates both PaCYP71AN24/PaCYP79D16 expression and prunasin accumulation, the immediate biosynthetic precursor of amygdalin. Overall, these findings provide mechanistic insights into the allelic variation underlying kernel bitterness and delineate the molecular cascade of amygdalin biosynthesis. The identified molecular markers and functional characterization establish a basis for marker-assisted breeding of low-amygdalin apricot cultivars, supporting the dual-purpose utilization of kernels in food and pharmaceutical industries.

Amygdalin

Algorithms to reconstruct past indels: The deletion-only parsimony problem.

Ancestral sequence reconstruction is an important task in bioinformatics, with applications ranging from protein engineering to the study of genome evolution. When sequences can only undergo substitutions, optimal reconstructions can be efficiently computed using well-known algorithms. However, accounting for indels in ancestral reconstructions is much harder. First, for biologically-relevant problem formulations, no polynomial-time exact algorithms are available. Second, multiple reconstructions are often equally parsimonious or likely, making it crucial to correctly display uncertainty in the results. Here, we consider a parsimony approach where only deletions are allowed, while addressing the aforementioned limitations. First, we describe an exact algorithm to obtain all the optimal solutions. The algorithm runs in polynomial time if only one solution is sought. Second, we show that all possible optimal reconstructions for a fixed node can be represented using a graph computable in polynomial time. While previous studies have proposed graph-based representations of ancestral reconstructions, this result is the first to offer a solid mathematical justification for this approach. Finally we provide arguments for the relevance of the deletion-only case for the general case.

Algorithms

Screening for dual sgRNAs with comparable indel efficiencies enhances CRISPR-mediated large-fragment deletion.

CRISPR-mediated large-fragment deletion provides a powerful approach for gene clusters, noncoding regions and structural variants, but its broader application is limited by low and variable deletion efficiency. Here, we systematically designed and evaluated 78 sgRNAs targeting nine representative gene clusters (ttn.1-ttn.2 cluster, 7 hox clusters and nppb-nppa cluster), containing 31 large fragments (5 kb-340 kb) to investigate the determinants of deletion efficiency. We found two key rules for achieving high deletion efficiency: (i) using dual sgRNAs with similar indel efficiencies, and (ii) applying a single sgRNA pair rather than multiple sgRNAs. Based on those rules, a 340 kb deletion is detected in the progenies of 95% of founders. Whereas the deletion size showed no significant linear correlation with deletion efficiency within the tested range. Implementing these rules resulted in an average of 70% of founders transmitting deletions across all tested sgRNA pairs. Therefore, screening sgRNAs can effectively enhance CRISPR utility in deletions, thereby facilitating the application of genomic manipulation in vertebrates and other species.

CRISPR

TriosCompass: a snakemake workflow for integrated detection of SNVs, indels, STRs, and structural de novo variants in parent-child trios.

MOTIVATION: The accurate and sensitive identification of de novo variants, which are unique to an individual and not found in the parents' germlines, is critical for understanding the genetic basis of rare diseases, developmental disorders, and evolutionary processes. Existing de novo variant detection pipelines often lack the flexibility to handle multiple variant types, struggle with speed and reproducibility across computational environments, demand extensive manual configuration, or require bioinformatics expertise for downstream curation and analysis, limiting their scalability and usability for large genomic studies. Accordingly, there is a pressing need to better address these challenges. RESULTS: We introduce TriosCompass, an open-source Snakemake workflow that addresses these challenges by providing a modular, accelerated, and environmentally-configurable end-to-end solution for comprehensive de novo variant discovery. It integrates state-of-the-art tools into a reproducible framework, empowering researchers to discover novel genetic insights with greater efficiency and reliability. AVAILABILITY: TriosCompass is implemented as a Snakemake workflow and is freely available at https://github.com/NCI-CGR/TriosCompass_v2 or on Zenodo (10.5281/zenodo.17981062). SUPPLEMENTARY INFORMATION: Supplementary data is available on GitHub at https://github.com/NCI-CGR/TriosCompass_v2/tree/manuscript/report_dashboards. Supplementary methods on DeepTrio benchmark runs can be viewed at: https://github.com/NCI-CGR/TriosCompass_v2/blob/manuscript/TriosCompass_Supp_Methods_deeptrio_benchmark.md.

Software

Comparisons Between Large-Scale Genomic Variants and SNPs in Driving Population Divergence and Local Adaptation.

Genomic variations, such as indels (2-49 bp) and structural variants (SVs, &#x2265;50 bp), are larger-scale mutations than single nucleotide polymorphisms (SNPs) and can substantially impact evolutionary processes, including speciation, adaptation, and phenotypes. Despite their functional importance, integrative population genetic analyses that jointly consider genome-wide SNPs, indels, and SVs remain under-explored. The ground tit (Pseudopodoces humilis), an endemic species to the Qinghai-Tibet Plateau (QTP), exhibits divergence across distinct glacial refugia, accompanied by habitat and morphological divergence, making it an excellent example for investigating how different types of genomic variants contribute to population divergence and local adaptation. Here, by retrieving 81 whole-genome sequence data, over 13 million SNPs, 2 million indels, and 22,101 SVs were identified. Variants were unevenly distributed across the genome, characterized by distinct hotspot regions. Indels and SVs revealed four genetic clusters consistent with previous SNP-based results, thereby validating the reliability of our variant datasets. FST and genotype-environment association (GEA) analyses independently revealed numerous candidate indels and SVs; each showed minimal overlap with previously identified SNPs, and were enriched in similar functional pathways such as signal transduction, skeletal muscle development, water transport, DNA repair, reproduction, nervous system development, and immunity. Collectively, our results demonstrated that indels and SVs could capture additional signatures besides SNPs. Furthermore, similar but distinct gene functions among different types of genomic variants collectively and complementarily drive genomic divergence across environmental gradients in such a high-elevation endemic species, underscoring its evolutionary relevance in local adaptation.

indels

Leveraging ONT move table values for signal aware variant calling.

Oxford Nanopore Technologies (ONT) sequencing enables long-range haplotype phasing and contiguous genome assembly but still exhibits elevated error rates that challenge small variant calling, particularly for insertions and deletions (Indels). While raw electrical signals contain rich information, existing signal-aware methods require computationally intensive processing of large signal files. Here, we present Clair3 v2, a method that leverages the ONT move table-a lightweight byproduct of basecalling that maps signal events to nucleotide positions-to improve variant calling accuracy. Clair3 v2 builds upon Clair3 and integrates signal-level dwelling time to significantly enhance variant calling performance. We also propose a genome position based circular buffer to incorporate dwelling time with minimal computational overhead. Benchmarking across six Genome in a Bottle samples demonstrates substantial improvements in variant calling accuracy. With HAC basecalling, Clair3 v2 achieves a mean SNP F1-score of 97.69% at 10 &#xd7; depth (compared to 96.45% for baseline Clair3), and Indel F1 scores improved from 64.27% to 76.70%, while gains persisted at higher depths. The benefits were most pronounced for longer Indels and in complex genomic regions, where Indel F1 scores in long homopolymer regions improved from 14.3% to 45.2%. Benchmark results across various basecalling modes, samples, and coverage settings outperformed Clair3 baselines and other methods, including DeepVariant and Dorado Variant, and demonstrate the significant benefits of Clair3 v2. Furthermore, Clair3 v2 incurs negligible runtime compared to standard Clair3, making it practical for routine use.

Sequence Analysis, DNA

Assessing the readiness of Oxford Nanopore sequencing for clinical genomics applications.

Long-read sequencing (LRS) technologies, namely, Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), have emerged as promising solutions to overcome the limitations of short-read sequencing (SRS). Nevertheless, the still higher sequencing error rates compared with SRS, need for customized pipelines, rapidly updating software, and incipient scalability continue to present challenges for adopting ONT in standard clinical practice. Here we assess the performance of ONT (R9 and R10 chemistries) in comparison to Illumina and MGI across 17 well-characterized reference samples with 11 clinical variants representing nine different genetic diseases. To enable this, we have implemented a production-ready pipeline including SNV, indel, STR, SV, and CNV detection, alongside reporting key summary metrics to ensure high-quality data at the production sequencing level. Our results show high accuracy of ONT across SNVs (F-score 0.978-0.983) and SVs (F-score = 0.75) but still weaknesses across indels (F-score 0.659-0.758). However, we highlight that ONT accurately detected all four pathogenic indels as well as the performance improvement in exons and with the newer R10 chemistry. We further demonstrated the importance of long reads to detect clinically impactful variants such as a FMR1 pathogenic expansion, often misclassified by SRS as being in the premutation range. Our multiplatform analysis and Sanger validation uncovered a 1 bp error in the Coriell annotation for a cystic fibrosis-causing indel in GM07829. This work underscores the growing readiness of ONT for clinical applications, highlighting both its advancements and its potential for broader adoption in clinical genomics and large-scale operations.

Humans

Molecular dissection and functional characterization of the liguleless1 gene for manipulation of leaf angle in maize.

Recessive liguleless1 (lg1) gene significantly reduces leaf angle in maize and has become the choice in breeding for high plant density. Here, we sequenced the entire lg1 gene (5560&#xa0;bp) among seven wild-type (Lg1) and one mutant (lg1) inbreds. The analysis revealed a total of 229 SNPs and 155 InDels within the Lg1 gene. The study also revealed the existence of three exons, with lg1-mutant having two exons. The lg1-mutant harboured an insertion of 130&#xa0;bp Tourist MITE transposable element in exon-2 at 1663rd base, which deleted 157 amino acids of C-terminal region of the mutant LG1 protein. The mutant LG1 protein was 247 amino acids in length, in contrast to 399-404 amino acids in wild-type protein. The analysis with 26 paralogues and 66 orthologues of Lg1 revealed conservation of the squamosa promoter-binding (SBP) domain. A PCR-based co-dominant InDel marker (MGU-lg1-Tourist) specific to insertion of 130&#xa0;bp was developed that differentiated the mutant allele (lg1) from the wild-type allele (Lg1). The marker was validated in two F2 populations, which showed a 1:2:1 ratio. F2 plants showed a 3 (wide angle: 41.58&#xb0;) :1 (narrow angle: 5.88&#xb0;) segregation for leaf angle. A set of 11 gene-based InDel markers (MGU-InDel1 to MGU-InDel11) specific to Lg1 was also developed, and along with MGU-lg1-Tourist, they classified a diverse set of 48 inbreds into 35 distinct haplotypes (hap1 to hap35) with lg1-based inbreds possessing hap1. This is the first report of the development and validation of a co-dominant gene-based marker specific to lg1, and information generated assumes great significance in maize breeding aimed to tailor the plant architecture suitable for high plant density..

Zea mays

Genomic and genetic dissection underlying seedling drought resilience in oats.

Drought threatens global crop yields, and common oat, a vital nutritional source for food and feed, is particularly constrained in the semi&#x2011;arid regions where it is widely cultivated. Here, we report two high-quality genome assemblies for drought-resilient (Borris37) and drought-sensitive (XymC06) oat accessions with distinct seedling survival rates and genome sizes of 10.92&#x2009;Gb and 10.96&#x2009;Gb, and construct comprehensive landscapes of insertion&#x2011;deletions (InDels) and structural variants (SVs). Integrating population-level genomic, transcriptomic and phenotypic (seedling survival rate), we demonstrate that InDels and SVs underpin divergent drought resilience and identify 52 candidate genes associated with drought resistance whose expression is significantly modulated by these variants. Borris37 accumulates 36 favorable alleles of these genes. An InDel in the AsNF-YB3 promoter enhances binding to AsARF1, upregulating AsNF&#x2011;YB3 under drought, and overexpression of AsNF&#x2011;YB3 reduces ROS accumulation. Our findings provide resources and targets for drought&#x2011;resistance breeding in oat, thereby supporting global food security.

Drought Resistance

SpacerScope: binary-vectorized, genome-wide off-target profiling for RNA-guided nucleases without prior candidate-site bias.

The precision of CRISPR/Cas systems is fundamental to their application in plant and animal biotechnology. However, comprehensive sequence-based off-target candidate discovery remains a computational bottleneck, particularly in large and complex genomes. Here we developed SpacerScope, an off-target candidate discovery framework that enables unbiased, genome-wide discovery by leveraging binary vectorization, bitwise filtering, and right-end-anchored alignment. Benchmarking against human CIRCLE-seq data demonstrated that SpacerScope recovered 100% of validated off-target sites (6142/6142), matching the sensitivity of exhaustive algorithms. Crucially, SpacerScope achieved this maximum candidate recovery while substantially reducing computational overhead. In large-genome evaluations, SpacerScope maintained low peak memory usage of 2.20 GiB and achieved substantial runtime improvements over indel-aware comparator tools, including more than 50-fold speedup relative to Cas-OFFinder 3 (544&#xa0;s versus 29&#xa0;185&#xa0;s). Furthermore, comparative analyses in polyploid species, such as the octoploid strawberry, revealed that SpacerScope identified larger sequence-compatible candidate burdens than standard web-based design platforms. Our results establish SpacerScope as a high-speed framework for sequence-based genome-wide off-target candidate discovery across diverse and highly repetitive genomic landscapes. The source code and program was publicly available at https://github.com/charlesqu666/SpacerScope. Short Abstract CRISPR/Cas sequence-based off-target candidate discovery remains computationally challenging in large, repetitive, and polyploid genomes. Existing tools either miss indel-containing candidate sites or incur prohibitive runtime and memory costs. We developed SpacerScope, a binary-vectorized framework that enables unbiased, genome-wide off-target candidate discovery without pre-selected candidate sites. By integrating bitwise filtering with right-end-anchored alignment, SpacerScope recovered 100% of validated off-target sites in human CIRCLE-seq data while using only 2.20 GiB of memory and achieving more than 10-fold speedup over indel-aware alternatives. Evaluation in plant genomes, including rice and octoploid strawberry, further demonstrated SpacerScope's capacity to identify larger sequence-compatible candidate burdens overlooked by standard tools. SpacerScope thus provides a high-speed framework for sequence-based genome-wide off-target candidate discovery across diverse and highly repetitive genomic landscapes, supporting downstream prioritization.

CRISPR-Cas Systems

A novel insertion/deletion in APC promotor 1B is associated with both gastric and colon polyposis.

Pathogenic variants in the APC gene are classically associated with autosomal dominant familial adenomatous polyposis (FAP), characterized by tens-to-thousands of colonic adenomatous polyps and a high-penetrance predisposition to colorectal cancer. More recently, specific PVs in the YY1 binding motif of APC promoter 1B have been associated with autosomal dominant gastric adenocarcinoma and proximal polyposis of the stomach (GAPPS), characterized by tens-to-thousands of fundic gland polyps and a predisposition to gastric cancer but which are only rarely associated with features consistent with FAP. Although management guidelines currently treat FAP and GAPPS as mutually exclusive conditions, the extent of phenotypic overlap is not well-characterized. Here, we present a multi-clinic and -laboratory collaboration reporting a previously undescribed APC promoter 1B insertion/deletion likely pathogenic variant in a family with mixed GAPPS and FAP phenotype. The family proband is a female of unspecified white ancestry. She was diagnosed with GAPPS at age 30 and, after developing gastric cancer at age 39, underwent curative gastrectomy. She is now 61 with a cumulative history of between 50 and 100 colon adenomas and recently completed subtotal colectomy. Her multi-gene panel testing in 2022 demonstrated a likely pathogenic insertion/deletion (indel) within the APC promoter 1B YY1 binding motif (APC c.-192_-191delATinsTAGCAAGGG). Review of a four-generation pedigree revealed the ages of gastric cancer presentation in the family ranged from 39-60's, with advanced gastric polyposis and prophylactic gastrectomy as early as ages 11 and 13 in the proband's daughter and nephew, respectively. Six of 10 (60%) family members known or presumed to carry the APC likely pathogenic variant underwent colectomy or hemicolectomy due to colon polyposis. The youngest known carrier in the family is a 12-year-old female, and the oldest living carrier is the proband's brother, age 66. A novel APC indel causes concomitant GAPPS and FAP presentations in this previously unreported large kindred. Mixed gastric and colon phenotypes have been rarely described in GAPPS families and the ages of presentation of gastric polyposis are strikingly young in the current family with prophylactic gastrectomies completed as early as age 11 and 13. These ages are significantly younger than the 15 years of age at which national guidelines currently recommend initiation of EGD for screening in GAPPS. Although the mechanism for this combined GAPPS-FAP phenotype is unclear, patients in this family and those with similar APC promoter 1B variants should be offered both gastric and colon cancer risk management.

Adult

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

With hundreds of copies of rDNA, it is unknown whether they possess sequence variations that form different types of ribosomes. Here, we developed an algorithm for long-read variant calling, termed RGA, which revealed that variations in human rDNA loci are predominantly insertion-deletion (indel) variants. We developed full-length rRNA sequencing (RIBO-RT) and in situ sequencing (SWITCH-seq), which showed that translating ribosomes possess variation in rRNA. Over 1,000 variants are lowly expressed. However, tens of variants are abundant and form distinct rRNA subtypes with different structures near indels as revealed by long-read rRNA structure probing coupled to dimethyl sulfate sequencing. rRNA subtypes show differential expression in endoderm/ectoderm-derived tissues, and in cancer, low-abundance rRNA variants can become highly expressed. Together, this study identifies the diversity of ribosomes at the level of rRNA variants, their chromosomal location, and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Humans