Search PubMedSearch

SEARCH · Search PubMed

Results for “indel mutation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Indel mutation in transcription factor PabHLH2 regulates amygdalin accumulation and kernel bitterness in apricot.

Amygdalin, the phytochemical responsible for the characteristic bitterness of apricot (Prunus armeniaca L.) kernels, also exhibits significant bioactive properties and therapeutic potential. Genetic regulation of amygdalin content is therefore a key objective in apricot breeding programs aimed at quality improvement. In this study, we conducted quantitative trait loci (QTL) mapping to uncover the genetic basis of sweet-bitter differentiation in apricot kernels. We identified a 15-bp insertion/deletion (indel) polymorphism strongly related to kernel bitterness, with marker validation achieving 100% concordance across 601 apricot germplasm accessions. Notably, this polymorphic site is located within the helix-loop-helix (HLH) domain of the basic HLH (bHLH) transcription factor PabHLH2. Protein interaction analyses revealed that the 15-bp deletion variant impaired dimerization capacity, reducing transcriptional activation of downstream targets. Using yeast one-hybrid screening and dual-luciferase reporter assays, we identified PaCYP71AN24 and PaCYP79D16 as direct transcriptional targets of PabHLH2. Functional characterization further indicated that the PabHLH2a variant (harboring the 15-bp insertion) significantly enhanced the promoter activity of these cytochrome P450 genes compared with the deletion variant. Transient overexpression and silencing experiments in apricot kernels further confirmed that the 15-bp insertion positively regulates both PaCYP71AN24/PaCYP79D16 expression and prunasin accumulation, the immediate biosynthetic precursor of amygdalin. Overall, these findings provide mechanistic insights into the allelic variation underlying kernel bitterness and delineate the molecular cascade of amygdalin biosynthesis. The identified molecular markers and functional characterization establish a basis for marker-assisted breeding of low-amygdalin apricot cultivars, supporting the dual-purpose utilization of kernels in food and pharmaceutical industries.

Amygdalin

Whole-Exome and Whole-Genome Sequencing of Candidate Pharmacogenomic and Schizophrenia-Related Genes in Sudanese Families with Schizophrenia.

BACKGROUND: Schizophrenia is considered a neuro-developmental disorder leading to disastrous lifelong disability of the patients and their families. There is a lack of data regarding pharmacogenomics of schizophrenia in Sudan. This study aimed to identify different genes affecting the treatment outcomes in Sudanese patients with schizophrenia. METHODS: A case-control study was conducted on seven families having more than one member diagnosed with schizophrenia. This was a small exploratory family-based sequencing study involving 18 affected individuals and 8 controls from seven families. Ethical clearance and informed consent were obtained. Demographic data were collected using a standardized data collection sheet. DNA was extracted from blood samples collected from patients and control groups. Then, whole-exome and genome sequencing were performed. Sixty-six genes associated with schizophrenia, treatment, and treatment resistance were selected from the variant calling file. Variants showing single-nucleotide polymorphisms (SNPs) were identified. These variants were then classified based on their impact on the protein-coding sequence into high- and moderate-impact. Moreover, indel mutations were also identified. RESULTS: Twelve variants of seven genes (COMT, FMO1, LPL, CYP2E1, ABCC1, GRM3, CYP2C9) were identified as genes with impact and potential association with schizophrenia (p-value=0.006632). Forty-three genes had a moderate impact, and they showed a potential association with schizophrenia (p-value=0.0004436). Two variants were indel mutations (CYP2D6, DTNBP1) and showed association with schizophrenia (p-value=0.004741). The p-values were generated from different databases. CONCLUSION: This exploratory family-based sequencing study identified several potentially relevant pharmacogenomic and schizophrenia-associated variants in Sudanese families, warranting validation in larger and ethnically diverse cohorts.

antipsychotics

Efficient and precise programmable DNA knock-in without double-strand breaks.

Programmable gene knock-in holds substantial promise for treating genetic diseases and advancing cell therapies. However, achieving precise and efficient kilobase-scale DNA fragment integration remains challenging1,2. Here we report CRISPR kilobase-scale nickase-targeting (KNIT) editing for efficient, precise and programmable kilobase-scale DNA insertion without double-strand DNA cleavage, which is enabled through the coupling of a Cas9 nickase with a DNA donor recruiting system. KNIT editing facilitates programmable integration of DNA fragments from 0.7 kb to more than 10 kb and is effective across genomic loci and cell types. It achieves up to 89% efficiency and markedly reduces unintended insertion-deletion mutation (indels) rates, translocations and off-target editing. The system supports repeated insertion editing and multiloci gene knock-in with minimal translocations. Its enhanced version, KNIT editor 2, further improves efficiency via a single transfection. Moreover, in mutant cells with a pathological mutation, KNIT editing restores normal gene expression by inserting a therapeutic gene into a safe harbour locus or its native locus. Notably, KNIT editing enables non-viral and programmable chimeric antigen receptor T cell (CAR-T cell) engineering without double-strand breaks and with clinically relevant efficiencies. Moreover, the engineered CAR-T cells exhibit effective antitumour activity in vitro and in mouse models. Therefore, by achieving programmable and site-specific kilobase-scale DNA insertions without double-strand breaks while reducing unintended outcomes, KNIT editing provides a versatile platform for advancing personalized medicine.

Animals

The mutation landscape of Daphnia obtusa reveals evolutionary forces shaping genome stability.

Spontaneous mutations are the primary source of genetic variation and play a central role in shaping evolutionary processes. To investigate mutational dynamics in Daphnia obtusa, we generated a chromosome-level genome assembly spanning 129.4 Mb across 12 chromosomes, encompassing 15,321 predicted protein-coding genes. Leveraging whole-genome sequencing of eight mutation accumulation (MA) lines propagated for an average of 482 generations (spanning over 20 years), we estimated a spontaneous single nucleotide mutation (SNM) rate of 2.23 × 10-9 and an indel mutation rate of 2.75 × 10-10 per site per generation. The SNM spectrum was strongly biased toward C:G > T:A transitions. Comparative analyses with natural population data revealed that exonic mutations observed in the MA lines were significantly less likely to be present in standing variation than intronic or intergenic mutations, suggesting that purifying selection in natural populations acts to remove deleterious alleles. We also identified 48 de novo loss-of-heterozygosity (LOH) events, comprising 8 heterozygous deletions and 40 gene conversion events. The genome-wide gene conversion rate was estimated at 2.62 × 10-5 per heterozygous site per generation. These findings provide a comprehensive view of the mutation spectrum, selective pressures, and mechanisms underlying genome stability in D. obtusa.

Daphnia obtusa

A quick guide to evaluating prime editing efficiency in mammalian cells.

According to the Clinvar database, modeling the diseases associated with pathogenic mutations requires the installation of base substitutions, small insertions or deletions. Prime editor (PE) was recently developed to precisely install any base substitutions and/or small insertions/deletions (indels) in mammalian cells and animals without requiring DSBs or donor DNA templates. PE also offers greater editing and targeting flexibility compared to other precision CRISPR editing methods because the versatile editing information is encoded in the reverse-transcription template of its prime editing guide RNA. However, optimal PE system selection and experimental design can be complex, and there are various factors that can affect PE efficiency. This chapter serves as a rapid entry-level guideline for the application of PE, providing an experimental framework for using PE at a specific genomic locus. RUNX1 was selected as a representative target site to illustrate the detailed methodology for constructing PE plasmids and the process of transfecting these plasmids into 293FT cells. We further examined the efficiency of PE-mediated genome editing in mammalian cells by using next-generation sequencing.

Gene Editing

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Molecular dissection and functional characterization of the liguleless1 gene for manipulation of leaf angle in maize.

Recessive liguleless1 (lg1) gene significantly reduces leaf angle in maize and has become the choice in breeding for high plant density. Here, we sequenced the entire lg1 gene (5560 bp) among seven wild-type (Lg1) and one mutant (lg1) inbreds. The analysis revealed a total of 229 SNPs and 155 InDels within the Lg1 gene. The study also revealed the existence of three exons, with lg1-mutant having two exons. The lg1-mutant harboured an insertion of 130 bp Tourist MITE transposable element in exon-2 at 1663rd base, which deleted 157 amino acids of C-terminal region of the mutant LG1 protein. The mutant LG1 protein was 247 amino acids in length, in contrast to 399-404 amino acids in wild-type protein. The analysis with 26 paralogues and 66 orthologues of Lg1 revealed conservation of the squamosa promoter-binding (SBP) domain. A PCR-based co-dominant InDel marker (MGU-lg1-Tourist) specific to insertion of 130 bp was developed that differentiated the mutant allele (lg1) from the wild-type allele (Lg1). The marker was validated in two F2 populations, which showed a 1:2:1 ratio. F2 plants showed a 3 (wide angle: 41.58°) :1 (narrow angle: 5.88°) segregation for leaf angle. A set of 11 gene-based InDel markers (MGU-InDel1 to MGU-InDel11) specific to Lg1 was also developed, and along with MGU-lg1-Tourist, they classified a diverse set of 48 inbreds into 35 distinct haplotypes (hap1 to hap35) with lg1-based inbreds possessing hap1. This is the first report of the development and validation of a co-dominant gene-based marker specific to lg1, and information generated assumes great significance in maize breeding aimed to tailor the plant architecture suitable for high plant density..

Zea mays

A novel insertion/deletion in APC promotor 1B is associated with both gastric and colon polyposis.

Pathogenic variants in the APC gene are classically associated with autosomal dominant familial adenomatous polyposis (FAP), characterized by tens-to-thousands of colonic adenomatous polyps and a high-penetrance predisposition to colorectal cancer. More recently, specific PVs in the YY1 binding motif of APC promoter 1B have been associated with autosomal dominant gastric adenocarcinoma and proximal polyposis of the stomach (GAPPS), characterized by tens-to-thousands of fundic gland polyps and a predisposition to gastric cancer but which are only rarely associated with features consistent with FAP. Although management guidelines currently treat FAP and GAPPS as mutually exclusive conditions, the extent of phenotypic overlap is not well-characterized. Here, we present a multi-clinic and -laboratory collaboration reporting a previously undescribed APC promoter 1B insertion/deletion likely pathogenic variant in a family with mixed GAPPS and FAP phenotype. The family proband is a female of unspecified white ancestry. She was diagnosed with GAPPS at age 30 and, after developing gastric cancer at age 39, underwent curative gastrectomy. She is now 61 with a cumulative history of between 50 and 100 colon adenomas and recently completed subtotal colectomy. Her multi-gene panel testing in 2022 demonstrated a likely pathogenic insertion/deletion (indel) within the APC promoter 1B YY1 binding motif (APC c.-192_-191delATinsTAGCAAGGG). Review of a four-generation pedigree revealed the ages of gastric cancer presentation in the family ranged from 39-60's, with advanced gastric polyposis and prophylactic gastrectomy as early as ages 11 and 13 in the proband's daughter and nephew, respectively. Six of 10 (60%) family members known or presumed to carry the APC likely pathogenic variant underwent colectomy or hemicolectomy due to colon polyposis. The youngest known carrier in the family is a 12-year-old female, and the oldest living carrier is the proband's brother, age 66. A novel APC indel causes concomitant GAPPS and FAP presentations in this previously unreported large kindred. Mixed gastric and colon phenotypes have been rarely described in GAPPS families and the ages of presentation of gastric polyposis are strikingly young in the current family with prophylactic gastrectomies completed as early as age 11 and 13. These ages are significantly younger than the 15 years of age at which national guidelines currently recommend initiation of EGD for screening in GAPPS. Although the mechanism for this combined GAPPS-FAP phenotype is unclear, patients in this family and those with similar APC promoter 1B variants should be offered both gastric and colon cancer risk management.

Adult

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

With hundreds of copies of rDNA, it is unknown whether they possess sequence variations that form different types of ribosomes. Here, we developed an algorithm for long-read variant calling, termed RGA, which revealed that variations in human rDNA loci are predominantly insertion-deletion (indel) variants. We developed full-length rRNA sequencing (RIBO-RT) and in situ sequencing (SWITCH-seq), which showed that translating ribosomes possess variation in rRNA. Over 1,000 variants are lowly expressed. However, tens of variants are abundant and form distinct rRNA subtypes with different structures near indels as revealed by long-read rRNA structure probing coupled to dimethyl sulfate sequencing. rRNA subtypes show differential expression in endoderm/ectoderm-derived tissues, and in cancer, low-abundance rRNA variants can become highly expressed. Together, this study identifies the diversity of ribosomes at the level of rRNA variants, their chromosomal location, and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Humans

ONCOLINER: A new solution for monitoring, improving, and harmonizing somatic variant calling across genomic oncology centers.

The characterization of somatic genomic variation associated with the biology of tumors is fundamental for cancer research and personalized medicine, as it guides the reliability and impact of cancer studies and genomic-based decisions in clinical oncology. However, the quality and scope of tumor genome analysis across cancer research centers and hospitals are currently highly heterogeneous, limiting the consistency of tumor diagnoses across hospitals and the possibilities of data sharing and data integration across studies. With the aim of providing users with actionable and personalized recommendations for the overall enhancement and harmonization of somatic variant identification across research and clinical environments, we have developed ONCOLINER. Using specifically designed mosaic and tumorized genomes for the analysis of recall and precision across somatic SNVs, insertions or deletions (indels), and structural variants (SVs), we demonstrate that ONCOLINER is capable of improving and harmonizing genome analysis across three state-of-the-art variant discovery pipelines in genomic oncology.

Humans

Genomic and genetic dissection underlying seedling drought resilience in oats.

Drought threatens global crop yields, and common oat, a vital nutritional source for food and feed, is particularly constrained in the semi‑arid regions where it is widely cultivated. Here, we report two high-quality genome assemblies for drought-resilient (Borris37) and drought-sensitive (XymC06) oat accessions with distinct seedling survival rates and genome sizes of 10.92 Gb and 10.96 Gb, and construct comprehensive landscapes of insertion‑deletions (InDels) and structural variants (SVs). Integrating population-level genomic, transcriptomic and phenotypic (seedling survival rate), we demonstrate that InDels and SVs underpin divergent drought resilience and identify 52 candidate genes associated with drought resistance whose expression is significantly modulated by these variants. Borris37 accumulates 36 favorable alleles of these genes. An InDel in the AsNF-YB3 promoter enhances binding to AsARF1, upregulating AsNF‑YB3 under drought, and overexpression of AsNF‑YB3 reduces ROS accumulation. Our findings provide resources and targets for drought‑resistance breeding in oat, thereby supporting global food security.

Drought Resistance

Whole-Genome Sequencing of 54 Dengchuan Cattle (Bos taurus) from Southwest China.

Domestic cattle (Bos taurus) play a significant role in human society as they provide abundant food resources and contribute to the development of agriculture and traditional culture. Dengchuan cattle, a local breed from Yunnan, Southwest China, are known for their high-quality milk and are at risk of extinction due to crossbreeding. To preserve the superior genetic resources of Dengchuan cattle, this study conducted whole-genome sequencing of 54 Dengchuan cattle using blood DNA samples, generating approximately 3.56 TB of clean data with an average sequencing depth of 32.78X. The sequencing data were aligned to the bovine reference genome (ARS-UCD2.0), achieving an average alignment rate of 99.85%. A total of 9,950,420 SNPs and 2,476,207 indels were detected using variant calling workflow. These data were utilized to characterize genomic profile of this unique cattle breed. The data generated in this study can be incorporated into the global cattle genomic diversity database, providing valuable information for comparative studies on cattle.

Animals

NanoFilter: enhancing phasing performance by utilizing highly consistent INDELs and SNVs in nanopore sequencing.

MOTIVATION: Nanopore sequencing data offer longer reads compared to other technologies, which is beneficial for phasing and genome assembly. INDELs provide valuable haplotype information and have significant potential to improve phasing performance. However, accurately identifying INDELs with variant callers is challenging, and incorporating INDELs into phasing remains a complex task. To address these issues, we developed NanoFilter, a novel filtering strategy designed to filter out INDELs that contain wrong phasing information based on their consistency. RESULTS: Our assessment using Nanopore R10 simplex data shows that filtering out low-consistency INDELs increases their precision from 88.3% to 98.8%, nearly matching the precision of SNVs. In the phasing results of Margin, incorporating these filtered INDELs leads to a 12.77% increase in N50 length and fewer switch errors. Furthermore, we found that SNVs filtered by NanoFilter will enhance assembly performance. When NanoFilter is integrated into the HapDup assembly pipeline, NanoFilter reduces the Hamming error rate and increases N50 length by 7.8%. AVAILABILITY AND IMPLEMENTATION: NanoFilter is available at https://github.com/Chenshanming-repo/NanoFilter (DOI: 10.5281/zenodo.16777826) and HapDup-NanoFilter is available at https://github.com/Chenshanming-repo/HapDup-NanoFilter (DOI: 10.5281/zenodo.16777890).

Nanopore Sequencing

TriosCompass: a snakemake workflow for integrated detection of SNVs, indels, STRs, and structural de novo variants in parent-child trios.

MOTIVATION: The accurate and sensitive identification of de novo variants, which are unique to an individual and not found in the parents' germlines, is critical for understanding the genetic basis of rare diseases, developmental disorders, and evolutionary processes. Existing de novo variant detection pipelines often lack the flexibility to handle multiple variant types, struggle with speed and reproducibility across computational environments, demand extensive manual configuration, or require bioinformatics expertise for downstream curation and analysis, limiting their scalability and usability for large genomic studies. Accordingly, there is a pressing need to better address these challenges. RESULTS: We introduce TriosCompass, an open-source Snakemake workflow that addresses these challenges by providing a modular, accelerated, and environmentally-configurable end-to-end solution for comprehensive de novo variant discovery. It integrates state-of-the-art tools into a reproducible framework, empowering researchers to discover novel genetic insights with greater efficiency and reliability. AVAILABILITY: TriosCompass is implemented as a Snakemake workflow and is freely available at https://github.com/NCI-CGR/TriosCompass_v2 or on Zenodo (10.5281/zenodo.17981062). SUPPLEMENTARY INFORMATION: Supplementary data is available on GitHub at https://github.com/NCI-CGR/TriosCompass_v2/tree/manuscript/report_dashboards. Supplementary methods on DeepTrio benchmark runs can be viewed at: https://github.com/NCI-CGR/TriosCompass_v2/blob/manuscript/TriosCompass_Supp_Methods_deeptrio_benchmark.md.

Software

Leveraging ONT move table values for signal aware variant calling.

Oxford Nanopore Technologies (ONT) sequencing enables long-range haplotype phasing and contiguous genome assembly but still exhibits elevated error rates that challenge small variant calling, particularly for insertions and deletions (Indels). While raw electrical signals contain rich information, existing signal-aware methods require computationally intensive processing of large signal files. Here, we present Clair3 v2, a method that leverages the ONT move table-a lightweight byproduct of basecalling that maps signal events to nucleotide positions-to improve variant calling accuracy. Clair3 v2 builds upon Clair3 and integrates signal-level dwelling time to significantly enhance variant calling performance. We also propose a genome position based circular buffer to incorporate dwelling time with minimal computational overhead. Benchmarking across six Genome in a Bottle samples demonstrates substantial improvements in variant calling accuracy. With HAC basecalling, Clair3 v2 achieves a mean SNP F1-score of 97.69% at 10 × depth (compared to 96.45% for baseline Clair3), and Indel F1 scores improved from 64.27% to 76.70%, while gains persisted at higher depths. The benefits were most pronounced for longer Indels and in complex genomic regions, where Indel F1 scores in long homopolymer regions improved from 14.3% to 45.2%. Benchmark results across various basecalling modes, samples, and coverage settings outperformed Clair3 baselines and other methods, including DeepVariant and Dorado Variant, and demonstrate the significant benefits of Clair3 v2. Furthermore, Clair3 v2 incurs negligible runtime compared to standard Clair3, making it practical for routine use.

Sequence Analysis, DNA

Assessing the readiness of Oxford Nanopore sequencing for clinical genomics applications.

Long-read sequencing (LRS) technologies, namely, Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), have emerged as promising solutions to overcome the limitations of short-read sequencing (SRS). Nevertheless, the still higher sequencing error rates compared with SRS, need for customized pipelines, rapidly updating software, and incipient scalability continue to present challenges for adopting ONT in standard clinical practice. Here we assess the performance of ONT (R9 and R10 chemistries) in comparison to Illumina and MGI across 17 well-characterized reference samples with 11 clinical variants representing nine different genetic diseases. To enable this, we have implemented a production-ready pipeline including SNV, indel, STR, SV, and CNV detection, alongside reporting key summary metrics to ensure high-quality data at the production sequencing level. Our results show high accuracy of ONT across SNVs (F-score 0.978-0.983) and SVs (F-score = 0.75) but still weaknesses across indels (F-score 0.659-0.758). However, we highlight that ONT accurately detected all four pathogenic indels as well as the performance improvement in exons and with the newer R10 chemistry. We further demonstrated the importance of long reads to detect clinically impactful variants such as a FMR1 pathogenic expansion, often misclassified by SRS as being in the premutation range. Our multiplatform analysis and Sanger validation uncovered a 1 bp error in the Coriell annotation for a cystic fibrosis-causing indel in GM07829. This work underscores the growing readiness of ONT for clinical applications, highlighting both its advancements and its potential for broader adoption in clinical genomics and large-scale operations.

Humans

Pangenomes aid accurate detection of large insertions and deletions from targeted sequencing: the case of cardiomyopathies.

BACKGROUND: Gene panels represent a widely used strategy for genetic testing in a vast range of Mendelian disorders. While this approach aids reliable bioinformatic detection of short coding variants, it often fails to detect many larger variants. Recent studies have recommended the adoption of pangenome references (as opposed to linear reference genomes like GRCh38) to augment detection of large variants from targeted sequencing, potentially providing diagnostic laboratories with the possibility to streamline diagnostic work-ups and reduce costs. METHODS: Here, we analyze 1969 cardiomyopathy cases and 1805 controls sequenced with the Illumina Trusight Cardio panel using a pangenome-based workflow (GRAF) and five conventional orthogonal methodologies (GATK HaplotypeCaller, GATK-gCNV, ExomeDepth, Manta and Lumpy-SV) to detect variants ≥ 20 bp in size. RESULTS: Following lab-based variant validation by means of PCR and Sanger sequencing, we show that GRAF conjugates higher precision and recall (F1 score 0.86) compared with other methods (F1 0-0.57) in detecting potentially pathogenic variants ≥ 20 bp from short-read panel data. Results were complemented by a comparison of the tools' performance in detecting ground truth variants on reference sample HG002 from Genome In A Bottle, which confirmed GRAF to outperform other tools also on exome sequencing (F1 0.97 vs. 0-0.94). Notably, in the HG002 benchmark dataset, GRAF also showed slightly improved performance compared to GATK HaplotypeCaller in the identification of small variants (1-19 bp; F1 0.975 vs. 0.968). CONCLUSIONS: Our results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.

Humans

Algorithms to reconstruct past indels: The deletion-only parsimony problem.

Ancestral sequence reconstruction is an important task in bioinformatics, with applications ranging from protein engineering to the study of genome evolution. When sequences can only undergo substitutions, optimal reconstructions can be efficiently computed using well-known algorithms. However, accounting for indels in ancestral reconstructions is much harder. First, for biologically-relevant problem formulations, no polynomial-time exact algorithms are available. Second, multiple reconstructions are often equally parsimonious or likely, making it crucial to correctly display uncertainty in the results. Here, we consider a parsimony approach where only deletions are allowed, while addressing the aforementioned limitations. First, we describe an exact algorithm to obtain all the optimal solutions. The algorithm runs in polynomial time if only one solution is sought. Second, we show that all possible optimal reconstructions for a fixed node can be represented using a graph computable in polynomial time. While previous studies have proposed graph-based representations of ancestral reconstructions, this result is the first to offer a solid mathematical justification for this approach. Finally we provide arguments for the relevance of the deletion-only case for the general case.

Algorithms