Search PubMedSearch

SEARCH · Search PubMed

Results for “Repeats”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Pentatricopeptide repeat protein targeting CUG repeat RNA ameliorates RNA toxicity in a myotonic dystrophy type 1 mouse model.

Myotonic dystrophy type 1 (DM1) is an autosomal dominant multisystemic disorder caused by the expansion of a CTG-triplet repeat in the 3' untranslated region of the dystrophia myotonica protein kinase (DMPK) gene. It results in the transcription of toxic RNAs that contain expanded CUG repeats (CUGexp). Splicing factors, such as muscleblind-like 1 (MBNL1), are sequestered by CUGexp, thereby disrupting the normal splicing program that is essential for various cellular functions. Pentatricopeptide repeat (PPR) proteins, originally found in plants, regulate RNA in organelles by binding in a sequence-specific manner. Here, we designed PPR proteins that specifically bind to the hexamer of CUG repeat RNAs (CUG-PPRs) and showed that CUG-PPR1 could ameliorate RNA toxicity induced by CUGexp in cell models of DM1. A single systemic recombinant adeno-associated virus (AAV9) vector-mediated gene delivery of CUG-PPR1 demonstrated long-term therapeutic effects on myotonia and restored splicing activity in a mouse model of DM1. These results highlight the potential of PPR molecules to target pathogenic RNA sequences in DM1 and potentially other RNA-mediated disorders.

Animals

Multiomic approaches identify a rare CCG repeat expansion in BCLAF3 in neurodevelopmental disorders.

BACKGROUND: Tandem repeat expansions have been implicated in various neurological conditions. Here, we present a novel hypermethylated CCG repeat expansion on Xp22 in the 5'UTR of BCLAF3 in males with neurodevelopmental disorders. METHODS: We used patient-derived fibroblasts and neuronal models from a family with BCLAF3 repeat expansions to generate multiomic data and investigate downstream molecular consequences of the repeat expansion. To identify additional affected individuals with BCLAF3 repeat expansions, we screened methylation arrays (n = 12,375) and short-read genomes (n = 15,963) from probands with neurodevelopmental presentations. We also characterized BCLAF3 repeat expansions in the general population using long-read sequencing data (n = 793) and population-level short-read sequencing data (n = 410,076). RESULTS: Long-read sequencing validated hypermethylation of expanded repeats. Patient-derived cells showed repressed BCLAF3 RNA and protein expression. We show that the BCLAF3 CCG repeat expansion constitutes a previously uncharacterized fragile site (FRAXG) that shifts the surrounding chromatin compartment from open euchromatin to closed heterochromatin. Using our multiomic screening approaches, we identified three additional unrelated males and one related male cousin with long-read sequencing validated (n = 2) or short-read sequencing predicted (n = 2) repeat expansions. In one family, the BCLAF3 repeats segregate with more severe phenotypes than expected for the primary diagnoses. Long-read sequencing in three carrier mothers showed skewed X-inactivation against the repeat expansion, highlighting the potential deleterious effect of an allele with an expansion. Expansions were absent in long-read sequencing data from control populations. Assessment of the BCLAF3 repeat expansion in the UK Biobank indicates that it may be ~ 20X rarer than FMR1 repeat expansions. CONCLUSIONS: CCG repeat expansions in the 5'UTR of BCLAF3 likely constitute a novel genetic etiology associated with X-linked neurodevelopmental phenotypes in males. Future work will be essential to delineate the phenotypic spectrum and determine a disease pathomechanism.

BCLAF3

Intermediate FMR1 cytosine‒guanine‒guanine repeats do not impair assisted reproductive technology outcomes in a large real-world cohort.

RESEARCH QUESTION: Does the presence of moderately elevated FMR1 cytosine‒guanine‒guanine (CGG) repeat numbers (40-70 repeats), identified through routine pre-pregnancy screening, adversely affect assisted reproductive technology (ART) outcomes in a real-world population? DESIGN: Retrospective cohort study including 760 first ART cycles conducted between 2010 and 2021 at a university-affiliated centre. FMR1 CGG repeat testing was conducted independently of infertility evaluation. Patients were categorized by repeat status in both alleles using two thresholds: 40 or more repeats (primary analysis) and 34 or more repeats (secondary analysis). Ovarian reserve markers, stimulation characteristics, oocyte yield, embryologic outcomes, positive beta-HCG and live birth rates were compared across groups. RESULTS: Among 760 patients, 669 (88%) had no allele of 40 or more repeats, 85 (11%) had one allele of 40 or more repeats and six (0.8%) had two alleles of 40 or more repats. The maximum observed repeat length was 71. Baseline demographics and ovarian reserve markers were similar between groups. No differences were observed in ovarian response, oocyte yield, fertilization or embryo development by FMR1 repeat category. Pregnancy and live birth rates were comparable between controls and patients with one expanded allele. Although elevated pregnancy and live birth rates were observed in patients with two expanded alleles, this subgroup was small, limiting interpretation. Analyses using the 34 or more repeat threshold yielded similar findings. CONCLUSIONS: Moderately elevated FMR1 CGG repeat numbers are not associated with impaired ART outcomes. Standard ART protocols remain appropriate, and FMR1 repeat length alone should not guide treatment modification in the absence of clinical ovarian insufficiency.

Humans

Replication of DNA Containing Trinucleotide Repeats by the Bacteriophage T7 Replisome.

Trinucleotide repeats in the human genome are implicated in various neurodegenerative diseases. The tendency of these repetitive DNA sequences to form non-B DNA structures can cause abnormal replication, leading to genomic instability. This instability contributes to disease progression, though the underlying mechanisms are not fully understood. We investigated the replication of DNA containing CAG and CTG trinucleotide repeats using individual components of the T7 bacteriophage replication machinery, as well as the complete replisome. Our results show that repeats in linear single-stranded DNA (ssDNA) inhibit the activity of T7 DNA polymerase and ssDNA-binding proteins, with a more pronounced effect observed in CTG repeats compared to CAG repeats. Direct unwinding assays showed that the T7 gene 4 helicase unwound forked substrates containing CAG or CTG repeats at least as efficiently as random-sequence substrates; however, the displaced repeat strands were recovered predominantly as compact, structured species rather than as unstructured single-stranded DNA, providing direct evidence that secondary structure forms immediately upon unwinding. Minicircle templates containing CTG repeats exhibited robust DNA synthesis on both the leading and lagging strands, though synthesis was not enhanced by the T7 gene 2.5 ssDNA-binding protein. The lagging strand products generated from the CTG repeat minicircle were significantly longer than those from random sequence templates, and their lengths were not extended by the presence of T7 gene 2.5 protein. When the repeated sequences were incorporated into the T7 phage genome, heterogeneity was observed downstream of the repeats, depending on their length. We propose that aberrant extension occurs predominantly in the lagging strand, driven by dynamic interactions between the repeated sequences and the DNA replisome. This study may provide a foundation for understanding the mechanisms underlying the extension or deletion of repetitive genomic regions.

DNA repeats

Dissecting the relationship between haplotypes around ATXN2 CAG repeats and the number of CAA interruptions by long-read sequencing.

BACKGROUND: CAG repeat expansions in ATXN2 are implicated as risk factors for several neurological diseases, including spinocerebellar ataxia type 2 (SCA2) when >=33 CAG repeats are present, and amyotrophic lateral sclerosis (ALS) when 27-33 CAG repeats are present. However, how haplotypes around the repeats and CAA interruptions within the repeats are associated with disease phenotypes remains poorly understood. Previous studies on haplotypes around ATXN2 were limited to SNPs very close to the repeats (<5kb) or were based on statistical inference only. METHODS: Here, we used long-read sequencing on the Oxford Nanopore Technologies (ONT) platform to simultaneously infer haplotypes around ATXN2, the number of CAG repeats, and the number of CAA interruptions, along with NYGC ALS Consortium NGS dataset. We further sequenced 41 individuals (EUR = 39) with neurological diseases with intermediate repeats by ONT. RESULTS: We found that haplotypes around ATXN2 and the number of interruptions show ethnicity-specific and ALS-specific distribution. Three CAA interruptions are present at low prevalence (~1%) in control populations in multiple ancestry groups, but high prevalence (~55%) in ALS individuals with intermediate repeats. Furthermore, we examined 159 individuals with ALS (~90% European ancestry) with intermediate ATXN2 repeats and found a unique haplotype in ALS individuals with three CAA interruptions, which can be tagged by an SNV, rs148019457. We also validated that the rs148019457-G allele is only present in haplotypes with three CAA interruptions. CONCLUSIONS: In summary, our study shows that 3 CAA interruptions are rarely seen in healthy controls but are common in those with expanded ATXN2 CAG repeats who have neurological disorders, and that rs148019457 tags a specific haplotype with 3 CAA interruptions within expanded ATXN2 CAG repeats in individuals of European ancestry. These results have implications for the development of precision genomic medicine for neurological disorders, and the tag SNP may help identify those with interruptions from existing population genotyping data.

ATXN2

Repeated drought induces a reproducible DNA methylation response associated with gene expression in Quercus lobata.

UNLABELLED: Long-lived trees must continually adjust to environmental change and face sustained climatic shifts over their lifetimes. One increasingly important challenge is the rising frequency of drought caused by climate change. Environmentally responsive DNA methylation is widespread in plants, but whether it contributes to gene expression during environmental stress remains unclear, particularly in long-lived trees. Here, we integrated long read methylomes and transcriptomes from valley oak ( Quercus lobata ) seedlings exposed to repeated drought and well-watered treatments. Repeated drought induced a reproducible DNA methylation response that repeatedly targeted the same genomic regions despite turnover of individual methylated sites. These repeatedly targeted regions were transposable elements (TEs) located near genes. Genes adjacent to CHH-methylated TEs were enriched for core drought-response pathways, including abscisic acid signaling, osmotic adjustment and cell-wall remodeling, and remained transcriptionally activated under drought. However, higher CHH methylation levels were associated with progressively smaller transcriptional responses, suggesting that environmentally responsive DNA methylation influences how strongly drought- response genes are activated rather than simply switching them on or off. At the same time, greater CHH methylation was associated with continued repression of nearby TEs, suggesting that this response may simultaneously regulate gene activity while maintaining genome stability. Together, these findings identify a reproducible genome- regulatory response associated with repeated environmental stress in a long-lived tree. By repeatedly targeting the same genomic regions despite turnover of individual sites, this response provides a framework for how long-lived trees repeatedly adjust gene expression while maintaining genome stability during environmental change. SIGNIFICANCE STATEMENT: Plants cannot escape environmental change, and trees must repeatedly respond to stresses, such as drought, over lifetimes spanning decades to centuries. Yet little is known about the molecular mechanisms that make this remarkable resilience possible. Using a widespread California oak, we show that repeated drought repeatedly induced the same DNA methylation pattern in the same parts of the genome, even though the differentially methylated individual sites changed between drought events. This pattern was linked to how strongly drought-response genes were activated, suggesting that trees repeatedly deploy the same molecular program to respond to environmental stress. Our findings provide a new framework for understanding how long-lived organisms repeatedly adjust to changing climates.

Journal Article

Accurate detection of tandem repeats exposes ubiquitous reuse of biological sequences.

Tandem repetition is one of the major processes underlying genome evolution and phenotypic diversification. While newly formed tandem repeats are often easy to identify, it is more challenging to detect repeat copies as they diverge over evolutionary timescales. Existing programs for finding tandem repeats return markedly different results, and it is unclear which predictions are more correct and how much room remains for improvement. Here, we introduce DetectRepeats, a new method that uses empirical information about structural repeats to improve the accuracy of repeat detection. We show that DetectRepeats advances the state-of-the-art by finding highly divergent repeats with relatively few false positive detections. We apply DetectRepeats to genomes across the tree of life to discover an enrichment of detectable tandem repeats within different genes, genome regions, and taxa. Furthermore, we use phylogenetic reconciliation to determine that some tandem repeats continue to evolve through intra-repeat unit replacement. In this manner, tandem repeats serve as a renewable genetic resource offering a bountiful source of alternative genetic material. Our work unlocks the confident detection of ancient tandem repeats, opening a doorway to future discoveries. DetectRepeats is part of the DECIPHER package for the R programming language and available via Bioconductor.

Tandem Repeat Sequences

Characterising the motif composition and allele length distribution of ZFHX3 GGC repeat expansions in amyotrophic lateral sclerosis.

A pathogenic GGC repeat expansion in zinc finger homeobox 3 (ZFHX3), encoding a pure polyglycine (polyG) tract, causes spinocerebellar ataxia type 4 (SCA4). Intermediate expansions of other SCA loci have been implicated in amyotrophic lateral sclerosis (ALS), while repeat motif composition is recognised to influence pathogenicity in neurodegenerative diseases. Given the genetic pleiotropy between ALS and SCA, we evaluated whether ZFHX3 GGC expansions are associated with ALS and characterised repeat motif composition. ZFHX3 GGC repeat sizes were genotyped using ExpansionHunter in short-read whole-genome sequencing data from ALS cases and healthy controls of European ancestry. Repeat sizes were visually inspected using REViewer, and motif configurations were manually derived from a subset. Receiver operating characteristic analysis and Youden's J statistic identified a candidate repeat size threshold. Logistic regression tested associations of repeat length and motif composition with ALS, while regression models assessed clinical phenotypes. Across 5785 ALS cases and 7982 controls, no association was observed between ZFHX3 expansions and ALS risk. Longer alleles showed a nominal association with later disease onset, however this did not remain significant after Bonferroni correction. Among 802 ALS cases and 800 controls, 50 distinct motif compositions were identified, including 11 encoding pure polyG tracts characteristic of pathogenic SCA4 expansions; none were associated with ALS. Although no association with ALS was observed, this study established the dynamic nature of ZFHX3 repeat motif composition and configuration. Variation within and between repeat sizes, including pure polyG repeats, supports consideration of motif composition alongside allele length when evaluating neurodegenerative disease risk.

Journal Article

The Spatial and Temporal Repeatability of Genomic Responses to Natural Selection as Demonstrated in Stickleback Populations Experiencing Highly Dynamic Environments.

The evolution of genotypic parallelism under shared environmental conditions provides strong evidence for the role of natural selection. However, analyses typically examine genomic signatures of selection long after the putative selection event and only assess the repeatability of responses across spatial population replicates. This impedes our ability to attribute a particular response to a given selection pressure and to distinguish non-parallel responses caused by stochastic processes from those caused by local selection. As such, the consistency of natural selection over space and time is unknown, and the role of persistent local selection pressures is unclear. Here, we leveraged the natural bar-built estuary system of Santa Cruz, California, to examine the repeatability of seasonal genomic change in threespine stickleback (Gasterosteus aculeatus) over space and time. By comparing allele-frequency shifts that are shared across locations (spatial repeatability) with those that are shared across years within locations (temporal repeatability), we identified both spatially shared and local components of putative selection. We found that repeated seasonal outlier responses occurred more often than expected under a neutral null model. Although repeatability declined as the number of estuaries sharing an outlier increased, enrichment above neutral expectations increased with broader spatial sharing, particularly for outliers repeated across both years. While the precise outlier SNPs varied across years, estuary-specific patterns of responses were broadly consistent, suggesting an important role for local conditions. Together, our findings show that temporal sampling can reveal components of putative selection that would be missed from spatial comparisons alone. More broadly, they highlight the importance of examining repeatability over both space and time to understand the parallel and non-parallel components of adaptive genomic change.

Animals

First clinical diagnosis of FAME3 via commercial Long-Read sequencing reveals mosaic repeat expansion in MARCHF6 gene.

Familial Adult Myoclonic Epilepsy type 3 (FAME3) is a rare autosomal dominant disorder characterized by cortical tremor and epilepsy, caused by a noncoding pentanucleotide repeat expansion (TTTTA/TTTCA)n in the MARCHF6 gene. Conventional genetic testing often fails to detect this expansion due to its repetitive structure and intronic location. We evaluated a 61-year-old woman with refractory myoclonic and generalized tonic-clonic seizures, whose prior genetic testing-including exome and genome sequencing-was non-diagnostic. Using PacBio HiFi long-read whole-genome sequencing and the tandem repeat genotyping tool TRGT, we identified a pathogenic MARCHF6 intronic expansion. The proband harbored one allele with 15 TTTTA repeats and a second allele with a compound expansion of 661 TTTTA and 12 TTTCA repeats. Three affected relatives shared similarly expanded alleles, but with increasing repeat size in the latter generations. Importantly, analysis using TRGT-instability revealed repeat mosaicism in all affected individuals, reflected by variability in motif counts across individual sequencing reads. This somatic heterogeneity may contribute to the phenotypic penetrance, variable expressivity and pleiotropism seen in FAME3 disease expression. To our knowledge, this is the first clinical diagnosis of FAME3 using a commercially available long-read sequencing platform, underscoring its diagnostic utility in resolving complex repeat expansion disorders and uncovering biologically relevant mosaicism.

Humans

AniAnn's: alignment-free annotation of tandem repeat arrays using fast average nucleotide identity estimates.

MOTIVATION: Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of tandem repeat arrays, are limited in their ability to annotate divergent or novel sequences. RESULTS: In this work, we introduce AniAnn's, an algorithm for annotating large blocks of tandemly repeating DNAs. AniAnn's exploits the high Average Nucleotide Identity (ANI) shared between repeat units of the same array to quickly and accurately infer the boundaries of such arrays. We show that AniAnn's improves the annotation of satellites and other tandem repeats within a variety of plant and animal genomes, while requiring only a fraction of the runtime compared to previous approaches. We conclude by exploring several use cases of AniAnn's as a lightweight method for masking repeats prior to whole-genome alignment as well as the de novo annotation and classification of satellite repeats. AVAILABILITY: AniAnn's is open source software and available at github.com/marbl/anianns.

Algorithms

The CGG triplet repeat binding protein 1 counteracts R-loop induced transcription-replication stress.

The CGG triplet repeat binding protein 1 (CGGBP1) binds to CGG repeats and has several important cellular functions, but how this DNA sequence-specific binding factor affects transcription and replication processes is an open question. Here, we show that CGGBP1 binds human gene promoters containing short (<&#x2009;5) CGG-repeat tracts prone to R-loop formation. Loss of CGGBP1 leads to deregulated transcription, transcription-replication-conflicts (TRCs) and accumulation of Serine-5 phosphorylated RNA polymerase II (RNAPII), indicative of promoter-proximal stalling and a defect in transcription elongation. Consistently, an episomal CGG-repeat-containing model locus as well as endogenous genes show deregulated transcription, R-loop accumulation and increased RNAPII chromatin occupancy in CGGBP1-depleted cells. We identify the DEAD-box RNA:DNA helicases DDX41 and DHX15 as interaction partners specifically recruited by CGGBP1. Co-depletion experiments show that DDX41 and CGGBP1 work in the same pathway to unwind R-loops and avoid TRCs. Together, our work shows that short trinucleotide repeats are a source of genome-destabilizing secondary structures, and cells rely on specific DNA-binding factors to maintain proper transcription and replication coordination at short CGG repeats.

Humans

tidk: a toolkit to rapidly identify telomeric repeats from genomic datasets.

SUMMARY: "tidk" (short for telomere identification toolkit) uses a simple, fast algorithm to scan long DNA reads for the presence of short tandemly repeated DNA in runs, and to aggregate them based on canonical DNA string representation. These are telomeric repeat candidates. Our algorithm is shown to be accurate in genomes for which the telomeric repeat unit is known and is tested across a wide variety of newly assembled genomes to uncover new telomeric repeat units. Tools are provided to identify telomeric repeats de novo, scan genomes for known telomeric repeats, and to visualize telomeric repeats on the assembly. "tidk" is implemented in Rust and is available as a command line tool which can be compiled using the Rust toolchain or downloaded as a binary from bioconda. AVAILABILITY AND IMPLEMENTATION: The "tidk" Rust crate is freely available under the MIT license (https://crates.io/crates/tidk), and the source code is available at https://github.com/tolkit/telomeric-identifier.

Telomere

Parallel Analysis of Repeat Expansions: An Updated Clinical Nanopore Cas9-Targeted Sequencing Workflow for Nanopore R10 Flow Cells.

Hereditary ataxias, caused by expansions of short tandem repeats, are difficult to diagnose using traditional PCR and Southern blot methods, which struggle to detect complex repeat expansions and cannot assess repeat interruptions or methylation. An updated Clinical Nanopore Cas9-Targeted Sequencing workflow is presented for analyzing repeat expansions, now compatible with the Oxford Nanopore Technologies R10 flow cell. The workflow incorporates the Oxford Nanopore Technologies wf-human-variation Epi2Me workflow, including the Straglr tool to analyze base-called reads, ensuring compatibility with past, current, and future sequencing chemistries. It expands the number of genes analyzed from 10 to 27 and introduces new gene panels for ataxia, myopathy, neurodegeneration, and amyotrophic lateral sclerosis/motor neuron disease. Validated with Coriell reference and clinical samples, this method improves the analysis of pathogenic repeat expansions, providing deeper insights into repeat structures while addressing the limitations of traditional approaches. In this work, the use of multiplexing, Flongle flow cells, and single-gene targeting were explored as alternatives to panel-based approaches in the Clinical Nanopore Cas9-Targeted Sequencing workflow, finding that only single-gene targeting provides compatibility and reliable performance.

Journal Article

Financial incentives and social messaging for repeat SARS-CoV-2 antibody testing among the underserved: A randomized trial.

Financial incentives may influence health behavior beyond their expected monetary value, and their effectiveness may depend on how the behavior is framed. Behavioral theories of decision making suggest that individuals may value protection against small-stakes losses more than expected utility predicts, while theories of family-centered health behavior suggest that messages emphasizing benefits to family members may strengthen participation in preventive health activities. We tested these ideas in a 2&#xd7;2 factorial randomized trial involving 625 households recruited from a Federally Qualified Health Center serving low-income Latino/Hispanic communities. Participants completed repeat SARS-CoV-2 antibody testing. The trial crossed two messaging strategies (Family vs. Personal) with two incentive structures (Loss Protection vs. Lottery) that offered equivalent expected monetary value. Family Messaging emphasized protecting one's family from COVID-19, whereas Personal Messaging emphasized protecting oneself. Loss Protection allowed participants to secure an at-risk reward through repeat testing, whereas the Lottery condition offered a chance of a large reward. Repeat testing was approximately 8 percentage points higher under Family Messaging and 7 percentage points higher under Loss Protection. Baseline trust in medical providers, financial barriers to vaccination, and risk aversion were associated with initial testing, whereas household characteristics were not associated with repeat testing. Incentive design may matter beyond expected monetary value and that framing health behaviors in terms of family welfare may increase participation in repeated healthy activities. Broadly, the results support behavioral theories emphasizing loss aversion, anticipated regret, and family-centered motivations, and suggest practical approaches for improving engagement in repeat health behaviors. CLINICALTRIALS.GOV REGISTRATION NUMBER:: NCT01901624.

Adult

Prevalence of intronic repeat expansions in the RFC1 gene in Polish patients with cerebellar syndrome.

Cerebellar ataxia with neuropathy and vestibular areflexia syndrome (CANVAS) is a recessively inherited neurodegenerative ataxic disorder, which has been associated with intronic biallelic repeat expansions in the RFC1 gene. Our objective was to assess retrospectively the prevalence of CANVAS in Polish population. We screened 2523 Polish patients in whom other repeat expansions were excluded. To determine the repeat expansions in the RFC1 gene in patients, we performed RFC1-flanking PCR and repeat primed PCR (RP-PCR) and to measure the size of the expansion we used Southern blotting and optical genome mapping to compare the results. We have observed the biallelic pathogenic motif/unit AAGGG expansions in 4.6% and expansions of non-pathogenic motifs AAAAG, AAAGG in 25% patients of our studied population. This is the first large-scale cohort study that confirms the relatively frequent occurrence of the CANVAS in Polish population. To increase the current diagnostics of late-onset ataxias within an unexplained molecular background, we suggest involving the RFC1 repeat expansions analysis to the routine diagnostic workflow.

Humans

Sequencing the orthologs of human autosomal forensic short tandem repeats provides individual- and species-level identification in African great apes.

BACKGROUND: Great apes are a global conservation concern, with anthropogenic pressures threatening their survival. Genetic analysis can be used to assess the effects of reduced population sizes and the effectiveness of conservation measures. In humans, autosomal short tandem repeats (aSTRs) are widely used in population genetics and for forensic individual identification and kinship testing. Traditionally, genotyping is length-based via capillary electrophoresis (CE), but there is an increasing move to direct analysis by massively parallel sequencing (MPS). An example is the ForenSeq DNA Signature Prep Kit, which amplifies multiple loci including 27 aSTRs, prior to sequencing via Illumina technology. Here we assess the applicability of this human-based kit in African great apes. We ask whether cross-species genotyping of the orthologs of these loci can provide both individual and (sub)species identification. RESULTS: The ForenSeq kit was used to amplify and sequence aSTRs in 52 individuals (14 chimpanzees; 4 bonobos; 16 western lowland, 6 eastern lowland, and 12 mountain gorillas). The orthologs of 24/27 human aSTRs amplified across species, and a core set of thirteen loci could be genotyped in all individuals. Genotypes were individually and (sub)species identifying. Both allelic diversity and the power to discriminate (sub)species were greater when considering STR sequences rather than allele lengths. Comparing human and African great-ape STR sequences with an orangutan outgroup showed general conservation of repeat types and allele size ranges. Variation in repeat array structures and a weak relationship with the known phylogeny suggests stochastic origins of mutations giving rise to diverse imperfect repeat arrays. Interruptions within long repeat arrays in African great apes do not appear to reduce allelic diversity. CONCLUSIONS: Orthologs of most human aSTRs in the ForenSeq DNA Signature Prep Kit can be analysed in African great apes. Primer redesign would reduce observed variability in amplification across some loci. MPS of the orthologs of human loci provides better resolution for both individual and (sub)species identification in great apes than standard CE-based approaches, and has the further advantage that there is no need to limit the number and size ranges of analysed loci.

Animals

GeomeTRe: accurate calculation of geometrical descriptors of tandem repeat proteins.

MOTIVATION: Structured tandem repeat proteins (STRPs) are characterized by preserved structural motifs arranged in a modular way. The structural and functional diversity of STRPs makes them particularly important for studying evolution and novel structure-function relationships, and ultimately for designing new synthetic proteins with specific functions. One crucial aspect of their classification is the estimation of geometrical parameters, which can provide better insight into their properties and the relationship between the spatial arrangement of repeated units and protein function. Calculating geometric descriptors for STRPs is challenging because naturally occurring repeats are not "perfect" and often contain insertions and deletions. Existing tools for predicting structural symmetry work well on simple cases but often fail for most natural proteins. RESULTS: Here, we present GeomeTRe, an algorithm that calculates geometrical descriptors such as curvature (yaw), twist (roll), and pitch for a protein structure with known repeat unit positions. The algorithm simulates the movement of consecutive units, identifies rotational axes, and calculates the corresponding Tait-Bryan angles. GeomeTRe's parameters can enhance STRP annotation and classification by identifying variations in geometric arrangements among different functional groups. The package is fast and suitable for processing large protein structure datasets when repeat region information (e.g. from RepeatsDB) is available. AVAILABILITY AND IMPLEMENTATION: GeomeTRe is available as a Python package; source code and documentation can be found at https://github.com/BioComputingUP/GeomeTRe.

Algorithms