Search PubMedSearch

SEARCH · Search PubMed

Results for “genome assembly error”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

24 records · Page 2Linked to original sources

CREAT: A CRISPR-Based Genome Trimming Strategy for Systematic Identification of Dispensable Regions and Rapid Genome Reduction.

The construction of minimal-genome microbes offers an ideal platform for understanding fundamental biological processes and synthetic biology, yet the research is hindered by incomplete lists of essential genes in microbes and by multiple rounds of genome trimming with a trial-and-error nature. To address this, we introduce CREAT (CRISPR-based genome trimming with a multi-homology-arm template)-a streamlined approach that integrates CRISPR-targeted genome cleavage and homology arm walking to classify essential from non-essential genomic subregions, thus providing the basis for predicting essential genes in a given organism. These essential genes were then assembled into synthetic gene cassettes for one-step replacement of the targeted non-deletable genomic regions for further genome trimming. Eight consecutive rounds of CREAT genome trimming achieved a 20.8% reduction in genome size in Saccharolobus islandicus. Furthermore, Cas9-based CREAT genome trimming was developed for Bacillus subtilis and Escherichia coli, with efficiency greatly enhanced by the λ-Red recombinase in the latter. Together, this iterative application of CREAT provides a scalable and generally applicable strategy for rapidly constructing minimal genomes across diverse microorganisms.

CRISPR-Cas Systems

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis

A Restriction-Free Cloning Approach for Molecular Engineering of Plasmids.

Molecular cloning by PCR amplification using a highly processive, high-fidelity DNA polymerase represents a robust and versatile technique for the precise manipulation of nucleic acid sequences. This approach enables the insertion, replacement, or modification of specific DNA fragments within a cloning vector, thereby generating an accurate copy of a gene or viral segment for downstream applications, such as protein expression, site-directed mutagenesis, and structural or functional analyses. The use of processive, high-fidelity polymerases significantly reduces the occurrence of base substitution errors, ensuring sequence integrity throughout the amplification process. Traditionally, restriction enzymes have been employed to facilitate directional cloning; however, alternative methods allow for mutagenesis without the need for unique and specific restriction sites and can be applied to virtually any cloning or seamless DNA assembly strategy. In this chapter, we describe a restriction enzyme-free and ligation-free PCR-based protocol widely applicable to any circular vector. This method enables targeted mutagenesis of the chikungunya virus (CHIKV) genome, offering a fast, efficient, and reliable strategy for generating mutant constructs suitable for virological and molecular studies.

Cloning, Molecular

NanoFilter: enhancing phasing performance by utilizing highly consistent INDELs and SNVs in nanopore sequencing.

MOTIVATION: Nanopore sequencing data offer longer reads compared to other technologies, which is beneficial for phasing and genome assembly. INDELs provide valuable haplotype information and have significant potential to improve phasing performance. However, accurately identifying INDELs with variant callers is challenging, and incorporating INDELs into phasing remains a complex task. To address these issues, we developed NanoFilter, a novel filtering strategy designed to filter out INDELs that contain wrong phasing information based on their consistency. RESULTS: Our assessment using Nanopore R10 simplex data shows that filtering out low-consistency INDELs increases their precision from 88.3% to 98.8%, nearly matching the precision of SNVs. In the phasing results of Margin, incorporating these filtered INDELs leads to a 12.77% increase in N50 length and fewer switch errors. Furthermore, we found that SNVs filtered by NanoFilter will enhance assembly performance. When NanoFilter is integrated into the HapDup assembly pipeline, NanoFilter reduces the Hamming error rate and increases N50 length by 7.8%. AVAILABILITY AND IMPLEMENTATION: NanoFilter is available at https://github.com/Chenshanming-repo/NanoFilter (DOI: 10.5281/zenodo.16777826) and HapDup-NanoFilter is available at https://github.com/Chenshanming-repo/HapDup-NanoFilter (DOI: 10.5281/zenodo.16777890).

Nanopore Sequencing

Complete chromosome 21 centromere sequencing of families with Down syndrome reveals centromere size asymmetry.

Down syndrome, the most common form of human intellectual disability, is caused by nondisjunction and chromosome 21 trisomy (T21). Small centromeres have been hypothesized to contribute to its aetiology and studies on mammals suggest that larger centromeres are more efficiently transmitted, yet complete sequencing of chromosome 21 (chr21) centromeres has been particularly challenging. Using long-read sequencing, we sequenced and assembled the centromeres from eight families that include a child with free T21 (1 trio, 6 child-mother duos, and 1 singleton) all resulting from maternal meiosis I errors. Two of these families carry the smallest chr21 centromeres (143 and 181 kbp) observed in female individuals to date, exhibiting a ~10.7- and ~19.4-fold centromeric α-satellite higher-order repeat array size difference between the maternally inherited homologs, respectively. In both cases, the longer centromere harbors a poorly defined centromere dip region, marked by DNA hypomethylation, in the proband but not in the mother. A comparison of all proband chr21 centromeres (n=24) to those of controls (n=261) shows that small centromeres are not enriched in families with T21 (p-value=0.73); contrarily, chr21 extreme centromere size asymmetry (>10-fold) is unique of T21 (p-value=0.003), suggesting that this feature may represent a genetic risk factor for a subset of families with free T21. Additionally, phylogenetic reconstruction reveals that human chr21 has been particularly prone to such variation with some of the biggest size differences occurring over the last ~17 thousand years of human evolution.

Down syndrome

Limitations of encapsidation of recombinant self-complementary adeno-associated viral genomes in different serotype capsids and their quantitation.

We previously reported that self-complementary adeno-associated virus (scAAV) type 2 genomes of up to 3.3 kb can be successfully encapsidated into AAV2 serotype capsids. Here we report that such oversized AAV2 genomes fail to undergo packaging in other AAV serotype capsids, such as AAV1, AAV3, AAV6, and AAV8, as determined by Southern blot analyses of the vector genomes, although hybridization signals on quantitative DNA slot-blots could still be obtained. Recently, it has been reported that quantitative real-time PCR assays may result in substantial differences in determining titers of scAAV vectors depending on the distance between the primer sets and the terminal hairpin structure in the scAAV genomes. We also observed that the vector titers determined by the standard DNA slot-blot assays were highly dependent on the specific probe being used, with probes hybridizing to the ends of viral genomes being significantly overrepresented compared with the probes hybridizing close to the middle of the viral genomes. These differences among various probes were not observed using Southern blot assays. This overestimation of titer is a systemic error during scAAV genome quantification, regardless of viral genome sequences and capsid serotypes. Furthermore, different serotypes capsid and modification of capsid sequence may affect the ability of packaging intact, full-length AAV genomes. Although the discrepancy is modest with wild-type serotype capsid and short viral genomes, the measured titer could be as much as fivefold different with capsid mutant vectors and large genomes. Thus, based on our data, we suggest that Southern blot analyses should be performed routinely to more accurately determine the titers of recombinant AAV vectors. At the very least, the use of probes/primers hybridizing close to the mutant inverted terminal repeat in scAAV genomes is recommended to avoid possible overestimation of vector titers.

Blotting, Southern