Search PubMedSearch

SEARCH · Search PubMed

Results for “DNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Primate repetitive DNAs: evidence for new satellite DNAs and similarities in non-satellite repetitive DNA sequence properties.

Repetitious DNA sequences have been isolated from a number of the primates in in both Suborders Anthropoidea and Prosimii by hydroxy-apatite chromatography at a Cot of 10. In addition to finding previously unreported possible AT-rich satellite DNAs in Orangutan, Gibbon, Rhesus and Slow Loris a clear similarity to human DNA was found in the nonsatellite repetitious DNA sequence properties of the primates in the Suborder Anthropoidea. This is based on the presence of the hydroxyapatitie isolated 1.703 and 1.714 g/cm3 DNA families in CsCl gradients in the analytical ultracentrifuge following renaturation and extensive DNA hyperpolymer network formation. Within the superfamily Hominoidea the amount of the 1.714 g/cm3 DNA family was greater than that of the 1.703 g/cm3 DNA family while the reverse situation was true within the Superfamily Cercopithecoidea. The orangutan 1.703 and 1.714 g/cm3 DNA families were shown to exhibit the same differential reassociation behavior demonstrated previously in human DNA (Marx et al., 1976a). These data are interpreted as preliminary evidence for a similar sequence organization in the Order Primates Suborder Anthropoidea.

Animals

Embed-Search-Align: DNA sequence alignment using Transformer models.

MOTIVATION: DNA sequence alignment, an important genomic task, involves assigning short DNA reads to the most probable locations on an extensive reference genome. Conventional methods tackle this challenge in two steps: genome indexing followed by efficient search to locate likely positions for given reads. Building on the success of Large Language Models in encoding text into embeddings, where the distance metric captures semantic similarity, recent efforts have encoded DNA sequences into vectors using Transformers and have shown promising results in tasks involving classification of short DNA sequences. Performance at sequence classification tasks does not, however, guarantee sequence alignment, where it is necessary to conduct a genome-wide search to align every read successfully, a significantly longer-range task by comparison. RESULTS: We bridge this gap by developing a "Embed-Search-Align" (ESA) framework, where a novel Reference-Free DNA Embedding (RDE) Transformer model generates vector embeddings of reads and fragments of the reference in a shared vector space; read-fragment distance metric is then used as a surrogate for sequence similarity. ESA introduces: (i) Contrastive loss for self-supervised training of DNA sequence representations, facilitating rich reference-free, sequence-level embeddings, and (ii) a DNA vector store to enable search across fragments on a global scale. RDE is 99% accurate when aligning 250-length reads onto a human reference genome of 3 gigabases (single-haploid), rivaling conventional algorithmic sequence alignment methods such as Bowtie and BWA-Mem. RDE far exceeds the performance of six recent DNA-Transformer model baselines such as Nucleotide Transformer, Hyena-DNA, and shows task transfer across chromosomes and species. AVAILABILITY AND IMPLEMENTATION: Please see https://anonymous.4open.science/r/dna2vec-7E4E/readme.md.

Sequence Analysis, DNA

Use of 3D chaos game representation to quantify DNA sequence similarity with applications for hierarchical clustering.

A 3D chaos game is shown to be a useful way for encoding DNA sequences. Since matching subsequences in DNA converge in space in 3D chaos game encoding, a DNA sequence's 3D chaos game representation can be used to compare DNA sequences without prior alignment and without truncating or padding any of the sequences. Two proposed methods inspired by shape-similarity comparison techniques show that this form of encoding can perform as well as alignment-based techniques for building phylogenetic trees. The first method uses the volume overlap of intersecting spheres and the second uses shape signatures by summarizing the coordinates, oriented angles, and oriented distances of the 3D chaos game trajectory. The methods are tested using: (1) the first exon of the beta-globin gene for 11 species, (2) mitochondrial DNA from four groups of primates, and (3) a set of synthetic DNA sequences. Simulations show that the proposed methods produce distances that reflect the number of mutation events; additionally, on average, distances resulting from deletion mutations are comparable to those produced by substitution mutations.

Animals

esloco: simulation-based estimation of local coverage in long-read DNA sequencing.

SUMMARY: Long-read DNA sequencing is increasingly applied for whole-genome studies, yet experimental planning often lacks reliable estimates of target region coverage, leading to costly and time-consuming pilot studies and replicates. We present esloco, a Monte Carlo-based simulation framework for estimating local coverage in long-read sequencing experiments, including scenarios with unknown target regions (e.g. viral integration, CRISPR-Cas9) or PCR-free designs (e.g. base modifications). By modeling coverage as a function of sequencing depth and read length distribution, esloco enables informed predictions of local sequencing outcomes. Benchmarking across a 45-gene panel demonstrated close agreement with empirical data, underscoring the framework's reliability. AVAILABILITY AND IMPLEMENTATION: esloco is a Python package available on PyPI (https://pypi.org/project/esloco/), GitHub (https://github.com/aweich/esloco), and Zenodo (https://doi.org/10.5281/zenodo.17776161).

Sequence Analysis, DNA

Prediction and functional interpretation of inter-chromosomal genome architecture from DNA sequence with TwinC.

Three-dimensional nuclear DNA architecture comprises well-studied intra-chromosomal (cis) folding and less characterized inter-chromosomal (trans) interfaces. Current predictive models of 3D genome folding can effectively infer pairwise cis-chromatin interactions from the primary DNA sequence but generally ignore trans contacts. There is an unmet need for robust models of trans-genome organization that provide insights into their underlying principles and functional relevance. We present TwinC, an interpretable convolutional neural network model that reliably predicts trans contacts measurable through proximity ligation-dependent (in situ and intact Hi-C) and independent (DNA SPRITE) genome-wide chromatin conformation assays. . TwinC uses a paired sequence design from replicate Hi-C experiments to learn single base pair relevance in trans interactions across two stretches of DNA. The method achieves high predictive accuracy (AUROC=0.80) on a cross-chromosomal test set from in situ and intact Hi-C experiments in heart tissue. Furthermore, we train TwinC using in situ Hi-C data from the widely used GM12878 cell line and validate its performance with orthogonal DNA SPRITE assay in the same cell type. Mechanistically, the neural network learns the importance of compartments, chromatin accessibility, clustered transcription factor binding and G-quadruplexes in forming trans contacts. In summary, TwinC models and interprets trans genome architecture, shedding light on this poorly understood aspect of gene regulation.

Journal Article

Evaluating the analytical validity of circulating tumor DNA sequencing assays for precision oncology.

Circulating tumor DNA (ctDNA) sequencing is being rapidly adopted in precision oncology, but the accuracy, sensitivity and reproducibility of ctDNA assays is poorly understood. Here we report the findings of a multi-site, cross-platform evaluation of the analytical performance of five industry-leading ctDNA assays. We evaluated each stage of the ctDNA sequencing workflow with simulations, synthetic DNA spike-in experiments and proficiency testing on standardized, cell-line-derived reference samples. Above 0.5% variant allele frequency, ctDNA mutations were detected with high sensitivity, precision and reproducibility by all five assays, whereas, below this limit, detection became unreliable and varied widely between assays, especially when input material was limited. Missed mutations (false negatives) were more common than erroneous candidates (false positives), indicating that the reliable sampling of rare ctDNA fragments is the key challenge for ctDNA assays. This comprehensive evaluation of the analytical performance of ctDNA assays serves to inform best practice guidelines and provides a resource for precision oncology.

Circulating Tumor DNA

ELYS associates with distinct DNA sequence environments during post-mitotic nuclear pore reassembly.

Nuclear pore complexes (NPCs) contribute to genome organization and cell identity, yet how post-mitotic NPC assembly is coordinated with chromatin architecture remains unclear. Here, we show that the nucleoporin ELYS preferentially associates with chromatin regions displaying distinct intrinsic DNA sequence features that are not explained by the repressive histone marks examined here. ELYS-bound regions are enriched for AT-rich sequences, whereas ELYS binding at super-enhancer-associated loci shift toward GC-rich sequence composition, revealing distinct sequence environments. These findings indicate that ELYS localization is associated with distinct intrinsic DNA sequence features and suggest a mechanism by which nuclear pore-associated architecture restores transcriptional programs after mitosis.

Journal Article

Accumulation of numerous cellular T-DNA sequences in the genus Diospyros by multiple rounds of natural transformation.

Horizontal gene transfer (HGT) is an important phenomenon in the evolutionary history of plants. Natural transformation by Agrobacterium is a special case of HGT and leads to the insertion of cellular T-DNA (cT-DNA) sequences, for example, in Diospyros lotus. The genus Diospyros contains about 795 species with economically important members, like different types of persimmon (D. kaki, D. lotus, and D. virginiana) and ebony (e.g., D. ebenum). Whole genome sequences (WGS) from D. kaki, D. oleifera, D. lotus, and D. virginiana were investigated for cT-DNAs. These four species belong to one clade and contain 15 different cT-DNAs (DiTA to DiTO). The hexaploid species D. kaki cv. "Xiaoguo-tianshi" contains seven types of cT-DNA (DiTA to DiTG) on 27 of 42 homeologs, adding up to 628 kb of cT-DNA. Five of these seven cT-DNAs are non-fixed, as shown by empty chromosomal insertion sites. The evolutionary history of the Diospyros cT-DNAs was reconstructed using the divergence of their inverted repeats. Insert age varied from 3 to 12 million years. Partial cT-DNA sequences were detected in 35 additional species from five Diospyros clades. Our data highlight the unexpectedly large scale of natural Agrobacterium transformation in Diospyros and demonstrate the necessity of whole genome approaches for studies on the origin and evolution of cT-DNAs.

Diospyros

MCALIGN: stochastic alignment of noncoding DNA sequences based on an evolutionary model of sequence evolution.

A method is described for performing global alignment of noncoding DNA sequences based on an evolutionary model parameterized by the frequency distribution of lengths of insertion/deletion events (indels) and their rate relative to nucleotide substitutions. A stochastic hill-climbing algorithm is used to search for the most probable alignment between a pair of sequences or three sequences of known phylogenetic relationship. The performance of the procedure, parameterized according to the empirical distribution of indel lengths in noncoding DNA of Drosophila species, is investigated by simulation. We show that there is excellent agreement between true and estimated alignments over a wide range of sequence divergences, and that the method outperforms other available alignment methods.

Algorithms

sedimix: a workflow for the analysis of hominin nuclear DNA sequences from sediments.

SUMMARY: Sediment DNA-the recovery of genetic material from archaeological sediments-is an exciting new frontier in ancient DNA research, offering the potential to study individuals at a given archaeological site without destructive sampling. In recent years, several studies have demonstrated the promise of this approach by extracting hominin DNA from prehistoric sediments, including those dating back to the Middle or Late Pleistocene. However, a lack of open-source workflows for analysis of hominin sediment DNA samples poses a challenge for data processing and reproducibility of findings across studies. Here, we introduce a snakemake workflow, sedimix, for processing genomic sequences from archaeological sediment DNA samples to identify hominin sequences and generate relevant summary statistics to assess the reliability of the pipeline. By performing simulations and comparing our results to two published studies with human DNA from ∼25,000 years ago (including shotgun data from a sediment sample and capture data from touch DNA recovered from a deer tooth pendant) we demonstrate that sedimix yields accurate and reliable inferences. sedimix offers a reliable and adaptable framework to aid in the analysis of sediment DNA datasets and improve reproducibility across studies. AVAILABILITY AND IMPLEMENTATION: sedimix is available as an open-source software with the associated code, example data, and user manual with installation instructions available at https://github.com/jierui-cell/sedimix. A permanent archived version of this release is available via Zenodo: https://doi.org/10.5281/zenodo.17244854.

Animals

RadiSeq: a single- and bulk-cell whole-genome DNA sequencing simulator for radiation-damaged cell models.

Objective.To build and validate a simulation framework to perform single-cell and bulk-cell whole genome sequencing simulation of radiation-exposed Monte Carlo (MC) cell models to assist radiation genomics studies.Approach.Sequencing the genomes of radiation-damaged cells can provide useful insight into radiation action for radiobiology research. However, carrying out post-irradiation sequencing experiments can often be challenging, expensive, and time-consuming. Although computational simulations have the potential to provide solutions to these experimental challenges, and aid in designing optimal experiments, the absence of tools currently limits such application. MC toolkits exist to simulate radiation exposures of cell models but there are no tools to simulate single- and bulk-cell sequencing of cell models containing radiation-damaged DNA. Therefore, we aimed to develop a MC simulation framework to address this gap by designing a tool capable of simulating sequencing processes for radiation-damaged cells. Main results.We developed RadiSeq-a multi-threaded whole-genome DNA sequencing simulator written in C++. RadiSeq can be used to simulate Illumina sequencing of radiation-damaged cell models produced by MC simulations. RadiSeq has been validated through comparative analysis, where simulated data were matched against experimentally obtained data, demonstrating reasonable agreement between the two. Additionally, it comes with numerous features designed to closely resemble actual whole-genome sequencing. RadiSeq is also highly customizable with a single input parameter file.Significance.RadiSeq enables the research community to perform complex simulations of radiation-exposed DNA sequencing, supporting the optimization, planning, and validation of costly and time-intensive radiation biology experiments. This framework provides a powerful tool for advancing radiation genomics research.

Monte Carlo Method

Sassy: fuzzy searching DNA sequences using SIMD.

MOTIVATION: Approximate string matching (ASM) is the problem of finding all occurrences of a pattern in a text while allowing up to k errors. Many modern methods use seed-chain-extend, which is fast in practice, but does not guarantee finding all matches with ≤k errors. However, applications such as CRISPR off-target detection require exhaustive results. RESULTS: We introduce Sassy, a library and tool for ASM of short patterns in long texts. Sassy splits the text into four parts that are searched in parallel, and uses bitvectors in the text direction rather than the pattern direction. This has complexity O(k⌈n/W⌉) when searching a random text of length n, where W=256 is the SIMD width, and provides significant speedups for small k. Separately, we allow matches of the pattern to extend beyond the text for an overhang cost of, e.g. α=0.5 per character, to find matches near contig or read ends.Sassy is 4× to 15× faster than Edlib for patterns ≤1000 bp, and can search text with a throughput near 2 Gbp/s. Likewise, Sassy is over 100× faster than parasail. We apply Sassy to CRISPR off-target detection by searching 61 guide sequences in a human genome. Sassy is 100× faster than SWOffinder and only slightly slower (for k≤3) than CHOPOFF, for which building its index takes 20 min. Sassy also scales well to larger k, unlike CHOPOFF whose index took over 10 h to build for k=5. AVAILABILITY AND IMPLEMENTATION: Sassy is available as library and binary at https://github.com/RagnarGrootKoerkamp/sassy, and archived at swh:1:dir:e884758dce5777a441bc2799dc8824e563c5f97b.

Sequence Analysis, DNA

A comparison of the circular dichroism spectra of synthetic DNA sequences of the homopurine . homopyrimidine and mixed purine- pyrimidine types.

We have obtained the ultraviolet circular dichroism spectra of two repeating trinucleotide DNAs, poly [d(A-G-G).d(C-C-T)] and poly[d(A-A-G).d(C-T-T)], that have all purines on one strand and all pyrimidines on the other. These spectra, together with spectra of other synthetic polymers, can be combined to give 3 first-neighbor calculations of the spectrum of poly[d(A).d(T)] and 2 first-neighbor calculations of the spectrum of poly [d(G).d(C)]. The results show (1) that first-neighbor calculations utilizing only spectra of homopurine.homopyrimidine DNA sequences are no more accurate than are similar calculations that involve spectra of mixed purine-pyrimidine sequences, demonstrating that double-stranded homopurine.homopyrimidine sequences do not obviously belong to a special class of secondary conformations, and (2) that the wavelength region above 250 nm in the CD spectra of synthetic DNAs is least predictable from first-neighbor equations, probably because this region is especially sensitive to sequence-dependent conformational differences.

Base Sequence

Long-read DNA sequencing resolves a rare case of alloimmune hemolysis mimicking autoimmune hemolysis.

BACKGROUND: Immune hemolytic anemia poses a significant challenge in transfusion medicine, as identification of underlying alloantibodies can be masked by warm and/or cold autoantibodies. This increases the risk of transfusing incompatible blood, which can precipitate or exacerbate hemolysis. Identifying alloantibodies in the presence of autoantibodies remains difficult with standard serologic and genotypic methods, often delaying accurate diagnosis and appropriate transfusion strategies. CASE REPORT: We describe a 63-year-old woman with autoimmune hemolytic anemia who suffered near-fatal hemolysis following transfusion. Despite extensive serologic and genotypic testing, the cause of her hemolytic transfusion reactions remained elusive. Given her clinical course and transfusion history, we hypothesized that her acute hemolytic transfusion reactions could be due to immune sensitization to a high-incidence RBC antigen. Research whole-genome long-read sequencing (LRS) revealed homozygosity for a rare KEL*02N.16 allele, consistent with a rare Ko phenotype, which was validated by Sanger sequencing. Retrospective serologic testing with Ko RBCs further confirmed alloimmunization within the Kell system. CONCLUSION: This case highlights the limitations of conventional serologic and genotypic methods in detecting rare blood group phenotypes, and emphasizes the diagnostic power of long-read sequencing in transfusion medicine. Early molecular testing in complex hemolytic cases can facilitate targeted transfusion strategies, reduce the risk of severe hemolysis, and improve patient outcomes. As sequencing technologies become more accessible, they have the potential to revolutionize blood group typing and alloimmunization risk assessment in clinical practice.

Humans

The N protein of bacteriophage lambda, defined by its DNA sequence, is highly basic.

Nucleotide sequence has been determined for the restriction fragments and cloned DNA from the pL-N-tL1 region of bacteriophage lambda. A unique reading frame for the N gene is defined by the absence of natural nonsense codons and by the presence of seven nonsense codons generated by mutations in N. This reading frame is initiated at two alternative ATG codons, the second of which is probably the in vivo translation start. Reading is stopped at a single TAG codon. The protein coded is therefore 133 or, more probably, 107 amino acids long, rich in lysine, arginine and proline.

Bacteriophage lambda

Characterizing the regulatory logic of transcriptional control at the DNA sequence level by ensembles of thermodynamic models.

MOTIVATION: Understanding how the genome encodes the regulatory logic of transcription is a main challenge of the post-genomic era, and can be overcome with the aid of customized computational tools. RESULTS: We report an automated framework for analyzing an ensemble of fits to data of a thermodynamics-based sequence-level model for transcriptional regulation. The fits are clustered accordingly with their intrinsic regulatory logic. A multiscale analysis enables visualization of quantitative features resulting from the deconvolution of the regulatory profile provided by multiple transcription factors interacting with the locus of a gene. Quantitative experimental data on reporters driven by the whole locus of the even-skipped gene in the blastoderm of Drosophila embryos was used for validating our approach. A few clusters of highly active DNA binding sites within the enhancers collectively modulate even-skipped gene transcription. Analysis of variable enhancers' length shows the importance of bound protein-protein interactions for transcriptional regulation. The interplay between activation and quenching enables function conservation of enhancers despite length variations. AVAILABILITY AND IMPLEMENTATION: The transcription factor level data used for performing the reported study is accessible in the input files in Zenodo and GitHub as well the full code. Additional data from formerly FlyEx database will be available under request.

Thermodynamics