Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Genetic evidence supports demic diffusion of Han culture.

The spread of culture and language in human populations is explained by two alternative models: the demic diffusion model, which involves mass movement of people; and the cultural diffusion model, which refers to cultural impact between populations and involves limited genetic exchange between them. The mechanism of the peopling of Europe has long been debated, a key issue being whether the diffusion of agriculture and language from the Near East was concomitant with a large movement of farmers. Here we show, by systematically analysing Y-chromosome and mitochondrial DNA variation in Han populations, that the pattern of the southward expansion of Han culture is consistent with the demic diffusion model, and that males played a larger role than females in this expansion. The Han people, who all share the same culture and language, exceed 1.16 billion (2000 census), and are by far the largest ethnic group in the world. The expansion process of Han culture is thus of great interest to researchers in many fields.

Agriculture↗

Word organization in coding DNA: a mathematical model.

This article deals with the relationship between vocabulary (total number of distinct oligomers or "words") and text-length (total number of oligomers or "words") for a coding DNA sequence (CDS). For natural human languages, Heaps established a mathematical formula known as Heaps' law, which relates vocabulary to text-length. Our analysis shows that Heaps' law fails to model this relationship for CDSs. Here we develop a mathematical model to establish the relationship between the number of type of words (vocabulary) and the number of words sampled (text-length) for CDSs, when non-overlapping nucleotide strings with the same length are treated as words. We use tangent-hyperbolic function, which captures the saturation property of vocabulary. Based on the parameters of the model, we formulate a mathematical equation, known as "equation of word organization", whose parameters essentially indicate that nucleotide organization of coding sequences are different from one another. We also compare the word organization of CDSs with the random word distribution and conclude that a CDS is neither similar to a natural human language nor to a random one. Moreover, these sequences have their unique nucleotide organization and it is completely structured for specific biological functioning.

Base Composition↗

Embed-Search-Align: DNA sequence alignment using Transformer models.

MOTIVATION: DNA sequence alignment, an important genomic task, involves assigning short DNA reads to the most probable locations on an extensive reference genome. Conventional methods tackle this challenge in two steps: genome indexing followed by efficient search to locate likely positions for given reads. Building on the success of Large Language Models in encoding text into embeddings, where the distance metric captures semantic similarity, recent efforts have encoded DNA sequences into vectors using Transformers and have shown promising results in tasks involving classification of short DNA sequences. Performance at sequence classification tasks does not, however, guarantee sequence alignment, where it is necessary to conduct a genome-wide search to align every read successfully, a significantly longer-range task by comparison. RESULTS: We bridge this gap by developing a "Embed-Search-Align" (ESA) framework, where a novel Reference-Free DNA Embedding (RDE) Transformer model generates vector embeddings of reads and fragments of the reference in a shared vector space; read-fragment distance metric is then used as a surrogate for sequence similarity. ESA introduces: (i) Contrastive loss for self-supervised training of DNA sequence representations, facilitating rich reference-free, sequence-level embeddings, and (ii) a DNA vector store to enable search across fragments on a global scale. RDE is 99% accurate when aligning 250-length reads onto a human reference genome of 3 gigabases (single-haploid), rivaling conventional algorithmic sequence alignment methods such as Bowtie and BWA-Mem. RDE far exceeds the performance of six recent DNA-Transformer model baselines such as Nucleotide Transformer, Hyena-DNA, and shows task transfer across chromosomes and species. AVAILABILITY AND IMPLEMENTATION: Please see https://anonymous.4open.science/r/dna2vec-7E4E/readme.md.

Sequence Analysis, DNA↗

An open benchmark and language models for AI in aging biology.

Over the past two decades, human aging has been characterized across DNA methylation, transcriptomic, proteomic, and clinical modalities, yet no benchmark evaluates whether AI systems can interpret these heterogeneous data types in the context of aging biology. We introduce LongevityBench, an open suite of 17 tasks spanning five biodata domains, and use it to assess 18 frontier AI systems from six developer teams. Despite recent advances in AI, no single model dominates all tasks, with omics-based age prediction being the hardest task regardless of scale. To test whether these gaps can be closed without frontier-scale resources, we fine-tuned a family of five multitask Longevity-LLMs on domain-specific aging data. The compact (0.6B-9B parameters) Longevity-LLMs matched or exceeded far larger frontier systems on LongevityBench, showing that general-purpose language models can be adapted to structured-omics tasks. We publicly release the benchmark, models, and Longevity Claw, an agentic research interface for aging researchers.

Aging↗

Calculating the exact probability of language-like patterns in biomolecular sequences.

We present algorithms for the exact computation of the probability that a random string of a certain length matches a given regular expression. These algorithms can be used to determine statistical significance in a variety of pattern searches such as motif searches and gene-finding. This work improves upon work of Kleffe and Langebacker (Kleffe & Langbecker 1990) and of Sewell and Durbin (Sewell & Durbin 1995) in several ways. First, in many cases of interest, the algorithms presented here are faster. In addition, the type of pattern considered here strictly includes those of both previous works but also allows, for instance, arbitrary length gaps. Also, the type of probability model which can be used is more general than that of Sewell and Durbin, allowing for Markov chains. The problem solved in this work is in fact in the class of NP-hard problems which are believed to be intractable. However, the problem is fixed-parameter tractable, meaning that it is tractable for small patterns. The is problem is also computationally feasible for many patterns which occur in practice. As a sample application, we consider calculating the statistical significance of most of the PROSITE patterns as in Sewell and Durbin. Whereas their method was only fast enough to exactly compute the probabilities for sequences of length 13 larger than the pattern length, we calculate these probabilities for sequences of up to length 2000. In addition, we calculate most of these probabilities using a first order Markov chain. Most of the PROSITE patterns have high significance at length 2000 under both the i.i.d. and Markov chain models. For further applications, we demonstrate the calculation of the probability of a PROSITE pattern occurring on either strand of a random DNA sequence of up to 500 kilo-bases and the probability of a simple gene model occurring in a random sequence of up to 1 megabase.

Algorithms↗

Feature expressions: creating and manipulating sequence datasets.

Annotation of features, such as introns, exons and protein coding regions in GenBank/EMBL/DDBJ entries is now standardized through use of the Features Table (FT) language. The essence of the FT language is described by the relation 'expression-->sequence', meaning that each FT expression evaluates to a sequence. For example, the expression M74750:1..50 evaluates to the first 50 bases of the sequence with accession number M74750. Because FT is intrinsic to the database definition, it can serve as a software- and platform-independent lingua franca for sequence manipulation. The XYLEM package makes it possible to create and manipulate sequence datasets using FT expressions. FEATURES is a program that resolves FT expressions into their corresponding sequences. Annotated features can be retrieved either by feature key or by expression. Even unannotated portions of a sequence can be retrieved by user-generated FT expressions. Applications of the FT language include retrieval of subsequences from large sequence entries, generation of chromosome models or artificial DNA constructs, and representation of restriction maps or mutants.

Base Sequence↗

Accelerating inference in genomic and proteomic foundation models via speculative decoding.

MOTIVATION: Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose inference is relatively slow, long-sequence generation quickly becomes costly. RESULTS: In this work we adapt speculative decoding to a representative GFM: the DNA model DNAGPT and two representative PFMs: ProGen2 and ProtGPT2. We implement a probabilistic variant of speculative decoding, in which a lightweight draft model proposes short token spans and a larger target model verifies or corrects them in parallel, while preserving the target model's sampling distribution. Across all three models we systematically study the effect of speculation window length, temperature, draft architecture and prompt length, and we benchmark tokens per second over multiple runs per configuration. Speculative decoding yields consistent speedups over standard key-value cached decoding, with maximum observed speedup reaching 100% increase, while average gains across models ranging between 20% and 40% (e.g. 1.2×-1.4×), without changing the underlying target model predictions. Our results show that speculative decoding is a practical and model-agnostic strategy for accelerating genomic and proteomic sequence generation without sacrificing prediction quality. AVAILABILITY AND IMPLEMENTATION: All code and results are freely available at https://github.com/Georgakopoulos-Soares-lab/BioSpecDec.

Genomics↗

easyLINKAGE: a PERL script for easy and automated two-/multi-point linkage analyses.

UNLABELLED: We have generated the program easyLINKAGE that combines automated setup and performance of linkage analyses and simulation under an easy to handle graphical user interface for Microsoft Windows 2000/XP and standard UNIX systems. The program package supports two-point linkage analyses (FastLink v4.1 and SPLink v1.09), multi-point linkage analyses [GENEHUNTER v2.1, GENEHUNTER-PLUS with the emendation by Kong and Cox v1.2 (allele sharing modelling)] and the simulation package SLINK v2.65, and provides genome-wide as well as chromosomal postscript plots of LOD scores, NPL scores, P-values and other parameters. AVAILABILITY: http://www.uni-wuerzburg.de/nephrologie/molecular_genetics/molecular_genetics.htm SUPPLEMENTARY INFORMATION: Supplementary information is available on the website.

Chromosome Mapping↗

libSRES: a C library for stochastic ranking evolution strategy for parameter estimation.

SUMMARY: Estimation of kinetic parameters in a biochemical pathway or network represents a common problem in systems studies of biological processes. We have implemented a C library, named libSRES, to facilitate a fast implementation of computer software for study of non-linear biochemical pathways. This library implements a (mu, lambda)-ES evolutionary optimization algorithm that uses stochastic ranking as the constraint handling technique. Considering the amount of computing time it might require to solve a parameter-estimation problem, an MPI version of libSRES is provided for parallel implementation, as well as a simple user interface. libSRES is freely available and could be used directly in any C program as a library function. We have extensively tested the performance of libSRES on various pathway parameter-estimation problems and found its performance to be satisfactory. AVAILABILITY: The source code (in C) is free for academic users at http://csbl.bmb.uga.edu/~jix/science/libSRES/

Algorithms↗

A measure of DNA sequence dissimilarity based on Mahalanobis distance between frequencies of words.

A number of algorithms exist for searching genetic databases for biologically significant similarities in DNA sequences. Past research has shown that word-based search tools are computationally efficient and can find similarities or dissimilarities invisible to other algorithms like FASTA. We characterize a family of word-based dissimilarity measures that define distance between two sequences by simultaneously comparing the frequencies of all subsequences of n adjacent letters (i.e., n-words) in the two sequences. Applications to real data demonstrate that currently used word-based methods that rely on Euclidean distance can be significantly improved by using Mahalanobis distance, which accounts for both variances and covariances between frequencies of n-words. Furthermore, in those cases where Mahalanobis distance may be too difficult to compute, using standardized Euclidean distance, which only corrects for the variances of frequencies of n-words, still gives better performance than the Euclidean distance. Also, a simple way of combining distances obtained at different n-words is considered. The goal is to obtain a single measure of dissimilarity between two DNA sequences. The performance ranking of the preceding three distances still holds for their combined counterparts. All results obtained in this paper are applicable to amino acid sequences with minor modifications.

Algorithms↗

Segmentation of yeast DNA using hidden Markov models.

MOTIVATION: Compositionally homogeneous segments of genomic DNA often correspond to meaningful biological units. Simple sliding window analysis is usually insufficient for compositional segmentation of natural sequences. Hidden Markov models (HMM) with a small number of states are a natural language for description of compositional properties of chromosome-size DNA sequences. RESULTS: The algorithms were applied to yeast Saccharomyces cerevisiae chromosomes (YC) I, III, IV, VI and IX. The optimal number of HMM states is found to be four. The optimal four-state HMMs for all chromosomes are very similar, as well as the reconstructed segmentations. In most cases the models with k + 1 states are obtained by 'splitting' one of the states in the model with k states, and the corresponding increase of the level of detail in segmentation. The high AT states usually correspond to intergenic regions. We also explore the model's likelihood landscape and analyze the dynamics of the optimization process, thus addressing the problem of reliability of the obtained optima and efficiency of the algorithms.

Algorithms↗

Klinefelter's syndrome as a model of anomalous cerebral laterality: testing gene dosage in the X chromosome pseudoautosomal region using a DNA microarray.

Consistent handedness and language laterality are two of the most striking behavioral and cognitive asymmetries observed in humans. Alterations in the typical pattern of cerebral laterality, termed "anomalous dominance," is observed in left-handers and some patients with verbal learning disabilities. We undertook the study of a genetically distinct group of subjects, XXY males (Klinefelter's syndrome; KS), who demonstrate anomalous dominance in a variety of testing paradigms in order to begin to elucidate the molecular basis of anomalous dominance in this population. KS subjects manifest specific verbal learning disability, evidence of altered functional laterality for phonologic processing, and an increase in left-handedness when measured by skill. It is proposed that an alteration in gene dosage in the pseudoautosomal region (PAR) of the sex chromosomes is the most likely explanation for anomalous dominance in these patients. This is especially intriguing in light of previously described genetic models of cerebral laterality that suggest a contributing locus in the PAR, or adjacent high homology regions of the X chromosome. We have developed an ordered DNA microarray covering the X chromosome PAR at high resolution for hybridization with two-color fluorescently labeled probes. We demonstrate the ability to detect changes in hybridization signal that will facilitate efficient large-scale screening of this region for alterations in gene dosage associated with features of anomalous dominance and other cognitive or behavioral phenotypes.

Case-Control Studies↗

Rational and computation-assisted engineering of a compact and efficient CRISPR-Cas12f genome editor.

The CRISPR-Cas12f system is an ultracompact genome-editing platform, yet only a few orthologs exhibit robust activity in mammalian cells. Here, we systematically screened 23 Cas12f orthologs and identified two active nucleases, PspCas12f1 and TcCas12f1, capable of genome editing in human cells. Single guide RNA (sgRNA) scaffold optimization enhanced the basal activity of PspCas12f1. To further improve its performance, we combined structure-guided rational design with protein language model-assisted filtering. Candidate mutations predicted by SaProt were further screened based on structural proximity to the DNA-binding interface and electrostatic compatibility. This integrative strategy identified Q100R and E293R, whose combination yielded the optimized variant enPspCas12f1. enPspCas12f1 achieved genome-editing efficiencies comparable to SpCas9 across multiple endogenous loci while maintaining high specificity. Collectively, our results demonstrate that integrating protein language model-assisted filtering with structure-guided rational design provides an effective strategy for engineering PspCas12f1 and may facilitate the optimization of additional compact CRISPR nucleases.

CRISPR-Cas12f↗

Using interval logic for order assembly.

Temporal logic, in particular, interval logic has been used to represent genome maps and to assist genome map constructions. However, interval logic itself appears to be limited in its expressive power because genome mapping requires various information such as partial order, distance and local orientation. In this paper, we first propose an integrated formalism based on a spatial-temporal logic where the concepts of metric information, local orientation and uncertainty are merged. Then, we present and discuss a deductive and object-oriented data model based on this formalism for a genetic deductive database, and the inference rules required. The formalism supports the maintenance of coarser knowledge of unordered, partially ordered and completely ordered genetic data in a relational hierarchy. We believe that this integrated formalism also provides a formal basis for designing a declarative query language.

Animals↗

The Longue Durée of genetic ancestry: multiple genetic marker systems and Celtic origins on the Atlantic facade of Europe.

Celtic languages are now spoken only on the Atlantic facade of Europe, mainly in Britain and Ireland, but were spoken more widely in western and central Europe until the collapse of the Roman Empire in the first millennium a.d. It has been common to couple archaeological evidence for the expansion of Iron Age elites in central Europe with the dispersal of these languages and of Celtic ethnicity and to posit a central European "homeland" for the Celtic peoples. More recently, however, archaeologists have questioned this "migrationist" view of Celtic ethnogenesis. The proposition of a central European ancestry should be testable by examining the distribution of genetic markers; however, although Y-chromosome patterns in Atlantic Europe show little evidence of central European influence, there has hitherto been insufficient data to confirm this by use of mitochondrial DNA (mtDNA). Here, we present both new mtDNA data from Ireland and a novel analysis of a greatly enlarged European mtDNA database. We show that mtDNA lineages, when analyzed in sufficiently large numbers, display patterns significantly similar to a large fraction of both Y-chromosome and autosomal variation. These multiple genetic marker systems indicate a shared ancestry throughout the Atlantic zone, from northern Iberia to western Scandinavia, that dates back to the end of the last Ice Age.

Base Sequence↗

Molecular modeling information transfer with VRML: from small molecules to large systems in bioscience.

The suitability of the Virtual Reality Modeling Language (VRML) for the communication of scientists via the internet is demonstrated with recent results from computer assisted cancer research: I. Substrate channels in cytochrome P450 enzymes. II. Binding properties of the wild type and mutated p53 tumor suppressor protein. Complex 3D molecular models were used to visualize new insights in the active site access of cytochrome P450 enzymes and in the p53 protein-DNA binding achieved by the use of computational methods. These 3D models of biomolecular systems were transferred into VRML scenarios. Additional implemented features allow users to receive related information interactively. With these examples it is shown that VRML provides an efficient method for scientific information exchange by the use of complex 3D molecular models.

Binding Sites↗