Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The Longue Durée of genetic ancestry: multiple genetic marker systems and Celtic origins on the Atlantic facade of Europe.

Celtic languages are now spoken only on the Atlantic facade of Europe, mainly in Britain and Ireland, but were spoken more widely in western and central Europe until the collapse of the Roman Empire in the first millennium a.d. It has been common to couple archaeological evidence for the expansion of Iron Age elites in central Europe with the dispersal of these languages and of Celtic ethnicity and to posit a central European "homeland" for the Celtic peoples. More recently, however, archaeologists have questioned this "migrationist" view of Celtic ethnogenesis. The proposition of a central European ancestry should be testable by examining the distribution of genetic markers; however, although Y-chromosome patterns in Atlantic Europe show little evidence of central European influence, there has hitherto been insufficient data to confirm this by use of mitochondrial DNA (mtDNA). Here, we present both new mtDNA data from Ireland and a novel analysis of a greatly enlarged European mtDNA database. We show that mtDNA lineages, when analyzed in sufficiently large numbers, display patterns significantly similar to a large fraction of both Y-chromosome and autosomal variation. These multiple genetic marker systems indicate a shared ancestry throughout the Atlantic zone, from northern Iberia to western Scandinavia, that dates back to the end of the last Ice Age.

Base Sequence↗

Molecular modeling information transfer with VRML: from small molecules to large systems in bioscience.

The suitability of the Virtual Reality Modeling Language (VRML) for the communication of scientists via the internet is demonstrated with recent results from computer assisted cancer research: I. Substrate channels in cytochrome P450 enzymes. II. Binding properties of the wild type and mutated p53 tumor suppressor protein. Complex 3D molecular models were used to visualize new insights in the active site access of cytochrome P450 enzymes and in the p53 protein-DNA binding achieved by the use of computational methods. These 3D models of biomolecular systems were transferred into VRML scenarios. Additional implemented features allow users to receive related information interactively. With these examples it is shown that VRML provides an efficient method for scientific information exchange by the use of complex 3D molecular models.

Binding Sites↗

BIOZON: a system for unification, management and analysis of heterogeneous biological data.

BACKGROUND: Integration of heterogeneous data types is a challenging problem, especially in biology, where the number of databases and data types increase rapidly. Amongst the problems that one has to face are integrity, consistency, redundancy, connectivity, expressiveness and updatability. DESCRIPTION: Here we present a system (Biozon) that addresses these problems, and offers biologists a new knowledge resource to navigate through and explore. Biozon unifies multiple biological databases consisting of a variety of data types (such as DNA sequences, proteins, interactions and cellular pathways). It is fundamentally different from previous efforts as it uses a single extensive and tightly connected graph schema wrapped with hierarchical ontology of documents and relations. Beyond warehousing existing data, Biozon computes and stores novel derived data, such as similarity relationships and functional predictions. The integration of similarity data allows propagation of knowledge through inference and fuzzy searches. Sophisticated methods of query that span multiple data types were implemented and first-of-a-kind biological ranking systems were explored and integrated. CONCLUSION: The Biozon system is an extensive knowledge resource of heterogeneous biological data. Currently, it holds more than 100 million biological documents and 6.5 billion relations between them. The database is accessible through an advanced web interface that supports complex queries, "fuzzy" searches, data materialization and more, online at http://biozon.org.

Animals↗

Efficient decoding algorithms for generalized hidden Markov model gene finders.

BACKGROUND: The Generalized Hidden Markov Model (GHMM) has proven a useful framework for the task of computational gene prediction in eukaryotic genomes, due to its flexibility and probabilistic underpinnings. As the focus of the gene finding community shifts toward the use of homology information to improve prediction accuracy, extensions to the basic GHMM model are being explored as possible ways to integrate this homology information into the prediction process. Particularly prominent among these extensions are those techniques which call for the simultaneous prediction of genes in two or more genomes at once, thereby increasing significantly the computational cost of prediction and highlighting the importance of speed and memory efficiency in the implementation of the underlying GHMM algorithms. Unfortunately, the task of implementing an efficient GHMM-based gene finder is already a nontrivial one, and it can be expected that this task will only grow more onerous as our models increase in complexity. RESULTS: As a first step toward addressing the implementation challenges of these next-generation systems, we describe in detail two software architectures for GHMM-based gene finders, one comprising the common array-based approach, and the other a highly optimized algorithm which requires significantly less memory while achieving virtually identical speed. We then show how both of these architectures can be accelerated by a factor of two by optimizing their content sensors. We finish with a brief illustration of the impact these optimizations have had on the feasibility of our new homology-based gene finder, TWAIN. CONCLUSIONS: In describing a number of optimizations for GHMM-based gene finders and making available two complete open-source software systems embodying these methods, it is our hope that others will be more enabled to explore promising extensions to the GHMM framework, thereby improving the state-of-the-art in gene prediction techniques.

Algorithms↗

The efficiency of different search strategies in estimating parsimony jackknife, bootstrap, and Bremer support.

BACKGROUND: For parsimony analyses, the most common way to estimate confidence is by resampling plans (nonparametric bootstrap, jackknife), and Bremer support (Decay indices). The recent literature reveals that parameter settings that are quite commonly employed are not those that are recommended by theoretical considerations and by previous empirical studies. The optimal search strategy to be applied during resampling was previously addressed solely via standard search strategies available in PAUP*. The question of a compromise between search extensiveness and improved support accuracy for Bremer support received even less attention. A set of experiments was conducted on different datasets to find an empirical cut-off point at which increased search extensiveness does not significantly change Bremer support and jackknife or bootstrap proportions any more. RESULTS: For the number of replicates needed for accurate estimates of support in resampling plans, a diagram is provided that helps to address the question whether apparently different support values really differ significantly. It is shown that the use of random addition cycles and parsimony ratchet iterations during bootstrapping does not translate into higher support, nor does any extension of the search extensiveness beyond the rather moderate effort of TBR (tree bisection and reconnection branch swapping) plus saving one tree per replicate. Instead, in case of very large matrices, saving more than one shortest tree per iteration and using a strict consensus tree of these yields decreased support compared to saving only one tree. This can be interpreted as a small risk of overestimating support but should be more than compensated by other factors that counteract an enhanced type I error. With regard to Bremer support, a rule of thumb can be derived stating that not much is gained relative to the surplus computational effort when searches are extended beyond 20 ratchet iterations per constrained node, at least not for datasets that fall within the size range found in the current literature. CONCLUSION: In view of these results, calculating bootstrap or jackknife proportions with narrow confidence intervals even for very large datasets can be achieved with less expense than often thought. In particular, iterated bootstrap methods that aim at reducing statistical bias inherent to these proportions are more feasible when the individual bootstrap searches require less time.

Bayes Theorem↗

Cell cycle simulation for flow cytometry.

A program has been implemented on a VAX computer that simulates the progression of cells through the cell cycle and generates data similar to those obtained in flow cytometry with different techniques. Features of the program are general applicability and flexibility, options including consideration of (a) mean duration of cell cycle phases, (b) their inter-cell distribution, (c) a first order commitment from G0 into G1 phase, and (d) a total or partial block of the output from any phase. Examples are given of simulated flow cytometric experiments of drug-induced cell cycle perturbation and of bromodeoxyuridine pulse labeling. This program should help to acquire a correct understanding of the relationship between kinetic features and flow cytometric data.

Animals↗

Repair-FunMap: a functional database of proteins of the DNA repair systems.

UNLABELLED: Repair-FunMap is a functional database of the DNA repair systems. This database contains not only the proteins directly involved in DNA repair, but also the proteins that interact with the DNA repair proteins. A protein interaction network associated with the human DNA repair processes was established according to the functional relationship between proteins in the database. This network represents the current knowledge on the intrinsic signaling pathways related to DNA repair. The Repair-FunMap could become an essential resource center for cancer research, providing clues to understanding the inter-relationship between proteins in the network, and to building scientific models of the DNA repair processes. AVAILABILITY: http://astro.temple.edu/~feng/Servers/BioinformaticServers.htm

Abstracting and Indexing↗

An empirical analysis of training protocols for probabilistic gene finders.

BACKGROUND: Generalized hidden Markov models (GHMMs) appear to be approaching acceptance as a de facto standard for state-of-the-art ab initio gene finding, as evidenced by the recent proliferation of GHMM implementations. While prevailing methods for modeling and parsing genes using GHMMs have been described in the literature, little attention has been paid as of yet to their proper training. The few hints available in the literature together with anecdotal observations suggest that most practitioners perform maximum likelihood parameter estimation only at the local submodel level, and then attend to the optimization of global parameter structure using some form of ad hoc manual tuning of individual parameters. RESULTS: We decided to investigate the utility of applying a more systematic optimization approach to the tuning of global parameter structure by implementing a global discriminative training procedure for our GHMM-based gene finder. Our results show that significant improvement in prediction accuracy can be achieved by this method. CONCLUSIONS: We conclude that training of GHMM-based gene finders is best performed using some form of discriminative training rather than simple maximum likelihood estimation at the submodel level, and that generalized gradient ascent methods are suitable for this task. We also conclude that partitioning of training data for the twin purposes of maximum likelihood initialization and gradient ascent optimization appears to be unnecessary, but that strict segregation of test data must be enforced during final gene finder evaluation to avoid artificially inflated accuracy measurements.

Algorithms↗

Deduction of probable events of lateral gene transfer through comparison of phylogenetic trees by recursive consolidation and rearrangement.

BACKGROUND: When organismal phylogenies based on sequences of single marker genes are poorly resolved, a logical approach is to add more markers, on the assumption that weak but congruent phylogenetic signal will be reinforced in such multigene trees. Such approaches are valid only when the several markers indeed have identical phylogenies, an issue which many multigene methods (such as the use of concatenated gene sequences or the assembly of supertrees) do not directly address. Indeed, even when the true history is a mixture of vertical descent for some genes and lateral gene transfer (LGT) for others, such methods produce unique topologies. RESULTS: We have developed software that aims to extract evidence for vertical and lateral inheritance from a set of gene trees compared against an arbitrary reference tree. This evidence is then displayed as a synthesis showing support over the tree for vertical inheritance, overlaid with explicit lateral gene transfer (LGT) events inferred to have occurred over the history of the tree. Like splits-tree methods, one can thus identify nodes at which conflict occurs. Additionally one can make reasonable inferences about vertical and lateral signal, assigning putative donors and recipients. CONCLUSION: A tool such as ours can serve to explore the reticulated dimensionality of molecular evolution, by dissecting vertical and lateral inheritance at high resolution. By this, we mean that individual nodes can be examined not only for congruence, but also for coherence in light of LGT. We assert that our tools will facilitate the comparison of phylogenetic trees, and the interpretation of conflicting data.

Animals↗

Using Tcl for molecular visualization and analysis.

Reading and manipulating molecular structure data is a standard task in every molecular visualization and analysis program, but is rarely available in a form readily accessible to the user. Instead, the development of new methods for analysis, display, and interaction is often achieved by writing a new program, rather than building on pre-existing software. We present the Tcl-based script language used in our molecular modeling program, VMD, and show how it can access information about the molecular structure, perform analysis, and graphically display and animate the results. The commands are available to the user and make VMD a useful environment for studying biomolecules.

Computer Simulation↗

Linking cDNA-AFLP-based gene expression patterns and ESTs.

Massive amounts of DNA sequence data, generated from expressed sequence tag (EST) and genome sequencing projects, require efficient methods to link sequence databases with temporal and spatial expression profiles. To meet this need, we have developed a powerful computer program (GenEST), which links cDNA sequence data (including EST sequences) with transcript profiles revealed by cDNA-amplified fragment length polymorphism (AFLP). cDNA-AFLP is a highly reproducible differential display method based on restriction enzyme digests and selective amplification under high stringency conditions. GenEST predicts the sizes of virtual transcript derived fragments (TDFs) from cDNA sequences digested in silico. The resulting virtual TDFs could be traced back among the thousands of TDFs displayed on cDNA-AFLP gels. As a consequence, cDNA sequence databases can be screened very efficiently to identify genes with relevant expression profiles. Vice versa, using the restriction enzyme recognition sites, the primer extensions and the estimated TDF size as identifiers, the DNA sequence(s) corresponding to a TDF with an interesting expression pattern can be identified.

Automation↗

Vestige: maximum likelihood phylogenetic footprinting.

BACKGROUND: Phylogenetic footprinting is the identification of functional regions of DNA by their evolutionary conservation. This is achieved by comparing orthologous regions from multiple species and identifying the DNA regions that have diverged less than neutral DNA. Vestige is a phylogenetic footprinting package built on the PyEvolve toolkit that uses probabilistic molecular evolutionary modelling to represent aspects of sequence evolution, including the conventional divergence measure employed by other footprinting approaches. In addition to measuring the divergence, Vestige allows the expansion of the definition of a phylogenetic footprint to include variation in the distribution of any molecular evolutionary processes. This is achieved by displaying the distribution of model parameters that represent partitions of molecular evolutionary substitutions. Examination of the spatial incidence of these effects across regions of the genome can identify DNA segments that differ in the nature of the evolutionary process. RESULTS: Vestige was applied to a reference dataset of the SCL locus from four species and provided clear identification of the known conserved regions in this dataset. To demonstrate the flexibility to use diverse models of molecular evolution and dissect the nature of the evolutionary process Vestige was used to footprint the Ka/Ks ratio in primate BRCA1 with a codon model of evolution. Two regions of putative adaptive evolution were identified illustrating the ability of Vestige to represent the spatial distribution of distinct molecular evolutionary processes. CONCLUSION: Vestige provides a flexible, open platform for phylogenetic footprinting. Underpinned by the PyEvolve toolkit, Vestige provides a framework for visualising the signatures of evolutionary processes across the genome of numerous organisms simultaneously. By exploiting the maximum-likelihood statistical framework, the complex interplay between mutational processes, DNA repair and selection can be evaluated both spatially (along a sequence alignment) and temporally (for each branch of the tree) providing visual indicators to the attributes and functions of DNA sequences.

Algorithms↗

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis↗

The application of molecular genetic approaches to the study of human evolution.

The past decade of advances in molecular genetic technology has heralded a new era for all evolutionary studies, but especially the science of human evolution. Data on various kinds of DNA variation in human populations have rapidly accumulated. There is increasing recognition of the importance of this variation for medicine and developmental biology and for understanding the history of our species. Haploid markers from mitochondrial DNA and the Y chromosome have proven invaluable for generating a standard model for evolution of modern humans. Conclusions from earlier research on protein polymorphisms have been generally supported by more sophisticated DNA analysis. Co-evolution of genes with language and some slowly evolving cultural traits, together with the genetic evolution of commensals and parasites that have accompanied modern humans in their expansion from Africa to the other continents, supports and supplements the standard model of genetic evolution. The advances in our understanding of the evolutionary history of humans attests to the advantages of multidisciplinary research.

Animals↗

An improved FORTRAN 77 recombinant DNA database management system with graphic extensions in GKS.

We have improved an existing clone database management system written in FORTRAN 77 and adapted it to our software environment. Improvements are that the database can be interrogated for any type of information, not just keywords. Also, recombinant DNA constructions can be represented in a simplified 'shorthand', whereafter a program assembles the full nucleotide sequence from the contributing fragments, which may be obtained from nucleotide sequence databases. Another improvement is the replacement of the database manager by programs, running in batch to maintain the databank and verify its consistency automatically. Finally, graphic extensions are written in Graphical Kernel System, to draw linear and circular restriction maps of recombinants. Besides restriction sites, recombinant features can be presented from the feature lines of recombinant database entries, or from the feature tables of nucleotide databases. The clone database management system is fully integrated into the sequence analysis software package from the Pasteur Institute, Paris, and is made accessible through the same menu. As a result, recombinant DNA sequences can directly be analysed by the sequence analysis programs.

Algorithms↗

TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders.

UNLABELLED: We describe two new Generalized Hidden Markov Model implementations for ab initio eukaryotic gene prediction. The C/C++ source code for both is available as open source and is highly reusable due to their modular and extensible architectures. Unlike most of the currently available gene-finders, the programs are re-trainable by the end user. They are also re-configurable and include several types of probabilistic submodels which can be independently combined, such as Maximal Dependence Decomposition trees and interpolated Markov models. Both programs have been used at TIGR for the annotation of the Aspergillus fumigatus and Toxoplasma gondii genomes. AVAILABILITY: Source code and documentation are available under the open source Artistic License from http://www.tigr.org/software/pirate

Algorithms↗