Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

SNAP: Combine and Map modules for multilocus population genetic analysis.

We have added two software tools to our Suite of Nucleotide Analysis Programs (SNAP) for working with DNA sequences sampled from populations. SNAP Map collapses DNA sequence data into unique haplotypes, extracts variable sites and manipulates output into multiple formats for input into existing software packages for evolutionary analyses. Map collapses DNA sequence data into unique haplotypes, extracts variable sites and manipulates output into multiple formats for input into existing software packages for evolutionary analyses. Map includes novel features such as recoding insertions or deletions, including or excluding variable sites that violate an infinite-sites model and the option of collapsing sequences with corresponding phenotypic information, important in testing for significant haplotype-phenotype associations. SNAP Combine merges multiple DNA sequence alignments into a single multiple alignment file. The resulting file can be the union or intersection of the input files. SNAP Combine currently reads from and writes to several sequence alignment file formats including both sequential and interleaved formats. Combine also keeps track of the start and end positions of each separate alignment file allowing the user to exclude variable sites or taxa, important in creating input files for multilocus analyses.

Algorithms↗

Advanced query mechanisms for biological databases.

Existing query interfaces for biological databases are either based on fixed forms or textual query languages. Users of a fixed form-based query interface are limited to performing some pre-defined queries providing a fixed view of the underlying database, while users of a free text query language-based interface have to understand the underlying data models, specific query languages and application schemas in order to formulate queries. Further, operations on application-specific complex data (e.g., DNA sequences, proteins), which are usually provided by a variety of software packages with their own format requirements and peculiarities, are not available as part of, nor integrated with biological query interfaces. In this paper, we describe generic tools that provide powerful and flexible support for interactively exploring biological databases in a uniform and consistent way, that is via common data models, formats, and notations, in the framework of the Object-Protocol Model (OPM). These tools include (i) a Java graphical query construction tool with support for automatic generation of Web query forms that can be either used for further specifying conditions, or can be saved and customized; (ii) query processors for interpreting and executing queries that may involve complex application-specific objects, and that could span multiple heterogeneous databases and file systems; and (iii) utilities for automatic generation of HTML pages containing query results, that can be browsed using a Web browser. These tools avoid the restrictions imposed by traditional fixed-form query interfaces, while providing users with simple and intuitive facilities for formulating ad-hoc queries across heterogeneous databases, without the need to understand the underlying data models and query languages.

Animals↗

The cell as the smallest DNA-based molecular computer.

The pioneering work of Adleman (1994) demonstrated that DNA molecules in test tubes can be manipulated to perform a certain type of mathematical computation. This has stimulated a theoretical interest in the possibility of constructing DNA-based molecular computers. To gauge the practicality of realizing such microscopic computers, it was thought necessary to learn as much as possible from the biology of the living cell--presently the only known DNA-based molecular computer in existence. Here the recently developed theoretical model of the living cell (the Bhopalator) and its associated theories (e.g. cell language), principles, laws and concepts (e.g. conformons, IDS's) are briefly reviewed and summarized in the form of a set of five laws of 'molecular semiotics' (synonyms include 'microsemiotics', 'cellular semiotics', or 'cytosemiotics') the study of signs mediating measurement, computation, and communication on the cellular and molecular levels. Hopefully, these laws will find practical applications in designing DNA-based computing systems.

Animals↗

Grammatical inference in bioinformatics.

Bioinformatics is an active research area aimed at developing intelligent systems for analyses of molecular biology. Many methods based on formal language theory, statistical theory, and learning theory have been developed for modeling and analyzing biological sequences such as DNA, RNA, and proteins. Especially, grammatical inference methods are expected to find some grammatical structures hidden in biological sequences. In this article, we give an overview of a series of our grammatical approaches to biological sequence analyses and related researches and focus on learning stochastic grammars from biological sequences and predicting their functions based on learned stochastic grammars.

Algorithms↗

EvoSNR-Prom: Predicting promoters at single-nucleotide resolution with label-aware transfer learning of the pretrained EVO model.

The precise identification of promoters is crucial for understanding gene regulation. Deep learning methods have achieved considerable success in promoter prediction, yet most operate at the sequence level with coarse-grained labels. This means they label an entire DNA segment as either a "promoter" or "non-promoter," which results in a lack of the nucleotide-level resolution in prediction. In this study, we propose EvoSNR-Prom, a model designed for promoter prediction at single-nucleotide resolution. EvoSNR-Prom is built on the Evo foundation model and formulates promoter identification as a token-level sequence labeling problem, analogous to named entity recognition in natural language processing. To address the limited contextual information available in single-nucleotide tokenization, we introduce a lexicon-enhanced embedding strategy that incorporates biologically meaningful DNA lexicons, enriching contextual representations and improving the model's ability to capture complex sequence motifs. Furthermore, to enhance predictive performance on small size datasets, we integrate a label-aware transfer learning framework to leverage knowledge from well-annotated source species to a target organism. The results across various prokaryotic datasets show that EvoSNR-Prom achieves excellent performance. This work provides a valuable computational framework for the high-precision analysis of gene regulatory elements, contributing to the advancement of promoter prediction at single-nucleotide resolution.

Promoter Regions, Genetic↗

Gene flow across linguistic boundaries in Native North American populations.

Cultural and linguistic groups are often expected to represent genetic populations. In this article, we tested the hypothesis that the hierarchical classification of languages proposed by J. Greenberg [(1987) Language in the Americas (Stanford Univ. Press, Stanford, CA)] also represents the genetic structure of Native North American populations. The genetic data are mtDNA sequences for 17 populations gleaned from literature sources and public databases. The hypothesis was rejected. Further analysis showed that departure of the genetic structure from the linguistic classification was pervasive and not due to an outlier population or a problematic language group. Therefore, Greenberg's language groups are at best an imperfect approximation to the genetic structure of these populations. Moreover, we show that the genetic structure among these Native North American populations departs significantly from the best-fitting hierarchical models. Analysis of median joining networks for mtDNA haplotypes provides strong evidence for gene flow across linguistic boundaries. In principle, the language of a population can be replaced more rapidly than its genes because language can be transmitted both vertically from parents to children and horizontally between unrelated people. However, languages are part of a cultural complex, and there may be strong pressure to maintain a language in place whereas genes are free to flow.

DNA, Mitochondrial↗

Ontologies for molecular biology.

Molecular biology has a communication problem. There are many databases using their own labels and categories for storing data objects and some using identical labels and categories but with a different meaning. A prominent example is the concept "gene" which is used with different semantics by major international genomic databases. Ontologies are one means to provide a semantic repository to systematically order relevant concepts in molecular biology and to bridge the different notions in various databases by explicitly specifying the meaning of and relation between the fundamental concepts in an application domain. Here, the upper level and a database branch of a prospective ontology for molecular biology (OMB) is presented and compared to other ontologies with respect to suitability for molecular biology (http:/(/)igd.rz-berlin.mpg.de/approximately www/oe/mbo.html).

Chromosome Mapping↗

SplitTester: software to identify domains responsible for functional divergence in protein family.

BACKGROUND: Many protein families have undergone functional divergence after gene duplications such that current subgroups of the family carry out overlapping but distinct biological roles. For the protein families with known functional subtypes (a functional split), we developed the software, SplitTester, to identify potential regions that are responsible for the observed distinct functional subtypes within the same protein family. RESULTS: Our software, SplitTester, takes a multiple protein sequences alignment as input, generated from protein members of two subgroups with known functional divergence. SplitTester was designed to construct the neighbor joining tree (a split cluster) from variable-sized sliding windows across the alignment in a process called split-clustering. SplitTester identifies the regions, whose split cluster is consistent with the functional split, but may be inconsistent with the phylogeny of the protein family. We hypothesize that at least some number of these identified regions, which are not following a random mutation process, are responsible for the observed functional split. To test our method, we used reverse transcriptase from a group of Pseudoviridae retrotransposons: to identify residues specific for diverged primer recognition. Candidate regions were then mapped onto the three dimensional structures of reverse transcriptase. The locations of these amino acids within the enzyme are consistent with their biological roles. CONCLUSION: SplitTester aims to identify specific domain sequences responsible for functional divergence of subgroups within a protein family. From the analysis of retroelements reverse transcriptase family, we successfully identified the regions splitting this family according to the primer specificity, implying their functions in the specific primer selection.

Algorithms↗

In silico identification of functional regions in proteins.

MOTIVATION: In silico prediction of functional regions on protein surfaces, i.e. sites of interaction with DNA, ligands, substrates and other proteins, is of utmost importance in various applications in the emerging fields of proteomics and structural genomics. When a sufficient number of homologs is found, powerful prediction schemes can be based on the observation that evolutionarily conserved regions are often functionally important, typically, only the principal functionally important region of the protein is detected, while secondary functional regions with weaker conservation signals are overlooked. Moreover, it is challenging to unambiguously identify the boundaries of the functional regions. METHODS: We present a new methodology, called PatchFinder, that automatically identifies patches of conserved residues that are located in close proximity to each other on the protein surface. PatchFinder is based on the following steps: (1) Assignment of conservation scores to each amino acid position on the protein surface. (2) Assignment of a score to each putative patch, based on its likelihood to be functionally important. The patch of maximum likelihood is considered to be the main functionally important region, and the search is continued for non-overlapping patches of secondary importance. RESULTS: We examined the accuracy of the method using the IGPS enzyme, the SH2 domain and a benchmark set of 112 proteins. These examples demonstrated that PatchFinder is capable of identifying both the main and secondary functional patches. AVAILABILITY: The PatchFinder program is available at: http://ashtoret.tau.ac.il/~nimrodg/

Algorithms↗

A dynamic programming algorithm for binning microbial community profiles.

MOTIVATION: A number of community profiling approaches have been widely used to study the microbial community composition and its variations in environmental ecology. Automated Ribosomal Intergenic Spacer Analysis (ARISA) is one such technique. ARISA has been used to study microbial communities using 16S-23S rRNA intergenic spacer length heterogeneity at different times and places. Owing to errors in sampling, random mutations in PCR amplification, and probably mostly variations in readings from the equipment used to analyze fragment sizes, the data read directly from the fragment analyzer should not be used for down stream statistical analysis. No optimal data preprocessing methods are available. A commonly used approach is to bin the reading lengths of the 16S-23S intergenic spacer. We have developed a dynamic programming algorithm based binning method for ARISA data analysis which minimizes the overall differences between replicates from the same sampling location and time. RESULTS: In a test example from an ocean time series sampling program, data preprocessing identified several outliers which upon re-examination were found to be because of systematic errors. Clustering analysis of the ARISA from different times based on the dynamic programming algorithm binned data revealed important features of the biodiversity of the microbial communities.

Algorithms↗

EMMA: an efficient massive mapping algorithm using improved approximate mapping filtering.

Efficient massive mapping algorithm (EMMA), an algorithm on efficiently mapping massive cDNAs onto genomic sequences, has recently been developed. The process of mapping massive cDNAs onto genomic sequences has been improved using more approximate mapping filtering based on an enhanced suffix array coupled with a pruned fast hash table, algorithms of block alignment extensions, and k-longest paths. When compared with the classical BLAT software in this field, the computing of EMMA ranges from two to forty-one times faster under similar prediction precisions.

Algorithms↗

Evolutionary sequence analysis of complete eukaryote genomes.

BACKGROUND: Gene duplication and gene loss during the evolution of eukaryotes have hindered attempts to estimate phylogenies and divergence times of species. Although current methods that identify clusters of orthologous genes in complete genomes have helped to investigate gene function and gene content, they have not been optimized for evolutionary sequence analyses requiring strict orthology and complete gene matrices. Here we adopt a relatively simple and fast genome comparison approach designed to assemble orthologs for evolutionary analysis. Our approach identifies single-copy genes representing only species divergences (panorthologs) in order to minimize potential errors caused by gene duplication. We apply this approach to complete sets of proteins from published eukaryote genomes specifically for phylogeny and time estimation. RESULTS: Despite the conservative criterion used, 753 panorthologs (proteins) were identified for evolutionary analysis with four genomes, resulting in a single alignment of 287,000 amino acids. With this data set, we estimate that the divergence between deuterostomes and arthropods took place in the Precambrian, approximately 400 million years before the first appearance of animals in the fossil record. Additional analyses were performed with seven, 12, and 15 eukaryote genomes resulting in similar divergence time estimates and phylogenies. CONCLUSION: Our results with available eukaryote genomes agree with previous results using conventional methods of sequence data assembly from genomes. They show that large sequence data sets can be generated relatively quickly and efficiently for evolutionary analyses of complete genomes.

Animals↗

Wildfire: distributed, Grid-enabled workflow construction and execution.

BACKGROUND: We observe two trends in bioinformatics: (i) analyses are increasing in complexity, often requiring several applications to be run as a workflow; and (ii) multiple CPU clusters and Grids are available to more scientists. The traditional solution to the problem of running workflows across multiple CPUs required programming, often in a scripting language such as perl. Programming places such solutions beyond the reach of many bioinformatics consumers. RESULTS: We present Wildfire, a graphical user interface for constructing and running workflows. Wildfire borrows user interface features from Jemboss and adds a drag-and-drop interface allowing the user to compose EMBOSS (and other) programs into workflows. For execution, Wildfire uses GEL, the underlying workflow execution engine, which can exploit available parallelism on multiple CPU machines including Beowulf-class clusters and Grids. CONCLUSION: Wildfire simplifies the tasks of constructing and executing bioinformatics workflows.

Algorithms↗

Design and implementation of microarray gene expression markup language (MAGE-ML).

BACKGROUND: Meaningful exchange of microarray data is currently difficult because it is rare that published data provide sufficient information depth or are even in the same format from one publication to another. Only when data can be easily exchanged will the entire biological community be able to derive the full benefit from such microarray studies. RESULTS: To this end we have developed three key ingredients towards standardizing the storage and exchange of microarray data. First, we have created a minimal information for the annotation of a microarray experiment (MIAME)-compliant conceptualization of microarray experiments modeled using the unified modeling language (UML) named MAGE-OM (microarray gene expression object model). Second, we have translated MAGE-OM into an XML-based data format, MAGE-ML, to facilitate the exchange of data. Third, some of us are now using MAGE (or its progenitors) in data production settings. Finally, we have developed a freely available software tool kit (MAGE-STK) that eases the integration of MAGE-ML into end users' systems. CONCLUSIONS: MAGE will help microarray data producers and users to exchange information by providing a common platform for data exchange, and MAGE-STK will make the adoption of MAGE easier.

Computer Simulation↗

Genetic structure of the indigenous populations of Siberia.

This study explores the genetic structure of Siberian indigenous populations on the basis of standard blood group and protein markers and DNA variable number of tandem repeats (VNTR) variation. Four analytical methods were utilized in this study: Harpending and Jenkin's R-matrix; Harpending and Ward's method of correlating genetic heterozygosity (H) to the distance from the centroid of the gene frequency array (rii); spatial autocorrelation, and Mantel tests. Because of the underlying assumptions of the various methods, the numbers of populations used in the analyses varied from 15 to 62. Since spatial autocorrelation is based upon separate correlations between alleles, a larger number of standard blood markers and populations were used. Fewest Siberian populations have been sampled for VNTRs, thus, only a limited comparison was possible. The four analytical procedures employed in this study yielded complementary results suggestive of the effects of unique historical events, evolutionary forces, and geography on the distribution of alleles in Siberian indigenous populations. The principal components analysis of the R-matrix demonstrated the presence of populational clusters that reflect their phylogenetic relationship. Mantel comparisons of matrices indicate that an intimate relationship exists between geography, languages, and genetics of Siberian populations. Spatial autocorrelation patterns reflect the isolation-by-distance model of Malecot and the possible effects of long-distance migration.

ABO Blood-Group System↗

Genetic clues to dispersal in human populations: retracing the past from the present.

Ongoing debate about proper interpretation of DNA sequence polymorphisms and their ability to reconstruct human population history illustrates a important change in perspective that we have achieved in the past 20 years of population genetics. To what extent does the history of a locus represent the history of a population? Tools originally developed for molecular systematics, where genetic lineages have been separated by speciation events, are routinely applied to the analysis of variation within our species, with conflicting results. Because of automated technologies and linkage analysis, we are poised to harvest a wealth of information about our past, if we are successful in moving beyond a current polarization regarding models of human evolution. Rather than just suggesting that true resolution will only come by considering fossil or archaeological evidence, the realistic and appropriate application of genetic models for analysis of population structure is also necessary. Three examples from different dispersal events are highlighted here.

Animals↗

Efficient combination of multiple word models for improved sequence comparison.

MOTIVATION: Studies of efficient and sensitive sequence comparison methods are driven by a need to find homologous regions of weak similarity between large genomes. RESULTS: We describe an improved method for finding similar regions between two sets of DNA sequences. The new method generalizes existing methods by locating word matches between sequences under two or more word models and extending word matches into high-scoring segment pairs (HSPs). The method is implemented as a computer program named DDS2. Experimental results show that DDS2 can find more HSPs by using several word models than by using one word model. AVAILABILITY: The DDS2 program is freely available for academic use in binary code form at http://bioinformatics.iastate.edu/aat/align/align.html and in source code form from the corresponding author.

Algorithms↗

Characterization of a novel cation transporter ATPase gene (ATP13A4) interrupted by 3q25-q29 inversion in an individual with language delay.

Specific language impairment (SLI) is defined as failure to acquire normal language skills despite adequate intelligence and environmental stimulation. Although SLI disorders are often heritable, the genetic basis is likely to involve a number of risk factors. This study describes a 7-year-old girl carrying an inherited paracentric inversion of the long arm of chromosome 3 [46XX, inv(3)(q25.32-q29)] having clinically defined expressive and receptive language delay. Fluorescence in situ hybridization (FISH) with locus-specific bacterial artificial chromosome clones (BACs) as probes was used to characterize the inverted chromosome 3. The proximal and distal inversion breakpoint was found to reside between markers D3S3692/D3S1553 and D3S3590/D3S2305, respectively. ATP13A4, a novel gene coding for a cation-transporting P-type ATPase, was found to be disrupted by the distal breakpoint. The ATP13A4 gene was shown to comprise a 3591-bp transcript encompassing 30 exons spanning 152 kb of the genomic DNA. This study discusses the characterization of ATP13A4 and its possible involvement in speech-language disorder.

Adenosine Triphosphatases↗