Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

A plant dialect of the histone language.

The genome contains all the information needed to build an organism. However, during differentiation and development, additional epigenetic information determines the functional state of cells and tissues. This epigenetic information can be introduced by cytosine methylation and by marking nucleosomal histones. The code written on histones consists of post-translational modifications, including acetylation and methylation. In contrast to the universal nature of the DNA code, the histone language and its decoding machinery differ among animals, plants and fungi. Plant cells have retained totipotency to generate the entire plant and maintained the ability to dedifferentiate, which suggests that the establishment and maintenance of epigenetic information differs from animals. Here, I aim to summarize the histone code and plant-specific aspects of setting and translating the code.

Acetylation↗

Syntactic recognition of regulatory regions in Escherichia coli.

MOTIVATION: One of the most common methodologies to identify cis-regulatory sites in regulatory regions in the DNA is that of weight matrices, as testified by several articles in this issue. An alternative to strengthen the computational predictions in regulatory regions is to develop methods that incorporate more biological properties present in such DNA regions. The grammatical implementation presented in this paper provides a concrete example in this direction. RESULTS: On the basis of the analysis of an exhaustive collection of regulatory regions in Escherichia coli, a grammatical model for the regulatory regions of sigma 70 promoters has been developed. The terminal symbols of the grammar represent individual sites for the binding of activator and repressor proteins, and include the precise position of sites in relation to transcription initiation. Combining these symbols, the grammar generates a large number of different sentences, each of which can be searched for matching against a collection of regulatory regions by means of weight matrices specific for each set of sites for individual proteins. On the basis of this grammatical model, a Prolog syntactic recognizer is presented here. Specific subgrammars for ArgR, LexA and TyrR were implemented. When parsing a collection of 128 sigma 70 promoter regions, the syntactic recognizer produces a much lower number of false-positive sites than the standard search using weight matrices.

Algorithms↗

DNA computing based on splicing: universality results.

The paper extends some of the most recently obtained results on the computational universality of specific variants of H systems (e.g. with regular sets of rules) and proves that we can construct universal computers based on various types of H systems with a finite set of splicing rules as well as a finite set of axioms, i.e. we show the theoretical possibility to design programmable universal DNA computers based on the splicing operation. For H systems working in the multiset style (where the numbers of copies of all available strings are counted) we elaborate how a Turing machine computing a partial recursive function can be simulated by an equivalent H system computing the same function; in that way, from a universal Turning machine we obtain a universal H system. Considering H systems as language generating devices we have to add various simple control mechanisms (checking the presence/absence of certain symbols in the spliced strings) to systems with a finite set of splicing rules as well as with a finite set of axioms in order to obtain the full computational power, i.e. to get a characterization of the family of recursively enumerable languages. We also introduce test tube systems, where several H systems work in parallel in their tubes and from time to time the contents of each tube are redistributed to all tubes according to certain separation conditions. By the construction of universal test tube systems we show that also such systems could serve as the theoretical basis for the development of biological (DNA) computers.

Alternative Splicing↗

Gene structure prediction using an orthologous gene of known exon-intron structure.

Given the availability of complete genome sequences from related organisms, sequence conservation can provide important clues for predicting gene structure. In particular, one should be able to leverage information about known genes in one species to help determine the structures of related genes in another. Such an approach is appealing in that high-quality gene prediction can be achieved for newly sequenced species, such as mouse and puffer fish, using the extensive knowledge that has been accumulated about human genes. This article reports a novel approach to predicting the exon-intron structures of mouse genes by incorporating constraints from orthologous human genes using techniques that have previously been exploited in speech and natural language processing applications. The approach uses a context-free grammar to parse a training corpus of annotated human genes. A statistical training procedure produces a weighted recursive transition network (RTN) intended to capture the general features of a mammalian gene. This RTN is expanded into a finite state transducer (FST) and composed with an FST capturing the specific features of the human orthologue. This model includes a trigram language model on the amino acid sequence as well as exon length constraints. A final stage uses the free software package ClustalW to align the top n candidates in the search space. For a set of 98 orthologous human-mouse pairs, we achieved 96% sensitivity and 97% specificity at the exon level on the mouse genes, given only knowledge gleaned from the annotated human genome.

Algorithms↗

Contrasting patterns of Y chromosome and mtDNA variation in Africa: evidence for sex-biased demographic processes.

To investigate associations between genetic, linguistic, and geographic variation in Africa, we type 50 Y chromosome SNPs in 1122 individuals from 40 populations representing African geographic and linguistic diversity. We compare these patterns of variation with those that emerge from a similar analysis of published mtDNA HVS1 sequences from 1918 individuals from 39 African populations. For the Y chromosome, Mantel tests reveal a strong partial correlation between genetic and linguistic distances (r=0.33, P=0.001) and no correlation between genetic and geographic distances (r=-0.08, P>0.10). In contrast, mtDNA variation is weakly correlated with both language (r=0.16, P=0.046) and geography (r=0.17, P=0.035). AMOVA indicates that the amount of paternal among-group variation is much higher when populations are grouped by linguistics (Phi(CT)=0.21) than by geography (Phi(CT)=0.06). Levels of maternal genetic among-group variation are low for both linguistics and geography (Phi(CT)=0.03 and 0.04, respectively). When Bantu speakers are removed from these analyses, the correlation with linguistic variation disappears for the Y chromosome and strengthens for mtDNA. These data suggest that patterns of differentiation and gene flow in Africa have differed for men and women in the recent evolutionary past. We infer that sex-biased rates of admixture and/or language borrowing between expanding Bantu farmers and local hunter-gatherers played an important role in influencing patterns of genetic variation during the spread of African agriculture in the last 4000 years.

Africa↗

The language of covalent histone modifications.

Histone proteins and the nucleosomes they form with DNA are the fundamental building blocks of eukaryotic chromatin. A diverse array of post-translational modifications that often occur on tail domains of these proteins has been well documented. Although the function of these highly conserved modifications has remained elusive, converging biochemical and genetic evidence suggests functions in several chromatin-based processes. We propose that distinct histone modifications, on one or more tails, act sequentially or in combination to form a 'histone code' that is, read by other proteins to bring about distinct downstream events.

Acetylation↗

Algebraic properties of DNA operations.

Any DNA strand can be identified with a word in the language X* where X=¿A, C, G, T¿. By encoding A as 000, C as 010, G as 101, and T as 111, we treat the DNA operations concatenation, union, reverse, complement, annealing and melting, from the algebraic point of view. The concatenation and union play the roles of multiplication and addition over some algebraic structures, respectively. Then the rest of the operations turn out to be the homomorphisms or anti-homomorphisms of these algebraic structures. Using this technique, we find the relationship among these DNA operations.

Animals↗

Prospects for a vaccine against human papillomavirus.

OBJECTIVE: To summarize existing data regarding the feasibility of developing strategies for prophylactic and therapeutic vaccination against human papillomavirus (HPV) infection. DATA SOURCES: We used the Medline data base and reference lists of articles to identify English-language papers that evaluate strategies for prophylactic and therapeutic vaccination against HPV infection. METHODS OF STUDY SELECTION: Our search uncovered several reports of systems that produce recombinant HPV major capsid proteins as antigens for biochemical, molecular, and immunologic studies and investigations that evaluate cell-mediated immune responses to HPV-induced, tumor-associated peptides. DATA EXTRACTION AND SYNTHESIS: Recombinant HPV major capsid proteins, which self-assemble into virus-like particles, are produced in quantity, mimic the conformation of native virions, react with neutralizing antibodies, and are type-specific. Human papillomavirus early viral peptides induce cytotoxic T lymphocyte responses that retard tumor progression and protect against tumor development after challenge in animal models. CONCLUSIONS: Recombinant papillomavirus virus-like particles are highly antigenic, protective in animal models, lack potentially carcinogenic viral DNA, and are, therefore, ideal candidates for a prophylactic vaccine against HPV infection. Immunization with HPV tumor peptides may be beneficial in tumor prevention, regression, and rejection. Vaccines against HPV infection can be important in reducing the incidence of cervical dysplasia and carcinoma worldwide, particularly in developing countries.

Animals↗

Significance of nucleotide sequence alignments: a method for random sequence permutation that preserves dinucleotide and codon usage.

The similarity of two nucleotide sequences is often expressed in terms of evolutionary distance, a measure of the amount of change needed to transform one sequence into the other. Given two sequences with a small distance between them, can their similarity be explained by their base composition alone? The nucleotide order of these sequences contributes to their similarity if the distance is much smaller than their average permutation distance, which is obtained by calculating the distances for many random permutations of these sequences. To determine whether their similarity can be explained by their dinucleotide and codon usage, random sequences must be chosen from the set of permuted sequences that preserve dinucleotide and codon usage. The problem of choosing random dinucleotide and codon-preserving permutations can be expressed in the language of graph theory as the problem of generating random Eulerian walks on a directed multigraph. An efficient algorithm for generating such walks is described. This algorithm can be used to choose random sequence permutations that preserve (1) dinucleotide usage, (2) dinucleotide and trinucleotide usage, or (3) dinucleotide and codon usage. For example, the similarity of two 60-nucleotide DNA segments from the human beta-1 interferon gene (nucleotides 196-255 and 499-558) is not just the result of their nonrandom dinucleotide and codon usage.

Base Sequence↗

Construction of molecular evolutionary phylogenetic trees from DNA sequences based on minimum complexity principle.

Ever since the discovery of a molecular clock, many methods have been developed to reconstruct the molecular evolutionary phylogenetic trees. In this paper, we deal with the problem from the viewpoint of an inductive inference and apply Rissanen's minimum description length principle to extract the minimum complexity phylogenetic tree. Our method describes the complexity of the molecular phylogenetic tree by three terms which are related to the tree topology, the sum of the branch lengths and the difference between the model and the data measured by logarithmic likelihood. Five mitochondrial DNA sequences, from the human, the common chimpanzee, the pygmy chimpanzee, the gorilla and the orangutan, are used for investigating the validity of this method. It is suggested that this method might be superior to the traditional method in that it still shows good accuracy even near the root of phylogenetic trees.

Algorithms↗

The preferential mode analysis of DNA sequence.

After reviewing approaches to the nucleotide correlation of DNA sequences the preferential mode analysis method is emphasized and discussed in detail. The preferred modes and poor modes in coding regions, as well as in introns, 5'-caps and 3'-tails are found through the statistical analysis of sequence data of all kinds of species in GenBank. The relation between the preferential mode analysis and informational parameter method is deduced. It is discovered that in higher species the coding sequences preferentially use the strong-weak bond (strong bond=C,G; weak bond=A, T) language and many noncoding regions (introns, 5'-caps, 3'-tails) use purine-pyrimidine language. The application of different languages in coding and noncoding sequences is a result of evolution, and it may be related to the functional differences in these two regions. Furthermore, we find that many preferential triplets in coding sequences can be expressed in a form of (* W S) (W=A,T; S=C,G), which may be explained by its relation to t-RNA abundance. The systematic change of some mode contents with evolution has also been found.

Animals↗

Computer program for calculating the melting temperature of degenerate oligonucleotides used in PCR or hybridization.

Degenerate primers or probes have been widely used in molecular cloning, but the calculation of their melting temperatures could not simply be done using thermodynamic parameters because of degeneracy and the lack of a computer program. We present here a simple computer program named dPrimer for the calculation of melting temperature of degenerate oligonucleotides based on the nearest-neighbor model. The program was written in C+2 computer language and implemented in Macintosh with a Symantec C+2 compiler. The degenerate sequencing data were read into a graph data structure. All possible oligonucleotide sequences were then determined by a depth-first search algorithm. Their melting temperature (Tm) values were individually calculated, and output was given as Tm range, mean and standard deviation. These data could help one in the selection of PCR annealing and hybridization temperatures as well as in the design of degenerate oligonucleotides with a desired range of Tm.

Algorithms↗

Empirical bayes microarray ANOVA and grouping cell lines by equal expression levels.

In the exploding field of gene expression techniques such as DNA microarrays, there are still few general probabilistic methods for analysis of variance. Linear models and ANOVA are heavily used tools in many other disciplines of scientific research. The usual F-statistic is unsatisfactory for microarray data, which explore many thousand genes in parallel, with few replicates. We present three potential one-way ANOVA statistics in a parametric statistical framework. The aim is to separate genes that are differently regulated across several treatment conditions from those with equal regulation. The statistics have different features and are evaluated using both real and simulated data. Our statistic B1 generally shows the best performance, and is extended for use in an algorithm that groups cell lines by equal expression levels for each gene. An extension is also outlined for more general ANOVA tests including several factors. The methods presented are implemented in the freely available statistical language R. They are available at http://www.math.uu.se/staff/pages/?uname=ingrid.

Journal Article↗

The hidden code in genomics: a tool for gene discovery.

Among new insights coming from the completion of sequencing of the human genome, reported in Nature and Science, are clues of how evolution has increased the complexity of species, and in particular how the genetic code has enabled this process. It is clear that life has not only evolved by increasing the number of genes, but also by ingeniously evolving an efficient code for expressing diversity in the building blocks (i.e. the amino acids). The rules of nucleic acid base pairing and the classification of amino acids according to hydrophobicity/hydrophilicity relationships define a binary DNA code, which determines the general biophysical characteristics of proteins. Sense and antisense strands can encode protein segments having inverted and complementary hydropathy. The underlying binary code controls association and dissociation of proteins and presumably represents a primordial code that might have emerged in the early stages of self-organizing biochemical cycles. It is the purpose of this communication to provide a perspective of the code in the context of a binary language from its primordial origin to its present day format and to propose to use this code as a genomic mining tool.

Expressed Sequence Tags↗

Lack of biological significance in the 'linguistic features' of noncoding DNA--a quantitative analysis.

Recently, the application of two statistical methods (related to Zipf's distribution and Shannon's redundancy), called 'linguistic' tests, to the primary structure of DNA sequences of living organisms has excited considerable interest. Of particular importance is the claim that noncoding DNA sequences in eukaryotes display specific 'linguistic' features, being reminiscent of natural languages. Furthermore, this implies that noncoding regions of DNA may carry some new, thus far unknown, biological information which is revealed by these tests. In this paper these claims are tested quantitatively. With the aid of computer simulations of natural DNA sequences, and by applying the same 'linguistic' tests to both natural and artificial sequences, we investigate in detail the reasons of the appearance of the claimed 'linguistic' features and the associated differences between coding and noncoding DNAs. The presented results show quantitatively that the 'linguistic' tests failed to reveal any new biological information in (noncoding or coding) DNA.

Base Composition↗

[Differentiation of closely related and distant human populations by multilocus DNA fingerprinting].

Using multilocus DNA fingerprinting with phage M13 DNA as a probe, we have investigated a heterogeneous group of four human populations from Eastern Europe and Northeastern Asia. These populations belong to two language families: Indo-European (Eastern Slavonic branch: Russians, Belarussians) and Altaian (Turkic branch: Yakuts). The experimental results were treated by different statistical techniques: cluster analysis, multidimensional scaling, and multiple correspondence analysis. Coefficients of genetic differentiation were estimated using similarity indices and heterozygosities. The results of our study demonstrated similarity of Belarussian populations and significant differences between the group of Slavonic populations and Yakuts.

Asian People↗

Mechanisms of disease: neurogenetics of MeCP2 deficiency.

Rett syndrome (RTT) is unique among genetic, chromosomal and other developmental disorders because of its extreme female gender bias, early normal development, and subsequent developmental regression with loss of motor and language skills. RTT is caused by heterozygosity for mutations in the X-linked gene MECP2, which encodes methyl-CpG binding protein 2. MeCP2 is a multifunctional protein that can act as an architectural chromatin-binding protein, a function that is unrelated to its ability to bind methyl-CpG and to attract chromatin modification complexes. Inactivating mutations that cause RTT in females are not prenatally lethal in males, but lead to profound congenital encephalopathy. Molecular diagnoses of RTT, through demonstration of a MECP2 mutation, made at an early stage of the disorder, usually confirm the sporadic nature and very low recurrence risk of the condition. A positive DNA test result, however, also predicts the inevitable clinical course, given the lack of effective intervention. Initial hypotheses indicating that the MeCP2 protein acts as a genome-wide transcriptional repressor were not confirmed by global gene expression studies in various tissues of individuals with RTT and mouse models of MeCP2 deficiency. Rather, recent evidence points to low-magnitude effects of a small number of genes--including the brain--derived neurotrophic factor pathway and glucocorticoid response genes-that might affect formation and maturation of synapses or synaptic function in postmitotic neurons.

Animals↗

Data-driven computer simulation of human cancer cell.

Using the Diagrammatic Cell Language trade mark, Gene Network Sciences (GNS) has created a network model of interconnected signal transduction pathways and gene expression networks that control human cell proliferation and apoptosis. It includes receptor activation and mitogenic signaling, initiation of cell cycle, and passage of checkpoints and apoptosis. Time-course experiments measuring mRNA abundance and protein activity are conducted on Caco-2 and HCT 116 colon cell lines. These data were used to constrain unknown regulatory interactions and kinetic parameters via sensitivity analysis and parameter optimization methods contained in the DigitalCell computer simulation platform. FACS, RNA knockdown, cell growth, and apoptosis data are also used to constrain the model and to identify unknown pathways, and cross talk between known pathways will also be discussed. Using the cell simulation, GNS tested the efficacy of various drug targets and performed validation experiments to test computer simulation predictions. The simulation is a powerful tool that can in principle incorporate patient-specific data on the DNA, RNA, and protein levels for assessing efficacy of therapeutics in specific patient populations and can greatly impact success of a given therapeutic strategy.

Apoptosis↗