Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Using credibility intervals instead of hypothesis tests in SAGE analysis.

MOTIVATION: Statistical methods usually used to perform Serial Analysis of Gene Expression (SAGE) analysis are based on hypothesis testing. They answer the biologist's question: 'what are the genes with differential expression greater than r with P-value smaller than P?'. Another useful and not yet explored question is: 'what is the uncertainty in differential expression ratio of a gene?'. RESULTS: We have used Bayesian model for SAGE differential gene expression ratios as a more informative alternative to hypothesis tests since it provides credibility intervals. AVAILABILITY: The model is implemented in R statistical language script and is available under GNU/GLP copyleft at supplemental web site. SUPPLEMENTARY INFORMATION: http://www.ime.usp.br/~rvencio/SAGEci/

Algorithms↗

Network dynamics and cell physiology.

Complex assemblies of interacting proteins carry out most of the interesting jobs in a cell, such as metabolism, DNA synthesis, movement and information processing. These physiological properties play out as a subtle molecular dance, choreographed by underlying regulatory networks. To understand this dance, a new breed of theoretical molecular biologists reproduces these networks in computers and in the mathematical language of dynamical systems.

CDC2 Protein Kinase↗

Who wrote the book of life? Information and the transformation of molecular biology, 1945-55.

This paper focuses on the opening of a discursive space: the emergence of informational and scriptural representations of life and hereditiy and their self-negating consequences for the construction of biological meaning. It probes the notion of writing and the book of life and shows how molecular biology's claims to a status of language and texuality undermines its own objective of control. These textual significations were historically contingent. The informational representations of heredity and life were not an outcome of the internal cognitive momentum of molecular biology; they were not a logical necessity of the unravelling of the base-pairing of the DNA double-helix. They were transported into molecular biology still within the protein paradigm of the gene in the 1940s and permeated nearly every discipline in the life and social sciences. These information-based models, metaphors, linguistic, and semiotic tools which were central to the formulation of the genetic code were transported into molecular biology from cybernetics, information theory, electronic computing, and control and communication systems--technosciences that were deeply embedded with the military experiences of World War II and the Cold War. The information discourse thus became fixed in molecular biology not because it worked in the narrow epistemic sense (it did not), but because it positioned molecular biology within postwar discourse and culture, perhaps within the transition to a post-modern information-based society.

History, 20th Century↗

[Origin of Hakka and Hakkanese: a genetics analysis].

Hakka is a distinctive Han Chinese population in Southern China speaking Hakkanese. The origin of Hakka has been controversial. In this report, we analyzed Y chromosomal markers in 148 Hakka males. Principle component analysis of Y-SNP haplotype distribution shows Hakka is clusteed strongly with the Han in Northern China, and is also close to She, a Hmong-Mien-speaking population, while the general Southern Han is fairly close to Daic populations. Admixture analysis revealed that the relative genetic contribution 80.2% (Han), 13% (She) and 6.8% (Kam) in Hakka. The network of Y-STR haplotype of M7 individuals in all concerned populations suggested two possible origins of Hmong-Mien contribution in Hakka: One is from Hubei and the other is from Canton. The Kam contribution in Hakka is likely from Kan-Yue, the ancient aborigine of Kiangsi (Jiangxi). The frequency of 9bp-deletion in Region V of mitochondrial DNA of Hakka is 19.7%, which is quite close to She but far from Han. We therefore concluded that genetically the majority of Hakka gene pool shall come from North Han with She contributing the most among all non-Han groups. Regarding the Hmong-Mien character of Hakkanese, the genetic structure of Hakka shows their core may be Kim-man, the ancient Hmong-Mien. We hypothesized that a great number of Han people from North China join this population in succession. Southern Chinese dialects, such as Hakkanese may also be those languages of Southern aborigines at first, and turn to extant appearance under the continuance effect of Northern Chinese.

China↗

A complex control circuit. Regulation of immunity in temperate bacteriophages.

Temperate bacteriophages can display in a stable way two essentially different behaviours. In the immune state, a gene (cI) produces a repressor which prevents expression of all the other viral genes; in the non-immune state the typically viral functions are expressed. The choice between the two pathways and the establishment of one of them have much in common with cell determination and differentiation. This choice depends on a complex control system, in fact one of the most intricate nets of regulation known in some detail. Our paper provides a formal description and partial analysis of this regulatory net. It is shown that even for relatively simple known models, this kind of analysis uncovers predictions which had previously remained hidden. Some of these predictions were checked experimentally. The experimental part chiefly deals with the efficiency of lysogenization by thermoinducible lambda phage carrying mutations in one or more of the regulatory genes, N, cro and cII. Although N- mutations are widely known for preventing efficient integration, and both N- and cII mutations for preventing efficient establishment of immunity, it is shown that, as predicted by a simple model, both N- and cII- phage efficiently lysogenize at low temperature if they are in addition cro-. In contrast with lambda N- cro+, lambda N- cro- is not propagated as a plasmid at low temperature, precisely because it establishes immunity too efficiently. Genetic control circuits are described in terms of sets of logic equations, which relate the state of expression of genes or of chemical reactions (functions) to input (genetic and environmental) variables and to the presence of gene and reaction products (internal, or memorization varibles). From the set of equations, one derives a matrix which shows the stable stationary states (if any) of the system, and from which one can derive the pathways (temporal sequences of states) consistent with the model. This kind of analysis is complementary to the more widely used analysis based on differential equations; it allows one to analyze in less detail more complex systems. The language might be used as well, mutatis mutandis, in fields very different from genetics. The last part of the discussion deals with the role of positive feedback loops in our specific problem (establishment and maintenance of immunity in temperate bacteriophages) and in developmental genetics in general. As a generalization of an old idea, it is suggested that cell determination (for a given character) depends on a set of genes whose interaction constitutes a positive feedback loop. Such a system has two stable stationary states: which one is chosen will usually depend on additional controls grafted on the loop.

Bacteriophages↗

The language of genes.

Linguistic metaphors have been woven into the fabric of molecular biology since its inception. The determination of the human genome sequence has brought these metaphors to the forefront of the popular imagination, with the natural extension of the notion of DNA as language to that of the genome as the 'book of life'. But do these analogies go deeper and, if so, can the methods developed for analysing languages be applied to molecular biology? In fact, many techniques used in bioinformatics, even if developed independently, may be seen to be grounded in linguistics. Further interweaving of these fields will be instrumental in extending our understanding of the language of life.

Computational Biology↗

Fractals for multicyclic synthesis conditions of biopolymers. Examples of oligonucleotide synthesis measured by high-performance capillary electrophoresis and ion-exchange high-performance liquid chromatography.

We have developed models of patterns for nucleotide chain growth. These patterns are measurable by high-performance capillary electrophoresis and ion-exchange high-performance liquid chromatography in crude products of solid-phase synthesized 30mer and 65mer oligodeoxyribonucleotide target sequences N. We introduce mathematical methods for finding characteristic values d(o) and p(o) for constant chemical modes of growth as well as d and p for non-constant chemical modes of growth (d = probability of propagation, p = probability of termination). These methods are employed by presenting the accompanying computer software developed by us in C code, Mathematica R languages, and Fortran. Characteristic values of the parameters d, p, and the target nucleotide length N describe the complete composition of the crude product. From this we have developed the relation 2 - [N/(N - 1)]/Da, measurable(N,d) as a universal quantitative measure for multicyclic synthesis conditions (D, fractal dimension and similarity exponent, respectively). We use this mathematical treatment to compare the efficiency of oligodeoxyribonucleotide syntheses of different target length N on polymer support materials. Further, we analyze selected syntheses of short and long oligodeoxyribonucleotides as well as single-stranded DNA sequences by well-known empirical autocorrelation, fast Fourier transformation, and embedding dimension techniques.

Base Sequence↗

Supervised cluster analysis for microarray data based on multivariate Gaussian mixture.

MOTIVATION: Grouping genes having similar expression patterns is called gene clustering, which has been proved to be a useful tool for extracting underlying biological information of gene expression data. Many clustering procedures have shown success in microarray gene clustering; most of them belong to the family of heuristic clustering algorithms. Model-based algorithms are alternative clustering algorithms, which are based on the assumption that the whole set of microarray data is a finite mixture of a certain type of distributions with different parameters. Application of the model-based algorithms to unsupervised clustering has been reported. Here, for the first time, we demonstrated the use of the model-based algorithm in supervised clustering of microarray data. RESULTS: We applied the proposed methods to real gene expression data and simulated data. We showed that the supervised model-based algorithm is superior over the unsupervised method and the support vector machines (SVM) method. AVAILABILITY: The program written in the SAS language implementing methods I-III in this report is available upon request. The software of SVMs is available in the website http://svm.sdsc.edu/cgi-bin/nph-SVMsubmit.cgi

Algorithms↗

Racial and ethnic disparity in participation in DNA collection at the Atlanta site of the National Birth Defects Prevention Study.

Genetic risk factors are a critical component of many epidemiologic studies; however, concerns about genetic research might affect participants' willingness to enroll. The authors assessed factors associated with completion of mailed buccal-cell collection kits following telephone interviews at the Atlanta, Georgia, study site of the National Birth Defects Prevention Study. Pregnant women who were interviewed after June 30, 1999, and had an estimated delivery date of December 31, 2002, or earlier were included (n = 1,606). For this time period, overall interview participation was 71.9%. Among those interviewed, 47.6% completed the buccal-cell collection kit (61.1% of non-Hispanic Whites, 34.9% of non-Hispanic Blacks, and 39.1% of Hispanics). Non-Hispanic White race/ethnicity, an English-language (vs. Spanish) interview, receipt of a redesigned mailing packet and an additional $20 incentive, and consumption of folic acid were associated with higher buccal-cell kit participation. Among non-Hispanic White mothers, higher education, intending to become pregnant, and having a child with a birth defect were associated with increased participation. Among non-Hispanic Black mothers, receipt of the redesigned packet and $20 incentive was associated with increased participation. Among Hispanic mothers, an English-language interview, higher education, and receipt of the redesigned packet and $20 incentive were associated with increased participation. At this study site, minority groups were less likely to participate in DNA collection. Factors associated with participation varied by race/ethnicity.

Attitude to Health↗

Speech sound disorder influenced by a locus in 15q14 region.

Despite a growing body of evidence indicating that speech sound disorder (SSD) has an underlying genetic etiology, researchers have not yet identified specific genes predisposing to this condition. The speech and language deficits associated with SSD are shared with several other disorders, including dyslexia, autism, Prader-Willi Syndrome (PWS), and Angelman's Syndrome (AS), raising the possibility of gene sharing. Furthermore, we previously demonstrated that dyslexia and SSD share genetic susceptibility loci. The present study assesses the hypothesis that SSD also shares susceptibility loci with autism and PWS. To test this hypothesis, we examined linkage between SSD phenotypes and microsatellite markers on the chromosome 15q14-21 region, which has been associated with autism, PWS/AS, and dyslexia. Using SSD as the phenotype, we replicated linkage to the 15q14 region (P=0.004). Further modeling revealed that this locus influenced oral-motor function, articulation and phonological memory, and that linkage at D15S118 was potentially influenced by a parent-of-origin effect (LOD score increase from 0.97 to 2.17, P=0.0633). These results suggest shared genetic determinants in this chromosomal region for SSD, autism, and PWS/AS.

Autistic Disorder↗

Image library of biological macromolecules.

An Image Library of Biological Macromolecules is described, which contains image and text files related to structures of biological macromolecules. Currently, the Library has approximately 3000 image files of approximately 300 structures of biological macromolecules whose coordinates are available in the Protein Data Bank and in the Nucleic Acid Database. The entries include all RNA structures, approximately 70 DNA structures, 150 proteins and a few carbohydrates. The Library contains further images of amino acids, of standard and modified nucleotides and of nucleic acid model structures. Each entry consists of an annotation file with bibliographic and sequence information and possibly comments, of a color-coded distance plot and of structure images. Almost all of the images are available both in a mono and in a stereo representation. Standard procedures for generating these images were strictly avoided. Therefore, mixed rendering, coloring and labeling techniques were used extensively. Since May 1995 the Library has a growing division of images in the new Virtual Reality Modeling Language (VRML) format. The Image Library of Biological Macromolecules can be accessed via the World-Wide Web (http://www.imb-jena.de/IMAGE.html). There is a large number of structures determined by experimental and/or modeling techniques which are not intended to be included into the Protein Data Bank or Nucleic Acid Database for some reason. The Image Library could be a repository of these structures and of images of these and other structures of biological macromolecules including structures which are not known at atomic detail. Authors who are willing to make available images or coordinates to the scientific community via the Image Library of Biological Macromolecules are requested to contact the author.

Computer Communication Networks↗

A universal time-varying distributed H-system of degree 2.

A time-varying distributed H system is a splicing system which has the following feature: at different moments one uses different sets of splicing rules. The number of these sets is called the degree of the system. The passing from a set of rules to another one is specified in a cycle. It is a well known fact that any formal language can be generated by a time-varying distributed H-system of degree at least 7. Here we prove that there are universal time-varying distributed H-systems of degree 2. The question of whether or not there are universal time-varying distributed H-systems of degree 1 remains open.

Animals↗

Haplotype parsing: methods for extracting information from human genetic variations.

While the shared consensus genetic sequence of our species contains a great deal of information about our common biology, there is also much to be learned from the subtle genetic variations across our species. These variations are believed to be generally of little or no direct functional significance and predominantly reflect the chance accumulation of small genetic changes since our emergence as a species. Therefore, they carry little useful information when observed in a single individual. When tallied across a whole population though, these chance mutations can teach us a great deal about our evolutionary history and the patterns of inheritance in particular individuals. In particular, frequently observed patterns of single nucleotide polymorphisms (SNPs) in a population can identify segments of chromosome that have been passed down largely intact through long stretches of our evolution. Finding these frequently conserved chromosomal segments, or haplotypes, and developing methods to identify haplotype patterns in particular individuals, will in turn help us to identify those particular segments that carry genetic factors influencing risk for many common human diseases. To make the best use of this data, we will need to develop new models for the encoding of information in genome variations--the "language of genetic variation"--and new algorithms for fitting datasets to those models. This article surveys past work by the author and colleagues on this problem, utilising computational methods for locating frequent patterns in haploid sequence data, and "parsing" sequences so as to optimally explain them given the knowledge of the general population structure. The author's recent work in this area has been compiled into a set of computational tools available at http://www-2.cs.cmu.edu/~russells/software/hapmotif.html.

Algorithms↗

DNAskew: statistical analysis of base compositional asymmetry and prediction of replication boundaries in the genome sequences.

Sueoka and Lobry declared respectively that, in the absence of bias between the two DNA strands for mutation and selection, the base composition within each strand should be A=T and C=G (this state is called Parity Rule type 2, PR2). However, the genome sequences of many bacteria, vertebrates and viruses showed asymmetries in base composition and gene direction. To determine the relationship of base composition skews with replication orientation, gene function, codon usage biases and phylogenetic evolution, in this paper a program called DNAskew was developed for the statistical analysis of strand asymmetry and codon composition bias in the DNA sequence. In addition, the program can also be used to predict the replication boundaries of genome sequences. The method builds on the fact that there are compositional asymmetries between the leading and the lagging strand for replication. DNAskew was written in Perl script language and implemented on the LINUX operating system. It works quickly with annotated or unannotated sequences in GBFF (GenBank flatfile) or fasta format. The source code is freely available for academic use at http://www.epizooty.com/pub/stat/DNAskew.

Algorithms↗

Object-oriented knowledge bases for the analysis of prokaryotic and eukaryotic genomes.

The amount of biological sequences introduced in the general collections, and the growing complexity of the biological knowledge require the construction of models to formalize this knowledge and particularly the relationships between several data types. Two examples of such situations are presented here, they result from the biological research lead in our team in the field of molecular evolution. ColiGene is a modelling of E. coli genetics devoted to the analysis of relationships between genomic sequences and gene expressivity. MultiMap implements a new formalization of genome maps allowing manipulation of "maps of maps" in two species. Application of ColiGene and MultiMap are not restricted to molecular evolution and, for instance, MultiMap offers new capabilities for infering data on a genome from knowledge on another species. This could be essential for many mapping projects (human, mouse but also other mammals like pig). Development and implementation of those models have been done using an object-oriented knowledge base management system (SHIRKA) interfaced with a dedicated genomic data base management system (ACNUC). Graphical interfaces have been designed to give an environment similar to the biological representations used by biologists.

Animals↗

A model for the mechanism of polymerase translocation.

A general mechanism for polymerase translocation is elaborated. The central feature of this mechanism is that a rapid translocational equilibrium is established after each cycle of nucleoside monophosphate incorporation such that the polymerase distributes itself by diffusional sliding between all accessible positions on the template with relative occupancy determined by relative free energy. While alternative models for translocation have not been fully developed, much of the language currently used to describe this step suggests an active mechanism coupled to conformational transitions in the polymerase. For example, a recent study of force generation by Escherichia coli RNA polymerase during transcription suggests that it is a mechanoenzyme analogous to kinesin of myosin motor proteins. While the proposed mechanism does not rule out conformational transitions during polymerase translocation, it suggests that they may be unnecessary and that translocation can be explained in terms of the affinity of the active site for nucleoside triphosphate and the relative free energies of the polymerase bound at different positions on the template. This mechanism makes specific predictions which are borne out experimentally with polymerases as distinct as E. coli DNAP I, phage T7 RNAP, and E. coli RNAP.

Bacteriophage T7↗

Knowledge representation of signal transduction pathways.

MOTIVATIONS: Signal transduction is the common term used to define a diverse topic that encompasses a large body of knowledge about the biochemical mechanisms. Since most of the knowledge of signal transduction resides in scientific articles and is represented by texts in natural language or by diagrams, there is the need of a knowledge representation model for signal transduction pathways that can be as readily processed by a computer as it is easily understood by humans. RESULTS: A signal transduction pathway representation model is presented. It is based on a compound graph structure and is designed to handle the diversity and hierarchical structure of pathways. A prototype knowledge base was implemented on a deductive database and a number of biological queries are demonstrated on it.

Amino Acid Motifs↗

A plant dialect of the histone language.

The genome contains all the information needed to build an organism. However, during differentiation and development, additional epigenetic information determines the functional state of cells and tissues. This epigenetic information can be introduced by cytosine methylation and by marking nucleosomal histones. The code written on histones consists of post-translational modifications, including acetylation and methylation. In contrast to the universal nature of the DNA code, the histone language and its decoding machinery differ among animals, plants and fungi. Plant cells have retained totipotency to generate the entire plant and maintained the ability to dedifferentiate, which suggests that the establishment and maintenance of epigenetic information differs from animals. Here, I aim to summarize the histone code and plant-specific aspects of setting and translating the code.

Acetylation↗