Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A genome scan for developmental dyslexia confirms linkage to chromosome 2p11 and suggests a new locus on 7q32.

Developmental dyslexia is a distinct learning disability with unexpected difficulty in learning to read despite adequate intelligence, education, and environment, and normal senses. The genetic aetiology of dyslexia is heterogeneous and loci on chromosomes 2, 3, 6, 15, and 18 have been repeatedly linked to it. We have conducted a genome scan with 376 markers in 11 families with 38 dyslexic subjects ascertained in Finland. Linkage of dyslexia to the vicinity of DYX3 on 2p was confirmed with a non-parametric linkage (NPL) score of 2.55 and a lod score of 3.01 for a dominant model, and a novel locus on 7q32 close to the SPCH1 locus was suggested with an NPL score of 2.77. The SPCH1 locus has previously been linked with a severe speech and language disorder and autism, and a mutation in exon 14 of the FOXP2 gene on 7q32 has been identified in one large pedigree. Because the language disorder associated with the SPCH1 locus has some overlap with the language deficits observed in dyslexia, we sequenced the coding region of FOXP2 as a candidate gene for our observed linkage in six dyslexic subjects. No mutations were identified. We conclude that DYX3 appears to be important for dyslexia susceptibility in many Finnish families, and a suggested linkage of dyslexia to chromosome 7q32 will need verification in other data sets.

Chromosome Mapping↗

Iberia: population genetics, anthropology, and linguistics.

Basques, Portuguese, Spaniards, and Algerians have been studied for HLA and mitochondrial DNA markers, and the data analysis suggests that pre-Neolithic gene flow into Iberia came from ancient white North Africans (Hamites). The Basque language has also been used to translate the Iberian-Tartesian language and also Etruscan and Minoan Linear A. Physical anthropometry of Iberian Mesolithic and Neolithic skeletons does not support the demic replacement in Iberia of preexisting Mesolithic people by Neolithic people bearing new farming technologies from Europe and the Middle East. Also, the presence of cardial impressed pottery in western Mediterranean Europe and across the Maghreb (North Africa) coasts at the beginning of the Neolithic provides good evidence of pre-Neolithic circum-Mediterranean contacts by sea. In addition, pre-dynastic Egyptian El-Badari culture (4,500 years ago) is similar to southern Iberian Neolithic settlements with regard to pottery and animal domestication. Taking the genetic, linguistic, anthropological, and archeological evidence together with the documented Saharan area desiccation starting about 10,000 years ago, we believe that it is possible that a genetic and cultural pre-Neolithic flow coming from southern Mediterranean coasts existed toward northern Mediterranean areas, including at least Iberia and some Mediterranean islands. This model would substitute for the demic diffusion model put forward to explain Neolithic innovations in Western Europe.

Anthropology↗

Non-similarity combinatorial problems.

Similarity problems intensively investigated in computational molecular biology have the following two stringology models: find the longest string included in any string of a given finite language, and find the shortest string including every string of a given finite language. These two problems are exemplified by the two well-known pairs of problems, the longest common subsequence (or substring) problem and the shortest common supersequence (or superstring) problem, interpretations. In this paper we consider opposite problems connected with string non-inclusion relations: find the shortest string included in no string of a given finite language and find the longest string including no string of a given finite language. The predicate "string alpha is not included in string beta" is interpreted either as "alpha is not a subsequence of beta" or as "alpha is not a substring of beta". The main purpose is to determine the complexity status of the non-similarity problems. Using graph approaches, we present NP-hardness proofs for the first interpretation and polynomial-time algorithms for the second one. Special cases of the problems, and related issues are discussed.

Algorithms↗

Developmental aspects of infant's cry melody and formants.

This paper deals with the analysis of cry melodies (time variations of the fundamental frequency) as well as vocal tract resonance frequencies (formants) from infant cry signals. The increase of complexity of cry melodies is a good indicator for neuro-muscular maturation as well as for the evaluation of pre-speech development. The variation of formant frequencies allows an estimation of articulatory activity during pre-speech vocalization. Subjects are three pairs of healthy identical twins (monocygozity determined by DNA-fingerprint). Spontaneous cries of these six children were recorded at different ages: 8th-9th week, 15th-17th week and 23rd-24th week. Analysis of 136 cry melodies and intensity contours was made using KAY-CSL 4300/MDVP. For formant estimation a spectral parametric technique was applied, which was based on autoregressive models (Digital spectral analysis with applications, 1987) whose order is adaptively estimated on subsequent signal frames by means of a new method (Med. Eng. Phys. 20 (1998) 432; Utras. Med. Biol. 21 (1995) 793). Cry melodies exhibited an increasing complexity during the observation period. Beginning with the second observation period (15th-17th week) an increasing coupling and tuning between melody and resonance frequencies was observed, which was interpreted as "intentional" articulatory activity. Possible applications are in cry diagnosis as well as in the evaluation of pre-speech development.

Crying↗

Mutually symmetric and complementary triplets: differences in their use distinguish systematically between coding and non-coding genomic sequences.

The general property of asymmetry in word use in meaningful texts written in a variety of languages, motivates a quantification of the differences in the use of mutually symmetric triplets in genomic sequences. When this is done in the three reading frames, high values found for one of them are used as indication that the sequence is coding for a protein. Moreover, a similar quantification of the differences in the use of complementary triplets is introduced, again with predictive power of the coding character of a sequence. This method reflects the non-equivalence between sense and anti-sense strand of a coding segment. In both approaches, "linguistic asymmetry" in coding sequences is related to the form of the genetic code and to the bias in codon usage and amino acid use skews.

Algorithms↗

Mild-onset presentation of Canavan's disease associated with novel G212A point mutation in aspartoacylase gene.

We describe two sisters with a mild-onset variant of Canavan's disease who presented at age 50 and 19 months with developmental delay but without macrocephaly, hypotonia, spasticity, or seizures. Remarkably, both patients had age-appropriate head control, gross motor development, and muscle tone. There were very mild deficits in fine motor skills, coordination, and gait. Both sisters had a history of strabismus, but otherwise vision was normal. The older child showed evidence of mild cognitive and social impairment, whereas language and behavior were normal for age in the infant. Both patients were found to be compound heterozygotes for C914A (A305E) and G212A (R71H) mutations in ASPA. Like all other known ASPA mutations, this previously unknown G212A mutation appears to have low absolute enzyme activity. Nevertheless, it is associated in these patients with an extremely benign phenotype that is highly atypical of Canavan's disease. Biochemical and clinical data were evaluated using a generalized linear mixed model generated from 25 other subjects with Canavan's disease. There were statistically significant differences in brain chemistry and clinical evaluations, supporting a distinct variant of Canavan's disease. Future studies of ASPA enzyme structure and gene regulation in these subjects could lead to a better understanding of Canavan's pathophysiology and improvements in ASPA gene therapy.

Adult↗

Similarities inferred from the studies of long range correlations among mitochondrial DNA sequences.

Existence of long range correlations within the DNA sequences of living organism has immense importance in understanding the language of DNA sequences. Recently it has been reported that long range correlations occur in DNA sequences. Some investigators claimed that these type of correlations occur only on intron containing DNA sequences. Some observers, however, have the opinion that long range correlations do not distinguish between the intron containing DNA sequences and intronless DNA sequences. The biological origin of long range correlations in the DNA sequences is not clearly known. In this paper we have demonstrated that long range correlations also occur on intronless mitochondrial DNA sequences, indicating that these special type of correlations are not the unique features for intron containing DNA sequences. We have also demonstrated that long range correlations simply originate in the region around which there is a large variation of pyrimidine and purine ratios. The similarities among the mitochondrial DNA sequences can be inferred by computing the fractal exponents in the region where there is a large variation of pyrimidine and purine ratio, as well as in the region where the ratio of pyrimidine and purine fluctuates in a nearly constant manner. In other words the similarities among the mitochondrial DNA sequences cannot be inferred by calculating the fractral exponents for the whole sequence.

Animals↗

SVM classification of human intergenic and gene sequences.

Despite constant improvement in prediction accuracy, gene-finding programs are still unable to provide automatic gene discovery with the desired correctness. This paper presents an analysis of gene and intergenic sequences from the point of view of language analysis, where gene and intergenic regions are regarded as two different subjects written in the four-letter alphabet {A,C,G,T}, and high frequency simple sequences are taken as keywords. A measurement alpha(l(tau)) was introduced to describe the relative repeat ratio of simple sequences. Threshold values were found for keyword selections. After eliminating 'noise', 178 short sequences were selected as keywords. DNA sequences are mapped to 178-dimensional Euclidean space, and SVM was used for prediction of gene regions. We showed by cross-validation that the program we developed could predict 93% of gene sequences with 7% false positives. When tested on a long genomic multi-gene sequence, our method improved nucleotide level specificity by 21%, and over 60% of predicted genes corresponded to actual genes.

Algorithms↗

Genetic counseling for fragile x syndrome: updated recommendations of the national society of genetic counselors.

These recommendations describe the minimum standard criteria for genetic counseling and testing of individuals and families with fragile X syndrome, as well as carriers and potential carriers of a fragile X mutation. The original guidelines (published in 2000) have been revised, replacing a stratified pre- and full mutation model of fragile X syndrome with one based on a continuum of gene effects across the full spectrum of FMR1 CGG trinucleotide repeat expansion. This document reviews the molecular genetics of fragile X syndrome, clinical phenotype (including the spectrum of premature ovarian failure and fragile X-associated tremor-ataxia syndrome), indications for genetic testing and interpretation of results, risks of transmission, family planning options, psychosocial issues, and references for professional and patient resources. These recommendations are the opinions of a multicenter working group of genetic counselors with expertise in fragile X syndrome genetic counseling, and they are based on clinical experience, review of pertinent English language articles, and reports of expert committees. These recommendations should not be construed as dictating an exclusive course of management, nor does use of such recommendations guarantee a particular outcome. The professional judgment of a health care provider, familiar with the facts and circumstances of a specific case, will always supersede these recommendations.

Alleles↗

PDA: Pooled DNA analyzer.

BACKGROUND: Association mapping using abundant single nucleotide polymorphisms is a powerful tool for identifying disease susceptibility genes for complex traits and exploring possible genetic diversity. Genotyping large numbers of SNPs individually is performed routinely but is cost prohibitive for large-scale genetic studies. DNA pooling is a reliable and cost-saving alternative genotyping method. However, no software has been developed for complete pooled-DNA analyses, including data standardization, allele frequency estimation, and single/multipoint DNA pooling association tests. This motivated the development of the software, 'PDA' (Pooled DNA Analyzer), to analyze pooled DNA data. RESULTS: We develop the software, PDA, for the analysis of pooled-DNA data. PDA is originally implemented with the MATLAB language, but it can also be executed on a Windows system without installing the MATLAB. PDA provides estimates of the coefficient of preferential amplification and allele frequency. PDA considers an extended single-point association test, which can compare allele frequencies between two DNA pools constructed under different experimental conditions. Moreover, PDA also provides novel chromosome-wide multipoint association tests based on p-value combinations and a sliding-window concept. This new multipoint testing procedure overcomes a computational bottleneck of conventional haplotype-oriented multipoint methods in DNA pooling analyses and can handle data sets having a large pool size and/or large numbers of polymorphic markers. All of the PDA functions are illustrated in the four bona fide examples. CONCLUSION: PDA is simple to operate and does not require that users have a strong statistical background. The software is available at http://www.ibms.sinica.edu.tw/%7Ecsjfann/first%20flow/pda.htm.

Algorithms↗

An automated comparative analysis of 17 complete microbial genomes.

MOTIVATION: As sequenced genomes become larger and sequencing becomes faster, there is a need to develop accurate automated genome comparison techniques and databases to facilitate derivation of genome functionality; identification of enzymes, putative operons and metabolic pathways; and to derive phylogenetic classification of microbes. RESULTS: This paper extends an automated pair-wise genome comparison technique (Bansal et al., Math. Model. Sci. Comput., 9, 1-23, 1998, Bansal and Bork, in First International Workshop of Declarative Languages, Springer, pp. 275-289, 1999) used to identify orthologs and gene groups to derive orthologous genes in a group of genomes and to identify genes with conserved functionality. Seventeen microbial genomes archived at ftp://ncbi.nlm.nih.gov/genbank/genomes have been compared using the automated technique. Data related to orthologs, gene groups, gene duplication, gene fusion, orthologs with conserved functionality, and genes specifically orthologous to Escherichia coli and pathogens has been presented and analyzed. AVAILABILITY: A prototype database is available at ftp://www.mcs.kent.edu/arvind/intellibio / orthos.html. The software is free for academic research under an academic license. The detailed database for every microbial genome in NCBI is commercially available through intellibio software and consultancy corporation (Web site: http://www.mcs.kent.edu/årvind/intellibio . html). CONTACT: arvind@mcs.kent.edu.

Algorithms↗

A linguistic representation of the regulation of transcription initiation. I. An ordered array of complex symbols with distinctive features.

The inadequacy of context-free grammars in the description of regulatory information contained in DNA gave the formal justification for a linguistic approach to the study of gene regulation. Based on that result, we have initiated a linguistic formalization of the regulatory arrays of 107 sigma 70 E. coli promoters. The complete sequences of promoter (Pr), operator (Op) and activator binding sites (I) have previously been identified as the smallest elements, or categories, for a combinatorial analysis of the range of transcription initiation of sigma 70 promoters. These categories are conceptually equivalent to phonemes of natural language. Several features associated with these categories are required in a complete description of regulatory arrays of promoters. We have to select the best way to describe the properties that are pertinent for the description of such regulatory regions. In this paper we define distinctive features of regulatory regions based on the following criteria: identification of subclasses of substitutable elements, simplicity, selection of the most directly related information, and distinction of one array among the whole set of promoters. Alternative ways to represent distances in between regulatory sites are discussed, permitting, together with a principle of precedence, the identification of an ordered set of complex symbols as a unique representation for a promoter and its associated regulatory sites. In the accompanying paper additional distinctive features of promoters and regulatory sites are identified.

Binding Sites↗

The language of methylation in genomics of eukaryotes.

Background studies have shown that 6-methylaminopurine (m6A) and 5-methylcytosine (m5C), detected in DNA, are products of its post-synthetic modification. At variance with bacterial genomes exhibiting both, eukaryotic genomes essentially carry only m5C in m5CpG doublets. This served to establish that, although a slight extra-S phase asymmetric methylation occurs de novo on 5'-CpC-3'/3'GpG-5', 5'-CpT-3'/3'-GpA-5', and 5'-CpA-3'/3'-GpT-5' dinucleotide pairs, a heavy methylation during S involves Okazaki fragments and thus semiconservatively newly made chains to guarantee genetic maintenance of -CH3 patterns in symmetrically dimethylated 5'-m5CpG-3'/3'-Gpm5C-5' dinucleotide pairs. On the other hand, whilst inverse correlation was observed between bulk DNA methylation, in S, and bulk RNA transcription, in G1 and G2, probes of methylated DNA helped to discover the presence of coding (exon) and uncoding (intron) sequences in the eukaryotic gene. These achievements led to the search for a language that genes regulated by methylation should have in common. Such a deciphering, initially providing restriction minimaps of hypermethylatable promoters and introns vs. hypomethylable exons, became feasible when bisulfite methodology allowed the direct sequencing of m5C. It emerged that, while in lymphocytes, where the transglutaminase gene (hTGc) is inactive, the promoter shows two fully methylated CpG-rich domains at 5 and one fully unmethylated CpG-rich domain at 3' (including the site +1 and a 5'-UTR), in HUVEC cells, where hTGc is active, in the first CpG-rich domain of its promoter four CpGs lack -CH3: a result suggesting new hypotheses on the mechanism of transcription, particularly in connection with radio-induced DNA demethylation.

Base Sequence↗

A worldwide analysis of AG molecular diversity inferred from serology.

Ten population samples from different geographic origins were tested serologically for the AG polymorphism of human beta-lipoproteins. Their haplotype frequencies were used with previously published data to perform a wide analysis of AG genetic differentiations throughout the world. Coancestry coefficients were computed from weighted F(ST)s among populations by using a matrix of molecular distances among AG haplotypes, which is here determined on the basis of DNA studies. Coancestry coefficients derived from unweighted F(ST)s and more classical Prevosti distances were computed on the same data and used for a comparison. In all cases a highly significant correlation was found between genetics and geography on a worldwide scale, while the significance of the correlation with linguistics differed. A test of significance of the pairwise F(ST)s among populations also gave different results depending on whether the molecular distance matrix among AG haplotypes was included. Globally, this study shows that in spite of being highly significantly correlated to each other, different genetic distance measures can lead to different interpretations of the same data set. Moreover, the elucidation of the molecular models related to the presently known serological polymorphisms may represent an additional tool for analyzing such polymorphisms in human population genetics studies.

Amino Acid Substitution↗

Unsupervised technique for robust target separation and analysis of DNA microarray spots through adaptive pixel clustering.

MOTIVATION: Microarray images challenge existing analytical methods in many ways given that gene spots are often comprised of characteristic imperfections. Irregular contours, donut shapes, artifacts, and low or heterogeneous expression impair corresponding values for red and green intensities as well as their ratio R/G. New approaches are needed to ensure accurate data extraction from these images. RESULTS: Herein we introduce a novel method for intensity assessment of gene spots. The technique is based on clustering pixels of a target area into foreground and background. For this purpose we implemented two clustering algorithms derived from k-means and Partitioning Around Medoids (PAM), respectively. Results from the analysis of real gene spots indicate that our approach performs superior to other existing analytical methods. This is particularly true for spots generally considered as problematic due to imperfections or almost absent expression. Both PX(PAM) and PX(KMEANS) prove to be highly robust against various types of artifacts through adaptive partitioning, which more correctly assesses expression intensity values. AVAILABILITY: The implementation of this method is a combination of two complementary tools Extractiff (Java) and Pixclust (free statistical language R), which are available upon request from the authors.

Algorithms↗

CGHPRO -- a comprehensive data analysis tool for array CGH.

BACKGROUND: Array CGH (Comparative Genomic Hybridisation) is a molecular cytogenetic technique for the genome wide detection of chromosomal imbalances. It is based on the co-hybridisation of differentially labelled test and reference DNA onto arrays of genomic BAC clones, cDNAs or oligonucleotides, and after correction for various intervening variables, loss or gain in the test DNA can be indicated from spots showing aberrant signal intensity ratios. Now that this technique is no longer confined to highly specialized laboratories and is entering the realm of clinical application, there is a need for a user-friendly software package that facilitates estimates of DNA dosage from raw signal intensities obtained by array CGH experiments, and which does not depend on a sophisticated computational environment. RESULTS: We have developed a user-friendly and versatile tool for the normalization, visualization, breakpoint detection and comparative analysis of array-CGH data. CGHPRO is a stand-alone JAVA application that guides the user through the whole process of data analysis. The import option for image analysis data covers several data formats, but users can also customize their own data formats. Several graphical representation tools assist in the selection of the appropriate normalization method. Intensity ratios of each clone can be plotted in a size-dependent manner along the chromosome ideograms. The interactive graphical interface offers the chance to explore the characteristics of each clone, such as the involvement of the clones sequence in segmental duplications. Circular Binary Segmentation and unsupervised Hidden Markov Model algorithms facilitate objective detection of chromosomal breakpoints. The storage of all essential data in a back-end database allows the simultaneously comparative analysis of different cases. The various display options facilitate also the definition of shortest regions of overlap and simplify the identification of odd clones. CONCLUSION: CGHPRO is a comprehensive and easy-to-use data analysis tool for array CGH. Since all of its features are available offline, CGHPRO may be especially suitable in situations where protection of sensitive patient data is an issue. It is distributed under GNU GPL licence and runs on Linux and Windows.

Algorithms↗

Pathway analysis of coronary atherosclerosis.

Large-scale gene expression studies provide significant insight into genes differentially regulated in disease processes such as cancer. However, these investigations offer limited understanding of multisystem, multicellular diseases such as atherosclerosis. A systems biology approach that accounts for gene interactions, incorporates nontranscriptionally regulated genes, and integrates prior knowledge offers many advantages. We performed a comprehensive gene level assessment of coronary atherosclerosis using 51 coronary artery segments isolated from the explanted hearts of 22 cardiac transplant patients. After histological grading of vascular segments according to American Heart Association guidelines, isolated RNA was hybridized onto a customized 22-K oligonucleotide microarray, and significance analysis of microarrays and gene ontology analyses were performed to identify significant gene expression profiles. Our studies revealed that loss of differentiated smooth muscle cell gene expression is the primary expression signature of disease progression in atherosclerosis. Furthermore, we provide insight into the severe form of coronary artery disease associated with diabetes, reporting an overabundance of immune and inflammatory signals in diabetics. We present a novel approach to pathway development based on connectivity, determined by language parsing of the published literature, and ranking, determined by the significance of differentially regulated genes in the network. In doing this, we identify highly connected "nexus" genes that are attractive candidates for therapeutic targeting and followup studies. Our use of pathway techniques to study atherosclerosis as an integrated network of gene interactions expands on traditional microarray analysis methods and emphasizes the significant advantages of a systems-based approach to analyzing complex disease.

Adult↗

Pegasys: software for executing and integrating analyses of biological sequences.

BACKGROUND: We present Pegasys--a flexible, modular and customizable software system that facilitates the execution and data integration from heterogeneous biological sequence analysis tools. RESULTS: The Pegasys system includes numerous tools for pair-wise and multiple sequence alignment, ab initio gene prediction, RNA gene detection, masking repetitive sequences in genomic DNA as well as filters for database formatting and processing raw output from various analysis tools. We introduce a novel data structure for creating workflows of sequence analyses and a unified data model to store its results. The software allows users to dynamically create analysis workflows at run-time by manipulating a graphical user interface. All non-serial dependent analyses are executed in parallel on a compute cluster for efficiency of data generation. The uniform data model and backend relational database management system of Pegasys allow for results of heterogeneous programs included in the workflow to be integrated and exported into General Feature Format for further analyses in GFF-dependent tools, or GAME XML for import into the Apollo genome editor. The modularity of the design allows for new tools to be added to the system with little programmer overhead. The database application programming interface allows programmatic access to the data stored in the backend through SQL queries. CONCLUSIONS: The Pegasys system enables biologists and bioinformaticians to create and manage sequence analysis workflows. The software is released under the Open Source GNU General Public License. All source code and documentation is available for download at http://bioinformatics.ubc.ca/pegasys/.

Computational Biology↗