Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Ultrasound-assisted lipoplasty.

BACKGROUND: Ultrasound-assisted lipoplasty (UAL) has been associated with particular types of complications and uncertain long-term effects arising from interactions between ultrasonic energy and living tissue. The present review seeks to address these issues. METHODS: Search strategy Three search strategies were devised to retrieve literature from Medline, Current Contents, Embase and Cochrane Library databases up until April 2000. Study selection Inclusion of papers was largely determined using a predetermined protocol. English language papers were selected. Acceptable study designs included randomized controlled trials, controlled clinical trials, case series or case reports. Data collection and analysis Thirty-six papers met the inclusion criteria. They were tabulated and critically appraised in terms of methodology and design, outcomes, and the possible influence of bias, confounding and chance. Other papers were also included to provide background material. RESULTS: There was little high-level evidence available comparing UAL and suction-assisted lipoplasty (SAL), with no conclusive evidence that UAL has a safety benefit, although low-quality evidence suggests that UAL is associated with reduced surgeon fatigue as well as increased operating times, slower aspiration rates and an increased learning curve. There is inadequate evidence to determine whether the theoretical potential for DNA damage from ultrasound is realized in the clinical setting. CONCLUSIONS: The evidence base for UAL is inadequate to determine the procedure's safety and efficacy. The potential for DNA damage must be investigated with appropriate in vivo animal models. Recommendations for the safe use of UAL are discussed.

Adipose Tissue↗

Structural analysis of disease-causing mutations in the P-subfamily of forkhead transcription factors.

Mutations in a number of forkhead transcription factors are associated with the development of inherited diseases in humans. Two closely related genes, FOXP2 and FOXP3, are implicated in two completely different human disorders. A point mutation in the forkhead domain of FOXP2 (R553H) is responsible for a severe speech and language disorder, while a series of missense mutations distributed over the forkhead domain of FOXP3 cause a fatal disorder called IPEX, characterized by immune dysregulation, polyendocrinopathy, and enteropathy. Homology model building techniques were used to generate atomic structures of FOXP2 and FOXP3, using the solution structures of the forkhead domain of the adipocyte-transcription factor FREAC-11 and AFX as templates. The impact of these disease-causing missense mutations on the three-dimensional structure, stability, and surface electrostatic charge distribution of the forkhead domains is examined here. The missense mutations R553H in FOXP2 and R397W in FOXP3 dramatically alter the electrostatic potentials of the molecular surface of their respective forkhead domains.

Amino Acid Sequence↗

What can developmental disorders tell us about the neurocomputational constraints that shape development? The case of Williams syndrome.

The uneven cognitive phenotype in the adult outcome of Williams syndrome has led some researchers to make strong claims about the modularity of the brain and the purported genetically determined, innate specification of cognitive modules. Such arguments have particularly been marshaled with respect to language. We challenge this direct generalization from adult phenotypic outcomes to genetic specification and consider instead how genetic disorders provide clues to the constraints on plasticity that shape the outcome of development. We specifically examine behavioral studies, brain imaging, and computational modeling of language in Williams syndrome but contend that our theoretical arguments apply equally to other cognitive domains and other developmental disorders. While acknowledging that selective deficits in normal adult patients might justify claims about cognitive modularity, we question whether similar, seemingly selective deficits found in genetic disorders can be used to argue that such cognitive modules are prespecified in infant brains. Cognitive modules are, in our view, the outcome of development, not its starting point. We note that most work on genetic disorders ignores one vital factor, the actual process of ontogenetic development, and argue that it is vital to view genetic disorders as proceeding under different neurocomputational constraints, not as demonstrations of static modularity.

Cognition Disorders↗

The evolution of information in the major transitions.

Maynard Smith and Szathmáry's analysis of the major transitions in evolution was based on changes in the way information is stored, transmitted and interpreted. With the exception of the transition to human linguistic societies, their discussion centred on changes in DNA and the genetic system. We argue that information transmitted by non-genetic means has played a key role in the major transitions, and that new and modified ways of transmitting non-DNA information resulted from them. We compare and attempt to categorise the major transitions, and suggest that the transition from RNA as both gene and enzyme to DNA as genetic material and proteins as enzymes may have been a double one. Unlike Maynard Smith and Szathmáry, we regard the emergence of the nervous system as a major transition. The evolution of a nervous system not only changed the way that information was transmitted between cells and profoundly altered the nature of the individuals in which it was present, it also led to a new type of heredity-social and cultural heredity-based on the transmission of behaviourally acquired information.

Animals↗

Informational parameters and randomness of mitochondrial DNA.

The informational content of genomes of nuclear and mitochondrial origin is examined. By using the parameters of Shannon's information theory the language of mitochondrial DNA is shown to be more similar to the language of bacterial DNA than to that of nuclear DNA in more evolutionarily advanced animals. Moreover, using the parameters of Kolmogorov's theory on randomness, genes of different organisms (Neurospora crassa and Saccharomyces cerevisiae) coding for the same protein (subunit 9 of ATPase) are shown to have, if both of mitochondrial origin, a similar degree of randomness, whereas genes coding for the same protein, both belonging to the same organisms, exhibit a quite different degree of randomness when one is of mitochondrial origin and the other of nuclear origin. These results are in favor of the symbiotic origin of mitochondria.

Adenosine Triphosphatases↗

MeSHer: identifying biological concepts in microarray assays based on PubMed references and MeSH terms.

UNLABELLED: MeSHer uses a simple statistical approach to identify biological concepts in the form of Medical Subject Headings (MeSH terms) obtained from the PubMed database that are significantly overrepresented within the identified gene set relative to those associated with the overall collection of genes on the underlying DNA microarray platform. As a demonstration, we apply this approach to gene lists acquired from a published study of the effects of angiotensin II (Ang II) treatment on cardiac gene expression and demonstrate that this approach can aid in the interpretation of the resulting 'significant' gene set. AVAILABILITY: The software is available at http://www.tm4.org. SUPPLEMENTARY INFORMATION: Results from the analysis of significant genes from the published Ang II study.

Artificial Intelligence↗

SeqHepB: a sequence analysis program and relational database system for chronic hepatitis B.

SeqHepB is a combination of a HBV genome sequence analysis program and a relational database that houses data collected from multiple data sources. Registered users can access the sequence analysis component of SeqHepB online for rapid and detailed interrogation of HBV genomic sequences. Its main function is to determine the HBV genotype, identify key mutations associated with antiviral resistance, and identify clinically important HBV mutants. All information generated is uploaded into a database and integrated with patient medical records, pathology laboratory tests, and supplemental virology results such as in vitro drug cross-resistance values. Combined with structured query language (SQL) queries developed in the database, it is possible to extract and correlate clinical, virological, and in vitro phenotypic data rapidly and efficiently. An important component of SeqHepB is its ability to integrate mutations detected within the reverse transcriptase (RT) and locate them onto a three-dimensional (3D) model of the HBV RT that can be viewed at any angle with known antiviral drug molecules in the catalytic pocket of the enzyme. SeqHepB will enable virologists and physicians to individualise patient management, cope with the explosion of antiviral associated HBV mutations, and to conduct cross-sectional retrospective or prospective studies on HBV-infected individuals during therapy.

DNA Mutational Analysis↗

cBrother: relaxing parental tree assumptions for Bayesian recombination detection.

UNLABELLED: Bayesian multiple change-point models accurately detect recombination in molecular sequence data. Previous Java-based implementations assume a fixed topology for the representative parental data. cBrother is a novel C language implementation that capitalizes on reduced computational time to relax the fixed tree assumption. We show that cBrother is 19 times faster than its predecessor and the fixed tree assumption can influence estimates of recombination in a medically-relevant dataset. AVAILABILITY: cBrother can be freely downloaded from http://www.biomath.org/dormanks/ and can be compiled on Linux, Macintosh and Windows operating systems. Online documentation and a tutorial are also available at the site.

Algorithms↗

Mixture models for assessing differential expression in complex tissues using microarray data.

MOTIVATION: The use of DNA microarrays has become quite popular in many scientific and medical disciplines, such as in cancer research. One common goal of these studies is to determine which genes are differentially expressed between cancer and healthy tissue, or more generally, between two experimental conditions. A major complication in the molecular profiling of tumors using gene expression data is that the data represent a combination of tumor and normal cells. Much of the methodology developed for assessing differential expression with microarray data has assumed that tissue samples are homogeneous. RESULTS: In this paper, we outline a general framework for determining differential expression in the presence of mixed cell populations. We consider study designs in which paired tissues and unpaired tissues are available. A hierarchical mixture model is used for modeling the data; a combination of methods of moments procedures and the expectation-maximization algorithm are used to estimate the model parameters. The finite-sample properties of the methods are assessed in simulation studies; they are applied to two microarray datasets from cancer studies. Commands in the R language can be downloaded from the URL http://www.sph.umich.edu/~ghoshd/COMPBIO/COMPMIX/.

Algorithms↗

Prediction of nucleoside-carcinogen reactivity. Alkylation of adenine, cytosine, guanine, and thymine and their deoxynucleosides by alkanediazonium ions.

MNDO semiempirical molecular orbital calculations for the SN2 alkylation of nucleic acid bases and deoxynucleosides by the methane-, ethane-, and propanediazonium ions are presented. An approximate correlation is demonstrated between the calculated relative activation enthalpies for attack at alternative base sites and the related experimental quantities for DNA modification by alkylnitrosoureas. The empirically observed shift from N- to O-alkylation with increasing complexity of the alkylating agent is reproduced by the calculations and rationalized by using an extension of a model worked out previously for the analogous reactions of simple nucleophiles. According to this model, the energetics of the related SN1 reactions, while not directly involved, have a profound influence on the SN2 transition-state geometries. For reactions in which the SN1 dissociation is unfavorable the forming bond to the incoming nucleophiles in the related SN2 transition state tends to be short and covalent interactions, which favor N-alkylation, play a significant role. When the SN1 reaction is more facile, the SN2 transition states are "looser" and the covalent interactions correspondingly smaller, leading to an overall shift away from N-alkylation. Consideration of the form of the electrostatic potential around the base, in conjunction with these ideas, provides a detailed explanation of the behavior of electrophiles toward the guanine N2-, 7-, and O6-positions. This model unifies much of the language already used in discussions of nucleic acid regiochemistry. At the same time it is consistent with the geometries and charge distributions in the transition states calculated for the gas-phase reaction processes.

Adenine↗

Exact computation of pattern probabilities in random sequences generated by Markov chains.

Observed patterns in macromolecular sequences are often considered as words and compared with their probabilities of occurring in random sequences. Calculation of these probabilities, however, often lacks rigour. We have developed an algorithm for exact computation of such probabilities for stochastic sequences that follow a Markov chain model. The method is applicable to the case that a random sequence contains one out of two given patterns P and Q, or both simultaneously. Another application yields the probability function P(x) that a sequence contains pattern P exactly x times. An application to patterns that include wild-card characters yields probabilities for homonucleotide clusters of a given length. We prove the probability of multiple runs of single nucleotides in the SV40 genome to be in accordance with the dinucleotide composition of the sequence, although it is in conflict with mononucleotide composition.

Algorithms↗

BNArray: an R package for constructing gene regulatory networks from microarray data by using Bayesian network.

UNLABELLED: BNArray is a systemized tool developed in R. It facilitates the construction of gene regulatory networks from DNA microarray data by using Bayesian network. Significant sub-modules of regulatory networks with high confidence are reconstructed by using our extended sub-network mining algorithm of directed graphs. BNArray can handle microarray datasets with missing data. To evaluate the statistical features of generated Bayesian networks, re-sampling procedures are utilized to yield collections of candidate 1st-order network sets for mining dense coherent sub-networks. AVAILABILITY: The R package and the supplementary documentation are available at http://www.cls.zju.edu.cn/binfo/BNArray/.

Algorithms↗

Mitochondrial DNA analysis of northwest African populations reveals genetic exchanges with European, near-eastern, and sub-Saharan populations.

Genetic studies have emphasized the contrast between North African and sub-Saharan populations, but the particular affinities of the North African mtDNA pool to that of Europe, the Near East, and sub-Saharan Africa have not previously been investigated. We have analysed 268 mtDNA control-region sequences from various Northwest African populations including several Senegalese groups and compared these with the mtDNA database. We have identified a few mitochondrial motifs that are geographically specific and likely predate the distribution and diversification of modern language families in North and West Africa. A certain mtDNA motif (16172C, 16219G), previously found in Algerian Berbers at high frequency, is apparently omnipresent in Northwest Africa and may reflect regional continuity of more than 20,000 years. The majority of the maternal ancestors of the Berbers must have come from Europe and the Near East since the Neolithic. The Mauritanians and West-Saharans, in contrast, bear substantial though not dominant mtDNA affinity with sub-Saharans.

Africa↗

Definite-clause grammars for the analysis of cis-regulatory regions in E. coli.

Based on an extensive collection of sigma 70 associated regulatory mechanisms, a grammatical model has been constructed that define the functional positions and combinations of sites within DNA regulatory regions. The syntactic rules and the dictionary implemented in a Prolog program were coupled to consensus matrices used as "sensors" to integrate a syntactic recognizer. A systematic comparison between the syntactic recognizer and the standard weight matrix methodology is presented using 12 regulatory proteins and the whole collection of about 130 sigma 70 DNA regulatory regions. On the average an increased sensitivity of 5 to 10 fold is obtained with this novel approach.

Binding Sites↗

Childhood-onset schizophrenia/autistic disorder and t(1;7) reciprocal translocation: identification of a BAC contig spanning the translocation breakpoint at 7q21.

Childhood-onset schizophrenia (COS) is defined by the development of first psychotic symptoms by age 12. While recruiting patients with COS refractory to conventional treatments for a trial of atypical antipsychotic drugs, we discovered a unique case who has a familial t(1;7)(p22;q21) reciprocal translocation and onset of psychosis at age 9. The patient also has symptoms of autistic disorder, which are usually transient before the first psychotic episode among 40-50% of the childhood schizophrenics but has persisted in him even after the remission of psychosis. Cosegregating with the translocation, among the carriers in the family available for the study, are other significant psychopathologies, including alcohol/drug abuse, severe impulsivity, and paranoid personality and language delay. This case may provide a model for understanding the genetic basis of schizophrenia or autism. Here we report the progress toward characterization of genomic organization across the translocation breakpoint at 7q21. The polymorphic markers, D7S630/D7S492 and D7S2410/D7S646, immediately flanking the breakpoint, may be useful for further confirming the genetic linkage for schizophrenia or autism in this region. Am. J. Med. Genet. (Neuropsychiatr. Genet.) 96:749-753, 2000. Published 2000 Wiley-Liss, Inc.

Autistic Disorder↗

Mice with truncated MeCP2 recapitulate many Rett syndrome features and display hyperacetylation of histone H3.

Mutations in the methyl-CpG binding protein 2 (MECP2) gene cause Rett syndrome (RTT), a neurodevelopmental disorder characterized by the loss of language and motor skills during early childhood. We generated mice with a truncating mutation similar to those found in RTT patients. These mice appeared normal and exhibited normal motor function for about 6 weeks, but then developed a progressive neurological disease that includes many features of RTT: tremors, motor impairments, hypoactivity, increased anxiety-related behavior, seizures, kyphosis, and stereotypic forelimb motions. Additionally, we show that although the truncated MeCP2 protein in these mice localizes normally to heterochromatic domains in vivo, histone H3 is hyperacetylated, providing evidence that the chromatin architecture is abnormal and that gene expression may be misregulated in this model of Rett syndrome.

Acetylation↗

Pvclust: an R package for assessing the uncertainty in hierarchical clustering.

SUMMARY: Pvclust is an add-on package for a statistical software R to assess the uncertainty in hierarchical cluster analysis. Pvclust can be used easily for general statistical problems, such as DNA microarray analysis, to perform the bootstrap analysis of clustering, which has been popular in phylogenetic analysis. Pvclust calculates probability values (p-values) for each cluster using bootstrap resampling techniques. Two types of p-values are available: approximately unbiased (AU) p-value and bootstrap probability (BP) value. Multiscale bootstrap resampling is used for the calculation of AU p-value, which has superiority in bias over BP value calculated by the ordinary bootstrap resampling. In addition the computation time can be enormously decreased with parallel computing option.

Algorithms↗

Development of bioinformatics resources for display and analysis of copy number and other structural variants in the human genome.

The discovery of an abundance of copy number variants (CNVs; gains and losses of DNA sequences >1 kb) and other structural variants in the human genome is influencing the way research and diagnostic analyses are being designed and interpreted. As such, comprehensive databases with the most relevant information will be critical to fully understand the results and have impact in a diverse range of disciplines ranging from molecular biology to clinical genetics. Here, we describe the development of bioinformatics resources to facilitate these studies. The Database of Genomic Variants (http://projects.tcag.ca/variation/) is a comprehensive catalogue of structural variation in the human genome. The database currently contains 1,267 regions reported to contain copy number variation or inversions in apparently healthy human cases. We describe the current contents of the database and how it can serve as a resource for interpretation of array comparative genomic hybridization (array CGH) and other DNA copy imbalance data. We also present the structure of the database, which was built using a new data modeling methodology termed Cross-Referenced Tables (XRT). This is a generic and easy-to-use platform, which is strong in handling textual data and complex relationships. Web-based presentation tools have been built allowing publication of XRT data to the web immediately along with rapid sharing of files with other databases and genome browsers. We also describe a novel tool named eFISH (electronic fluorescence in situ hybridization) (http://projects.tcag.ca/efish/), a BLAST-based program that was developed to facilitate the choice of appropriate clones for FISH and CGH experiments, as well as interpretation of results in which genomic DNA probes are used in hybridization-based experiments.

Algorithms↗