Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

The PowerAtlas: a power and sample size atlas for microarray experimental design and research.

BACKGROUND: Microarrays permit biologists to simultaneously measure the mRNA abundance of thousands of genes. An important issue facing investigators planning microarray experiments is how to estimate the sample size required for good statistical power. What is the projected sample size or number of replicate chips needed to address the multiple hypotheses with acceptable accuracy? Statistical methods exist for calculating power based upon a single hypothesis, using estimates of the variability in data from pilot studies. There is, however, a need for methods to estimate power and/or required sample sizes in situations where multiple hypotheses are being tested, such as in microarray experiments. In addition, investigators frequently do not have pilot data to estimate the sample sizes required for microarray studies. RESULTS: To address this challenge, we have developed a Microrarray PowerAtlas. The atlas enables estimation of statistical power by allowing investigators to appropriately plan studies by building upon previous studies that have similar experimental characteristics. Currently, there are sample sizes and power estimates based on 632 experiments from Gene Expression Omnibus (GEO). The PowerAtlas also permits investigators to upload their own pilot data and derive power and sample size estimates from these data. This resource will be updated regularly with new datasets from GEO and other databases such as The Nottingham Arabidopsis Stock Center (NASC). CONCLUSION: This resource provides a valuable tool for investigators who are planning efficient microarray studies and estimating required sample sizes.

Algorithms↗

Computer programs to assist in high resolution thermal denaturation and circular dichroism studies on nucleic acids.

Computer programs are described that direct the collection, processing, and graphical display of numerical data obtained from high resolution thermal denaturation (1-3) and circular dichroism (4) studies. Besides these specific applications, the programs may also be useful, either directly or as programming models, in other types of spectrophotometric studies employing computers, programming languages, or instruments similar to those described here (see Materials and Methods).

Circular Dichroism↗

Predicting bacterial transcription units using sequence and expression data.

MOTIVATION: A key aspect of elucidating gene regulation in bacterial genomes is identifying the basic units of transcription. We present a method, based on probabilistic language models, that we apply to predict operons, promoters and terminators in the genome of Escherichia coli K-12. Our approach has two key properties: (i) it provides a coherent set of predictions for related regulatory elements of various types and (ii) it takes advantage of both DNA sequence and gene expression data, including expression measurements from inter-genic probes. RESULTS: Our experimental results show that we are able to predict operons and localize promoters and terminators with high accuracy. Moreover, our models that use both sequence and expression data are more accurate than those that use only one of these two data sources.

Algorithms↗

Functional genetic analysis of mutations implicated in a human speech and language disorder.

Mutations in the FOXP2 gene cause a severe communication disorder involving speech deficits (developmental verbal dyspraxia), accompanied by wide-ranging impairments in expressive and receptive language. The protein encoded by FOXP2 belongs to a divergent subgroup of forkhead-box transcription factors, with a distinctive DNA-binding domain and motifs that mediate hetero- and homodimerization. Here we report the first direct functional genetic investigation of missense and nonsense mutations in FOXP2 using human cell-lines, including a well-established neuronal model system. We focused on three unusual FOXP2 coding variants, uniquely identified in cases of verbal dyspraxia, assessing expression, subcellular localization, DNA-binding and transactivation properties. Analysis of the R553H forkhead-box substitution, found in all affected members of a large three-generation family, indicated that it severely affects FOXP2 function, chiefly by disrupting nuclear localization and DNA-binding properties. The R328X truncation mutation, segregating with speech/language disorder in a second family, yields an unstable, predominantly cytoplasmic product that lacks transactivation capacity. A third coding variant (Q17L) observed in a single affected child did not have any detectable functional effect in the present study. In addition, we used the same systems to explore the properties of different isoforms of FOXP2, resulting from alternative splicing in human brain. Notably, one such isoform, FOXP2.10+, contains dimerization domains, but no DNA-binding domain, and displayed increased cytoplasmic localization, coupled with aggresome formation. We hypothesize that expression of alternative isoforms of FOXP2 may provide mechanisms for post-translational regulation of transcription factor function.

Alternative Splicing↗

Design of long oligonucleotide probes for functional gene detection in a microbial community.

MOTIVATION: Analysis of the functions of microorganisms and their dynamics in the environment is essential for understanding microbial ecology. For analysis of highly similar sequences of a functional gene family using microarrays, the previous long oligonucleotide probe design strategies have not been useful in generating probes. RESULTS: We developed a Hierarchical Probe Design (HPD) program that designs both sequence-specific probes and hierarchical cluster-specific probes from sequences of a conserved functional gene based on the clustering tree of the genes, specifically for analyses of functional gene diversity in environmental samples. HPD was tested on datasets for the nirS and pmoA genes. Our results showed that HPD generated more sequence-specific probes than several popular oligonucleotide design programs. With a combination of sequence-specific and cluster-specific probes, HPD generated a probe set covering all the sequences of each test set. AVAILABILITY: http://brcapp.kribb.re.kr/HPD/

Algorithms↗

Automatic annotation of BIND molecular interactions from three-dimensional structures.

Software to automate the process of extracting molecular interactions from three-dimensional (3D) structures has been developed that records these as Biomolecular Interaction Network Database (BIND) pairwise interaction records. Full annotation of BIND records is provided through a database processing tool called MMDBind, including detailed atom-atom and residue-residue level interaction information. BIND three-dimensional interaction annotation is synthesized by combining information from the Molecular Modeling Database (MMDB), and the HET (heterogen) group dictionary of small molecules in the macromolecular Crystallographic Information Format (mmCIF). Interactions are validated using the Protein Quaternary Structure (PQS) system. A total of 18,166 interactions were removed as being redundant or biologically irrelevant after PQS validation. This first pass MMDBind annotation creates two new divisions of BIND, 3D Biopolymers (BIND-3DBP) comprising 16,737 initial interaction records, and 3D Small Molecules (BIND-3DSM) comprising 48,219 records. Visualization of interacting residues and nucleotides within a macromolecular structure is possible directly from the BIND database owing to added 3D feature annotation within the BIND records that can be conveniently seen using Cn3D ("see-in-3D") after query from the BIND Data Manager. These interaction records provide a further demonstration of the completeness of the BIND data specification and its capabilities as storage and exchange format for all kinds of molecular interactions, including RNA, DNA, protein, and small molecules. Data from the 3DBP and 3DSM sets are available for downloading in Abstract Syntax Notation.1 (ASN.1) or Extensible Markup Language (XML) formats at ftp://ftp.bind.ca/DB/MMDBBind. Data from the 3DBP set is available for interactive query from the BIND Data Manager at www.bind.ca.

Crystallography, X-Ray↗

Estimating allele frequencies of hypervariable DNA systems.

Several polymorphisms of human DNA have been shown to be hypervariable due to the recurrence of a variable number of tandem repeats (VNTRs) in the lengths of allelic restriction fragments. The recurrence of allelic variants in this novel class of polymorphisms seems to comply well with a model of continuous random variables. Based on this assumption, we have compiled some simple algorithms for classification of continuous data and estimation of classes of relative frequencies and have implemented these routines for the management of databases storing hypervariable single locus DNA genetic systems. The algorithms are compiled in BASIC language and can be incorporated in task-oriented computer programs. Three procedures are discussed, based in turn on: (a) using predetermined, arbitrary classes; (b) point estimations of frequencies for single fragments using error measurements associated with the kilobase value assignment; (c) estimates of phenotype frequencies according to error measurements. Error measurements are obtained from a statistic of values pertaining to several restriction fragments (genomic controls) repeatedly tested in different experiments. Problems related to these approaches are discussed.

Algorithms↗

Ultraviolet A and melanoma: a review.

The incidence and mortality rates of melanoma have risen for many decades in the United States. Increased exposure to ultraviolet (UV) radiation is generally considered to be responsible. Sunburns, a measure of excess sun exposure, have been identified as a risk factor for the development of melanoma. Because sunburns are primarily due to UVB (280-320 nm) radiation, UVB has been implicated as a potential contributing factor to the pathogenesis of melanoma. The adverse role of UVA (320-400 nm) in this regard is less well studied, and currently there is a great deal of controversy regarding the relationship between UVA exposure and the development of melanoma. This article reviews evidence in the English-language literature that surrounds the controversy concerning a possible role for UVA in the origin of melanoma. Our search found that UVA causes DNA damage via photosensitized reactions that result in the production of oxygen radical species. UVA can induce mutations in various cultured cell lines. Furthermore, in two animal models, the hybrid Xiphophorus fish and the opossum (Mondelphis domestica), melanomas and melanoma precursors can be induced with UVA. UVA radiation has been reported to produce immunosuppression in laboratory animals and in humans. Some epidemiologic studies have reported an increase in melanomas in users of sunbeds and sunscreens and in patients exposed to psoralen and UVA (PUVA) therapy. There is basic scientific evidence of the harmful effects of UVA on DNA, cells and animals. Collectively, these data suggest a potential role for UVA in the pathogenesis of melanoma. To date evidence from epidemiologic studies and clinical observations are inconclusive but seem to be consistent with this hypothesis. Additional research on the possible role of UVA in the pathogenesis of melanoma is required.

Animals↗

ShiftDetector: detection of shift mutations.

MOTIVATION: Sequencing of a bi-allelic PCR product, which contains an allele with a deletion/insertion mutation results in a superimposed tracefile following the site of this shift mutation. A trace file of this type hampers the use of current computer programs for base calling. ShiftDetector analyses a sequencing trace file in order to discover if it is a superimposed sequence of two molecules that differ in a shift mutation of 1 to 25 bases. The program calculates a probability score for the existence of such a shift and reconstructs the sequence of the original molecule. AVAILABILITY: ShiftDetector is available from http://cowry.agri.huji.ac.il

Alleles↗

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans↗

Novel approaches and applications in identifying DNA methylation markers of cardio-kidney-metabolic disease.

Cardio-kidney-metabolic (CKM) diseases represent a major public health challenge, accounting for a large proportion of global burden of morbidity and mortality. These conditions share risk factors, including genetic predisposition, environmental exposures, and lifestyle influences, which collectively drive disease development and progression. Epigenetic modifications, particularly DNA methylation (DNAm), serve as key mediators and biomarkers between these risk factors and disease phenotypes by regulating gene expression without altering the DNA sequence. Epigenome-wide association studies have identified DNAm markers associated with CKM diseases and related phenotypes, highlighting both shared pathways and disease-specific epigenetic signatures in inflammation, metabolic dysfunction, and aging-related processes. Longitudinal studies further demonstrate the dynamic nature of DNAm changes over time, offering insights into disease trajectories. Additionally, methylation risk scores integrating multiple epigenetic markers show promise in improving disease prediction and risk stratification beyond traditional clinical factors. To synthesize the current evidence, we conducted a targeted literature search in PubMed for English-language, peer-reviewed articles published between 2014 and the present. Future research leveraging large, well-phenotyped cohorts, advanced statistical methods, and innovative study designs will be critical for uncovering novel biomarkers, refining risk prediction models, and developing targeted epigenetic therapies to mitigate the global burden.

Humans↗

MET Exon 14 Skipping Mutation in NSCLC: From Genomic Discovery to Biomarker-Guided Therapeutic Innovation.

INTRODUCTION: Non-small cell lung cancer (NSCLC) is the most common type of lung cancer, and the MET exon 14 skipping mutation is a key oncogenic driver, which promotes tumor progression and provides a new direction for precision therapy. METHODS: A systematic search of English-language literature and clinical trial data related to the MET exon 14 skipping mutation from 2020-2025 was performed to summarize the role of the mutation and therapeutic advances. RESULTS: DNA-based next-generation sequencing (NGS), RNA-based NGS, and RT-qPCR were employed as the main detection methods. Preclinical models confirmed that mutations promote tumor progression by activating the RAS/MAPK pathway. Clinical trials have reported objective remission rates (ORR) of 46-68% for first-line treatment with MET inhibitors in NSCLC patients harboring MET exon 14 skipping mutations. DISCUSSION: MET exon 14 skipping mutation as a therapeutic target for NSCLC has made significant progress, and MET inhibitors are more advantageous than chemotherapy and immunotherapy, and have been recommended by national and international guidelines as a first-line treatment option. Additionally, NGS technology has the potential to dynamically monitor tumor evolution and drugresistant mutations, thereby helping to realize precision medicine. CONCLUSION: The MET exon 14 skipping mutation is an important target for the precision treatment of NSCLC, and MET-TKIs have remarkable efficacy but a prominent problem with drug resistance. The construction of a precision medicine system encompassing diagnosis, treatment, and drug resistance management through multi-omics research, technological innovation, and international collaboration is a key direction for improving prognosis.

Humans↗

Examination of AVPR1a as an autism susceptibility gene.

Impaired reciprocal social interaction is one of the core features of autism. While its determinants are complex, one biomolecular pathway that clearly influences social behavior is the arginine-vasopressin (AVP) system. The behavioral effects of AVP are mediated through the AVP receptor 1a (AVPR1a), making the AVPR1a gene a reasonable candidate for autism susceptibility. We tested the gene's contribution to autism by screening its exons in 125 independent autistic probands and genotyping two promoter polymorphisms in 65 autism affected sibling pair (ASP) families. While we found no nonconservative coding sequence changes, we did identify evidence of linkage and of linkage disequilibrium. These results were most pronounced in a subset of the ASP families with relatively less severe impairment of language. Thus, though we did not demonstrate a disease-causing variant in the coding sequence, numerous nontraditional disease-causing genetic abnormalities are known to exist that would escape detection by traditional gene screening methods. Given the emerging biological, animal model, and now genetic data, AVPR1a and genes in the AVP system remain strong candidates for involvement in autism susceptibility and deserve continued scrutiny.

Autistic Disorder↗

CAGER: classification analysis of gene expression regulation using multiple information sources.

BACKGROUND: Many classification approaches have been applied to analyzing transcriptional regulation of gene expressions. These methods build models that can explain a gene's expression level from the regulatory elements (features) on its promoter sequence. Different types of features, such as experimentally verified binding motifs, motifs discovered by computer programs, or transcription factor binding data measured with Chromatin Immunoprecipitation (ChIP) assays, have been used towards this goal. Each type of features has been shown successful in modeling gene transcriptional regulation under certain conditions. However, no comparison has been made to evaluate the relative merit of these features. Furthermore, most publicly available classification tools were not designed specifically for modeling transcriptional regulation, and do not allow the user to combine different types of features. RESULTS: In this study, we use a specific classification method, decision trees, to model transcriptional regulation in yeast with features based on predefined motifs, automatically identified motifs, ChlP-chip data, or their combinations. We compare the accuracies and stability of these models, and analyze their capabilities in identifying functionally related genes. Furthermore, we design and implement a user-friendly web server called CAGER (Classification Analysis of Gene Expression Regulation) that integrates several software components for automated analysis of transcriptional regulation using decision trees. Finally, we use CAGER to study the transcriptional regulation of Arabidopsis genes in response to abscisic acid, and report some interesting new results. CONCLUSION: Models built with ChlP-chip data suffer from low accuracies when the condition under which gene expressions are measured is significantly different from the condition under which the ChIP experiment is conducted. Models built with automatically identified motifs can sometimes discover new features, but their modeling accuracies may have been over-estimated in previous studies. Furthermore, models built with automatically identified motifs are not stable with respect to noises. A combination of ChlP-chip data and predefined motifs can substantially improve modeling accuracies, and is effective in identifying true regulons. The CAGER web server, which is freely available at http://cic.cs.wustl.edu/CAGER/, allows the user to select combinations of different feature types for building decision trees, and interact with the models graphically. We believe that it will be a useful tool to facilitate the discovery of gene transcriptional regulatory networks.

Algorithms↗

A Hidden Markov model web application for analysing bacterial genomotyping DNA microarray experiments.

Whole genome DNA microarray genomotyping experiments compare the gene content of different species or strains of bacteria. A statistical approach to analysing the results of these experiments was developed, based on a Hidden Markov model (HMM), which takes adjacency of genes along the genome into account when calling genes present or absent. The model was implemented in the statistical language R and applied to three datasets. The method is numerically stable with good convergence properties. Error rates are reduced compared with approaches that ignore spatial information. Moreover, the HMM circumvents a problem encountered in a conventional analysis: determining the cut-off value to use to classify a gene as absent. An Apache Struts web interface for the R script was created for the benefit of users unfamiliar with R. The application may be found at http://hmmgd.cryst.bbk.ac.uk/hmmgd. The source code illustrating how to run R scripts from an Apache Struts-based web application is available from the corresponding author on request. The application is also available for local installation if required.

Algorithms↗

Conceptual data modelling for bioinformatics.

Current research in the biosciences depends heavily on the effective exploitation of huge amounts of data. These are in disparate formats, remotely dispersed, and based on the different vocabularies of various disciplines. Furthermore, data are often stored or distributed using formats that leave implicit many important features relating to the structure and semantics of the data. Conceptual data modelling involves the development of implementation-independent models that capture and make explicit the principal structural properties of data. Entities such as a biopolymer or a reaction, and their relations, eg catalyses, can be formalised using a conceptual data model. Conceptual models are implementation-independent and can be transformed in systematic ways for implementation using different platforms, eg traditional database management systems. This paper describes the basics of the most widely used conceptual modelling notations, the ER (entity-relationship) model and the class diagrams of the UML (unified modelling language), and illustrates their use through several examples from bioinformatics. In particular, models are presented for protein structures and motifs, and for genomic sequences.

Computational Biology↗

Deciphering the language of the genome.

The non-coding DNA in eukaryotic genomes encodes a language which programs organismal growth and development. We show that a linguistic and cryptographic approach can be used to deduce the syntax of this programming language for gene regulation and to compile a dictionary of enhancers which form its words.

Animals↗