Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

htSNPer1.0: software for haplotype block partition and htSNPs selection.

BACKGROUND: There is recently great interest in haplotype block structure and haplotype tagging SNPs (htSNPs) in the human genome for its implication on htSNPs-based association mapping strategy for complex disease. Different definitions have been used to characterize the haplotype block structure in the human genome, and several different performance criteria and algorithms have been suggested on htSNPs selection. RESULTS: A heuristic algorithm, generalized branch-and-bound algorithm, is applied to the searching of minimal set of haplotype tagging SNPs (htSNPs) according to different htSNPs performance criteria. We develop a software htSNPer1.0 to implement the algorithm, and integrate three htSNPs performance criteria and four haplotype block definitions for haplotype block partitioning. It is a software with powerful Graphical User Interface (GUI), which can be used to characterize the haplotype block structure and select htSNPs in the candidate gene or interested genomic regions. It can find the global optimization with only a fraction of the computing time consumed by exhaustive searching algorithm. CONCLUSION: htSNPer1.0 allows molecular geneticists to perform haplotype block analysis and htSNPs selection using different definitions and performance criteria. The software is a powerful tool for those focusing on association mapping based on strategy of haplotype block and htSNPs.

Algorithms↗

Using ESTs to improve the accuracy of de novo gene prediction.

BACKGROUND: ESTs are a tremendous resource for determining the exon-intron structures of genes, but even extensive EST sequencing tends to leave many exons and genes untouched. Gene prediction systems based exclusively on EST alignments miss these exons and genes, leading to poor sensitivity. De novo gene prediction systems, which ignore ESTs in favor of genomic sequence, can predict such "untouched" exons, but they are less accurate when predicting exons to which ESTs align. TWINSCAN is the most accurate de novo gene finder available for nematodes and N-SCAN is the most accurate for mammals, as measured by exact CDS gene prediction and exact exon prediction. RESULTS: TWINSCAN_EST is a new system that successfully combines EST alignments with TWINSCAN. On the whole C. elegans genome TWINSCAN_EST shows 14% improvement in sensitivity and 13% in specificity in predicting exact gene structures compared to TWINSCAN without EST alignments. Not only are the structures revealed by EST alignments predicted correctly, but these also constrain the predictions without alignments, improving their accuracy. For the human genome, we used the same approach with N-SCAN, creating N-SCAN_EST. On the whole genome, N-SCAN_EST produced a 6% improvement in sensitivity and 1% in specificity of exact gene structure predictions compared to N-SCAN. CONCLUSION: TWINSCAN_EST and N-SCAN_EST are more accurate than TWINSCAN and N-SCAN, while retaining their ability to discover novel genes to which no ESTs align. Thus, we recommend using the EST versions of these programs to annotate any genome for which EST information is available.TWINSCAN_EST and N-SCAN_EST are part of the TWINSCAN open source software package http://genes.cse.wustl.edu/distribution/download_TS.html.

Algorithms↗

CapsID: a web-based tool for developing parsimonious sets of CAPS molecular markers for genotyping.

BACKGROUND: Genotyping may be carried out by a number of different methods including direct sequencing and polymorphism analysis. For a number of reasons, PCR-based polymorphism analysis may be desirable, owing to the fact that only small amounts of genetic material are required, and that the costs are low. One popular and cheap method for detecting polymorphisms is by using cleaved amplified polymorphic sequence, or CAPS, molecular markers. These are also known as PCR-RFLP markers. RESULTS: We have developed a program, called CapsID, that identifies snip-SNPs (single nucleotide polymorphisms that alter restriction endonuclease cut sites) within a set or sets of reference sequences, designs PCR primers around these, and then suggests the most parsimonious combination of markers for genotyping any individual who is not a member of the reference set. The output page includes biologist-friendly features, such as images of virtual gels to assist in genotyping efforts. CapsID is freely available at http://bbc.botany.utoronto.ca/capsid. CONCLUSION: CapsID is a tool that can rapidly provide minimal sets of CAPS markers for molecular identification purposes for any biologist working in genetics, community genetics, plant and animal breeding, forensics and other fields.

Computational Biology↗

Algorithms for determining the fate of sites and domain boundaries in computer simulations of recombinant DNA procedures.

Structural and functional features within genomic sequences are best described by their position within the genomic structure. Cleavage sites can be conveniently described by single positions but genomic domains require the position of their two boundaries. The handling of these positions simultaneously to sequence manipulations in computer simulations of recombinant DNA procedures greatly improves the understanding of the resulting recombinant constructs. In addition, the algorithms describing the fate of domain boundaries can be used for the handling of nucleotide sequences in dynamic database environments handled by languages like Prolog which are particularly suitable for artificial intelligence implementations. This communication describes a set of algorithms for the automatic updating of single sites and double domain boundaries in linear and circular models for computer simulation of recombinant DNA procedures.

Algorithms↗

DNA splicing systems and post systems.

This paper concerns the formal study on the generative powers of extended splicing (H) systems. First, using a classical result by Post which characterizes the recursively enumerable languages in terms of his Post Normal systems, we establish several new characterizations of extended H systems which not only allow us to have very simple alternative proof methods for the previous results mentioned above, but also give a new insight into the relationships between families of extended H systems. We show a kind of normal form for extended H systems exactly characterizing the class of regular languages. We also show a new representation result for the family of context-free languages in terms of extended H systems.

Alternative Splicing↗

Integrating database homology in a probabilistic gene structure model.

We present an improved stochastic model of genes in DNA, and describe a method for integrating database homology into the probabilistic framework. A generalized hidden Markov model (GHMM) describes the grammar of a legal parse of a DNA sequence. Probabilities are estimated for gene features by using dynamic programming to combine information from multiple sensors. We show how matches to homologous sequences from a database can be integrated into the probability estimation by interpreting the likelihood of a sequence in terms of the bit-cost to encode a sequence given a homology match. We also demonstrate how homology matches in protein databases can be exploited to help identify splice sites. Our experiments show significant improvements in the sensitivity and specificity of gene structure identification when these new features are added to our gene-finding system, Genie. Experimental results in tests using a standard set of annotated genes showed that Genie identified 95% of coding nucleotides correctly with a specificity of 91%, and 77% of exons were identified exactly.

Algorithms↗

New computing paradigms suggested by DNA computing: computing by carving.

Inspired by the experiments in the emerging area of DNA computing, a somewhat unusual type of computation strategy was recently proposed by one of us: to generate a (large) set of candidate solutions of a problem, then remove the non-solutions such that what remains is the set of solutions. This has been called a computation by carving. This idea leads both to a speculation with possible important consequences--computing non-recursively enumerable languages--and to interesting theoretical computer science (formal language) questions.

Animals↗

MBEToolbox: a MATLAB toolbox for sequence data analysis in molecular biology and evolution.

BACKGROUND: MATLAB is a high-performance language for technical computing, integrating computation, visualization, and programming in an easy-to-use environment. It has been widely used in many areas, such as mathematics and computation, algorithm development, data acquisition, modeling, simulation, and scientific and engineering graphics. However, few functions are freely available in MATLAB to perform the sequence data analyses specifically required for molecular biology and evolution. RESULTS: We have developed a MATLAB toolbox, called MBEToolbox, aimed at filling this gap by offering efficient implementations of the most needed functions in molecular biology and evolution. It can be used to manipulate aligned sequences, calculate evolutionary distances, estimate synonymous and nonsynonymous substitution rates, and infer phylogenetic trees. Moreover, it provides an extensible, functional framework for users with more specialized requirements to explore and analyze aligned nucleotide or protein sequences from an evolutionary perspective. The full functions in the toolbox are accessible through the command-line for seasoned MATLAB users. A graphical user interface, that may be especially useful for non-specialist end users, is also provided. CONCLUSION: MBEToolbox is a useful tool that can aid in the exploration, interpretation and visualization of data in molecular biology and evolution. The software is publicly available at http://web.hku.hk/~jamescai/mbetoolbox/ and http://bioinformatics.org/project/?group_id=454

Algorithms↗

HDBStat!: a platform-independent software suite for statistical analysis of high dimensional biology data.

BACKGROUND: Many efforts in microarray data analysis are focused on providing tools and methods for the qualitative analysis of microarray data. HDBStat! (High-Dimensional Biology-Statistics) is a software package designed for analysis of high dimensional biology data such as microarray data. It was initially developed for the analysis of microarray gene expression data, but it can also be used for some applications in proteomics and other aspects of genomics. HDBStat! provides statisticians and biologists a flexible and easy-to-use interface to analyze complex microarray data using a variety of methods for data preprocessing, quality control analysis and hypothesis testing. RESULTS: Results generated from data preprocessing methods, quality control analysis and hypothesis testing methods are output in the form of Excel CSV tables, graphs and an Html report summarizing data analysis. CONCLUSION: HDBStat! is a platform-independent software that is freely available to academic institutions and non-profit organizations. It can be downloaded from our website http://www.soph.uab.edu/ssg_content.asp?id=1164.

Algorithms↗

[Mathematical methods in flow cytometry: the problem of evaluating DNA histograms of partially synchronous cell populations].

There is presented a procedure for determining the phase fractions in a cell population (i.e. the fractions of the G1-, S- and G2 + M-cells) from the corresponding DNA histogram obtained by flow cytometry. The evaluation procedure can regard arbitrary apriori information about the model parameters. It is based on a mathematical model of the DNA distribution for a growing cell population which contains the phase fractions and some further values as parameters. The parameter estimations rest on the Maximum-Likelihood-Principle by which the parameters of the theoretical model must be determined in such a way that the measured histogram corresponds to a maximum probability. As a consequence of the model flexibility especially the S-phase fraction which can be determined independently by labeling techniques, can be regarded as a fixed predetermined value. Such apriori information enhances the preciseness of the parameter estimation. The applications have shown that the determination of the S-phase fraction by means of flow cytometry can be very unprecise especially under partial synchronization. For the application of the evaluation programs (programming language ALGOL) the histograms must be free of cell detritus and clumping.

Animals↗

The DNA damage response and patient safety: engaging our molecular biology-oriented colleagues.

The imperative to improve patient safety is clear. Biomedical scientists, who account for a large proportion of medical school faculty, and clinicians tend to speak different languages. Biological systems are remarkable for their high robustness, flexibility, and efficiency. Biomedical scientists possess a profound understanding of the complex mechanisms that govern organisms. Their insights may inform the design of safer health care systems. We propose a model to assist in bi-directional communication between these disciplines. We use the principles and mechanisms of the DNA damage response to describe the central concepts of safety science and discuss similarities and differences between the systems of DNA repair and organizational approaches to safety in health care. We suggest that such biomedical scientists can and should be engaged in the effort to bring education about patient safety management into the medical school curriculum and to make patient care safer.

Biomedical Research↗

Genomic history of the Caucasus: A systematic review and meta-analysis of ancient DNA studies.

The Caucasus region represents a unique natural laboratory for paleogenetic research due to its complex topography, long-standing role as a migratory corridor and glacial refugium, and exceptional preservation conditions for ancient DNA. This review synthesizes recent genome-wide studies to reconstruct the demographic history shaping the distinctive genetic landscape of modern Caucasus populations. The analysis reveals a deep pattern of continuity, isolation, and periodic admixture. Early genetic differentiation emerged in the Neolithic and Chalcolithic, forming distinct steppe and mountain population clusters. The Bronze Age was a pivotal period marked by large-scale gene flow from the Eurasian Steppe, particularly linked to the Yamnaya expansion, and interactions with Iranian and Anatolian-related groups. Despite these influences, many populations demonstrate remarkable genetic continuity from the Bronze Age to the present day. Significant knowledge gaps persist, particularly for the Paleolithic, Mesolithic, and Neolithic of the North Caucasus, as well as for the Late Medieval and Early Modern periods across the entire region. Addressing these gaps through targeted archaeogenomic studies is crucial for understanding the fine-scale processes that formed the hierarchical structure and high linguistic diversity of Caucasus populations, offering a powerful model for studying human adaptation, interaction, and language-genetics dynamics in a mountainous environment.

Humans↗

Visualization methods for statistical analysis of microarray clusters.

BACKGROUND: The most common method of identifying groups of functionally related genes in microarray data is to apply a clustering algorithm. However, it is impossible to determine which clustering algorithm is most appropriate to apply, and it is difficult to verify the results of any algorithm due to the lack of a gold-standard. Appropriate data visualization tools can aid this analysis process, but existing visualization methods do not specifically address this issue. RESULTS: We present several visualization techniques that incorporate meaningful statistics that are noise-robust for the purpose of analyzing the results of clustering algorithms on microarray data. This includes a rank-based visualization method that is more robust to noise, a difference display method to aid assessments of cluster quality and detection of outliers, and a projection of high dimensional data into a three dimensional space in order to examine relationships between clusters. Our methods are interactive and are dynamically linked together for comprehensive analysis. Further, our approach applies to both protein and gene expression microarrays, and our architecture is scalable for use on both desktop/laptop screens and large-scale display devices. This methodology is implemented in GeneVAnD (Genomic Visual ANalysis of Datasets) and is available at http://function.princeton.edu/GeneVAnD. CONCLUSION: Incorporating relevant statistical information into data visualizations is key for analysis of large biological datasets, particularly because of high levels of noise and the lack of a gold-standard for comparisons. We developed several new visualization techniques and demonstrated their effectiveness for evaluating cluster quality and relationships between clusters.

Algorithms↗

Modelling the checkpoint response to telomere uncapping in budding yeast.

One of the DNA damage-response mechanisms in budding yeast is temporary cell-cycle arrest while DNA repair takes place. The DNA damage response requires the coordinated interaction between DNA repair and checkpoint pathways. Telomeres of budding yeast are capped by the Cdc13 complex. In the temperature-sensitive cdc13-1 strain, telomeres are unprotected over a specific temperature range leading to activation of the DNA damage response and subsequently cell-cycle arrest. Inactivation of cdc13-1 results in the generation of long regions of single-stranded DNA (ssDNA) and is affected by the activity of various checkpoint proteins and nucleases. This paper describes a mathematical model of how uncapped telomeres in budding yeast initiate the checkpoint pathway leading to cell-cycle arrest. The model was encoded in the Systems Biology Markup Language (SBML) and simulated using the stochastic simulation system Biology of Ageing e-Science Integration and Simulation (BASIS). Each simulation follows the time course of one mother cell keeping track of the number of cell divisions, the level of activity of each of the checkpoint proteins, the activity of nucleases and the amount of ssDNA generated. The model can be used to carry out a variety of in silico experiments in which different genes are knocked out and the results of simulation are compared to experimental data. Possible extensions to the model are also discussed.

Cell Cycle↗

Maximum likelihood estimates of allele frequencies and error rates from samples of related individuals by gene counting.

SUMMARY: Graphical modeling is used to extend the gene counting method to compute maximum likelihood estimates of allele frequencies for samples of individuals related in extended pedigrees. Genotypes may be missing or partially observed, and error rates can be simultaneously estimated. AVAILABILITY: The Java classes and Javadocs pages for \mathsf\hbox GeneCountAlleles can be obtained from bioinformatics.med.utah.edu/~alun, which also has more information on its use and file formats.

Biological Evolution↗

AnovArray: a set of SAS macros for the analysis of variance of gene expression data.

BACKGROUND: Analysis of variance is a powerful approach to identify differentially expressed genes in a complex experimental design for microarray and macroarray data. The advantage of the anova model is the possibility to evaluate multiple sources of variation in an experiment. RESULTS: AnovArray is a package implementing ANOVA for gene expression data using SAS statistical software. The originality of the package is 1) to quantify the different sources of variation on all genes together, 2) to provide a quality control of the model, 3) to propose two models for a gene's variance estimation and to perform a correction for multiple comparisons. CONCLUSION: AnovArray is freely available at http://www-mig.jouy.inra.fr/stat/AnovArray and requires only SAS statistical software.

Algorithms↗

Statistical analysis of real-time PCR data.

BACKGROUND: Even though real-time PCR has been broadly applied in biomedical sciences, data processing procedures for the analysis of quantitative real-time PCR are still lacking; specifically in the realm of appropriate statistical treatment. Confidence interval and statistical significance considerations are not explicit in many of the current data analysis approaches. Based on the standard curve method and other useful data analysis methods, we present and compare four statistical approaches and models for the analysis of real-time PCR data. RESULTS: In the first approach, a multiple regression analysis model was developed to derive DeltaDeltaCt from estimation of interaction of gene and treatment effects. In the second approach, an ANCOVA (analysis of covariance) model was proposed, and the DeltaDeltaCt can be derived from analysis of effects of variables. The other two models involve calculation DeltaCt followed by a two group t-test and non-parametric analogous Wilcoxon test. SAS programs were developed for all four models and data output for analysis of a sample set are presented. In addition, a data quality control model was developed and implemented using SAS. CONCLUSION: Practical statistical solutions with SAS programs were developed for real-time PCR data and a sample dataset was analyzed with the SAS programs. The analysis using the various models and programs yielded similar results. Data quality control and analysis procedures presented here provide statistical elements for the estimation of the relative expression of genes using real-time PCR.

Analysis of Variance↗