Cross-talk between steroids and NF-kappa B: what language?
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
BACKGROUND: The EMBL Nucleotide Sequence Database is a comprehensive database of DNA and RNA sequences and related information traditionally made available in flat-file format. Queries through tools such as SRS (Sequence Retrieval System) also return data in flat-file format. Flat files have a number of shortcomings, however, and the resources therefore currently lack a flexible environment to meet individual researchers' needs. The Object Management Group's common object request broker architecture (CORBA) is an industry standard that provides platform-independent programming interfaces and models for portable distributed object-oriented computing applications. Its independence from programming languages, computing platforms and network protocols makes it attractive for developing new applications for querying and distributing biological data. RESULTS: A CORBA infrastructure developed by EMBL-EBI provides an efficient means of accessing and distributing EMBL data. The EMBL object model is defined such that it provides a basis for specifying interfaces in interface definition language (IDL) and thus for developing the CORBA servers. The mapping from the object model to the relational schema in the underlying Oracle database uses the facilities provided by PersistenceTM, an object/relational tool. The techniques of developing loaders and 'live object caching' with persistent objects achieve a smart live object cache where objects are created on demand. The objects are managed by an evictor pattern mechanism. CONCLUSIONS: The CORBA interfaces to the EMBL database address some of the problems of traditional flat-file formats and provide an efficient means for accessing and distributing EMBL data. CORBA also provides a flexible environment for users to develop their applications by building clients to our CORBA servers, which can be integrated into existing systems.
UNLABELLED: Gene expression index calculations from Affymetrix GeneChips have been dominated by the Affymetrix MAS, dChip, and RMA methods. A new method to estimate the gene expression value utilizing the probe sequence information named position-dependent nearest-neighbor (PDNN) has been suggested by Zhang et al. (2003). Here we describe an open source implementation of the PDNN method for the statistical language R. AVAILABILITY: The package can be downloaded from http://www.bioconductor.org/repository/devel/package/html/affypdnn.html CONTACT: hbjorn@cbs.dtu.dk.
UNLABELLED: GeneRecon is a tool for fine-scale association mapping using a coalescence model. GeneRecon takes as input case-control data from phased or unphased SNP and microsatellite genotypes. The posterior distribution of disease locus position is obtained by Metropolis-Hastings sampling in the state space of genealogies. Input format, search strategy and the sampled statistics can be configured through the Guile Scheme programming language embedded in GeneRecon, making GeneRecon highly configurable. AVAILABILITY: The source code for GeneRecon, written in C++ and Scheme, is available under the GNU General Public License (GPL) at http://www.birc.au.dk/~mailund/GeneRecon CONTACT: mailund@birc.au.dk.
The aim of the RNA Operator Theory is to propose a new explanation for the mechanism of certain biological control functions, including morphogenesis and brain function. It assumes the existence of a natural computing language, the vocabulary of which, in machine language form, is constituted of bytes of nucleotide bits. These are observable empirically as, inter alia, repeated sequences in genetic nucleic acid. The assumed computing language possesses a range of operators which act directly when coded as RNS to effect the positional manipulation of matter on a molecular level.
The symbolic sequences of the exons that make human proteins are subjected to methods of statistical linguistics. The ideas developed for the natural languages by G. K. Zipf, when applied to these sequences, show significant promise. In particular, we argue, the Zipf's exponent differentiates, and hence, identifies disparate human sequences.
The convergence of clinical medicine and the Life Sciences, commencing with opportunities in clinical trials and clinically linked medical research, presents many novel challenges. The Genomic Messaging System (GMS) described here was originally developed as a tool for assembling clinical genomic records of individual and collective patients, and was then generalized to become a flexible workflow component that will link clinical records to a variety of computational biology research tools, for research and ultimately for a more personalized, focused, and preventative healthcare system. Prominent among the applications linked are protein science applications, including the rapid automated modeling of patient proteins with their individual structural polymorphisms. In an initial study, GMS formed the basis of a fully automated system for modeling patient proteins with structural polymorphisms as a basis for drug selection and ultimately design on an individual patient basis.
Nucleic acid sequences may be looked upon as words over the alphabet of nucleotides. Naturally occurring DNAs and RNAs form subsets of the set of all possible words. The use of formal languages is proposed to describe the structure of these subsets. Regular languages defined by finite automata are introduced to demonstrate the application of the concept on RNA-phages of group I. This approach permits a concise characterization of grammatical patterns in genetic information.
Chemical carcinogenesis is a process beginning with carcinogen absorption and ending with development of a malignant tumor. Individual elements of this process have been studied intensively but no comprehensive model has been developed. This report describes a comprehensive model which incorporates carcinogen pharmacokinetics, biochemical mechanism of action, and the resultant mutation of normal cells to malignancy. Model parameters correspond to specific physiological and biochemical structures and processes. The model was encoded in a simulation language and used to examined biochemical and cellular effects of exposure to an initiator and a promoter. With laboratory validation, the model should be useful for interpretation and design of studies on carcinogenic mechanisms and for risk assessment.
MOTIVATION: The analysis of gene expression data in its chromosomal context has been a recent development in cancer research. However, currently available methods fail to account for variation in the distance between genes, gene density and genomic features (e.g. GC content) in identifying increased or decreased chromosomal regions of gene expression. RESULTS: We have developed a model-based scan statistic that accounts for these aspects of the complex landscape of the human genome in the identification of extreme chromosomal regions of gene expression. This method may be applied to gene expression data regardless of the microarray platform used to generate it. To demonstrate the accuracy and utility of this method, we applied it to a breast cancer gene expression dataset and tested its ability to predict regions containing medium-to-high level DNA amplification (DNA ratio values >2). A classifier was developed from the scan statistic results that had a 10-fold cross-validated classification rate of 93% and a positive predictive value of 88%. This result strongly suggests that the model-based scan statistic and the expression characteristics of an increased chromosomal region of gene expression can be used to accurately predict chromosomal regions containing amplified genes. AVAILABILITY: Functions in the R-language are available from the author upon request. CONTACT: fcouples@umich.edu.
MOTIVATION: Accurate time series for biological processes are difficult to estimate due to problems of synchronization, temporal sampling and rate heterogeneity. Methods are needed that can utilize multi-dimensional data, such as those resulting from DNA microarray experiments, in order to reconstruct time series from unordered or poorly ordered sets of observations. RESULTS: We present a set of algorithms for estimating temporal orderings from unordered sets of sample elements. The techniques we describe are based on modifications of a minimum-spanning tree calculated from a weighted, undirected graph. We demonstrate the efficacy of our approach by applying these techniques to an artificial data set as well as several gene expression data sets derived from DNA microarray experiments. In addition to estimating orderings, the techniques we describe also provide useful heuristics for assessing relevant properties of sample datasets such as noise and sampling intensity, and we show how a data structure called a PQ-tree can be used to represent uncertainty in a reconstructed ordering. AVAILABILITY: Academic implementations of the ordering algorithms are available as source code (in the programming language Python) on our web site, along with documentation on their use. The artificial 'jelly roll' data set upon which the algorithm was tested is also available from this web site. The publicly available gene expression data may be found at http://genome-www.stanford.edu/cellcycle/ and http://caulobacter.stanford.edu/CellCycle/.
Many powerful tools have been created to detect and describe the similarities between nucleic acid or protein sequences. Frequently these take the form of a sequence consensus, expressing simple most popular positional identities, positional identities with allowances for varying positions or some type of statistical description of the positional frequency characteristics of the defining sequence family. Despite the fact that some provide intuitively interpretable descriptions of the consensuses themselves, they typically do not give the viewer any information about regions of the sequence that might have inter-positional dependencies, and that therefore do not obey a strict consensus behavior. Herein, we present MAVL (Multiple Alignment Variation Linker) and StickWRLD. MAVL is our web-based application for detecting and displaying both positive and negative inter-positional correlations in nucleic acid sequences. MAVL examines all positional pairs in each of a collection of pre-aligned sequences and determines any pairs that occur with either greater or lesser frequency than a positional frequency matrix would predict. These data are then composited into a StickWRLD representation and supplied back to the user as a VRML (virtual reality modeling language) file. MAVL and StickWRLD can be accessed at http://www.microbial-pathogenesis.org/stickwrld/. A tutorial that explains MAVL features and demonstrates typical user interactions with StickWRLD graphs is available at http://www.microbial-pathogenesis.org/stickwrld/tutorial/sticktut2.html. This tutorial is quite large; please be patient while it loads.
How can biological plasticity been added to a simulation of neuritic growth? Coming from this question, we have chosen a new access to simulate neuritic growth under the very aspect of meaningful and progredient development of single cells. Based on a specific description-language, we have set up a computer-program, to construct neurite-models and to simulate neuritic interaction during their development. Instead of using mathematical equations, we define various types of cytoskeletons by taking a specified graph grammar. Using this technique, we are able to define strings, combined with other influencing parameters, which allow the setting up of very naturally behaving artificial nervecells, in which distinct statistical variance and fixed rules as given in DNA operate together. In this paper, we want to discuss the underlying principles of the given grammar and to show some results from these computer-simulations, which enable us to study growth, development and other specific characteristics of neurites within a simulator in comparison to in vivo-experiments.
The first successes in cloning experiments and stem cell "reprogramming" have already demonstrated the primordial role of cellular working-space memory and regulatory mechanisms, which use the knowledge stored in the DNA database in read mode. We present an analogy between living systems and informatics systems by considering: 1) the cell cytoplasm as a memory device accessible as read/write; 2) the mechanisms of regulation as a programming language defined by a grammar, a molecular algebra; 3) biological processes as volatile programs which are executed without being written; 4) DNA as a database in read only mode. We also present applications to two biological algorithms: the immune response and glycogen metabolism.
SUMMARY: One of the significant challenges in gene expression analysis is to find unknown subtypes of several diseases at the molecular levels. This task can be addressed by grouping gene expression patterns of the collected samples on the basis of a large number of genes. Application of commonly used clustering methods to such a dataset however are likely to fail owing to over-learning, because the number of samples to be grouped is much smaller than the data dimension which is equal to the number of genes involved in the dataset. To overcome such difficulty, we developed a novel model-based clustering method, referred to as the mixed factors analysis. The ArrayCluster is a freely available software to perform the mixed factors analysis. It provides us some analytic tools for clustering DNA microarray experiments, data visualization and an automatic detector for module transcriptional of genes that are relevant to the calibrated molecular subtypes and so on.
INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.
The question of the origins of the Polynesians has, for over 200 years, been the subject of adventure science. Since Captain Cook's first speculations on these isolated Pacific islanders, their language affiliations have been seen as an essential clue to the solution. The geographic and numeric centre of gravity of the Austronesian language family is in island Southeast Asia, which was therefore originally seen as their dispersal homeland. However, another view has held sway for 15 years, the 'out of Taiwan' model, popularly known as the 'express train to Polynesia'. This model, based on the combined evidence of archaeology and linguistics, proposes a common origin for all Austronesian-speaking populations, in an expansion of rice agriculturalists from south China/Taiwan beginning around 6,000 years ago. However, it is becoming clear that there is, in fact, little supporting evidence in favour of this view. Alternative models suggest that the ancestors of the Polynesians achieved their maritime skills and horticultural Neolithic somewhere between island Southeast Asia and Melanesia, at an earlier date. Recent advances in human genetics now allow for an independent test of these models, lending support to the latter view rather than the former. Although local gene flow occurring between the bio-geographic regions may have been the means for the dramatic cultural spread out to the Pacific, the immediate genetic substrate for the Polynesian expansion came not from Taiwan, but from east of the Wallace line, probably in Wallacea itself.
MOTIVATION: Although a large amount of information on the structure, function and properties of biomolecules is becoming available, it is difficult to understand the relationship between them. Thus, we have attempted to create an integrated relational database, search and visualization tool, 3DinSight, to help researchers to gain insight into their relationship. RESULTS: We have gathered data on the structure, function and properties of biomolecules, and implemented them into a relational database system. The structural data contain several subset data such as protein homologues, protein-DNA complex, in order to enable searching within a specific class of data. The functional data include motif sequence and mutation data of proteins. Also, various amino acid properties are implemented as a relational table. The World Wide Web (WWW) interfaces enable users to carry out various kinds of searches among these data. The locations of motif sequences and mutations are automatically mapped on the structure, and visualized in three-dimensional (3D) space by interactive viewers, VRML (Virtual Reality Modeling Language) and RasMol. In the case of VRML, the mapped 3D objects are hyper-linked to the corresponding document data. Also, amino acid properties, linked with structure, functional and mutation sites, can be displayed as graph plots. AVAILABILITY: 3DinSight is freely accessible through the Internet (http://www.rtc.riken.go.jp/3DinSight.h tml). CONTACT: sarai@rtc.riken.go.jp