Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Protein language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

In silico identification of functional regions in proteins.

MOTIVATION: In silico prediction of functional regions on protein surfaces, i.e. sites of interaction with DNA, ligands, substrates and other proteins, is of utmost importance in various applications in the emerging fields of proteomics and structural genomics. When a sufficient number of homologs is found, powerful prediction schemes can be based on the observation that evolutionarily conserved regions are often functionally important, typically, only the principal functionally important region of the protein is detected, while secondary functional regions with weaker conservation signals are overlooked. Moreover, it is challenging to unambiguously identify the boundaries of the functional regions. METHODS: We present a new methodology, called PatchFinder, that automatically identifies patches of conserved residues that are located in close proximity to each other on the protein surface. PatchFinder is based on the following steps: (1) Assignment of conservation scores to each amino acid position on the protein surface. (2) Assignment of a score to each putative patch, based on its likelihood to be functionally important. The patch of maximum likelihood is considered to be the main functionally important region, and the search is continued for non-overlapping patches of secondary importance. RESULTS: We examined the accuracy of the method using the IGPS enzyme, the SH2 domain and a benchmark set of 112 proteins. These examples demonstrated that PatchFinder is capable of identifying both the main and secondary functional patches. AVAILABILITY: The PatchFinder program is available at: http://ashtoret.tau.ac.il/~nimrodg/

Algorithms↗

Linguistic approaches to biological sequences.

Biologists have long made use of linguistic metaphors in describing and naming cellular processes involving nucleic acid and protein sequences. Indeed, it is very natural to view the genetic 'text' and its sequential transliterations in these terms. However, a metaphor is not a tool, and it is necessary to ask whether the techniques used in analyzing other kinds of languages, such as human and computer languages, can in fact be of any use in tackling problems in molecular biology. This paper reviews the work of the author and others in applying the methods of computational linguistics to biological sequences.

Base Sequence↗

Integrating computation and visualization for biomolecular analysis: an example using python and AVS.

One of the challenges in biocomputing is to enable the efficient use of a wide variety of fast-evolving computational methods to simulate, analyze, and understand the complex properties and interactions of molecular systems. Our laboratory investigates several areas including molecular visualization, protein-ligand docking, protein-protein docking, molecular surfaces, and the derivation of phenomenological potentials. In this paper we present an approach based on the Python programming language to achieve a high level of integration between these different computational methods and our primary visualization system AVS. This approach removes many limitations of AVS while increasing dramatically the inter-operability of our computational tools. Several examples are shown to illustrate how this approach enables a high level of integration and inter-operability between different tools, while retaining modularity and avoiding the creation of a large monolithic package that is difficult to extend and maintain.

Capsid↗

Rapid numerical integration algorithm for finding the equilibrium state of a system of coupled binding reactions.

We have adapted a simple method of numerical integration to predict the equilibrium state of a population of components undergoing reversible association according to the Law of Mass Action. Its particular application is to populations of protein molecules in aqueous solution. The method is based on Euler integration but employs an adaptive step size: the time increment being reduced if it would make the concentration of any component negative and increased while the concentration of any component changes at greater than a specified rate. Parameters of the algorithm have been optimized empirically using a model set of binding equilibria with dissociation constants ranging from 10(-5) M to 10(-9) M. The method obtains the solution to a set of binding equilibria more rapidly than the conventional initial value methods (simple Euler, 4th order Runge-Kutta and variable-step Runge-Kutta methods were tested) for the same accuracy. A computer code in standard C is presented.

Algorithms↗

MOLECULAR DESIGNER: an interactive program for the display of protein structure on the IBM-PC.

A BASIC interactive graphics program has been developed for the IBM-PC which utilizes the graphics capabilities of that computer to display and manipulate protein structure from coordinates. Structures may be generated from typed files, or from Brookhaven National Laboratories' Protein Data Bank data tapes. Once displayed, images may be rotated, translated and expanded to any desired size. Figures may be viewed as ball-and-stick or space-filling models. Calculated multiple-point perspective may also be added to the display. Docking manipulations are possible since more than a single figure may be displayed and manipulated simultaneously. Further, stereo images and red/blue three-dimensional images may be generated using the accompanying DESIPLOT program and an HP-7475A plotter. A version of the program is also currently available for the Apple Macintosh. Full implementation on the Macintosh requires 512 K and at least one disk drive. Otherwise this version is essentially identical to the IBM-PC version described herein.

Computer Graphics↗

CLANS: a Java application for visualizing protein families based on pairwise similarity.

SUMMARY: The main source of hypotheses on the structure and function of new proteins is their homology to proteins with known properties. Homologous relationships are typically established through sequence similarity searches, multiple alignments and phylogenetic reconstruction. In cases where the number of potential relationships is large, for example in P-loop NTPases with many thousands of members, alignments and phylogenies become computationally demanding, accumulate errors and lose resolution. In search of a better way to analyze relationships in large sequence datasets we have developed a Java application, CLANS (CLuster ANalysis of Sequences), which uses a version of the Fruchterman-Reingold graph layout algorithm to visualize pairwise sequence similarities in either two-dimensional or three-dimensional space. AVAILABILITY: CLANS can be downloaded at http://protevo.eb.tuebingen.mpg.de/download.

Algorithms↗

Automatic detection of subsystem/pathway variants in genome analysis.

MOTIVATION: Proteins work together in pathways and networks, collectively comprising the cellular machinery. A subsystem (a generalization of pathway concept) is a group of related functional roles (such as enzymes) jointly involved in a specific aspect of the cellular machinery. Subsystems provide a natural framework for comparative genome analysis and functional annotation. A subsystem may be implemented in a number of different functional variants in individual species. In order to reliably project functional assignments across multiple genomes, we have to be able to identify the variants implemented in each genome. The analysis of such variants across diverse species is an interesting problem by itself and may provide new evolutionary insights. However, no computational techniques are presently available for an automated detection and analysis of subsystem variants. RESULTS: Here we formulate the subsystem variant detection problem as finding the minimum number of subgraphs of a subsystem, which is represented as a graph, and solve the optimization problem by integer programming approach. The performance of our method was tested on subsystems encoded in the SEED, a genomic integration platform developed by the Fellowship for Interpretation of Genomes as a component of a large-scale effort on comparative analysis and annotation of multiple diverse genomes. Here we illustrate the results obtained for two expert-encoded subsystems of the biosynthesis of Coenzyme A and FMN/FAD cofactors. Applications of variant detection, to support genomic annotations and to assess divergence of species, are briefly discussed in the context of these universally conserved and essential metabolic subsystems. SUPPLEMENTARY INFORMATION: The details of the variant detection results are available at http://ffas.burnham.org/svar/supp.html.

Animals↗

EMAN: semiautomated software for high-resolution single-particle reconstructions.

We present EMAN (Electron Micrograph ANalysis), a software package for performing semiautomated single-particle reconstructions from transmission electron micrographs. The goal of this project is to provide software capable of performing single-particle reconstructions beyond 10 A as such high-resolution data become available. A complete single-particle reconstruction algorithm is implemented. Options are available to generate an initial model for particles with no symmetry, a single axis of rotational symmetry, or icosahedral symmetry. Model refinement is an iterative process, which utilizes classification by model-based projection matching. CTF (contrast transfer function) parameters are determined using a new paradigm in which data from multiple micrographs are fit simultaneously. Amplitude and phase CTF correction is then performed automatically as part of the refinement loop. A graphical user interface is provided, so even those with little image processing experience will be able to begin performing reconstructions. Advanced users can directly use the lower level shell commands and even expand the package utilizing EMAN's extensive image-processing library. The package was written from scratch in C++ and is provided free of charge on our Web site. We present an overview of the package as well as several conformance tests with simulated data.

Algorithms↗

Simultaneous modelling of metabolic, genetic and product-interaction networks.

The creation of cell models from annotated genome information, as well as additional data from other databases, requires both a format and medium for its distribution. Standards are described for the representation of the data in the form of Document Type Definitions (DTDs) for XML files. Separate DTDs are detailed for genetic, metabolic and gene product-interaction networks, which can be used to hold information on individual subsystems, or which may be combined to create a whole cell DTD. In the execution of this work, a fifth DTD was also created for a metabolite thesaurus, which allows incorporation of metabolite synonyms and generic nomenclature data into the models. A gene-regulation classification scheme was also created, to facilitate incorporation of gene regulatory information in an efficient manner. The work is described with particular reference to the metabolic network of Escherichia coli, which contains 808 individual enzymes. The assignment of confidence levels to these data, through the use of Gene Ontology evidence codes, is highlighted. In silico investigations may now be performed using the mathematical simulation workbench, DBsolve, which incorporates the facility to introduce data directly from XML.

Computational Biology↗

CNplot: visualizing pre-clustered networks.

SUMMARY: CNplot is a simple technique for the visualization of global connectivity within pre-clustered network data. CNplot is easy to implement and in most cases produces informative and satisfactory summary of the data. AVAILABILITY: A Java implementation is available that allows users to modify graphics parameters and produces a LaTeX output. This software is free and is available at http://csb.stanford.edu/nbatada/VCN.html

Cluster Analysis↗

BioNetGen: software for rule-based modeling of signal transduction based on the interactions of molecular domains.

BioNetGen allows a user to create a computational model that characterizes the dynamics of a signal transduction system, and that accounts comprehensively and precisely for specified enzymatic activities, potential post-translational modifications and interactions of the domains of signaling molecules. The output defines and parameterizes the network of molecular species that can arise during signaling and provides functions that relate model variables to experimental readouts of interest. Models that can be generated are relevant for rational drug discovery, analysis of proteomic data and mechanistic studies of signal transduction.

Algorithms↗

BioLingua: a programmable knowledge environment for biologists.

UNLABELLED: BioLingua is an interactive, web-based programming environment that enables biologists to analyze biological systems by combining knowledge and data through direct end-user programming. BioLingua embeds a mature symbolic programming language in a frame-based knowledge environment, integrating genomic and pathway knowledge about a class of similar organisms. The BioLingua language provides interfaces to numerous state-of-the-art bioinformatic tools, making these available as an integrated package through the novel use of web-based programmability and an integrated Wiki-based community code and data store. The pilot instantiation of BioLingua, which has been developed in collaboration with several cyanobacteriologists, integrates knowledge about a subset of cyanobacteria with the Gene Ontology, KEGG and BioCyc knowledge bases. We introduce the BioLingua concept, architecture and language, and give several examples of its use in complex analyses. AVAILABILITY: Extensive documentation is available online at http://nostoc.stanford.edu/Docs/index.html CONTACT: JShrager@Stanford.edu

Artificial Intelligence↗

MESHI: a new library of Java classes for molecular modeling.

UNLABELLED: Adapting a modular and object-oriented approach in the design of molecular modeling packages may reduce the software development barrier between ideas and their programed applications. Towards this goal we developed MESHI, a new, strictly object-oriented, molecular modeling suite written in Java. MESHI provides a comprehensive library of extendable classes for all the essential components of molecular modeling: molecular and geometry elements, energy functions and optimization methods. AVAILABILITY: MESHI and its related documentation are freely available at http://www.cs.bgu.ac.il/~meshi; the MESHI API is available at http://www.cs.bgu.ac.il/~meshi/API CONTACT: keasar@cs.bgu.ac.il SUPPLEMENTARY INFORMATION: The Supplementary information includes (1) a detailed description of several key packages and classes, and (2) a brief presentation of results achieved by using the MESHI application--Beautify--in the CASP6 experiment.

Computer Simulation↗

COPASI--a COmplex PAthway SImulator.

MOTIVATION: Simulation and modeling is becoming a standard approach to understand complex biochemical processes. Therefore, there is a big need for software tools that allow access to diverse simulation and modeling methods as well as support for the usage of these methods. RESULTS: Here, we present COPASI, a platform-independent and user-friendly biochemical simulator that offers several unique features. We discuss numerical issues with these features; in particular, the criteria to switch between stochastic and deterministic simulation methods, hybrid deterministic-stochastic methods, and the importance of random number generator numerical resolution in stochastic simulation. AVAILABILITY: The complete software is available in binary (executable) for MS Windows, OS X, Linux (Intel) and Sun Solaris (SPARC), as well as the full source code under an open source license from http://www.copasi.org.

Algorithms↗

Design and implementation of a tool for translating SBML into the biochemical stochastic pi-calculus.

MOTIVATION: SBML is becoming a standard 'de-facto' to represent and store biological models. Although SBML is very useful in defining ways of exchanging and storing biological information, it is not formal enough to allow direct translation into non ambiguous formal representation languages to perform analysis and simulation of models. We here suggest to map SBML models into process calculi representations. RESULTS: We implemented and validated a tool that translates SBML descriptions into stochastic pi-calculus specifications. AVAILABILITY: Source code is freely available for academic use by contacting the authors.

Algorithms↗

Effective ambiguity checking in biosequence analysis.

BACKGROUND: Ambiguity is a problem in biosequence analysis that arises in various analysis tasks solved via dynamic programming, and in particular, in the modeling of families of RNA secondary structures with stochastic context free grammars. Several types of analysis are invalidated by the presence of ambiguity. As this problem inherits undecidability (as we show here) from the namely problem for context free languages, there is no complete algorithmic solution to the problem of ambiguity checking. RESULTS: We explain frequently observed sources of ambiguity, and show how to avoid them. We suggest four testing procedures that may help to detect ambiguity when present, including a just-in-time test that permits to work safely with a potentially ambiguous grammar. We introduce, for the special case of stochastic context free grammars and RNA structure modeling, an automated partial procedure for proving non-ambiguity. It is used to demonstrate non-ambiguity for several relevant grammars. CONCLUSION: Our mechanical proof procedure and our testing methods provide a powerful arsenal of methods to ensure non-ambiguity.

Algorithms↗

ROKU: a novel method for identification of tissue-specific genes.

BACKGROUND: One of the important goals of microarray research is the identification of genes whose expression is considerably higher or lower in some tissues than in others. We would like to have ways of identifying such tissue-specific genes. RESULTS: We describe a method, ROKU, which selects tissue-specific patterns from gene expression data for many tissues and thousands of genes. ROKU ranks genes according to their overall tissue specificity using Shannon entropy and detects tissues specific to each gene if any exist using an outlier detection method. We evaluated the capacity for the detection of various specific expression patterns using synthetic and real data. We observed that ROKU was superior to a conventional entropy-based method in its ability to rank genes according to overall tissue specificity and to detect genes whose expression pattern are specific only to objective tissues. CONCLUSION: ROKU is useful for the detection of various tissue-specific expression patterns. The framework is also directly applicable to the selection of diagnostic markers for molecular classification of multiple classes.

Algorithms↗