Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Protein language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Ruby-Helix: an implementation of helical image processing based on object-oriented scripting language.

Helical image analysis in combination with electron microscopy has been used to study three-dimensional structures of various biological filaments or tubes, such as microtubules, actin filaments, and bacterial flagella. A number of packages have been developed to carry out helical image analysis. Some biological specimens, however, have a symmetry break (seam) in their three-dimensional structure, even though their subunits are mostly arranged in a helical manner. We refer to these objects as "asymmetric helices". All the existing packages are designed for helically symmetric specimens, and do not allow analysis of asymmetric helical objects, such as microtubules with seams. Here, we describe Ruby-Helix, a new set of programs for the analysis of "helical" objects with or without a seam. Ruby-Helix is built on top of the Ruby programming language and is the first implementation of asymmetric helical reconstruction for practical image analysis. It also allows easier and semi-automated analysis, performing iterative unbending and accurate determination of the repeat length. As a result, Ruby-Helix enables us to analyze motor-microtubule complexes with higher throughput to higher resolution.

Algorithms↗

Sequential and parallel molecular mechanics calculations.

This article describes a gradient algorithm for the computational optimization of model molecular structures, and discusses the various compromises inherent in the practical expression of the algorithm in a Fortran computer program (VULCAN) for both sequential and parallel computers. Details are given of some previously undiscussed properties of gradient algorithms; various acceleration techniques are compared; and some traps for the unwary are highlighted.

Algorithms↗

Object-oriented design tools for supramolecular devices and biomedical nanotechnology.

Nanotechnology provides multifunctional agents for in vivo use that increasingly blur the distinction between pharmaceuticals and medical devices. Realization of such therapeutic nanodevices requires multidisciplinary effort that is difficult for individual device developers to sustain, and identification of appropriate collaborations outside ones own field can itself be challenging. Further, as in vivo nanodevices become increasingly complex, their design will increasingly demand systems level thinking. System engineering tools such as object-oriented analysis, object-oriented design (OOA/D) and unified modeling language (UML) are applicable to nanodevices built from biological components, help logically manage the knowledge needed to design them, and help identify useful collaborative relationships for device designers. We demonstrate the utility of these systems engineering tools by reverse engineering an existing molecular device (the bacmid molecular cloning system) using them, and illustrate how object-oriented approaches identify fungible components (objects) in nanodevices in a way that facilitates design of families of related devices, rather than single inventions. We also explore the utility of object-oriented approaches for design of another class of therapeutic nanodevices, vaccines. While they are useful for design of current nanodevices, the power of systems design tools for biomedical nanotechnology will become increasingly apparent as the complexity and sophistication of in vivo nanosystems increases. The nested, hierarchical nature of object-oriented approaches allows treatment of devices as objects in higher-order structures, and so will facilitate concatenation of multiple devices into higher-order, higher-function nanosystems.

Baculoviridae↗

ATID: a web-oriented database for collection of publicly available alternative translational initiation events.

SUMMARY: Alternative translational initiation is an important cellular mechanism contributing to the diversity of protein products and functions. We develop a database that provides a comprehensive collection of alternative translational initiation events. The purpose of this alternative translational initiation database (ATID) is to facilitate the systematic study of alternative translational initiation of genes. The current version of database contains 300 genes from Homo sapiens, Mus musculus and other species. Each of the genes has two or more isoforms due to alternative translational initiation. Resources in ATID, including gene information, alternative products of genes and domain structures of isoforms, are provided through a user-friendly web interface. AVAILABILITY: The ATID database is available for public use at http://bioinfo.au.tsinghua.edu.cn/atie/.

Amino Acid Sequence↗

An iterative statistical approach to the identification of protein phosphorylation motifs from large-scale data sets.

With the recent exponential increase in protein phosphorylation sites identified by mass spectrometry, a unique opportunity has arisen to understand the motifs surrounding such sites. Here we present an algorithm designed to extract motifs from large data sets of naturally occurring phosphorylation sites. The methodology relies on the intrinsic alignment of phospho-residues and the extraction of motifs through iterative comparison to a dynamic statistical background. Results show the identification of dozens of novel and known phosphorylation motifs from recently published serine, threonine and tyrosine phosphorylation studies. When applied to a linguistic data set to test the versatility of the approach, the algorithm successfully extracted hundreds of language motifs. This method, in addition to shedding light on the consensus sequences of identified and as yet unidentified kinases and modular protein domains, may also eventually be used as a tool to determine potential phosphorylation sites in proteins of interest.

Algorithms↗

Linguistic complexity of protein sequences as compared to texts of human languages.

A notion and a measure of linguistic complexity introduced earlier (Trifonov, 1990) were originally used for analysis of nucleotide sequences. This measure was shown to reflect multiplicity of codes (messages) of different natures superimposed in the sequences. Unlike human language texts, genetic texts are 'read' by cellular mechanisms in several different ways, each time using a different selection of the characters of the same text while skipping others (Trifonov, 1989). Human texts are read in one way only, sequentially and involving all characters (one code). The conceptual significance and essence of the idea on the multiplicity of overlapping codes in genetic sequences, as opposed to human languages, is discussed. The linguistic complexity technique allows a calculation to be made of the structural complexity of any linear sequence of characters irrespective of whether the text is cognized or presently undeciphered. The texts (sequences) are compared exclusively from the point of view of their structural complexity with no reference to the meaning of the texts which is beyond the scope of this article. Results of such a comparison of protein sequences with various texts, written in English, Italian and Welsh are presented. The human texts are found to be structurally simpler than genetic (protein) texts, reflecting, apparently, a difference in the reading modes: single code versus many codes.

Amino Acid Sequence↗

Evaluation of methods for predicting the topology of beta-barrel outer membrane proteins and a consensus prediction method.

BACKGROUND: Prediction of the transmembrane strands and topology of beta-barrel outer membrane proteins is of interest in current bioinformatics research. Several methods have been applied so far for this task, utilizing different algorithmic techniques and a number of freely available predictors exist. The methods can be grossly divided to those based on Hidden Markov Models (HMMs), on Neural Networks (NNs) and on Support Vector Machines (SVMs). In this work, we compare the different available methods for topology prediction of beta-barrel outer membrane proteins. We evaluate their performance on a non-redundant dataset of 20 beta-barrel outer membrane proteins of gram-negative bacteria, with structures known at atomic resolution. Also, we describe, for the first time, an effective way to combine the individual predictors, at will, to a single consensus prediction method. RESULTS: We assess the statistical significance of the performance of each prediction scheme and conclude that Hidden Markov Model based methods, HMM-B2TMR, ProfTMB and PRED-TMBB, are currently the best predictors, according to either the per-residue accuracy, the segments overlap measure (SOV) or the total number of proteins with correctly predicted topologies in the test set. Furthermore, we show that the available predictors perform better when only transmembrane beta-barrel domains are used for prediction, rather than the precursor full-length sequences, even though the HMM-based predictors are not influenced significantly. The consensus prediction method performs significantly better than each individual available predictor, since it increases the accuracy up to 4% regarding SOV and up to 15% in correctly predicted topologies. CONCLUSIONS: The consensus prediction method described in this work, optimizes the predicted topology with a dynamic programming algorithm and is implemented in a web-based application freely available to non-commercial users at http://bioinformatics.biol.uoa.gr/ConBBPRED.

Algorithms↗

BIOCHAM: an environment for modeling biological systems and formalizing experimental knowledge.

UNLABELLED: BIOCHAM (the BIOCHemical Abstract Machine) is a software environment for modeling biochemical systems. It is based on two aspects: (1) the analysis and simulation of boolean, kinetic and stochastic models and (2) the formalization of biological properties in temporal logic. BIOCHAM provides tools and languages for describing protein networks with a simple and straightforward syntax, and for integrating biological properties into the model. It then becomes possible to analyze, query, verify and maintain the model with respect to those properties. For kinetic models, BIOCHAM can search for appropriate parameter values in order to reproduce a specific behavior observed in experiments and formalized in temporal logic. Coupled with other methods such as bifurcation diagrams, this search assists the modeler/biologist in the modeling process. AVAILABILITY: BIOCHAM (v. 2.5) is a free software available for download, with example models, at http://contraintes.inria.fr/BIOCHAM/.

Algorithms↗

The Molecular Biology Toolkit (MBT): a modular platform for developing molecular visualization applications.

BACKGROUND: The large amount of data that are currently produced in the biological sciences can no longer be explored and visualized efficiently with traditional, specialized software. Instead, new capabilities are needed that offer flexibility, rapid application development and deployment as standalone applications or available through the Web. RESULTS: We describe a new software toolkit--the Molecular Biology Toolkit (MBT; http://mbt.sdsc.edu)--that enables fast development of applications for protein analysis and visualization. The toolkit is written in Java, thus offering platform-independence and Internet delivery capabilities. Several applications of the toolkit are introduced to illustrate the functionality that can be achieved. CONCLUSIONS: The MBT provides a well-organized assortment of core classes that provide a uniform data model for the description of biological structures and automate most common tasks associated with the development of applications in the molecular sciences (data loading, derivation of typical structural information, visualization of sequence and standard structural entities).

Algorithms↗

A path planning approach for computing large-amplitude motions of flexible molecules.

MOTIVATION: Motion is inherent in molecular interactions. Molecular flexibility must be taken into account in order to develop accurate computational techniques for predicting interactions. Energy-based methods currently used in molecular modeling (i.e. molecular dynamics, Monte Carlo algorithms) are, in practice, only able to compute local motions while accounting for molecular flexibility. However, large-amplitude motions often occur in biological processes. We investigate the application of geometric path planning algorithms to compute such large motions in flexible molecular models. Our purpose is to exploit the efficacy of a geometric conformational search as a filtering stage before subsequent energy refinements. RESULTS: In this paper two kinds of large-amplitude motion are treated: protein loop conformational changes (involving protein backbone flexibility) and ligand trajectories to deep active sites in proteins (involving ligand and protein side-chain flexibility). First studies performed using our two-stage approach (geometric search followed by energy refinements) show that, compared to classical molecular modeling methods, quite similar results can be obtained with a performance gain of several orders of magnitude. Furthermore, our results also indicate that the geometric stage can provide highly valuable information to biologists. AVAILABILITY: The algorithms have been implemented in the general-purpose motion planning software Move3D, developed at LAAS-CNRS. We are currently working on an optimized stand-alone library that will be available to the scientific community.

Algorithms↗

A combined approach for ab initio construction of low resolution protein tertiary structures from sequence.

An approach to construct low resolution models of protein structure from sequence information using a combination of different methodologies is described. All possible compact self-avoiding C alpha conformations (approximately 10 million) of a small protein chain were exhaustively enumerated on a tetrahedral lattice. The best scoring 10,000 conformations were selected using a lattice-based scoring function. All-atom structures were then generated by fitting an off-lattice four-state phi/psi model to the lattice conformations, using idealised helix and sheet values based on predicted secondary structure. The all-atom conformations were minimised using ENCAD and scored using a second hybrid scoring function. The best scoring 50, 100, and 500 conformations were input to a consensus-based distance geometry routine that used constraints from each the conformation sets and produced a single structure for each set (total of three). Secondary structures were again fitted to the three structures, and the resulting structures were minimised and scored. The lowest scoring conformation was taken to be the "correct" answer. The results of application of this method to twelve proteins are presented.

Amino Acid Sequence↗

Generating quantitative models describing the sequence specificity of biological processes with the stabilized matrix method.

BACKGROUND: Many processes in molecular biology involve the recognition of short sequences of nucleic-or amino acids, such as the binding of immunogenic peptides to major histocompatibility complex (MHC) molecules. From experimental data, a model of the sequence specificity of these processes can be constructed, such as a sequence motif, a scoring matrix or an artificial neural network. The purpose of these models is two-fold. First, they can provide a summary of experimental results, allowing for a deeper understanding of the mechanisms involved in sequence recognition. Second, such models can be used to predict the experimental outcome for yet untested sequences. In the past we reported the development of a method to generate such models called the Stabilized Matrix Method (SMM). This method has been successfully applied to predicting peptide binding to MHC molecules, peptide transport by the transporter associated with antigen presentation (TAP) and proteasomal cleavage of protein sequences. RESULTS: Herein we report the implementation of the SMM algorithm as a publicly available software package. Specific features determining the type of problems the method is most appropriate for are discussed. Advantageous features of the package are: (1) the output generated is easy to interpret, (2) input and output are both quantitative, (3) specific computational strategies to handle experimental noise are built in, (4) the algorithm is designed to effectively handle bounded experimental data, (5) experimental data from randomized peptide libraries and conventional peptides can easily be combined, and (6) it is possible to incorporate pair interactions between positions of a sequence. CONCLUSION: Making the SMM method publicly available enables bioinformaticians and experimental biologists to easily access it, to compare its performance to other prediction methods, and to extend it to other applications.

Algorithms↗

viwish: a visualization server for protein modelling and docking.

A visualization tool viwish for proteins based on the Tcl command language has been developed. The system is completely menu driven and can display arbitrary many proteins in arbitrary many windows. It isinstantly t o use, even for non computer experts and provides possibilities to modify menus, configurations, and windows. It may be used as a stand-alone molecular graphics package or as a graphics server for external programs. Communications with these client applications is established even across different machines (through the send command to Tk, an extension of Tcl). In addition, a wide rage of chemical data like molecular surfaces and 3D gridded samplings of chemical features can be displayed. Therefore the systmen is especially useful for the development of algorithms that need visual distributed freely, including the source code.

Binding Sites↗

PDBML: the representation of archival macromolecular structure data in XML.

SUMMARY: The Protein Data Bank (PDB) has recently released versions of the PDB Exchange dictionary and the PDB archival data files in XML format collectively named PDBML. The automated generation of these XML files is driven by the data dictionary infrastructure in use at the PDB. The correspondences between the PDB dictionary and the XML schema metadata are described as well as the XML representations of PDB dictionaries and data files.

Amino Acid Sequence↗

XtalView/Xfit--A versatile program for manipulating atomic coordinates and electron density.

Xfit is a model-building and map viewing program in XtalView that is used by the structural biology community including researchers in the fields of crystallography, molecular modeling, and electron microscopy. Among its distinguishing features are built-in fast Fourier transforms that allow users flexibility in map calculations including the creation of OMIT maps and the updating of structure factors to reflect model changes from within the program. Written in C and using the freely available XView toolkit, it is highly portable to almost any X-windows based workstation including Intel-based LINUX systems. Its user interface is designed to aid in facile model-building and contains a semiautomated fitting system that allows the user to interactively and rapidly build chain de novo into an electron density map. The program is highly optimized to allow such features as interactive contour levels and map calculations to be completed within a few seconds. Features in the latest version including phase-combination, solvent-flattening, automated water addition, and small-probe dot contact surfaces, as well as basic design features, are discussed.

Computing Methodologies↗

Algorithm for exhaustive and nonredundant organic stereoisomer generation.

Generation of organic stereoisomers with R/S, Z/E, and/or M/P configurations that may contain heteroatoms, multiple bonds, and any kind of cycle (isolated, spiro, condensed, and nested) is described. Inputs for processing are molecular structures in a N_tuple format resident on an automatic (canonical) or manual (non canonical) generated file which are processed by doing internal molecular graph construction, a weighted bipartite tree construction for all atoms and bonds to detect stereocenters, and symmetrical atom groups (SAG) with some specific SAG parameters that constitute a novel way for redundancy elimination of meso structures. Finally, determination of ligand CIP priorities allows for writing the output N_tuples with stereoisomer description. Several examples showing application of this methology to a wide number of structures are also presented.

Algorithms↗

AstexViewer: a visualisation aid for structure-based drug design.

AstexViewer is a Java molecular graphics program that can be used for visualisation in many aspects of structure-based drug design. This paper describes its functionality, implementation and examples of its use. The program can run as an Applet in a web browser allowing structures to be displayed without installing additional software. Applications of its use are described for visualisation and as part of a structure based design platform. The software is being made freely available to the community and may be downloaded from http://www.astex-technology.com/AstexViewer.

Computer Graphics↗