Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Protein language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Stratus not altocumulus: a new view of the yeast protein interaction network.

Systems biology approaches can reveal intermediary levels of organization between genotype and phenotype that often underlie biological phenomena such as polygenic effects and protein dispensability. An important conceptualization is the module, which is loosely defined as a cohort of proteins that perform a dedicated cellular task. Based on a computational analysis of limited interaction datasets in the budding yeast Saccharomyces cerevisiae, it has been suggested that the global protein interaction network is segregated such that highly connected proteins, called hubs, tend not to link to each other. Moreover, it has been suggested that hubs fall into two distinct classes: "party" hubs are co-expressed and co-localized with their partners, whereas "date" hubs interact with incoherently expressed and diversely localized partners, and thereby cohere disparate parts of the global network. This structure may be compared with altocumulus clouds, i.e., cotton ball-like structures sparsely connected by thin wisps. However, this organization might reflect a small and/or biased sample set of interactions. In a multi-validated high-confidence (HC) interaction network, assembled from all extant S. cerevisiae interaction data, including recently available proteome-wide interaction data and a large set of reliable literature-derived interactions, we find that hub-hub interactions are not suppressed. In fact, the number of interactions a hub has with other hubs is a good predictor of whether a hub protein is essential or not. We find that date hubs are neither required for network tolerance to node deletion, nor do date hubs have distinct biological attributes compared to other hubs. Date and party hubs do not, for example, evolve at different rates. Our analysis suggests that the organization of global protein interaction network is highly interconnected and hence interdependent, more like the continuous dense aggregations of stratus clouds than the segregated configuration of altocumulus clouds. If the network is configured in a stratus format, cross-talk between proteins is potentially a major source of noise. In turn, control of the activity of the most highly connected proteins may be vital. Indeed, we find that a fluctuation in steady-state levels of the most connected proteins is minimized.

Computational Biology↗

PDB file parser and structure class implemented in Python.

UNLABELLED: The biopython project provides a set of bioinformatics tools implemented in Python. Recently, biopython was extended with a set of modules that deal with macromolecular structure. Biopython now contains a parser for PDB files that makes the atomic information available in an easy-to-use but powerful data structure. The parser and data structure deal with features that are often left out or handled inadequately by other packages, e.g. atom and residue disorder (if point mutants are present in the crystal), anisotropic B factors, multiple models and insertion codes. In addition, the parser performs some sanity checking to detect obvious errors. AVAILABILITY: The Biopython distribution (including source code and documentation) is freely available (under the Biopython license) from http://www.biopython.org

Computer Simulation↗

The use of edge-betweenness clustering to investigate biological function in protein interaction networks.

BACKGROUND: This paper describes an automated method for finding clusters of interconnected proteins in protein interaction networks and retrieving protein annotations associated with these clusters. RESULTS: Protein interaction graphs were separated into subgraphs of interconnected proteins, using the JUNG implementation of Girvan and Newman's Edge-Betweenness algorithm. Functions were sought for these subgraphs by detecting significant correlations with the distribution of Gene Ontology terms which had been used to annotate the proteins within each cluster. The method was implemented using freely available software (JUNG and the R statistical package). Protein clusters with significant correlations to functional annotations could be identified and included groups of proteins know to cooperate in cell metabolism. The method appears to be resilient against the presence of false positive interactions. CONCLUSION: This method provides a useful tool for rapid screening of small to medium size protein interaction datasets.

Algorithms↗

MACiE: a database of enzyme reaction mechanisms.

SUMMARY: MACiE (mechanism, annotation and classification in enzymes) is a publicly available web-based database, held in CMLReact (an XML application), that aims to help our understanding of the evolution of enzyme catalytic mechanisms and also to create a classification system which reflects the actual chemical mechanism (catalytic steps) of an enzyme reaction, not only the overall reaction. AVAILABILITY: http://www-mitchell.ch.cam.ac.uk/macie/.

Catalysis↗

FPV: fast protein visualization using Java 3D.

MOTIVATION: Many tools have been developed to visualize protein structures. Tools that have been based on Java 3D((TM)) are compatible among different systems and they can be run remotely through web browsers. However, using Java 3D for visualization has some performance issues with it. The primary concerns about molecular visualization tools based on Java 3D are in their being slow in terms of interaction speed and in their inability to load large molecules. This behavior is especially apparent when the number of atoms to be displayed is huge, or when several proteins are to be displayed simultaneously for comparison. RESULTS: In this paper we present techniques for organizing a Java 3D scene graph to tackle these problems. We have developed a protein visualization system based on Java 3D and these techniques. We demonstrate the effectiveness of the proposed method by comparing the visualization component of our system with two other Java 3D based molecular visualization tools. In particular, for van der Waals display mode, with the efficient organization of the scene graph, we could achieve up to eight times improvement in rendering speed and could load molecules three times as large as the previous systems could. AVAILABILITY: EPV is freely available with source code at the following URL: http://www.cs.ucsb.edu/~tcan/fpv/

Computer Graphics↗

Clustering the annotation space of proteins.

BACKGROUND: Current protein clustering methods rely on either sequence or functional similarities between proteins, thereby limiting inferences to one of these areas. RESULTS: Here we report a new approach, named CLAN, which clusters proteins according to both annotation and sequence similarity. This approach is extremely fast, clustering the complete SwissProt database within minutes. It is also accurate, recovering consistent protein families agreeing on average in more than 97% with sequence-based protein families from Pfam. Discrepancies between sequence- and annotation-based clusters were scrutinized and the reasons reported. We demonstrate examples for each of these cases, and thoroughly discuss an example of a propagated error in SwissProt: a vacuolar ATPase subunit M9.2 erroneously annotated as vacuolar ATP synthase subunit H. CLAN algorithm is available from the authors and the CLAN database is accessible at http://maine.ebi.ac.uk:8000/cgi-bin/clan/ClanSearch.pl CONCLUSIONS: CLAN creates refined function-and-sequence specific protein families that can be used for identification and annotation of unknown family members. It also allows easy identification of erroneous annotations by spotting inconsistencies between similarities on annotation and sequence levels.

Adenosine Triphosphatases↗

Calculation of helix packing angles in protein structures.

UNLABELLED: Software is presented for the calculation of packing angles and geometry of helical secondary structure elements in protein structures. AVAILABILITY: C language source code and documentation is available from http://www.bioinformatics.leeds.ac.uk.

Algorithms↗

A top-level ontology of functions and its application in the Open Biomedical Ontologies.

MOTIVATION: A clear understanding of functions in biology is a key component in accurate modelling of molecular, cellular and organismal biology. Using the existing biomedical ontologies it has been impossible to capture the complexity of the community's knowledge about biological functions. RESULTS: We present here a top-level ontological framework for representing knowledge about biological functions. This framework lends greater accuracy, power and expressiveness to biomedical ontologies by providing a means to capture existing functional knowledge in a more formal manner. An initial major application of the ontology of functions is the provision of a principled way in which to curate functional knowledge and annotations in biomedical ontologies. Further potential applications include the facilitation of ontology interoperability and automated reasoning. A major advantage of the proposed implementation is that it is an extension to existing biomedical ontologies, and can be applied without substantial changes to these domain ontologies. AVAILABILITY: The Ontology of Functions (OF) can be downloaded in OWL format from http://onto.eva.mpg.de/. Additionally, a UML profile and supplementary information and guides for using the OF can be accessed from the same website.

Biomedical Engineering↗

Visualization of the entire surface of a protein by cartographic projection.

A minicomputer based system for the determination and schematic representation of protein surfaces is described. The algorithms are based on the atomic coordinates of globular protein molecules of the Brookhaven Protein Data Bank. Using a cartographic projection a normalized graphic representation is obtained of the amino acid residues located on the surface of the considered protein. The programs are written in FORTRAN IV (surface determination) and BASIC (graphic representation).

Algorithms↗

Atlas - a data warehouse for integrative bioinformatics.

BACKGROUND: We present a biological data warehouse called Atlas that locally stores and integrates biological sequences, molecular interactions, homology information, functional annotations of genes, and biological ontologies. The goal of the system is to provide data, as well as a software infrastructure for bioinformatics research and development. DESCRIPTION: The Atlas system is based on relational data models that we developed for each of the source data types. Data stored within these relational models are managed through Structured Query Language (SQL) calls that are implemented in a set of Application Programming Interfaces (APIs). The APIs include three languages: C++, Java, and Perl. The methods in these API libraries are used to construct a set of loader applications, which parse and load the source datasets into the Atlas database, and a set of toolbox applications which facilitate data retrieval. Atlas stores and integrates local instances of GenBank, RefSeq, UniProt, Human Protein Reference Database (HPRD), Biomolecular Interaction Network Database (BIND), Database of Interacting Proteins (DIP), Molecular Interactions Database (MINT), IntAct, NCBI Taxonomy, Gene Ontology (GO), Online Mendelian Inheritance in Man (OMIM), LocusLink, Entrez Gene and HomoloGene. The retrieval APIs and toolbox applications are critical components that offer end-users flexible, easy, integrated access to this data. We present use cases that use Atlas to integrate these sources for genome annotation, inference of molecular interactions across species, and gene-disease associations. CONCLUSION: The Atlas biological data warehouse serves as data infrastructure for bioinformatics research and development. It forms the backbone of the research activities in our laboratory and facilitates the integration of disparate, heterogeneous biological sources of data enabling new scientific inferences. Atlas achieves integration of diverse data sets at two levels. First, Atlas stores data of similar types using common data models, enforcing the relationships between data types. Second, integration is achieved through a combination of APIs, ontology, and tools. The Atlas software is freely available under the GNU General Public License at: http://bioinformatics.ubc.ca/atlas/

Computational Biology↗

PTGL--a web-based database application for protein topologies.

Protein Topology Graph Library (PTGL) is a database application for the representation and retrieval of protein topologies. Protein topologies are based on a graph-theoretical protein model at secondary structure level. Different views on protein topology are given by four linear notations for their characterization. Protein topologies can be derived at different description levels considering alpha- and beta-structures. The on-line search tool is based on an object-relational database and provides a query browser for data interrogation by string patterns, keyword queries and sequence similarity. Protein topologies are represented both as schematic diagrams and as three-dimensional images.

Computer Graphics↗

FOLD: integrated analysis and display of protein secondary structure.

FOLD, a computer program for the definition and analysis of protein secondary structure, is described. Algorithms implemented in the software are reviewed. These include methods for the identification of simple features such as hydrogen bonds, alpha helices, beta strands, beta bulges, and beta and psi turns. Techniques are also described for the definition and analysis of higher-order structures, such as beta hairpins, beta sheets and their topology, and beta barrels. In addition to considerable textual output the program supports visualization of protein secondary structure in either an atom-based display style or one reproducing the characteristics of a so-called ribbon drawing.

Computer Graphics↗

Comparing protein-ligand docking programs is difficult.

There is currently great interest in comparing protein-ligand docking programs. A review of recent comparisons shows that it is difficult to draw conclusions of general applicability. Statistical hypothesis testing is required to ensure that differences in pose-prediction success rates and enrichment rates are significant. Numerical measures such as root-mean-square deviation need careful interpretation and may profitably be supplemented by interaction-based measures and visual inspection of dockings. Test sets must be of appropriate diversity and of good experimental reliability. The effects of crystal-packing interactions may be important. The method used for generating starting ligand geometries and positions may have an appreciable effect on docking results. For fair comparison, programs must be given search problems of equal complexity (e.g. binding-site regions of the same size) and approximately equal time in which to solve them. Comparisons based on rescoring require local optimization of the ligand in the space of the new objective function. Re-implementations of published scoring functions may give significantly different results from the originals. Ostensibly minor details in methodology may have a profound influence on headline success rates.

Algorithms↗

CX, an algorithm that identifies protruding atoms in proteins.

MOTIVATION: A simple and fast algorithm is described that calculates a measure of protrusion (cx) for atoms in protein structures, directly useable with the common molecular graphics programs. RESULTS: A sphere of predetermined radius is centered around each non-hydrogen atom, and the volume occupied by the protein and the free volume within the sphere (internal and external volumes, respectively) are calculated. Atoms in protruding regions have a high ratio (cx) between the external and the internal volume. The program reads a PDB file, and writes the output in the same format, with cx values in the B factor field. Output structure files can be directly displayed with standard molecular graphics programs like RASMOL, MOLMOL, Swiss-PDB Viewer and colored according to cx values. We show the potential use of this program in the analysis of two protein-protein complexes and in the prediction of limited proteolysis sites in native proteins. AVAILABILITY: The algorithm is implemented in a standalone program written in C and its source is freely available at ftp.icgeb.trieste.it/pub/CX or on request from the authors.

Algorithms↗

Query3d: a new method for high-throughput analysis of functional residues in protein structures.

BACKGROUND: The identification of local similarities between two protein structures can provide clues of a common function. Many different methods exist for searching for similar subsets of residues in proteins of known structure. However, the lack of functional and structural information on single residues, together with the low level of integration of this information in comparison methods, is a limitation that prevents these methods from being fully exploited in high-throughput analyses. RESULTS: Here we describe Query3d, a program that is both a structural DBMS (Database Management System) and a local comparison method. The method conserves a copy of all the residues of the Protein Data Bank annotated with a variety of functional and structural information. New annotations can be easily added from a variety of methods and known databases. The algorithm makes it possible to create complex queries based on the residues' function and then to compare only subsets of the selected residues. Functional information is also essential to speed up the comparison and the analysis of the results. CONCLUSION: With Query3d, users can easily obtain statistics on how many and which residues share certain properties in all proteins of known structure. At the same time, the method also finds their structural neighbours in the whole PDB. Programs and data can be accessed through the PdbFun web interface.

Algorithms↗

Representation for discovery of protein motifs.

There are several dimensions and levels of complexity in which information on protein motifs may be available. For example, one-dimensional sequence motifs may be associated with secondary structure identifiers. Alternatively, three-dimensional information on polypeptide segments may be used to induce prototypical three-dimensional structure templates. This paper surveys various representations encountered in the protein motif discovery literature. Many of the representations are based on incompatible semantics, making difficult the comparison and combination of previous results. To make better use of machine learning techniques and to provide for an integrated knowledge representation framework, a general representation language--in which all types of motifs can be encoded and given a uniform semantics--is required. In this paper we propose such a model, called a spatial description logic, and present a machine learning approach based on the model.

Amino Acids↗

Finding association rules on heterogeneous genome data.

A novel approach for discovery of knowledge from genome data, which has been recently watched with interest in the research area of database, is applied to finding unified rules spreading over sequence, structure, and function of protein. As the result of experiments using data extracted from PDB, SWISS-PROT, and PROSITE, some association rules stating sequential/structural/functional aspects of two kinds of endopeptidases were found.

Computer Simulation↗