Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

PRINTS--a database of protein motif fingerprints.

PRINTS is a compendium of protein motif 'fingerprints'. A fingerprint is defined as a group of motifs excised from conserved regions of a sequence alignment, whose diagnostic power or potency is refined by iterative databasescanning (in this case the OWL composite sequence database). Generally, the motifs do not overlap, but are separated along a sequence, though they may be contiguous in 3D-space. The use of groups of independent, linearly- or spatially-distinct motifs allows protein folds and functionalities to be characterised more flexibly and powerfully than conventional single-component patterns or regular expressions. The current version of the database contains 200 entries (encoding 950 motifs), covering a wide range of globular and membrane proteins, modular polypeptides, and so on. The growth of the databaseis influenced by a number of factors; e.g. the use of multiple motifs; the maximisation of sequence information through iterative database scanning; and the fact that the database searched is a large composite. The information contained within PRINTS is distinct from, but complementary to the consensus expressions stored in the widely-used PROSITE dictionary of patterns.

Amino Acid Sequence↗

Sequence analysis of plasmid pCC5.2 from cyanobacterium Synechocystis PCC 6803 that replicates by a rolling circle mechanism.

The cyanobacterium Synechocystis sp. strain PCC 6803 contains several cryptic plasmids of 2.4, 5.2, and about 50 or 100 kbp. The complete nucleotide sequence of the 5.2-kbp plasmid, pCC5.2, has been analyzed and is reported here. This plasmid contains 5214 bp and 53.1% A+T. Six open reading frames, ORFs A-F (encoding peptides of larger than 90 amino acid residues), were located on both strands of pCC5.2. ORF B codes for a potential replication protein containing 971 amino acids, for which there are three homologous proteins encoded by other cyanobacterial plasmids. ORF C encodes a polypeptide of 93 amino acids which shares homologies with products of two ORFs found in the protein database. Counterparts of the products of ORF A, D, E, and F could not be found in the protein database. Detection of a single-stranded DNA intermediate during replication of pCC5.2 indicates that this plasmid may also replicate by a rolling circle mechanism, as has been reported for pCA2.4 and pCB2.4 from the same strain of Synechocystis (PCC 6803).

Amino Acid Sequence↗

The Universal Protein Resource (UniProt).

The Universal Protein Resource (UniProt) provides the scientific community with a single, centralized, authoritative resource for protein sequences and functional information. Formed by uniting the Swiss-Prot, TrEMBL and PIR protein database activities, the UniProt consortium produces three layers of protein sequence databases: the UniProt Archive (UniParc), the UniProt Knowledgebase (UniProt) and the UniProt Reference (UniRef) databases. The UniProt Knowledgebase is a comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase with extensive cross-references. This centrepiece consists of two sections: UniProt/Swiss-Prot, with fully, manually curated entries; and UniProt/TrEMBL, enriched with automated classification and annotation. During 2004, tens of thousands of Knowledgebase records got manually annotated or updated; we introduced a new comment line topic: TOXIC DOSE to store information on the acute toxicity of a toxin; the UniProt keyword list got augmented by additional keywords; we improved the documentation of the keywords and are continuously overhauling and standardizing the annotation of post-translational modifications. Furthermore, we introduced a new documentation file of the strains and their synonyms. Many new database cross-references were introduced and we started to make use of Digital Object Identifiers. We also achieved in collaboration with the Macromolecular Structure Database group at EBI an improved integration with structural databases by residue level mapping of sequences from the Protein Data Bank entries onto corresponding UniProt entries. For convenient sequence searches we provide the UniRef non-redundant sequence databases. The comprehensive UniParc database stores the complete body of publicly available protein sequence data. The UniProt databases can be accessed online (http://www.uniprot.org) or downloaded in several formats (ftp://ftp.uniprot.org/pub). New releases are published every two weeks.

Amino Acid Sequence↗

The optimization of protein secondary structure determination with infrared and circular dichroism spectra.

We have used the circular dichroism and infrared spectra of a specially designed 50 protein database [Oberg, K.A., Ruysschaert, J.M. & Goormaghtigh, E. (2003) Protein Sci. 12, 2015-2031] in order to optimize the accuracy of spectroscopic protein secondary structure determination using multivariate statistical analysis methods. The results demonstrate that when the proteins are carefully selected for the diversity in their structure, no smaller subset of the database contains the necessary information to describe the entire set. One conclusion of the paper is therefore that large protein databases, observing stringent selection criteria, are necessary for the prediction of unknown proteins. A second important conclusion is that only the comparison of analyses run on circular dichroism and infrared spectra independently is able to identify failed solutions in the absence of known structure. Interestingly, it was also found in the course of this study that the amide II band has high information content and could be used alone for secondary structure prediction in place of amide I.

Algorithms↗

Locality-aware pooling enhances protein language model performance across varied applications.

MOTIVATION: Protein language models (PLMs) are amongst the most exciting recent advances for characterizing protein sequences, and have enabled a diverse set of applications, including structure determination, functional property prediction, and mutation impact assessment, all from single protein sequences alone. State-of-the-art PLMs leverage transformer architectures originally developed for natural language processing, and are pre-trained on large protein databases to generate contextualized representations of individual amino acids. To harness the power of these PLMs to predict protein-level properties, these per-residue embeddings are typically "pooled" to fixed-size vectors that are further utilized in downstream prediction networks. Common pooling strategies include Cls-Pooling and Avg-Pooling, but neither of these approaches can capture the local substructures and long-range interactions observed in proteins. RESULTS: We propose the use of attention pooling, which can naturally capture these important features of proteins. To make the expensive attention operator (quadratic in the length of the input protein) feasible in practice, we introduce bag-of-mer pooling, or BoM-Pooling, a locality-aware hierarchical pooling technique that combines windowed average pooling with attention pooling. We empirically demonstrate that both full attention pooling and BoM-Pooling outperform previous pooling strategies on three important, diverse tasks: (i) predicting the activities of two proteins as they are varied; (ii) detecting remote homologs; and (iii) predicting signaling protein interactions with peptides. Overall, our work highlights the advantages of biologically inspired pooling techniques in protein sequence modeling and is a step toward more effective adaptations of language models in biological settings. AVAILABILITY AND IMPLEMENTATION: https://github.com/Singh-Lab/bom-pooling.

Natural Language Processing↗

Protein identification in DNA databases by peptide mass fingerprinting.

Proteins can be identified using a set of peptide fragment weights produced by a specific digestion to search a protein database in which sequences have been replaced by fragment weights calculated for various cleavage methods. We present a method using multidimensional searches that greatly increases the confidence level for identification, allowing DNA sequence databases to be examined. This method provides a link between 2-dimensional gel electrophoresis protein databases and genome sequencing projects. Moreover, the increased confidence level allows unknown proteins to be matched to expressed sequence tags, potentially eliminating the need to obtain sequence information for cloning. Database searching from a mass profile is offered as a free service by an automatic server at the ETH, Zürich. For information, send an electronic message to the address cbrg/inf.ethz.ch with the line: help mass search, or help all.

Animals↗

The information encrypted in accurate peptide masses-improved protein identification and assistance in glycopeptide identification and characterization.

Analytically useful information from accurate mass data for peptides with an error of </=20 ppm is discussed. The deltamass (= mass value following the decimal point) distribution of natural peptides is extracted from a protein database. Compared with the random peptide data, the natural data show a higher average deltamass value and a smaller width of the mass distribution. This deviation can be ascribed to the non-random abundances of the standard amino acids. In particular, accurate mass data for peptides located near the edges of the natural mass distribution contain analytical information. Mass data near the edges generate very few hits in a protein database search and are therefore highly specific for protein identification. Mass signals near the low-mass edge indicate either a high probability that the peptide contains one or several cysteine sites, or that the peptide is highly acidic due to the presence of several D and/or E residues or that it is a glycopeptide. Mass data near the high-mass edge indicate a non-polar peptide with a high abundance of the non-polar amino acids leucine, isoleucine and valine. An Internet page is introduced that analyzes the deviation of a peptide mass from the average deltamass value and that supports the characterization of glycopeptides found near the low-mass edge of the mass distribution.

Amino Acid Sequence↗

Grafting of discontinuous sites: a protein modeling strategy.

A strategy for modeling continuous as well as discontinuous sites in protein structures has been developed. Central to this modeling strategy is the search algorithm of FITSITE, a program to search a given target structure for suitable combinations of backbone positions mirroring as closely as possible the geometric relationships of a source structural motif of interest. All target sites detected by FITSITE are further refined to mimic the source geometry. The sidechain rotamer library concept fails to precisely describe side chains involved in coordinative bonding (e.g, metal binding sites). Therefore an algorithm using detailed data-base bonding parameter information was applied for the side-chain construction. The FITSITE program and the subsequent processing of the program output are presented in a test case. The Rop protein, a four-helix bundle structure, served as the target protein. It was searched for candidate sites to model a variety of metal binding sites, with structures extracted from Brookhaven Protein Database entries. The preliminary protein models were investigated for structural overlaps with neighboring residues by interactive computer graphics; if required, additional changes were performed. A set of parameters for energy minimization with AMBER (including metal ions) was developed, and the completed Rop variants were energy minimized. Finally, 12 potentially metal binding Rop variants were selected for production via genetic engineering.

Algorithms↗

EXProt--a database for EXPerimentally verified Protein functions.

EXProt (database for EXPerimentally verified Protein functions) is a new non-redundant database containing protein sequences for which the function has been experimentally verified. It is a selection of 3976 entries from the Prokaryotes section of the EMBL Nucleotide Sequence Database, Release 66, and 375 entries from the Pseudomonas Community Annotation Project (PseudoCAP). The entries in EXProt all have a unique ID number and provide information about the organism, protein sequence, functional annotation, link to entry in original database, and if known, gene name and link to references in PubMed/Medline. The EXProt web page (http://www.cmbi.nl/EXProt) provides further details of the database and a link to a BLAST search (blastp & blastx) of the database. The EXProt entries are indexed in SRS (http://www.cmbi.nl/srs/) and can be searched by means of keywords. Authors can be reached by email (exprot(cmbi.kun.nl).

Amino Acid Sequence↗

An algorithm to classify amino acid sequences into protein groups of Bothrops jararacussu venomous gland.

An algorithm for automatic clustering of database protein sequences from Bothrops jararacussu venomous gland, according to sequence similarities of their domains, is described. The program was written in C and Perl languages. This algorithm compares a domain with each ORF protein sequence in the database. Each nucleotide FASTA sequence generates six ORFs. As a result, the user has a list containing all sequences found in a specific domain and a display of the sequence, domain and number of hits. The algorithm lists only the sequences that present a minimum similarity of 30 hits and the best alignment. This limit was considered appropriate. The algorithm is available in the Internet (www.compbionet.org.br/cgi-domains/homesnake) and it can quickly and accurately organizes large database into classes.

Algorithms↗

Drug Adverse Reaction Target Database (DART) : proteins related to adverse drug reactions.

An adverse drug reaction (ADR) often results from interaction of a drug or its metabolites with specific protein targets important in normal cellular function. Knowledge about these targets is both important in facilitating the study of the mechanisms of ADRs and in new drug discovery. It is also useful in the development and testing of rational drug design and safety evaluation tools. The Drug Adverse Reaction Database (DART) is intended to provide comprehensive information about adverse effect targets of drugs described in the literature. Moreover, proteins involved in adverse effect targets of chemicals not yet confirmed as ADR targets are also included as potential targets. This database gives physiological function of each target, binding drugs/agonists/antagonists/activators/inhibitors, IC(50) values of the inhibitors, corresponding adverse effects, and type of ADR induced by drug binding to a target. Cross-links to other databases are also introduced to facilitate the access of information about the sequence, 3-dimensional structure, function, and nomenclature of each target along with drug/ligand binding properties, and related literature. The database currently contains entries for 147 ADR targets and 89 potential targets. A total of 187 adverse reaction conditions, 257 drugs, and 1080 ligands known to bind to each of these targets are also currently described. Each entry can be retrieved through multiple search methods including target name, target physiological function, adverse effect, ligand name, and biological pathways. A special page is provided for contribution of new or additional information. This database can be accessed at http://xin.cz3.nus.edu.sg/group/drt/dart.asp.

Adverse Drug Reaction Reporting Systems↗

Hydrogen/deuterium exchange for higher specificity of protein identification by peptide mass fingerprinting.

Genome sequencing projects produce large amounts of information that could be translated into potential protein sequences. Such amounts of material continuously increase protein database sizes. At present, 22 times more protein sequences are available in the SWISS-PROT and TrEMBL databases than 8 years ago in SWISS-PROT. One of the methods of choice for protein identification makes use of specific endoproteolytic cleavage followed by matrix-assisted laser desorption/ionisation mass spectrometric (MALDI-MS) analysis of the digested product. Since 1993, when this technique was first demonstrated, the conditions required for a correct identification have changed dramatically. Whilst 4-5 peptides with an uncertainty of 2-3 Da were sufficient for a correct identification in 1993, 10-13 peptides with less than 60 ppm mass error are now required for human and E. coli proteins. This evolution is directly related to the continuous increase in protein database sizes, which causes an increase in the number of false positive matches in identification results. Use of an information complement deduced from the primary protein sequence, in the process of identification by peptide mass fingerprints, can help to increase confidence in the identification results. In this article, we propose the exchange of labile hydrogen atoms with deuterium atoms to provide an alternative information complement. The exchange reaction with optimised techniques has shown an average 95% of hydrogen/deuterium (H/D) exchange on tryptic peptides. This level of exchange was sufficient to single out one or more peptides from a list of potential candidate proteins due to the dependence of H/D exchange on the peptide primary structure. This technique also has clear advantages in the identification of small proteins where direct protein identification is impaired by the limited number of endoproteolytic peptides. Then, information related to primary sequence obtained with this technique could help to identify proteins with high confidence without any expensive tandem mass spectrometry instruments.

Amino Acid Sequence↗

Compression of protein sequence databases.

We have created an algorithm for compressing a PIR database to assist individual researchers and software developers who utilize sequence database information but may not have huge storage space. The resulting compact databank contains compressed PIR information and an interface written in C which allows fast direct access to the stored information without extensive decompression of corresponding files. The databank files as well as the interface C-file can be used on both PC-compatibles and UNIX-based computers without any modifications. The interface supports all standard PIR Request Network queries (i.e. gets databank SEQ number by entry; for a defined databank SEQ number, gets specified information like: name, organism(s), keyword(s), sequence, sequence features with coordinates, etc.). In contrast with PIR Request Network, our package allows us to call PIR-contained information directly from the C programs, even on a personal computer not on a network. Our PIR-derived databank, SAGITTARIUS PIR, was implemented in the form of separate file sets. Each file set contains database information of independent types (i.e. sequences, entry indexes, organisms, etc.). On a particular computer, the available configuration of the PIR information (and storage space) can be easily changed as needed by the user without affecting retrievals of other types of stored information. Due to an original alignment-based algorithm, in the compression of protein sequences themselves, our package out-performs the well-known ZIP file compressor. For PC-compatibles, a dialogue shell is available which supports all standard PIR Request Network queries plus homology searches, alignments, etc.

Algorithms↗

Proteomic analysis of native metabotropic glutamate receptor 5 protein complexes reveals novel molecular constituents.

We used a proteomic approach to identify novel proteins that may regulate metabotropic glutamate receptor 5 (mGluR5) responses by direct or indirect protein interactions. This approach does not rely on the heterologous expression of proteins and offers the advantage of identifying protein interactions in a native environment. The mGluR5 protein was immunoprecipitated from rat brain lysates; co-immunoprecipitating proteins were analyzed by mass spectrometry and identified peptides were matched to protein databases to determine the correlating parent proteins. This proteomic approach revealed the interaction of mGluR5 with known regulatory proteins, as well as novel proteins that reflect previously unidentified molecular constituents of the mGluR5-signaling complex. Immunoblot analysis confirmed the interaction of high confidence proteins, such as phosphofurin acidic cluster sorting protein 1, microtubule-associated protein 2a and dynamin 1, as mGluR5-interacting proteins. These studies show that a proteomic approach can be used to identify candidate interacting proteins. This approach may be particularly useful for neurobiology applications where distinct protein interactions within a signaling complex can dramatically alter the outcome of the response to neurotransmitter release, or the disruption of normal protein interactions can lead to severe neurological and psychiatric disorders.

Algorithms↗