Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Application of a novel and fast information-theoretic method to the discovery of higher-order correlations in protein databases.

We present a fast, discrete data-mining approach to the problem of finding kappa-tuples of correlated amino acid residues in protein sequence data. When sets of sequence-distant sites display high mutual information, they may bespeak important structural or functional features. Our novel methodology overcomes the limitations of previous methods which examined only single-residue features or pairwise interactions.

AIDS Vaccines↗

Use of protein database for the computation of the dipole moments of normal and abnormal hemoglobins.

Previously, we discussed the calculation of the dipole moments of small proteins using the three-dimensional protein data-base. Our results demonstrate that the calculated dipole moments are in acceptable agreement with measured values. We, however, noted the difficulty of the calculation with larger proteins, in particular those consisting of several subunits. Hemoglobin (Hb) is a protein having a molecular weight of 64,000 that consists of four subunits, a typical case where the computation was found to be difficult. To circumvent the difficulties, we calculated the dipole moment of each subunit separately. The dipole moment of the whole protein was calculated by the vectorial summation of subunit moments. With this method, the calculated net dipole moment is in good agreement with the experimental value. Our calculation shows that the dipole moment vectors of subunits are, by and large, antiparallel in tetramers causing partial cancellation of the net dipole moment. In addition to normal HbA, the dipole moment of abnormal HbS was calculated using an approximate computational technique. Because of the loss of two negative changes as a result of the replacement of glutamic acid with valine in beta-chains, the dipole moment of HbS was found, experimentally and theoretically, to be significantly smaller than that of HbA.

Biophysical Phenomena↗

Construction of a two-dimensional gel electrophoresis protein database for the Nicotiana tabacum cv. Bright Yellow-2 cell suspension culture.

Using two-dimensional gel electrophoresis (2-DE) and electrospray-tandem mass spectrometry (ESI-MS/MS), we have started the proteome analysis of the cell line Nicotiana tabacum cv. Bright Yellow-2 (tobacco BY-2). The BY-2 cell suspension culture is widely used as a model system to study the growth and development of plant cells. We present a protocol describing the sample preparation and 2-DE, enabling us to separate and display more than 1000 proteins from this cell culture. A reference gel was generated, using immobilized pH gradient isoelectric focusing in a linear gradient from pH 3 to 10 and 12% Sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE). Although the tobacco genome is not sequenced yet, a range of protein spots from this reference map was identified by means of a semi-automated liquid chromatography-ESI-quadrupole time of flight-tandem MS (LC-ESI-QTOF-MS-MS) setup and cross-species matching. These data were integrated in a database, which can be accessed at http://tby2-www.uia.ac.be/tby2/. On the on-line reference map, the identified protein spots are hyperlinked to individual protein entries. Each protein entry contains all identification information, as well as links to relevant entries in other on-line databases. Comprehensive search functions are implemented. Especially for an unsequenced but widespread model organism like tobacco BY-2, such a reference database is a convenient source for protein information that brings protein identification within reach without the need for extensive MS. This publicly accessible database provides a solid basis for tobacco BY-2 proteomics in the future.

Databases as Topic↗

SALSA: improved protein database searching by a new algorithm for assembly of sequence fragments into gapped alignments.

MOTIVATION: Optimal sequence alignment based on the Smith-Waterman algorithm is usually too computationally demanding to be practical for searching large sequence databases. Heuristic programs like FASTA and BLAST have been developed which run much faster, but at the expense of sensitivity. RESULTS: In an effort to approximate the sensitivity of an optimal alignment algorithm, a new algorithm has been devised for the computation of a gapped alignment of two sequences. After scanning for high-scoring words and extensions of these to form fragments of similarity, the algorithm uses dynamic programming to build an accurate alignment based on the fragments initially identified. The algorithm has been implemented in a program called SALSA and the performance has been evaluated on a set of test sequences. The sensitivity was found to be close to the Smith-Waterman algorithm, while the speed was similar to FASTA (ktup = 2). AVAILABILITY: Searches can be performed from the SALSA homepage at http://dna.uio.no/salsa/ using a wide range of databases. Source code and precompiled executables are also available. CONTACT: torbjorn.rognes@labmed.uio.no

Algorithms↗

Towards establishing a protein database of Drosophila.

An improved method of high-resolution two-dimensional gel electrophoresis has been used to study the patterns of protein synthesis in wing imaginal discs of late instar larvae of Drosophila melanogaster. A total of one thousand and twenty five labelled polypeptides (787 acidic and 238 basic) have so far been separated and catalogued. For convenience, all these polypeptides have been numbered and their position fixed by its molecular weight and relative mobility. They are indicated on a reference protein map for further studies.

Animals↗

Attributable testing for abnormal prion protein, database linkage, and blood-borne vCJD risks.

CONTEXT: National prospective collection of tonsillar tissue to be tested anonymously for abnormal lymphoreticular accumulation of prion protein (PrP) was approved to begin in the UK in 2004. The UK is not, however, testing autopsy specimens attributably for abnormal PrP (PrP(SC)) so that recipients at risk after a blood transfusion from, or exposed to surgical instruments from, a deceased carrier of variant Creutzfeldt-Jakob disease (vCJD) can be followed up to quantify transmission risks. In Switzerland, surveillance for subclinical vCJD includes unconsented testing in autopsies: consented testing of tonsillar tissue is potentially attributable to interrupt human-to-human vCJD transmission or treat it. STARTING POINT: The UK announced its first case of probable blood-borne vCJD transmission in December, 2003, and first detected a case of probable blood-borne subclinical vCJD in July, 2004. To reduce the possible risk of onward transmission to other people, UK patients who had received vCJD-implicated plasma products are being contacted. They, and their general practitioner, are asked to inform anyone giving them medical, surgical, or dental treatment, and the patients must refrain from donating blood, tissues, or organs. WHERE NEXT? Prudent additional surveillance options for human PrP(SC)--particularly at autopsy or to sanction the release of quarantined operation sets pending effective decontamination--can be costed by reference to results for cattle and sheep. Some ethical or legal impediments to the UK's potentially-attributable testing for PrP(SC) may yet be rued.

Animals↗

Protein database, human retinal pigment epithelium.

The retinal pigment epithelium (RPE) is a single cell layer adjacent to the rod and cone photoreceptors that plays key roles in retinal physiology and the biochemistry of vision. RPE cells were isolated from normal adult human donor eyes, subcellular fractions were prepared, and proteins were fractionated by electrophoresis. Following in-gel proteolysis, proteins were identified by peptide sequencing using liquid chromatography tandem electrospray mass spectrometry and/or by peptide mass mapping using matrix-assisted laser desorption ionization time-of-flight mass spectrometry. Preliminary analyses have identified 278 proteins and provide a starting point for building a database of the human RPE proteome.

Chromatography, Liquid↗

Access to DNA and protein databases on the Internet.

During the past year, the number of biological databases that can be queried via Internet has dramatically increased. This increase has resulted from the introduction of networking tools, such as Gopher and WAIS, that make it easy for research workers to index databases and make them available for on-line browsing. Biocomputing in the nineties will see the advent of more client/server options for the solution of problems in bioinformatics.

Amino Acid Sequence↗

The alpha/beta fold family of proteins database and the cholinesterase gene server ESTHER.

ESTHER (for esterases, alpha/betahydrolase enzyme and relatives) is a database of sequences phylogenetically related to cholinesterases. These sequences define a homogeneous group of enzymes (carboxylesterases, lipases and hormone-sensitive lipases) sharing a similar structure of a central beta-sheet surrounded by alpha-helices. Among these proteins a wide range of functions can be found (hydrolases, adhesion molecules, hormone precursors). The purpose of ESTHER is to help comparison of structures and functions of members of the family. Since the last release, new features have been added to the server. A BLAST comparison tool allows sequence homology searches within the database sequences. New sections are available: kinetics and inhibitors of cholinesterases, fasciculin-acetylcholinesterase interaction and a gene structure review. The mutation analysis compilation has been improved with three-dimensional images. A mailing list has been created.

Amino Acid Sequence↗

EXProt: a database for proteins with an experimentally verified function.

EXProt is a non-redundant protein database containing a selection of entries from genome annotation projects and public databases, aimed at including only proteins with an experimentally verified function. In EXProt release 2.0 we have collected entries from the Pseudomonas aeruginosa community annotation project (PseudoCAP), the Escherichia coli genome and proteome database (GenProtEC) and the translated coding sequences from the Prokaryotes division of EMBL nucleotide sequence database, which are described as having an experimentally verified function. Each entry in EXProt has a unique ID number and contains information about the species, amino acid sequence, functional annotation and, in most cases, links to references in MEDLINE/PubMed and to the entry in the original database. EXProt is indexed in SRS at CMBI (http://www.cmbi.kun.nl/srs/) and can be searched with BLAST and FASTA through the EXProt web page (http://www.cmbi.kun.nl/EXProt/).

Animals↗

Bacterial identification by protein mass mapping combined with an experimentally derived protein mass database.

A protein mass mapping approach using mass spectrometry (MS) combined with an experimentally derived protein mass database is presented for rapid and effective identification of bacterial species. A prototype mass database from the protein extracts of nine bacterial species has been created by off-line high-performance liquid chromatography (HPLC) matrix-assisted laser desorption/ionization (MALDI) MS, in which the microbiological parameter of bacterial growth time is considered. A numerical method using a statistical weight factor algorithm is devised for matching the protein masses of an unknown bacterial sample against the database. The sum of these weight factors produces a corresponding summed weight factor score for each bacterial species listed in the database, and the database species producing the highest score represents the identity of the respective unknown bacterium. The applicability and reliability of this protein mass mapping approach has been tested with seven bacterial species in a single-blind study by both direct MALDI MS and HPLC electrospray ionization MS methods, and identification results with 100% accuracy are obtained. Our studies have demonstrated that the protein mass database can be rapidly established and readily adopted with relatively less dependency on experimental factors. Furthermore, it is shown that a number of proteins can be detected using a protein sample amount equivalent to an extract of less than 1000 cells, demonstrating that this protein mass mapping approach can potentially be highly sensitive for rapid bacterial identification.

Bacteria↗

The PIR-International Protein Sequence Database.

The Protein Information Resource (PIR; http://www-nbrf.georgetown. edu/pir/) supports research on molecular evolution, functional genomics, and computational biology by maintaining a comprehensive, non-redundant, well-organized and freely available protein sequence database. Since 1988 the database has been maintained collaboratively by PIR-International, an international association of data collection centers cooperating to develop this resource during a period of explosive growth in new sequence data and new computer technologies. The PIR Protein Sequence Database entries are classified into superfamilies, families and homology domains, for which sequence alignments are available. Full-scale family classification supports comparative genomics research, aids sequence annotation, assists database organization and improves database integrity. The PIR WWW server supports direct on-line sequence similarity searches, information retrieval, and knowledge discovery by providing the Protein Sequence Database and other supplementary databases. Sequence entries are extensively cross-referenced and hypertext-linked to major nucleic acid, literature, genome, structure, sequence alignment and family databases. The weekly release of the Protein Sequence Database can be accessed through the PIR Web site. The quarterly release of the database is freely available from our anonymous FTP server and is also available on CD-ROM with the accompanying ATLAS database search program.

Amino Acid Sequence↗

Constructing ontology-driven protein family databases.

MOTIVATION: Protein family databases provide a central focus for scientific communities as well as providing useful resources to aide research. However, such resources require constant curation and often become outdated and discontinued. We have developed an ontology-driven system for capturing and managing protein family data that addresses the problems of maintenance and sustainability. RESULTS: Using protein phosphatases and ABC transporters as model protein families, we constructed two protein family database resources around a central DAML+OIL ontology. Each resource contains specialist information about each protein family, providing specialized domain-specific resources based on the same template structure. The formal structure, combined with the extraction of biological data using GO terms, allows for automated update strategies. Despite the functional differences between the two protein families, the ontology model was equally applicable to both, demonstrating the generic nature of the system. AVAILABILITY: The protein phosphatase resource, PhosphaBase, is freely available on the internet (http://www.bioinf.man.ac.uk/phosphabase). The DAML+OIL ontology for the protein phosphatases and the ABC transporters is available on request from the authors. CONTACT: kwolstencroft@cs.man.ac.uk.

ATP-Binding Cassette Transporters↗

PFDB: a generic protein family database integrating the CATH domain structure database with sequence based protein family resources.

MOTIVATION: The PFDB (Protein Family Database) is a new database designed to integrate protein family-related data with relevant functional and genomic data. It currently manages biological data for three projects-the CATH protein domain database (Orengo et al., 1997; Pearl et al., 2001), the VIDA virus domains database (Albà et al., 2001) and the Gene3D database (Buchan et al., 2001). The PFDB has been designed to accommodate protein families identified by a variety of sequence based or structure based protocols and provides a generic resource for biological research by enabling mapping between different protein families and diverse biochemical and genetic data, including complete genomes. RESULTS: A characteristic feature of the PFDB is that it has a number of meta-level entities (for example aggregation, collection and inclusion) represented as base tables in the final design. The explicit representation of relationships at the meta-level has a number of advantages, including flexibility-both in terms of the range of queries that can be formulated and the ability to integrate new biological entities within the existing design. A potential drawback with this approach-poor performance caused by the number of joins across meta-level tables-is avoided by implementing the PFDB with materialized views using the mature relational database technology of Oracle 8i. The resultant database is both fast and flexible. This paper presents the principles on which the database has been designed and implemented, and describes the current status of the database and query facilities supported.

Database Management Systems↗