Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Two-dimensional gel protein database of Saccharomyces cerevisiae.

With the systematic sequencing of the yeast genome, yeast biology has entered a new era where novel challenges have to be faced. One challenge is the identification of the function of the several hundred novel genes discovered by genome sequencing. Another is to understand how all yeast genes act in concert to ensure and maintain cell organization. Two-dimensional (2-D) gel electrophoresis is the technique of choice to take up these challenges because it provides the opportunity of obtaining an overall view of genome expression. In prospect of these studies we have undertaken the construction of a yeast 2-D gel protein database that contains information on polypeptides of the yeast protein map. In this paper we report the information presently contained in this database. The reported information includes the identification of 250 protein spots and the characterization of polypeptides corresponding to N-terminal acetylated proteins, mitochondrial proteins, glucose-repressed proteins, heat shock induced proteins and proteins encoded by intron-containing genes. In all, 600 spots are annotated. These data can be accessed on the Yeast Protein Map server through the World Wide Web network.

Computer Communication Networks↗

Amino acid analysis and protein database compositional search as a rapid and inexpensive method to identify proteins.

The identification of protein samples in minute quantities of protein samples, e.g., from two-dimensional polyacrylamide gel electrophoresis analysis, is an everyday problem in biology laboratories. Here we show that computer-assisted amino acid analysis can fulfill this task. Amino acid analysis data can be used to compare the amino acid composition of an unknown protein with protein compositions in a database (compositional search). Routine amino acid analysis data can, despite a certain margin of error, be used to identify a protein. Compared to protein sequencing, amino analysis is much cheaper, faster, and allows higher sample throughput. Thus, the method may replace protein sequencing as a first attempt in identification, provided a homolog can be found in the database.

Amino Acid Sequence↗

Microorganism identification by mass spectrometry and protein database searches.

A method for rapid identification of microorganisms is presented, which exploits the wealth of information contained in prokaryotic genome and protein sequence databases. The method is based on determining the masses of a set of ions by MALDI TOF mass spectrometry of intact or treated cells. Subsequent correlation of each ion in the set to a protein, along with the organismic source of the protein, is performed by searching an Internet-accessible protein database. Convoluting the lists for all ions and ranking the organisms corresponding to matched ions results in the identification of the microorganism. The method has been successfully demonstrated on B. subtilis and E. coli, two organisms with completely sequenced genomes. The method has been also tested for identification from mass spectra of mixtures of microorganisms, from spectra of an organism at different growth stages, and from spectra originating at other laboratories. Experimental factors such as MALDI matrix preparation, spectral reproducibility, contaminants, mass range, and measurement accuracy on the database search procedure are addressed too. The proposed method has several advantages over other MS methods for microorganism identification.

Bacillus subtilis↗

Sequence patterns produced by incomplete enzymatic digestion or one-step Edman degradation of peptide mixtures as probes for protein database searches.

Mass spectrometric peptide mapping of proteins isolated by polyacrylamide gel electrophoresis is a rapid method for identifying proteins in sequence databases. A majority of tryptic peptide maps were found to contain pairs of peptide ion peaks separated by the molecular weight of the lysyl or arginyl residue. These peaks originate from amino acid sequence patterns such as Lys-Lys where trypsin has cleaved C-terminals to either one of the lysines. The peptide mass and the pattern define an N- or C-terminal sequence tag. Searching sequence databases by such a sequence tag results in only a moderate number of matches and significantly reduces the number of database matches when used in combination with a peptide mass map. Two N- or C-terminal sequence tags alone unambiguously identify a protein in most cases. The technique discussed here is simple, does not require additional measurements, and increases the percentage of protein samples that can be identified by their mass maps alone. N-Terminal peptide sequence tags for database searching can also be generated by manual one-step Edman degradation of the unseparated peptide mixture.

Databases, Factual↗

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein↗

Microsequences of 145 proteins recorded in the two-dimensional gel protein database of normal human epidermal keratinocytes.

Microsequencing of proteins recovered from two-dimensional (2-D) gels is being used systematically to identify proteins in the master human keratinocyte 2-D gel database. To date, about 250 protein spots recorded in human 2-D gel databases have been microsequenced and, of these, 145 are recorded in the keratinocyte database under the entry partial amino acid sequence. Coomassie Brilliant Blue-stained protein spots cut from several (up to 40) dry gels were concentrated by elution-concentration gel electrophoresis, electroblotted onto PVDF membranes and digested in situ with trypsin. Eluting peptides were separated by reversed-phase HPLC, collected individually and sequenced. Computer search using the FASTA and TFASTA programs from Genetics Computer Group indicated that 110 of the microsequenced polypeptides shared significant similarity with proteins contained in the PIR, Mipsx or GenEMBL databases. Only 35 polypeptides corresponded to hitherto unknown proteins. Peptide sequences of all 145 proteins are listed together with their coordinates (apparent molecular weight and pI) in the keratinocyte database.

Amino Acid Sequence↗

Proteomic analysis of the human colon carcinoma cell line (LIM 1215): development of a membrane protein database.

The proteomic definition of plasma membrane proteins is an important initial step in searching for novel tumor marker proteins expressed during the different stages of cancer progression. However, due to the charge heterogeneity and poor solubility of membrane-associated proteins this subsection of the cell's proteome is often refractory to two-dimensional electrophoresis (2-DE), the current paradigm technology for studying protein expression profiles. Here, we describe a non-2-DE method for identifying membrane proteins. Proteins from an enriched membrane preparation of the human colorectal carcinoma cell line LIM1215 were initially fractionated by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE, 4-20%). The unstained gel was cut into 16 x 3 mm slices, and peptide mixtures resulting from in-gel tryptic digestion of each slice were individually subjected to capillary-column reversed phase-high performance liquid chromatography (RP-HPLC) coupled with electrospray ionization-ion trap-mass spectrometry (ESI-IT-MS). Interrogation of genomic databases with the resulting collision-induced dissociation (CID) generated peptide ion fragment data was used to identify the proteins in each gel slice. Over 284 proteins (including 92 membrane proteins) were identified, including many integral membrane proteins not previously identified by 2-DE, many proteins seen at the genomic level only, as well as several proteins identified by expressed sequence tags (ESTs) only. Additionally, a number of peptides, identified by de novo MS sequence analysis, have not been described in the databases. Further, a "targeted" ion approach was used to unambiguously identify known low-abundance plasma membrane proteins, using the membrane-associated A33 antigen, a gastrointestinal-specific epithelial cell protein, as an example. Following localization of the A33 antigen in the gel by immunoblotting, ions corresponding to the theoretical A33 antigen tryptic peptide masses were selected using an "inclusion" mass list for automated sequence analysis. Six peptides corresponding to the A33 antigen, present at levels well below those accessible using the standard automated "nontargeted" approach, were identified. The membrane protein database may be accessed via the World Wide Web (WWW) at http://www.ludwig. edu.au/jpsl/jpslhome.html.

Colonic Neoplasms↗

Error-tolerant protein database searching using peptide product-ion spectra.

A method for matching proteins in databases with unknown samples using data derived from peptide product ion masses is described. The power of the method is due to the speed of analysis and to the ability to locate proteins in a database even when the experimental data show anomalies due to derivatization or post-translational modification.

Amino Acid Sequence↗

Quantitative determination of TCR cross-reactivity using peptide libraries and protein databases.

A single T cell clone can be activated by many different peptides in the context of a particular HLA molecule. To quantify the number of peptides that can be recognized by a CD4(+) T cell clone, we screened a one-bead-one-peptide synthetic peptide library and a protein database for peptides that stimulate an HLA-DR3-restricted, human glutamic acid decarboxylase (GAD65)-reactive CD4(+) T cell clone. Both the library screening and the database analysis indicated that this T cell clone is able to recognize approximately 10(6) 11-mer peptides at low nanomolar concentration. Furthermore, we determined that the frequency of cross-reactivity increased only 1.5-3 times when the peptide concentration increased 10 times, in the range of 0.01 - 1 microM. These data imply that there is a considerable potential for T cell cross-reactivity and are useful for studies on the role of molecular mimicry in the etiology of T cell-mediated disease.

Amino Acid Sequence↗

Dilated cardiomyopathy-associated proteins and their presentation in a WWW-accessible two-dimensional gel protein database.

High resolution two-dimensional electrophoresis (2-DE) and computer-assisted image analysis were used to screen 13 patients suffering from dilated cardiomyopathy (DCM) versus 15 control patients for quantitative and qualitative differences in their myocardial protein expression. Right atrial tissue samples were obtained from end-stage failing explanted hearts and control hearts. Fifty-two spots differed significantly in average intensity between the DCM and the control groups. Myosin light chain 2, ventricular (MLC2) and heat shock protein HSP 27 were identified by protein microsequencing and gel map comparison with other databases. These proteins were found to be characteristic protein markers for DCM in the right atrium. In DCM patients, the spot intensity (protein abundance) of MLC2 is increased to 336% and HSP 27 is decreased to 59%, compared to the control group. The HEART-2DPAGE, a World Wide Web-accessible 2-DE database, was used and extended for the presentation of these disease-associated proteins. Retrievable via Internet we present a list of disease-associated proteins, their altered level of expression in DCM, and their position on a right atrial protein pattern. The accession number to protein sequence databases confers a connection to databases like SWISS-PROT to obtain a detailed functional and structural description of disease-associated proteins. New DCM-associated proteins are detected and their presentation in a 2-DE gel protein database is described.

Cardiomyopathy, Dilated↗

Nearest neighbor classification in 3D protein databases.

In molecular databases, structural classification is a basic task that can be successfully approached by nearest neighbor methods. The underlying similarity models consider spatial properties such as shape and extension as well as thematic attributes. We introduce 3D shape histograms as an intuitive and powerful approach to model similarity for solid objects such as molecules. Errors of measurement, sampling, and numerical rounding may result in small displacements of atomic coordinates. These effects may be handled by using quadratic form distance functions. An efficient processing of similarity queries based on quadratic forms is supported by a filter-refinement architecture. Experiments on our 3D protein database demonstrate the high classification accuracy of more than 90% and the good performance of the technique.

Algorithms↗

Analysis of chaperonin-containing TCP-1 subunits in the human keratinocyte two-dimensional protein database: further characterisation of antibodies to individual subunits.

The chaperonin-containing TCP-1 (CCT), found in the eukaryotic cytosol, is currently the focus of extensive research. CCT consists of at least eight different subunit types encoded by independent but related genes, and a set of antibodies that recognise individual subunits has proved useful in the characterisation and functional analysis of CCT. These antibodies were used to identify subunits of CCT in the human keratinocyte two-dimensional protein database. Accurate values for the pI and molecular mass of human CCT subunits were determined from the database, and biological data was obtained regarding changes in subunit levels in response to extracellular agents and growth conditions. The second part of the study describes the characterisation of seven monoclonal antibodies raised against mouse TCP-1, also known as CCT alpha, using a combination of epitope mapping and immunoblot analysis of protein extracts from different species and tissue types. Some antibodies were not monospecific for TCP-1, and a number of epitope-related proteins were identified.

Animals↗

SCOP: a structural classification of proteins database for the investigation of sequences and structures.

To facilitate understanding of, and access to, the information available for protein structures, we have constructed the Structural Classification of Proteins (scop) database. This database provides a detailed and comprehensive description of the structural and evolutionary relationships of the proteins of known structure. It also provides for each entry links to co-ordinates, images of the structure, interactive viewers, sequence data and literature references. Two search facilities are available. The homology search permits users to enter a sequence and obtain a list of any structures to which it has significant levels of sequence similarity. The key word search finds, for a word entered by the user, matches from both the text of the scop database and the headers of Brookhaven Protein Databank structure files. The database is freely accessible on World Wide Web (WWW) with an entry point to URL http: parallel scop.mrc-lmb.cam.ac.uk magnitude of scop.

Amino Acid Sequence↗

Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.

The BLAST programs are widely used tools for searching protein and DNA databases for sequence similarities. For protein comparisons, a variety of definitional, algorithmic and statistical refinements described here permits the execution time of the BLAST programs to be decreased substantially while enhancing their sensitivity to weak similarities. A new criterion for triggering the extension of word hits, combined with a new heuristic for generating gapped alignments, yields a gapped BLAST program that runs at approximately three times the speed of the original. In addition, a method is introduced for automatically combining statistically significant alignments produced by BLAST into a position-specific score matrix, and searching the database using this matrix. The resulting Position-Specific Iterated BLAST (PSI-BLAST) program runs at approximately the same speed per iteration as gapped BLAST, but in many cases is much more sensitive to weak but biologically relevant sequence similarities. PSI-BLAST is used to uncover several new and interesting members of the BRCT superfamily.

Algorithms↗

The rat liver epithelial (RLE) cell protein database.

Computer databases of rat liver epithelial (RLE) cellular polypeptides have been established using high resolution two-dimensional gel electrophoresis and computer-assisted analysis. Databases have been constructed utilizing both [35S]methionine- and [32P]orthophosphate-labeled as well as silver-stained polypeptides from normal RLE cells. The RLE database, which contains both qualitative and quantitative annotations, includes experiments with normal, chemically and oncogene transformed as well as spontaneously transformed cell lines. A total of 2537 [35S]methionine-labeled polypeptides from whole cell lysates (1920 acidic and 617 basic, separated in the first dimension using isoelectric focusing and nonequilibrium pH gradient electrophoresis, respectively) were analyzed and databases constructed using the Elsie 5 gel analysis system. To increase the "viewing window" and hence the usefulness of the RLE database, subcellular fractionation of whole cell preparations was performed and high resolution two-dimensional maps of the individual subcellular components were constructed. Databases representing 1229 cytosolic, 1539 acidic and 674 basic nuclear, 1746 membrane-associated, 415 mitochondrial, 773 in vitro translated and 350 phosphoproteins were established from these maps. The RLE databases contain the Elsie 5 identification number, protein name (if known), molecular weight and pI information, quantitative and spot shape data, and specific information regarding transformation-sensitive, growth-related (exponentially proliferating versus confluent) cell populations as well as those polypeptides modulated by specific growth factors. The RLE databases represent initial efforts toward the establishment of comprehensive databases of rat liver proteins and serve as a vital resource for on-going as well as future studies regarding the regulation of growth and differentiation as well as transformation of RLE cells.

Animals↗

A comparative two-dimensional gel protein database of the intact and regenerating newt limbs.

In this paper we describe a two-dimensional gel database of the regenerating newt limb. Protein synthesis was compared in the intact limb, in the 1-week regenerating limb, representing the dedifferentiation stage, and in the 2-week regenerating limb, representing the formation of the blastema. This comparative database provided data on differential expression of about 800 proteins during the process of limb regeneration. In addition, a map has been generated for these proteins for future guidance in characterizing further new, unknown proteins. The overall expression patterns of the proteins indicated that the dedifferentiation stage was marked by down-regulation of most proteins, while the blastema formation was marked by the appearance of many new proteins. The potential use of such a database in isolating factors involved during limb regeneration is discussed.

Animals↗