Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

THoR: a tool for domain discovery and curation of multiple alignments.

We describe a tool, THoR, that automatically creates and curates multiple sequence alignments representing protein domains. This exploits both PSI-BLAST and HMMER algorithms and provides an accurate and comprehensive alignment for any domain family. The entire process is designed for use via a web-browser, with simple links and cross-references to relevant information, to assist the assessment of biological significance. THoR has been benchmarked for accuracy using the SMART and pufferfish genome databases.

Algorithms↗

The microbial proteome database--an automated laboratory catalogue for monitoring protein expression in bacteria.

Laboratories devoted to high-throughput characterisation of purified proteins arrayed via two-dimensional (2-D) gel electrophoresis face an arduous task in maintaining a centralised and constantly evolving record of information relating to the characterisation of proteins and their responses following biological challenges. The Microbial Proteome Database (MPD) has been conceived as an in-house resource for complementing the plethora of genomic databases available for such organisms. The database utilises commercially available software to provide an electronic 'lab book' of information obtained daily from 2-D electrophoresis gels, image analysis packages, protein characterisation methodologies, and biological experimentation. The MPD begins from a single 2-D gel image (a 2-D 'reference map') with clickable spots that link to a 'protein catalogue' (ProtCat) with spot information including protein identity, changes in expression determined under experimental conditions, cellular location, mass, and pI. The entry for each protein then contains further links to gel images corresponding to the presence of the particular protein within different subproteomes (as defined by the pH of narrow- and wide-range immobilised pH gradients or from differential extraction methods used to determine the location of the protein within a functional cell). The database currently contains information from strains of three microbial species (Escherichia coil, Pseudomonas aeruginosa and Staphylococcus aureus) and 32 master gel images. The rapid accessibility of information obtained from microbial proteomes is an essential step towards the integrated analysis of these organisms at the gene, transcript, protein and functional levels and will aid in reducing turnaround times between sample preparation and the discovery of molecules of biological significance.

Automation↗

Analysis of Human Proteome Organization Plasma Proteome Project (HUPO PPP) reference specimens using surface enhanced laser desorption/ionization-time of flight (SELDI-TOF) mass spectrometry: multi-institution correlation of spectra and identification of biomarkers.

We report on a multicenter analysis of HUPO reference specimens using SELDI-TOF MS. Eight sites submitted data obtained from serum and plasma reference specimen analysis. Spectra from five sites passed preliminary quality assurance tests and were subjected to further analysis. Intralaboratory CVs varied from 15 to 43%. A correlation coefficient matrix generated using data from these five sites demonstrated high level of correlation, with values >0.7 on 37 of 42 spectra. More than 50 peaks were differentially present among the various sample types, as observed on three chip surfaces. Additionally, peaks at approximately 9200 and approximately 15,950 m/z were present only in select reference specimens. Chromatographic fractionation using anion-exchange, membrane cutoff, and reverse phase chromatography, was employed for protein purification of the approximately 9200 m/z peak. It was identified as the haptoglobin alpha subunit after peptide mass fingerprinting and high-resolution MS/MS analysis. The differential expression of this protein was confirmed by Western blot analysis. These pilot studies demonstrate the potential of the SELDI platform for reproducible and consistent analysis of serum/plasma across multiple sites and also for targeted biomarker discovery and protein identification. This approach could be exploited for population-based studies in all phases of the HUPO PPP.

Biomarkers↗

The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1998.

SWISS-PROT (http://www.expasy.ch/) is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, its domains structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to two additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT.

Amino Acid Sequence↗

Characterization of wheat gliadin proteins by combined two-dimensional gel electrophoresis and tandem mass spectrometry.

A proteomics-based approach was used for characterizing wheat gliadins from an Italian common wheat (Triticum aestivum) cultivar. A two-dimensional gel electrophoresis (2-DE) map of roughly 40 spots was obtained by submitting the 70% alcohol-soluble crude protein extract to isoelectric focusing on immobilized pH gradient strips across two pH gradient ranges, i.e., 3-10 or pH 6-11, and to sodium dodecyl sulfate-polyacrylamide electrophoresis in the second dimension. The chymotryptic digest of each spot was characterized by matrix-assisted laser desorption/ionization-time of flight mass spectrometry and nano electrospray ionization-tandem mass spectrometry (MS/MS) analysis, providing a "peptide map" for each digest. The measured masses were subsequently sought in databases for sequences. For accurate identification of the parent protein, it was necessary to determine de novo sequences by MS/MS experiments on the peptides. By partial mass fingerprinting, we identified protein molecules such as alpha/beta-, gamma-, omega-gliadin, and high molecular weight-glutenin. The single spots along the 2-DE map were discriminated on the basis of their amino acid sequence traits. alpha-Gliadin, the most represented wheat protein in databases, was highly conserved as the relative N-terminal sequence of the components from the 2-DE map contained only a few silent amino acid substitutions. The other closely related gliadins were identified by sequencing internal peptide chains. The results gave insight into the complex nature of gliadin heterogeneity. This approach has provided us with sound reference data for differentiating gliadins amongst wheat varieties.

Amino Acid Sequence↗

Annotating the human proteome: the Human Proteome Survey Database (HumanPSD) and an in-depth target database for G protein-coupled receptors (GPCR-PD) from Incyte Genomics.

The Proteome Division of Incyte Genomics has released new volumes to the BioKnowledge Library to add human, mouse and rat protein information to its rich collection of model organism Proteome Databases. The Human Proteome Survey Database (HumanPSD) compiles the fundamental properties of more than 25 000 characterized mammalian proteins. HumanPSD includes clear, concise and current protein descriptions (Title Lines), the protein sequence, calculated physical properties, precomputed BLAST alignments, controlled-vocabulary protein properties and Gene Ontology terms, and a list of published references. Each report also contains expression data, Pfam domain information and an associated Mouse Mutant Phenotype section describing behavioral, physiological and cellular phenotypes for over 1500 mouse mutant phenotypes. GPCR-PD contains more than 3200 Protein Reports from the three mammalian species for G protein-coupled receptors, their protein ligands, associated G-proteins and their downstream signaling proteins. In addition to the features described above, each GPCR-PD Protein Report displays annotations of experimental findings from over 10 000 publications. These databases provide important new volumes of Proteome's BioKnowledge Library (http://www.incyte.com), integrating protein information from model organisms with the human proteome.

Amino Acid Sequence↗

Human cerebrospinal fluid protein database: edition 1992.

Two-dimensional electrophoresis maps of human cerebrospinal fluid proteins are presented in the form of labeled images. 931 protein spots are identified in spinal fluid from a normal volunteer. Distinct spots that represent variants of the same protein, especially posttranslational modifications, are estimated to reduce the 931 different spots to < 200 different proteins. 248 spots of 29 protein groups have been identified and are indicated on enlargements of specific gel regions. The distribution of protein abundance, mass, charge and shape characteristics of these normal 931 spinal fluid spots are graphically profiled. Analysis of the shape parameter "vertical height: width ratio" reveals that a ratio > 3.5 correlates with glycoproteins, enabling their identification simply by image analysis. Proteins that are not present on the normal map, but appear in spinal fluid in patients with schizophrenia and Creutzfeldt-Jakob disease are illustrated on additional maps.

Cerebrospinal Fluid Proteins↗

High-performance human myocardial two-dimensional electrophoresis database: edition 1996.

The master gel of the human myocardial two-dimensional electrophoresis (2-DE) gel database contains about 3300 protein spots characterized in terms of isoelectric point (pI) and molecular mass. A high-performance technique was applied, using large gels (23 x 30 cm). Isoelectric focusing with anodic sample preparation and nonequilibrium running conditions (NEPHGE) was combined with sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) in 15% acrylamide gels in the second dimension. The range of pI extends from pH 4.5 to 9.6. Seventy proteins were identified by combinations of amino acid analysis, N-terminal and internal sequencing, immunostaining, matrix assisted laser desorption/ionization-mass spectrometry (MALDI-MS) peptide mass fingerprinting, post-source decay MALDI-MS and ladder sequencing by carboxypeptidase P. The identification of additional proteins, not found in the master gel, was achieved by immunoblotting. Unequivocal identification with high sensitivity and good yield was obtained by combining internal sequencing and MALDI-MS. In-gel digestion, the concentration and purification of peptides in a peptide collecting device, and the improved FRAGMOD program for peptide mass fingerprinting have added to the security and sensitivity of identification. The high-performance human myocardial 2-DE database was built up with proteins detected by the TOPSPOT program. Spots within six sections of the whole pattern are clickable. Protein description includes detailed information about identification, characterization, and links to the related SWISS-PROT, other 2-DE databases and Medline entries. The database is constructed in accordance with four of the rules for a federated database.

Computer Communication Networks↗

Environment-specific amino acid substitution tables: tertiary templates and prediction of protein folds.

The local environment of an amino acid in a folded protein determines the acceptability of mutations at that position. In order to characterize and quantify these structural constraints, we have made a comparative analysis of families of homologous proteins. Residues in each structure are classified according to amino acid type, secondary structure, accessibility of the side chain, and existence of hydrogen bonds from the side chains. Analysis of the pattern of observed substitutions as a function of local environment shows that there are distinct patterns, especially for buried polar residues. The substitution data tables are available on diskette with Protein Science. Given the fold of a protein, one is able to predict sequences compatible with the fold (profiles or templates) and potentially to discriminate between a correctly folded and misfolded protein. Conversely, analysis of residue variation across a family of aligned sequences in terms of substitution profiles can allow prediction of secondary structure or tertiary environment.

Amino Acid Sequence↗

Protein kinase C isoforms from Giardia duodenalis: identification and functional characterization of a beta-like molecule during encystment.

Protein kinase C (PKC) is a family of serine/threonine kinases that regulate many different cellular processes such as cell growth and differentiation in eukaryotic cells. Using specific polyclonal antibodies raised against mammalian PKC isoforms, it was demonstrated here for the first time that Giardia duodenalis expresses several PKC isoforms (beta, delta, epsilon, theta and zeta). All PKC isoforms detected showed changes in their expression pattern during encystment induction. In addition, selective PKC inhibitors blocked the encystment in a dose-dependent manner, suggesting that PKC isozymes may play important roles during this differentiation process. We have characterized here the only conventional-type PKC member found so far in Giardia, which showed an increased expression and changes in its intracellular localization pattern during cyst formation. The purified protein obtained by chromatography on DEAE-cellulose followed by size-exclusion chromatography, displayed in vitro kinase activity using histone HI-IIIS as substrate, which was dependent on cofactors required by conventional PKCs, i.e., phospholipids and calcium. An open reading frame in the Giardia Genome Database that encodes a homolog of PKCbeta catalytic domain was identified and cloned. The expressed recombinant protein was also recognized by a mammalian anti-PKCbeta antibody and was referred as giardial PKCbeta on the basis of all these experimental evidence.

Animals↗

Exploring conformational space with a simple lattice model for protein structure.

We present a low resolution lattice model for which we can exhaustively generate all possible compact backbone conformations for small proteins. Using simple structural and energetic criteria, for a variety of proteins, we can select for lattice structures that have significant similarities with their known native structures. Our energetic parameters are based on pairwise amino acid contact frequencies in a database of experimentally determined structures. A key step in our method involves the threading of a sequence onto every lattice model, such that a locally optimal pattern of tertiary interactions is formed. We evaluate our results against statistics collected for structures covering all of conformational space, and against statistics collected for permuted sequences. Despite the low resolution of the model, our low energy structures contain many native features. These results indicate that the overall pattern of hydrophobicity of a sequence significantly constrains the range of folds that sequence is likely to adopt.

Algorithms↗

MEROPS: the peptidase database.

Peptidases (proteolytic enzymes) and their natural, protein inhibitors are of great relevance to biology, medicine and biotechnology. The MEROPS database (http://merops.sanger.ac.uk) aims to fulfil the need for an integrated source of information about these proteins. The organizational principle of the database is a hierarchical classification in which homologous sets of proteins of interest are grouped into families and the homologous families are grouped in clans. The most important addition to the database has been newly written, concise text annotations for each peptidase family. Other forms of information recently added include highlighting of active site residues (or the replacements that render some homologues inactive) in the sequence displays and BlastP search results, dynamically generated alignments and trees at the peptidase or inhibitor level, and a curated list of human and mouse homologues that have been experimentally characterized as active. A new way to display information at taxonomic levels higher than species has been devised. In the Literature pages, references have been flagged to draw attention to particularly 'hot' topics.

Animals↗

Recognition of transmembrane alpha-helical segments with environmental profiles.

A method for assessing the environmental properties of membrane-spanning alpha-helical peptides in proteins has been proposed. The algorithm employs a set of environmental preference parameters derived for amino acid residues based on the analysis of the 3-D structures of membrane domains in bacteriorhodopsin and photoreaction centers Rhodopseudomonas viridis and Rhodobacter sphaeroides. The resulting 3-D-1-D scores for transmembrane segments are significantly different from those derived for alpha-helices in globular proteins. The parameters obtained have been used to construct environmental profiles for membrane alpha-helices in bacteriorhodopsin and photoreaction centers. The profiles successfully recognize their own sequences in several specially designed large databases. The method has been applied to several membrane proteins with unknown spatial structures. Most of their membrane-spanning peptides were efficiently recognized by the profiles. The predicted environment of the residues in the membrane segments fits the experimental data well. The approach is independent of any homology data and can be employed to delineate the membrane segments of a protein with environmental characteristics close to those of bacteriorhodopsin and photoreaction centers. The alignment of these segments with the reference profiles provides a considerable amount of data about their lipid and protein exposure.

Algorithms↗

Establishing mathematical laws of genomic variation.

As the biological arm of the Rasch community, genomic measurement is concerned with asserting and testing hypotheses regarding the quantitative status of genomic variables, including alleles, genotypes, gene expression levels, and phenotypes, as well as DNA, RNA, and protein sequence information. The defining goal of this scientific paradigm, in contrast to the sample-dependent model-fitting and deterministic hypothesis testing of classical statistical genetics, is the identification, validation, and maintenance of a common unit of genomic measurement that maintains its magnitude and meaning, within an allowable range of error, regardless of the laboratory technology used to generate outcomes or the particular group of individuals or organisms under investigation. Such an invariant metric, the basis of a standard genometric scale and associated system of genomic metrology, can be identified, validated, and maintained through 1) routine implementation of the Rasch family of measurement models to construct sample- and scale-free measures from different types of genomic data and 2) cross-calibration of genomic measurement instruments between and among researchers, laboratories, universities, corporations, and databases. This manuscript provides an introductory overview of the guiding principles of fundamental measurement theory and the work of Rasch, connects these concepts to well-known tenets of population genetics, and highlights the potential benefits, both theoretical and applied, associated with achieving objectivity in genomic measurement.

Animals↗

FANTASIA suite: a reproducible and configurable framework for embedding-based functional annotation of proteins.

Embedding-based annotation transfer is increasingly used for protein function inference due to protein language models capture sequence, structural, and functional signals that may extend beyond conventional pairwise similarity. However, systematic application of these approaches requires control over model choice, reference composition, lookup parameters, evidence traceability, and output formats. We developed the FANTASIA suite, a configurable framework for embedding-based functional annotation of proteins. The suite combines a database-backed implementation for reproducible and extensible analyses with a portable flat-file implementation for rapid local annotation and pipeline integration. Using non-model and model-organism proteomes, we show that larger neighbourhood sizes remain practical for proteome-scale analyses and that taxonomy and sequence-identity filtering support leakage-aware benchmarking. We also compare the supported models with baseline methods through external CAFA5 evaluation and provide practical guidance based on empirical evidence variables. FANTASIA provides a controlled, scalable, and reproducible framework for extending functional annotation across the rapidly expanding diversity of sequenced organisms.

Software↗

The SWISS-PROT protein sequence data bank and its new supplement TREMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc), a minimal level of redundancy and a high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to seven additional databases; a variety of new documentation files; the creation of TREMBL, and unannotated supplement to SWISS-PROT. This supplement consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except CDS already included in SWISS-PROT.

Amino Acid Sequence↗

SuperHapten: a comprehensive database for small immunogenic compounds.

The immune system protects organisms from foreign proteins, peptide epitopes and a multitude of chemical compounds. Among these, haptens are small molecules, eliciting an immune response when conjugated with carrier molecules. Known haptens are xenobiotics or natural compounds, which can induce a number of autoimmune diseases like contact dermatitis or asthma. Furthermore, haptens are utilized in the development of biosensors, immunomodulators and new vaccines. Although hapten-induced allergies account for 6-10% of all adverse drug effects, the understanding of the correlation between structural and haptenic properties is rather fragmentary. We have developed a manually curated hapten database, SuperHapten, integrating information from literature and web resources. The current version of the database compiles 2D/3D structures, physicochemical properties and references for about 7500 haptens and 25,000 synonyms. The commercial availability is documented for about 6300 haptens and 450 related antibodies, enabling experimental approaches on cross-reactivity. The haptens are classified regarding their origin: pesticides, herbicides, insecticides, drugs, natural compounds, etc. Queries allow identification of haptens and associated antibodies according to functional class, carrier protein, chemical scaffold, composition or structural similarity. SuperHapten is available online at http://bioinformatics.charite.de/superhapten.

Cross Reactions↗

A two-dimensional electrophoresis proteomic reference map and systematic identification of 1367 proteins from a cell suspension culture of the model legume Medicago truncatula.

The proteome of a Medicago truncatula cell suspension culture was analyzed using two-dimensional electrophoresis and nanoscale HPLC coupled to a tandem Q-TOF mass spectrometer (QSTAR Pulsar i) to yield an extensive protein reference map. Coomassie Brilliant Blue R-250 was used to visualize more than 1661 proteins, which were excised, subjected to in-gel trypsin digestion, and analyzed using nanoscale HPLC/MS/MS. The resulting spectral data were queried against a custom legume protein database using the MASCOT search engine. A total of 1367 of the 1661 proteins were identified with high rigor, yielding an identification success rate of 83% and 907 unique protein accession numbers. Functional annotation of the M. truncatula suspension cell proteins revealed a complete tricarboxylic acid cycle, a nearly complete glycolytic pathway, a significant portion of the ubiquitin pathway with the associated proteolytic and regulatory complexes, and many enzymes involved in secondary metabolism such as flavonoid/isoflavonoid, chalcone, and lignin biosynthesis. Proteins were also identified from most other functional classes including primary metabolism, energy production, disease/defense, protein destination/storage, protein synthesis, transcription, cell growth/division, and signal transduction. This work represents the most extensive proteomic description of M. truncatula suspension cells to date and provides a reference map for future comparative proteomic and functional genomic studies of the response of these cells to biotic and abiotic stress.

Amino Acid Sequence↗