Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Development of a liquid chromatography-tandem mass spectrometry method using capillary liquid chromatography and nanoelectrospray ionization-quadrupole time-of-flight hybrid mass spectrometer for the detection of milk allergens.

Liquid chromatography-tandem mass spectrometry (LC-MS/MS) analysis of the tryptic digest of a cleaned-up food matrix extract was used for the detection of milk allergens. The emphasis of this study was on casein, which is the most abundant milk protein and is also considered the most allergenic. A sample cleanup method was developed using an ion exchange column and centriprep device. Cookies spiked with milk powder from 0 to 1250 ppm were extracted, cleaned up, and either digested directly by trypsin or further cleaned up by gel electrophoresis before digestion. The peptide mixture was analyzed on a capillary LC-quadrupole time-of-flight system. Two marker peptides from alphaS1-casein were identified and used for prescreening. The MS/MS data from the mass spectrometry system were processed with Masslynx v4.0 and submitted for database search using either ProteinLynx Global Server or Mascot for protein identification. The LC-MS/MS method, using casein enzyme-linked immunosorbent assay as a reference, was tested on the cookie matrix and was extended to other sample matrices. There were good agreements between the two. This LC-MS/MS method provides a valuable confirmatory method for the presence of casein. It also allows the simultaneous detection of other milk allergens.

Allergens↗

The EBI SRS server--recent developments.

MOTIVATION: The current data explosion is intractable without advanced data management systems. The numerous data sets become really useful when they are interconnected under a uniform interface--representing the domain knowledge. The SRS has become an integration system for both data retrieval and applications for data analysis. It provides capabilities to search multiple databases by shared attributes and to query across databases fast and efficiently. RESULTS: Here we present recent developments at the EBI SRS server (http://srs.ebi.ac.uk). The EBI SRS server contains today more than 130 biological databases and integrates more than 10 applications. It is a central resource for molecular biology data as well as a reference server for the latest developments in data integration. One of the latest additions to the EBI SRS server is the InterPro database-Integrated Resource of Protein Domains and Functional Sites. Distributed in XML format it became a turning point in low level XML-SRS integration. We present InterProScan as an example of data analysis applications, describe some advanced features of SRS6, and introduce the SRSQuickSearch JavaScript interfaces to SRS.

Computational Biology↗

Distance-scaled, finite ideal-gas reference state improves structure-derived potentials of mean force for structure selection and stability prediction.

The distance-dependent structure-derived potentials developed so far all employed a reference state that can be characterized as a residue (atom)-averaged state. Here, we establish a new reference state called the distance-scaled, finite ideal-gas reference (DFIRE) state. The reference state is used to construct a residue-specific all-atom potential of mean force from a database of 1011 nonhomologous (less than 30% homology) protein structures with resolution less than 2 A. The new all-atom potential recognizes more native proteins from 32 multiple decoy sets, and raises an average Z-score by 1.4 units more than two previously developed, residue-specific, all-atom knowledge-based potentials. When only backbone and C(beta) atoms are used in scoring, the performance of the DFIRE-based potential, although is worse than that of the all-atom version, is comparable to those of the previously developed potentials on the all-atom level. In addition, the DFIRE-based all-atom potential provides the most accurate prediction of the stabilities of 895 mutants among three knowledge-based all-atom potentials. Comparison with several physical-based potentials is made.

Computational Biology↗

At what scale should microarray data be analyzed?

INTRODUCTION: The hybridization intensities derived from microarray experiments, for example Affymetrix's MAS5 signals, are very often transformed in one way or another before statistical models are fitted. The motivation for performing transformation is usually to satisfy the model assumptions such as normality and homogeneity in variance. Generally speaking, two types of strategies are often applied to microarray data depending on the analysis need: correlation analysis where all the gene intensities on the array are considered simultaneously, and gene-by-gene ANOVA where each gene is analyzed individually. AIM: We investigate the distributional properties of the Affymetrix GeneChip signal data under the two scenarios, focusing on the impact of analyzing the data at an inappropriate scale. METHODS: The Box-Cox type of transformation is first investigated for the strategy of pooling genes. The commonly used log-transformation is particularly applied for comparison purposes. For the scenario where analysis is on a gene-by-gene basis, the model assumptions such as normality are explored. The impact of using a wrong scale is illustrated by log-transformation and quartic-root transformation. RESULTS: When all the genes on the array are considered together, the dependent relationship between the expression and its variation level can be satisfactorily removed by Box-Cox transformation. When genes are analyzed individually, the distributional properties of the intensities are shown to be gene dependent. Derivation and simulation show that some loss of power is incurred when a wrong scale is used, but due to the robustness of the t-test, the loss is acceptable when the fold-change is not very large.

Algorithms↗

A systematic review of protein and energy supplementation for hip fracture aftercare in older people.

OBJECTIVES: To evaluate whether protein and energy supplementation influences recovery after hip fracture. DESIGN: Systematic review of randomised and quasi-randomised trials in people aged 65 y and over. DATA SOURCES: We searched seven electronic databases from 1966 to April 2002, four journals and reference lists of relevant articles. We contacted trial investigators and experts for details of other trials. MAIN OUTCOME MEASURES: Mortality, complications and unfavourable outcome (mortality or survivors with complications) were the primary outcomes. We also sought data on length of hospital stay, functional status after hip fracture, quality of life and compliance with supplementation. RESULTS: In total, 12 randomised trials involving 898 participants were included. Nine trials evaluated protein and energy supplementation (five oral and four nasogastric feeding), and a further three trials tested oral protein supplementation. Potential biases resulting from inadequate allocation concealment and lack of assessor blinding and intention-to-treat analysis, as well as the limited outcome data, mean that the results must be interpreted with caution. Pooled data from eight of the nine trials evaluating protein and energy supplements showed no evidence for an effect on mortality (relative risk 0.92, 95% CI 0.56-1.50). Limited data from only three trials showed that oral protein and energy supplements may reduce unfavourable outcome (relative risk 0.52, 95% CI 0.32-0.84). CONCLUSION: Based on limited evidence, oral protein and energy supplementation after hip fracture may reduce unfavourable outcome. Further evidence from good-quality randomised trials is required to inform clinical practice.

Aged↗

High-Resolution Chromosome-Level Genome Assembly and Annotation of Triplophysa stewarti, an Endemic Plateau Loach from the Qinghai-Tibet Plateau.

The bottom-dwelling fish Triplophysa stewarti, endemic to the Qinghai-Tibet Plateau, is a valuable model for studying high-altitude adaptation in aquatic ecosystems. However, the lack of a high-quality reference genome has hindered comparative genomic and evolutionary studies within this genus. Here, we present a chromosome-level genome assembly for T. stewarti, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding. The 697.9 Mb assembly is highly continuous (scaffold N50 of 253.58 Mb) and encompasses 25 chromosomes, representing 92.65% of the genome. BUSCO analysis indicated a 98.4% completeness, supporting the high quality of the assembly. We annotated 28,009 protein-coding genes, with 97.04% being functionally assigned across multiple databases (NR, UniProt, KEGG, GO, Pfam and InterPro). Additionally, repetitive elements constituted 42.47% of the genome, and we identified 52,709 non-coding RNAs. This high-quality reference genome provides a fundamental resource for exploring the adaptive evolution, population structure, and conservation genetics of T. stewarti and related species on the Qinghai-Tibet Plateau.

Animals↗

A two-dimensional proteome map of Shigella flexneri.

Shigella flexneri is a Gram-negative facultatively intracellular pathogen responsible for bacillary dysentery in humans. In this study, extracellular proteins from the culture medium and whole cell proteins in cellular extracts of S. flexneri 2a strain 2457T were examined by two-dimensional (2-D) gel electrophoresis using immobilized pH gradient (IPG) technology. Proteins were identified by matrix-assisted laser desorption/ionization-mass spectrometry (MALDI-MS) in combination with Mascot search program. In total, among the 488 proteins spots processed, 388 proteins were identified. The identified proteins represented 169 genes. By comparing results of Mascot search against databases of Escherichia coli and genomes of S. flexneri 2a, one S. flexneri-specific protein was identified and one possible gap was found in 2457T genome sequences. Although this proteome map is still incomplete, it is already a useful reference for future studies involving pathogenicity, vaccine development, design of novel antibacterial drugs, etc. Proteome maps and a table of all identified proteins are available on the internet at www.proteomics.com.cn.

Bacterial Proteins↗

The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domains structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include format and content enhancements, cross-references to additional databases, new documentation files and improvements to TrEMBL, a computer-annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDSs) in the EMBL Nucleotide Sequence Database, except the CDSs already included in SWISS-PROT. We also describe the Human Proteomics Initiative (HPI), a major project to annotate all known human sequences according to the quality standards of SWISS-PROT. SWISS-PROT is available at: http://www.expasy.ch/sprot/ and http://www.ebi.ac.uk/swissprot/

Animals↗

Similarity metrics for ligands reflecting the similarity of the target proteins.

In this study we evaluate how far the scope of similarity searching can be extended to identify not only ligands binding to the same target as the reference ligand(s) but also ligands of other homologous targets without initially known ligands. This "homology-based similarity searching" requires molecular representations reflecting the ability of a molecule to interact with target proteins. The Similog keys, which are introduced here as a new molecular representation, were designed to fulfill such requirements. They are based only on the molecular constitution and are counts of atom triplets. Each triplet is characterized by the graph distances and the types of its atoms. The atom-typing scheme classifies each atom by its function as H-bond donor or acceptor and by its electronegativity and bulkiness. In this study the Similog keys are investigated in retrospective in silico screening experiments and compared with other conformation independent molecular representations. Studied were molecules of the MDDR database for which the activity data was augmented by standardized target classification information from public protein classification databases. The MDDR molecule set was split randomly into two halves. The first half formed the candidate set. Ligands of four targets (dopamine D2 receptor, opioid delta-receptor, factor Xa serine protease, and progesterone receptor) were taken from the second half to form the respective reference sets. Different similarity calculation methods are used to rank the molecules of the candidate set by their similarity to each of the four reference sets. The accumulated counts of molecules binding to the reference target and groups of targets with decreasing homology to it were examined as a function of the similarity rank for each reference set and similarity method. In summary, similarity searching based on Unity 2D-fingerprints or Similog keys are found to be equally effective in the identification of molecules binding to the same target as the reference set. However, the application of the Similog keys is more effective in comparison with the other investigated methods in the identification of ligands binding to any target belonging to the same family as the reference target. We attribute this superiority to the fact that the Similog keys provide a generalization of the chemical elements and that the keys are counted instead of merely noting their presence or absence in a binary form. The second most effective molecular representation are the occurrence counts of the public ISIS key fragments, which like the Similog method, incorporates key counting as well as a generalization of the chemical elements. The results obtained suggest that ligands for a new target can be identified by the following three-step procedure: 1. Select at least one target with known ligands which is homologous to the new target. 2. Combine the known ligands of the selected target(s) to a reference set. 3. Search candidate ligands for the new targets by their similarity to the reference set using the Similog method. This clearly enlarges the scope of similarity searching from the classical application for a single target to the identification of candidate ligands for whole target families and is expected to be of key utility for further systematic chemogenomics exploration of previously well explored target families.

Algorithms↗

Utility of electrophoretically derived protein mass estimates as additional constraints in proteome analysis of human serum based on MS/MS analysis.

The proteome of a HUPO human serum reference sample was analyzed using multidimensional separation techniques at both the protein and the peptide levels. To eliminate false-positive identifications from the search results, we employed a data filtering method using molecular weight (MW) correlations derived from denaturing 1-DE. First, the six most abundant serum proteins were removed from the sample using immunoaffinity chromatography. 1-DE was then used to fractionate the remaining serum proteins according to the MW. Gel bands were isolated and in-gel digested with trypsin, and the resulting peptides were analyzed by 2-D LC/ESI-MS/MS. A SEQUEST search using the MS/MS results identified 494 proteins. Of these, 202 were excluded formally using protein data filtering as they were single-assignment proteins and their theoretical and electrophoretically-derived MWs did not correlate at high confidence. To evaluate this method, the results were compared with those of 1-D LC/MALDI-TOF/TOF and HUPO Plasma Proteome Project analyses. Our data filtering approach proved valuable in analysis of complex, large-scale proteomes such as human serum.

Amino Acid Sequence↗

Enhanced detection of human immunodeficiency virus type 1-specific T-cell responses to highly variable regions by using peptides based on autologous virus sequences.

The antigenic diversity of human immunodeficiency virus type 1 (HIV-1) represents a significant challenge for vaccine design as well as the comprehensive assessment of HIV-1-specific immune responses in infected persons. In this study we assessed the impact of antigen variability on the characterization of HIV-1-specific T-cell responses by using an HIV-1 database to determine the sequence variability at each position in all expressed HIV-1 proteins and a comprehensive data set of CD8 T-cell responses to a reference strain of HIV-1 in infected persons. Gamma interferon Elispot analysis of HIV-1 clade B-specific T-cell responses to 504 overlapping peptides spanning the entire expressed HIV-1 genome derived from 57 infected subjects demonstrated that the average amino acid variability within a peptide (entropy) was inversely correlated to the measured frequency at which the peptide was recognized (P = 6 x 10(-7)). Subsequent studies in six persons to assess T-cell responses against p24 Gag, Tat, and Vpr peptides based on autologous virus sequences demonstrated that 29% (12 of 42) of targeted peptides were only detected with peptides representing the autologous virus strain compared to the HIV-1 clade B consensus sequence. The use of autologous peptides also allowed the detection of significantly stronger HIV-1-specific T-cell responses in the more variable regulatory and accessory HIV-1 proteins Tat and Vpr (P = 0.007). Taken together, these data indicate that accurate assessment of T-cell responses directed against the more variable regulatory and accessory HIV-1 proteins requires reagents based on autologous virus sequences. They also demonstrate that CD8 T-cell responses to the variable HIV-1 proteins are more common than previously reported.

Amino Acid Sequence↗

THoR: a tool for domain discovery and curation of multiple alignments.

We describe a tool, THoR, that automatically creates and curates multiple sequence alignments representing protein domains. This exploits both PSI-BLAST and HMMER algorithms and provides an accurate and comprehensive alignment for any domain family. The entire process is designed for use via a web-browser, with simple links and cross-references to relevant information, to assist the assessment of biological significance. THoR has been benchmarked for accuracy using the SMART and pufferfish genome databases.

Algorithms↗

The microbial proteome database--an automated laboratory catalogue for monitoring protein expression in bacteria.

Laboratories devoted to high-throughput characterisation of purified proteins arrayed via two-dimensional (2-D) gel electrophoresis face an arduous task in maintaining a centralised and constantly evolving record of information relating to the characterisation of proteins and their responses following biological challenges. The Microbial Proteome Database (MPD) has been conceived as an in-house resource for complementing the plethora of genomic databases available for such organisms. The database utilises commercially available software to provide an electronic 'lab book' of information obtained daily from 2-D electrophoresis gels, image analysis packages, protein characterisation methodologies, and biological experimentation. The MPD begins from a single 2-D gel image (a 2-D 'reference map') with clickable spots that link to a 'protein catalogue' (ProtCat) with spot information including protein identity, changes in expression determined under experimental conditions, cellular location, mass, and pI. The entry for each protein then contains further links to gel images corresponding to the presence of the particular protein within different subproteomes (as defined by the pH of narrow- and wide-range immobilised pH gradients or from differential extraction methods used to determine the location of the protein within a functional cell). The database currently contains information from strains of three microbial species (Escherichia coil, Pseudomonas aeruginosa and Staphylococcus aureus) and 32 master gel images. The rapid accessibility of information obtained from microbial proteomes is an essential step towards the integrated analysis of these organisms at the gene, transcript, protein and functional levels and will aid in reducing turnaround times between sample preparation and the discovery of molecules of biological significance.

Automation↗

Analysis of Human Proteome Organization Plasma Proteome Project (HUPO PPP) reference specimens using surface enhanced laser desorption/ionization-time of flight (SELDI-TOF) mass spectrometry: multi-institution correlation of spectra and identification of biomarkers.

We report on a multicenter analysis of HUPO reference specimens using SELDI-TOF MS. Eight sites submitted data obtained from serum and plasma reference specimen analysis. Spectra from five sites passed preliminary quality assurance tests and were subjected to further analysis. Intralaboratory CVs varied from 15 to 43%. A correlation coefficient matrix generated using data from these five sites demonstrated high level of correlation, with values >0.7 on 37 of 42 spectra. More than 50 peaks were differentially present among the various sample types, as observed on three chip surfaces. Additionally, peaks at approximately 9200 and approximately 15,950 m/z were present only in select reference specimens. Chromatographic fractionation using anion-exchange, membrane cutoff, and reverse phase chromatography, was employed for protein purification of the approximately 9200 m/z peak. It was identified as the haptoglobin alpha subunit after peptide mass fingerprinting and high-resolution MS/MS analysis. The differential expression of this protein was confirmed by Western blot analysis. These pilot studies demonstrate the potential of the SELDI platform for reproducible and consistent analysis of serum/plasma across multiple sites and also for targeted biomarker discovery and protein identification. This approach could be exploited for population-based studies in all phases of the HUPO PPP.

Biomarkers↗

The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1998.

SWISS-PROT (http://www.expasy.ch/) is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, its domains structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to two additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT.

Amino Acid Sequence↗

Characterization of wheat gliadin proteins by combined two-dimensional gel electrophoresis and tandem mass spectrometry.

A proteomics-based approach was used for characterizing wheat gliadins from an Italian common wheat (Triticum aestivum) cultivar. A two-dimensional gel electrophoresis (2-DE) map of roughly 40 spots was obtained by submitting the 70% alcohol-soluble crude protein extract to isoelectric focusing on immobilized pH gradient strips across two pH gradient ranges, i.e., 3-10 or pH 6-11, and to sodium dodecyl sulfate-polyacrylamide electrophoresis in the second dimension. The chymotryptic digest of each spot was characterized by matrix-assisted laser desorption/ionization-time of flight mass spectrometry and nano electrospray ionization-tandem mass spectrometry (MS/MS) analysis, providing a "peptide map" for each digest. The measured masses were subsequently sought in databases for sequences. For accurate identification of the parent protein, it was necessary to determine de novo sequences by MS/MS experiments on the peptides. By partial mass fingerprinting, we identified protein molecules such as alpha/beta-, gamma-, omega-gliadin, and high molecular weight-glutenin. The single spots along the 2-DE map were discriminated on the basis of their amino acid sequence traits. alpha-Gliadin, the most represented wheat protein in databases, was highly conserved as the relative N-terminal sequence of the components from the 2-DE map contained only a few silent amino acid substitutions. The other closely related gliadins were identified by sequencing internal peptide chains. The results gave insight into the complex nature of gliadin heterogeneity. This approach has provided us with sound reference data for differentiating gliadins amongst wheat varieties.

Amino Acid Sequence↗

Annotating the human proteome: the Human Proteome Survey Database (HumanPSD) and an in-depth target database for G protein-coupled receptors (GPCR-PD) from Incyte Genomics.

The Proteome Division of Incyte Genomics has released new volumes to the BioKnowledge Library to add human, mouse and rat protein information to its rich collection of model organism Proteome Databases. The Human Proteome Survey Database (HumanPSD) compiles the fundamental properties of more than 25 000 characterized mammalian proteins. HumanPSD includes clear, concise and current protein descriptions (Title Lines), the protein sequence, calculated physical properties, precomputed BLAST alignments, controlled-vocabulary protein properties and Gene Ontology terms, and a list of published references. Each report also contains expression data, Pfam domain information and an associated Mouse Mutant Phenotype section describing behavioral, physiological and cellular phenotypes for over 1500 mouse mutant phenotypes. GPCR-PD contains more than 3200 Protein Reports from the three mammalian species for G protein-coupled receptors, their protein ligands, associated G-proteins and their downstream signaling proteins. In addition to the features described above, each GPCR-PD Protein Report displays annotations of experimental findings from over 10 000 publications. These databases provide important new volumes of Proteome's BioKnowledge Library (http://www.incyte.com), integrating protein information from model organisms with the human proteome.

Amino Acid Sequence↗