Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Plant genome resources at the national center for biotechnology information.

The National Center for Biotechnology Information (NCBI) integrates data from more than 20 biological databases through a flexible search and retrieval system called Entrez. A core Entrez database, Entrez Nucleotide, includes GenBank and is tightly linked to the NCBI Taxonomy database, the Entrez Protein database, and the scientific literature in PubMed. A suite of more specialized databases for genomes, genes, gene families, gene expression, gene variation, and protein domains dovetails with the core databases to make Entrez a powerful system for genomic research. Linked to the full range of Entrez databases is the NCBI Map Viewer, which displays aligned genetic, physical, and sequence maps for eukaryotic genomes including those of many plants. A specialized plant query page allow maps from all plant genomes covered by the Map Viewer to be searched in tandem to produce a display of aligned maps from several species. PlantBLAST searches against the sequences shown in the Map Viewer allow BLAST alignments to be viewed within a genomic context. In addition, precomputed sequence similarities, such as those for proteins offered by BLAST Link, enable fluid navigation from unannotated to annotated sequences, quickening the pace of discovery. NCBI Web pages for plants, such as Plant Genome Central, complete the system by providing centralized access to NCBI's genomic resources as well as links to organism-specific Web pages beyond NCBI.

Biotechnology↗

A novel sequence similarity searching and visualization method based on overlappingly translated nucleic acids: the blastNP.

Sequence data are stored in nucleic acid and protein databases. Searching the nucleic acid databases is very specific but rather insensitive method. Searching protein databases is sensitive but not very specific procedure. It was expected that the combination of these methods might provide an optimal approach. Therefore an alternative method to TblastX has been developed, known as blastNP. Nucleic acids in database and query sequences were translated into overlapping protein-like sequences (overlappingly translated sequences or OTSs) before searching with blastP. Thus, each nucleic acid sequence is represented by a single "protein like" sequence (instead of three hypothetical proteins in different reading frames). The blastNP method is defined as a blastP that is performed on an overlappingly translated nucleic acid database using a similarly converted nucleic acid query. The specificity and sensitivity of blastNP and TblastX is very similar, however blastNP is more sensitive to detect short sequence similarities (less than 50 residues). BlastNP combines the advantages of nucleotide and protein blasts and bypasses many difficulties: (1). it is more sensitive to weak sequence similarities than blastN, (2). codon redundancy is eliminated, (3). the sensitivity to single nucleotide polymorphism, mutation and sequencing errors are reduced, (4). it is insensitive to frame shifts. This novel method was proved to find significant sequence similarities which remained hidden for other methods and is a promising tool for further understanding (and annotating) the function of many old and new sequences.

Amino Acid Sequence↗

The iProClass integrated database for protein functional analysis.

Increasingly, scientists have begun to tackle gene functions and other complex regulatory processes by studying organisms at the global scales for various levels of biological organization, ranging from genomes to metabolomes and physiomes. Meanwhile, new bioinformatics methods have been developed for inferring protein function using associative analysis of functional properties to complement the traditional sequence homology-based methods. To fully exploit the value of the high-throughput system biology data and to facilitate protein functional studies requires bioinformatics infrastructures that support both data integration and associative analysis. The iProClass database, designed to serve as a framework for data integration in a distributed networking environment, provides comprehensive descriptions of all proteins, with rich links to over 50 databases of protein family, function, pathway, interaction, modification, structure, genome, ontology, literature, and taxonomy. In particular, the database is organized with PIRSF family classification and maps to other family, function, and structure classification schemes. Coupled with the underlying taxonomic information for complete genomes, the iProClass system (http://pir.georgetown.edu/iproclass/) supports associative studies of protein family, domain, function, and structure. A case study of the phosphoglycerate mutases illustrates a systematic approach for protein family and phylogenetic analysis. Such studies may serve as a basis for further analysis of protein functional evolution, and its relationship to the co-evolution of metabolic pathways, cellular networks, and organisms.

Amino Acid Sequence↗

Secondary structure-based profiles: use of structure-conserving scoring tables in searching protein sequence databases for structural similarities.

The profile method, for detecting distantly related proteins by sequence comparison, has been extended to incorporate secondary structure information from known X-ray structures. The sequence of a known structure is aligned to sequences of other members of a given folding class. From the known structure, the secondary structure (alpha-helix, beta-strand or "other") is assigned to each position of the aligned sequences. As in the standard profile method, a position-dependent scoring table, termed a profile, is calculated from the aligned sequences. However, rather than using the standard Dayhoff mutation table in calculating the profile, we use distinct amino acid mutation tables for residues in alpha-helices, beta-strands or other secondary structures to calculate the profile. In addition, we also distinguish between internal and external residues. With this new secondary structure-based profile method, we created a profile for eight-stranded, antiparallel beta barrels of the insecticyanin folding class. It is based on the sequences of retinol-binding protein, insecticyanin and beta-lactoglobulin. Scanning the sequence database with this profile, it was possible to detect the sequence of avidin. The structure of streptavidin is known, and it appears to be distantly related to the antiparallel beta barrels. Also detected is the sequence of complement component C8, which we therefore predict to be a member of this folding class.

Amino Acid Sequence↗

REFOLD: an analytical database of protein refolding methods.

The expression and harvesting of proteins from insoluble inclusion bodies by solubilization and refolding is a technique commonly used in the production of recombinant proteins. To bring clarity to the large and widespread quantity of published protein refolding data, we have recently established the REFOLD database (http://refold.med.monash.edu.au), which is a freely available, open repository for protocols describing the refolding and purification of recombinant proteins. Refolding methods are currently published in many different formats and resources--REFOLD provides a standardized system for the structured reporting and presentation of these data. Furthermore, data in REFOLD are readily accessible using a simple search function, and the database also enables analyses which identify and highlight particular trends between suitable refolding and purification conditions and specific protein properties. This information may in turn serve to facilitate the rational design and development of new refolding protocols for novel proteins. There are approximately 200 proteins currently listed in REFOLD, and it is anticipated that with the continued contribution of data by researchers this number will grow significantly, thus strengthening the emerging trends and patterns and making this database a valuable tool for the scientific community.

Databases, Protein↗

PDZBase: a protein-protein interaction database for PDZ-domains.

SUMMARY: PDZBase is a database that aims to contain all known PDZ-domain-mediated protein-protein interactions. Currently, PDZBase contains approximately 300 such interactions, which have been manually extracted from > 200 articles. The database can be queried through both sequence motif and keyword-based searches, and the sequences of interacting proteins can be visually inspected through alignments (for the comparison of several interactions), or as residue-based diagrams including schematic secondary structure information (for individual complexes).

Database Management Systems↗

CluSTr: a database of clusters of SWISS-PROT+TrEMBL proteins.

The CluSTr (Clusters of SWISS-PROT and TrEMBL proteins) database offers an automatic classification of SWISS-PROT and TrEMBL proteins into groups of related proteins. The clustering is based on analysis of all pairwise comparisons between protein sequences. Analysis has been carried out for different levels of protein similarity, yielding a hierarchical organisation of clusters. The database provides links to InterPro, which integrates information on protein families, domains and functional sites from PROSITE, PRINTS, Pfam and ProDom. Links to the InterPro graphical interface allow users to see at a glance whether proteins from the cluster share particular functional sites. CluSTr also provides cross-references to HSSP and PDB. The database is available for querying and browsing at http://www.ebi.ac.uk/clustr.

Animals↗

Large-scale open bioinformatics data resources.

The data explosion in bioinformatics is relentless. More and more genomes are being sequenced and many new types of datasets are being generated in large-scale projects. Integration and true open access to the data are still difficult issues, although they are gradually being addressed. Notably, certain fields have good standardization and interoperability, while others lag behind. This review summarizes the latest developments in genome and sequences databases, transcriptomics data (ESTs, ORESTES, full-length cDNAs), proteomics data (protein databases, protein structures, family and domain classification) as well as loosely integrated fields, such as microarray experiments, mutation databases and databases of regulatory regions and elements. The review attempts to resist simply summarizing what data are available, and aims to provide a critical look at some of the integration and access issues associated with several of these resources.

Computational Biology↗

RPG: the Ribosomal Protein Gene database.

RPG (http://ribosome.miyazaki-med.ac.jp/) is a new database that provides detailed information about ribosomal protein (RP) genes. It contains data from humans and other organisms, including Drosophila melanogaster, Caenorhabditis elegans, Saccharo myces cerevisiae, Methanococcus jannaschii and Escherichia coli. Users can search the database by gene name and organism. Each record includes sequences (genomic, cDNA and amino acid sequences), intron/exon structures, genomic locations and information about orthologs. In addition, users can view and compare the gene structures of the above organisms and make multiple amino acid sequence alignments. RPG also provides information on small nucleolar RNAs (snoRNAs) that are encoded in the introns of RP genes.

Amino Acid Sequence↗

A novel scoring schema for peptide identification by searching protein sequence databases using tandem mass spectrometry data.

BACKGROUND: Tandem mass spectrometry (MS/MS) is a powerful tool for protein identification. Although great efforts have been made in scoring the correlation between tandem mass spectra and an amino acid sequence database, improvements could be made in three aspects, including characterization ofpeaks in spectra, adoption of effective scoring functions and access to thereliability of matching between peptides and spectra. RESULTS: A novel scoring function is presented, along with criteria to estimate the performance confidence of the function. Through learning the typesof product ions and the probability of generating them, a hypothetic spectrum was generated for each candidate peptide. Then relative entropy was introduced to measure the similarity between the hypothetic and the observed spectra. Based on the extreme value distribution (EVD) theory, a threshold was chosen to distinguish a true peptide assignment from a random one. Tests on a public MS/MS dataset demonstrated that this method performs better than the well-known SEQUEST. CONCLUSION: A reliable identification of proteins from the spectra promises a more efficient application of tandem mass spectrometry to proteomes with high complexity.

Algorithms↗

Mitoproteome: human heart mitochondrial protein sequence database.

The human mitochondrial proteome database has been developed by deriving data from a combination of public repositories and experimental and computational prediction methods. The experimental data is derived from highly purified mitochondria from human heart tissue, whereas predictions have been performed by MITOPRED, a genome-scale method for the prediction of nucleus-encoded mitochondrial proteins. Mitochondrial protein sequences from different sources have been clustered to generate a nonredundant dataset. Annotations related to the protein function, structure, disease association, pathways, and so on are collected from a number of public databases using commonly used UNIX and Perl scripts. This chapter provides a detailed description of various data sources and methods used to download, curate, parse, and generate meaningful annotations from primary as well as derived databases.

Computational Biology↗

The HSSP database of protein structure-sequence alignments and family profiles.

HSSP (http: //www.sander.embl-ebi.ac.uk/hssp/) is a derived database merging structure (3-D) and sequence (1-D) information. For each protein of known 3D structure from the Protein Data Bank (PDB), we provide a multiple sequence alignment of putative homologues and a sequence profile characteristic of the protein family, centered on the known structure. The list of homologues is the result of an iterative database search in SWISS-PROT using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed putative homologues are very likely to have the same 3D structure as the PDB protein to which they have been aligned. As a result, the database not only provides aligned sequence families, but also implies secondary and tertiary structures covering 33% of all sequences in SWISS-PROT.

Computer Communication Networks↗

The PRINTS database of protein fingerprints: a novel information resource for computational molecular biology.

PRINTS is a compendium of protein motif fingerprints derived from the OWL composite sequence database. Fingerprints are groups of motifs within sequence alignments whose conserved nature allows them to be used as signatures of family membership. Fingerprints inherently offer improved diagnostic reliability over single motif methods by virtue of the mutual context provided by motif neighbors. To date, 650 fingerprints have been constructed and stored in PRINTS, the size of which has doubled in the last 2 years. The current version, 14.0, encodes 3500 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. The database is now accessible via the UCL Bioinformatics Server on http:@ www.biochem.ucl.ac.uk/bsm/dbbrowser/. We describe here progress with the database, its compilation and interrogation software, and its Web interface.

Amino Acid Sequence↗

A protein expression database for the molecular pharmacology of cancer.

In the last six years, the Developmental Therapeutics Program (DTP) of the US National Cancer Institute (NCI) has screened over 60,000 chemical compounds and a larger number of natural product extracts for their ability to inhibit growth of 60 different cancer cell lines representing different organs of origin. Whereas inhibition of the growth of one cancer cell type gives no information on drug specificity, the relative growth inhibitory activities against 60 different cells constitute patterns that encode detailed information on mechanisms of action and resistance (as reviewed in Boyd and Paull, Drug Devel. Res. 1995, 34, 19-109 and Weinstein et al., Science 1997, 275, 343-349). In order to correlate the patterns of activity with properties of the cells, we and other laboratories are characterizing the cells with respect to a large number of factors at the DNA, mRNA, and protein levels. As part of that effort, we have developed a two-dimensional gel electrophoresis (2-DE) protein expression database covering all 60 cell types (Buolamwini et al., submitted). Here we present analyses of the correlations among protein spots (i) in terms of their patterns of expression and (ii) in terms of their apparent relationships to the pharmacology of a set of 3989 screened compounds. The correlations tend to be stronger for the latter than for the former, suggesting that the spots have more robust signatures in terms of the pharmacology than in terms of expression levels. Links to pertinent databases and tools of analysis will be updated progressively at http:@www.nci.nih.gov/intra/lmp/jnwbio.htm and http:@epnwsl.ncifcrf.gov:2345/dis3d/dtp.++ +html.

Antineoplastic Agents↗

Prolinks: a database of protein functional linkages derived from coevolution.

The advent of whole-genome sequencing has led to methods that infer protein function and linkages. We have combined four such algorithms (phylogenetic profile, Rosetta Stone, gene neighbor and gene cluster) in a single database--Prolinks--that spans 83 organisms and includes 10 million high-confidence links. The Proteome Navigator tool allows users to browse predicted linkage networks interactively, providing accompanying annotation from public databases. The Prolinks database and the Proteome Navigator tool are available for use online at http://dip.doe-mbi.ucla.edu/pronav.

ATP Synthetase Complexes↗

Nanoliter chemistry combined with mass spectrometry for peptide mapping of proteins from single mammalian cell lysates.

A nanoliter-chemistry station combined with matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectrometry was developed to characterize proteins at the attomole level. Chemical reactions including protein digestion were carried out in nanoliter or subnanoliter volumes, followed by microspot sample deposition of the digest to a MALDI-TOF mass spectrometer. Accurate mass determination of the peptides from the enzyme digest, in conjunction with protein database searching, allowed the identification of the proteins in the protein database. This method is particularly useful for handling small-volume samples such as in single-cell analysis. The high sensitivity and specificity of this method were demonstrated by peptide mapping and identifying hemoglobin variants of sickle cell disease from a single red blood cell. The approach of combining nanoliter chemistry with highly sensitive mass spectrometric analysis should find general use in characterizing proteins from biological systems where only a limited amount of material is available for interrogation.

Animals↗

Identification of microbial mixtures by capillary electrophoresis/selective tandem mass spectrometry.

In this paper, we propose a new strategy for identifying specific bacteria in bacterial mixtures by using CE-selective MS/MS of peptide marker ions associated with the bacteria of interest. We searched the CE-MS/MS spectra acquired from the proteolytic digests of pure bacterial cell extracts against protein databases. The identified peptides that match the protein associated with the corresponding species were selected as marker ions for bacterial identification. Specific peptide marker ions were obtained for each of the following three pathogens: Pseudomonas aeruginasa, Staphylococcus aureus, and Staphylococcus epidermidis. To identify a bacterial species in a sample, we performed CE-MS/MS analysis of the selected marker ions in the proteolytic digest of the cell extract and then performed protein database searches. The selected peptides that we identified correctly from Xcorr values ranking at the top of the search results allowed us to identify the corresponding bacterial species present in the sample. We have applied this method successfully to the identification of various mixtures of the three pathogens. Even minor bacterial species present at a concentration of 1% can be identified with great confidence. This method for CE-MS/MS analysis of bacteria-specific marker peptides provides excellent selectivity and high accuracy when identifying bacterial species in complex systems. In addition, we have used this approach to identify P. aeruginasa in a saliva sample spiked with E.coli and P. aeruginasa.

Amino Acid Sequence↗

StructSorter: a method for continuously updating a comprehensive protein structure alignment database.

Advances in protein crystallography and homology modeling techniques are producing vast amounts of high resolution protein structure data at ever increasing rates. As such, the ability to quickly and easily extract structural similarities is a key tool in discovering important functional relationships. We report on an approach for creating and maintaining a database of pairwise structure alignments for a comprehensive database comprising the PDB and homology models for the human and select pathogen genomes. Our approach consists of a novel, multistage method for determining pairwise structural similarity coupled with an efficient clustering protocol that approximates a full NxN assessment in a fraction of the time. Since biologists are commonly interested in recently released structures, and the homology models built from them, an automatically updating database of structural alignments has great value. Our approach yields a querying system that allows scientists to retrieve databank-wide protein structure similarities as easily as retrieving protein sequence similarities via BLAST or PSI-BLAST. Basic, noncommercial access to the database can be requested at https://tip.eidogen-sertanty.com/.

Databases, Protein↗