Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

PSSM-based prediction of DNA binding sites in proteins.

BACKGROUND: Detection of DNA-binding sites in proteins is of enormous interest for technologies targeting gene regulation and manipulation. We have previously shown that a residue and its sequence neighbor information can be used to predict DNA-binding candidates in a protein sequence. This sequence-based prediction method is applicable even if no sequence homology with a previously known DNA-binding protein is observed. Here we implement a neural network based algorithm to utilize evolutionary information of amino acid sequences in terms of their position specific scoring matrices (PSSMs) for a better prediction of DNA-binding sites. RESULTS: An average of sensitivity and specificity using PSSMs is up to 8.7% better than the prediction with sequence information only. Much smaller data sets could be used to generate PSSM with minimal loss of prediction accuracy. CONCLUSION: One problem in using PSSM-derived prediction is obtaining lengthy and time-consuming alignments against large sequence databases. In order to speed up the process of generating PSSMs, we tried to use different reference data sets (sequence space) against which a target protein is scanned for PSI-BLAST iterations. We find that a very small set of proteins can actually be used as such a reference data without losing much of the prediction value. This makes the process of generating PSSMs very rapid and even amenable to be used at a genome level. A web server has been developed to provide these predictions of DNA-binding sites for any new protein from its amino acid sequence. AVAILABILITY: Online predictions based on this method are available at http://www.netasa.org/dbs-pssm/

Algorithms↗

euHCVdb: the European hepatitis C virus database.

The hepatitis C virus (HCV) genome shows remarkable sequence variability, leading to the classification of at least six major genotypes, numerous subtypes and a myriad of quasispecies within a given host. A database allowing researchers to investigate the genetic and structural variability of all available HCV sequences is an essential tool for studies on the molecular virology and pathogenesis of hepatitis C as well as drug design and vaccine development. We describe here the European Hepatitis C Virus Database (euHCVdb, http://euhcvdb.ibcp.fr), a collection of computer-annotated sequences based on reference genomes. The annotations include genome mapping of sequences, use of recommended nomenclature, subtyping as well as three-dimensional (3D) molecular models of proteins. A WWW interface has been developed to facilitate database searches and the export of data for sequence and structure analyses. As part of an international collaborative effort with the US and Japanese databases, the European HCV Database (euHCVdb) is mainly dedicated to HCV protein sequences, 3D structures and functional analyses.

Databases, Protein↗

Web-based two-dimensional database of Saccharomyces cerevisiae proteins using immobilized pH gradients from pH 6 to pH 12 and matrix-assisted laser desorption/ionization-time of flight mass spectrometry.

An image based two-dimensional (2-D) reference map of very alkaline yeast cell proteins was established by using immobilized pH gradients (IPG) up to pH 12 (IPG 6-12, IPG 9-12 and IPG 10-12) for 2-D electrophoresis and by using matrix-assisted laser desorption/ionization-time of flight mass spectrometry peptide mass fingerprinting for spot identification. Up to now 106 proteins with theoretical isoelectric points up to pH 11.15 and molecular mass between 7.5 and 115 kDa were localized and identified. Additionally, due to the improved resolution of steady-state isoelectric focussing with IPGs, even low copy number proteins with codon bias below 0.02 were detected and identified.

Databases as Topic↗

Database of mutations within the adenovirus 5 E1A oncogene.

The Ad5 E1A database is a listing of mutations affecting the early region 1A (E1A) proteins of human adenovirus type 5. The database contains the name of the mutation, the nucleic acid sequence changes, the resulting alterations in amino acid sequence and reference. Additional notes and references are provided on the effect of each mutation on E1A function. The database is contained within the Adenovirus 5 E1A page on the World Wide Web at: http://www.geocities.com/CapeCanaveral/Hangar /2541/

Adenovirus E1A Proteins↗

ESTAnnotator: A tool for high throughput EST annotation.

In high throughput sequence analysis, it is often necessary to combine the results of contemporary bioinformatics tools, because no individual tool alone computes all the requested information. ESTAnnotator is a tool for the high throughput annotation of expressed sequence tags (ESTs) by automatically running a collection of bioinformatics applications. In the first step, a quality check is performed and repeats, vector parts and low quality sequences are masked. Then successive steps of database searching and EST clustering are performed. Already known transcripts present within mRNA and genomic DNA reference databases are identified. Subsequently, tools for the clustering of anonymous ESTs, and for further database searches at the protein level, are applied. Finally, the outputs of each individual tool are gathered and the relevant results presented in a descriptive summary. ESTAnnotator was already successfully applied for the systematic identification and characterisation of novel human genes involved in cartilage/bone formation, growth, differentiation and homeostasis. ESTAnnotator is available at http://genome.dkfz-heidelberg.de, contact: genome@dkfz.de.

Cartilage↗

Mutational data integration in gene-oriented files of the Hermansky-Pudlak Syndrome database.

Hermansky-Pudlak Syndrome (HPS) is a genetically heterogeneous disorder characterized by oculocutaneous albinism and prolonged bleeding due to abnormal vesicle trafficking to lysosomes and related organelles such as melanosomes and platelet dense granules. This HPS database (HPSD; http://liweilab.genetics.ac.cn/HPSD/) provides integrated, annotatory, and curative data that is distributed in a variety of public databases or predicted by bioinformatics servers for the recently cloned human and mouse HPS genes, as well as for the genes responsible for HPSrelated syndromes, such as ChediakHigashi Syndrome (CHS), Griscelli syndrome (GS), oculocutaneous albinism (OCA), Usher syndrome type 1B (USH1B), and ocular albinism (OA). The HPSD is designed by using a unique GeneOriented File (GOF) format. Seven blocks (genomic, transcript, protein, function, mutation, phenotype, and reference) are carefully annotated in each userfriendly GOF entry. The HPSD emphasizes paired human and mouse GOF entries. The genes included in this database (currently 58 in total) are arbitrarily divided into four categories: 1) Human and Mouse HPS, 2) Mouse HPS Only, 3) Putative Mouse or Human HPS, and 4) HPS Related Syndromes. All the mutations in these genes are integrated in the GOFs. We expect that these very informative and peerreviewed GOFs will be shortcuts to utilize the webbased information for the emerging interdisciplinary studies of HPS.

Animals↗

High-throughput screening of historic collections: observations on file size, biological targets, and file diversity.

At Pfizer Central Research, high-throughput screening has been an important source of new leads for drug discovery for a decade. Our experience with over 150 high-throughput screens can address questions about necessary file size, how well particular biological targets fare (with particular reference to protein-protein interactions), and what file diversity means in practice.

Databases, Factual↗

GlycoSuiteDB: a curated relational database of glycoprotein glycan structures and their biological sources. 2003 update.

GlycoSuiteDB is an annotated and curated relational database of glycan structures reported in the literature. It contains information on the glycan type, core type, linkages and anomeric configurations, mass, composition and the analytical methods used by the researchers to determine the glycan structure. Native and recombinant sources are detailed, including species, tissue and/or cell type, cell line, strain, life stage, disease, and if known the protein to which the glycan structures are attached. There are links to SWISS-PROT/TrEMBL and PubMed where applicable. Recent developments include the implementation of searching by 2D structure and substructure, disease and reference. The database is updated twice a year, and now contains over 7650 entries. Access to GlycoSuiteDB is available at http://www.glycosuite.com.

Animals↗

TCDB: the Transporter Classification Database for membrane transport protein analyses and information.

The Transporter Classification Database (TCDB) is a web accessible, curated, relational database containing sequence, classification, structural, functional and evolutionary information about transport systems from a variety of living organisms. TCDB is a curated repository for factual information compiled from >10,000 references, encompassing approximately 3000 representative transporters and putative transporters, classified into >400 families. The transporter classification (TC) system is an International Union of Biochemistry and Molecular Biology approved system of nomenclature for transport protein classification. TCDB is freely accessible at http://www.tcdb.org. The web interface provides several different methods for accessing the data, including step-by-step access to hierarchical classification, direct search by sequence or TC number and full-text searching. The functional ontology that underlies the database structure facilitates powerful query searches that yield valuable data in a quick and easy way. The TCDB website also offers several tools specifically designed for analyzing the unique characteristics of transport proteins. TCDB not only provides curated information and a tool for classifying newly identified membrane proteins, but also serves as a genome transporter-annotation tool.

Databases, Protein↗

Human plasma proteome analysis by reversed sequence database search and molecular weight correlation based on a bacterial proteome analysis.

In shotgun proteomics, proteins can be fractionated by 1-D gel electrophoresis and digested into peptides, followed by liquid chromatography to separate the peptide mixture. Mass spectrometry generates hundreds of thousands of tandem mass spectra from these fractions, and proteins are identified by database searching. However, the search scores are usually not sufficient to distinguish the correct peptides. In this study, we propose a confident protein identification method for high-throughput analysis of human proteome. To build a filtering protocol in database search, we chose Pseudomonas putida KT2440 as a reference because this bacterial proteome contains fewer modifications and is simpler than the human proteome. First, the P. putida KT2440 proteome was filtered by reversed sequence database search and correlated by the molecular weight in 1-D-gel band positions. The characterization protocol was then applied to determine the criteria for clustering of the human plasma proteome into three different groups. This protein filtering method, based on bacterial proteome data analysis, represents a rapid way to generate higher confidence protein list of the human proteome, which includes some of heavily modified and cleaved proteins.

Blood Proteins↗

DIALIGN-T: an improved algorithm for segment-based multiple sequence alignment.

BACKGROUND: We present a complete re-implementation of the segment-based approach to multiple protein alignment that contains a number of improvements compared to the previous version 2.2 of DIALIGN. This previous version is superior to Needleman-Wunsch-based multi-alignment programs on locally related sequence sets. However, it is often outperformed by these methods on data sets with global but weak similarity at the primary-sequence level. RESULTS: In the present paper, we discuss strengths and weaknesses of DIALIGN in view of the underlying objective function. Based on these results, we propose several heuristics to improve the segment-based alignment approach. For pairwise alignment, we implemented a fragment-chaining algorithm that favours chains of low-scoring local alignments over isolated high-scoring fragments. For multiple alignment, we use an improved greedy procedure that is less sensitive to spurious local sequence similarities. To evaluate our method on globally related protein families, we used the well-known database BAliBASE. For benchmarking tests on locally related sequences, we created a new reference database called IRMBASE which consists of simulated conserved motifs implanted into non-related random sequences. CONCLUSION: On BAliBASE, our new program performs significantly better than the previous version of DIALIGN and is comparable to the standard global aligner CLUSTAL W, though it is outperformed by some newly developed programs that focus on global alignment. On the locally related test sets in IRMBASE, our method outperforms all other programs that we evaluated.

Algorithms↗

Analysis of peptide MS/MS spectra from large-scale proteomics experiments using spectrum libraries.

A widespread proteomics procedure for characterizing a complex mixture of proteins combines tandem mass spectrometry and database search software to yield mass spectra with identified peptide sequences. The same peptides are often detected in multiple experiments, and once they have been identified, the respective spectra can be used for future identifications. We present a method for collecting previously identified tandem mass spectra into a reference library that is used to identify new spectra. Query spectra are compared to references in the library to find the ones that are most similar. A dot product metric is used to measure the degree of similarity. With our largest library, the search of a query set finds 91% of the spectrum identifications and 93.7% of the protein identifications that could be made with a SEQUEST database search. A second experiment demonstrates that queries acquired on an LCQ ion trap mass spectrometer can be identified with a library of references acquired on an LTQ ion trap mass spectrometer. The dot product similarity score provides good separation of correct and incorrect identifications.

Amino Acid Sequence↗

The human keratinocyte two-dimensional gel protein database (update 1992): towards an integrated approach to the study of cell proliferation, differentiation and skin diseases.

The master two-dimensional gel database of human keratinocytes currently lists 2980 cellular proteins (2098 isoelectric focusing, IEF; and 882 nonequilibrium pH gradient electrophoresis, NEPHGE) many of which correspond to posttranslational modifications. About 20% of all recorded proteins have been identified (protein name, organelle components, etc.) and they are listed in alphabetical order together with their M(r), pI, cellular localization and credit to the investigator(s) that aided in the identification. Also, we have listed 145 microsequenced proteins that are recorded in this database. As an aid in localizing the polypeptides we have included blow-ups of the master images (IEF, NEPHGE) displaying all the protein numbers. In the long run, the master keratinocyte database is expected to link protein and DNA sequencing and mapping information (Human Genome Program) and to provide an integrated picture of the expression levels and properties of the thousands of proteins that orchestrate various keratinocyte functions both in health and disease.

Cell Differentiation↗

Global transcript analysis of rice leaf and seed using SAGE technology.

We have compiled two comprehensive gene expression profiles from mature leaf and immature seed tissue of rice (Oryza sativa ssp. japonica cultivar Nipponbare) using Serial Analysis of Gene Expression (SAGE) technology. Analysis revealed a total of 50 519 SAGE tags, corresponding to 15 131 unique transcripts. Of these, the large majority (approximately 70%) occur only once in both libraries. Unexpectedly, the most abundant transcript (approximately 3% of the total) in the leaf library was derived from a type 3 metallothionein gene. The overall frequency profiles of the abundant tag species from both tissues differ greatly and reveal seed tissue as exhibiting a non-typical pattern of gene expression characterized by an over abundance of a small number of transcripts coding for storage proteins. A high proportion ( approximately 80%) of the abundant tags (> or = 9) matched entries in our reference rice EST database, with many fewer matches for low abundant tags. Singleton transcripts that are common to both tissues were collated to generate a summary of low abundant transcripts that are expressed constitutively in rice tissues. Finally and most surprisingly, a significant number of tags were found to code for antisense transcripts, a finding that suggests a novel mechanism of gene regulation, and may have implications for the use of antisense constructs in transgenic technology.

Journal Article↗

Soy isoflavone analysis: quality control and a new internal standard.

Development of a database of the soy isoflavone content of foods requires accurate and precise evaluation of different food matrixes. To evaluate accuracy, we estimated recoveries of both internal and external standards in 5 different soyfoods weekly. Standards were evaluated daily for system quality assurance. To evaluate sample precision, we analyzed soybeans and soymilk bimonthly for within-day precision and over 4 d for day-to-day precision. CVs should be < or = 8%. We validated our methods for single and multiple recovery concentrations by using our new internal standard, 2,4,4'-trihydroxydeoxybenzoin, and the external standards daidzein, genistein, and genistin. Concentrations of 12 isoflavone isomers, 3 aglycones (daidzein, genistein, and glycitein), and 9 glucosides (daidzin, genistin, glycitin, acetyldaidzin, acetylgenistin, acetylglycitin, malonyldaidzin, malonylgenistin, and malonylglycitin) were measured in a variety of soybeans and soyfoods. The extraction methods used depended on soyfood type. The HPLC conditions for soy isoflavone analysis were improved, leading to good separation with a short analysis time (60 min/sample). A data bank of concentration and distribution of isoflavones in different soybean products was assembled. A wide range of isoflavone concentrations, from < 50 microg/g to > 20,000 microg/g, was found in different soy products. The glucoside forms are almost twice the molecular weight of the aglycones; reported isoflavone concentrations should be normalized to the aglycone mass (or an isoflavonoid equivalent) rather than a simple sum of all isomers.

Chromatography, High Pressure Liquid↗