Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Study of human laryngeal muscle protein using two-dimensional electrophoresis and mass spectrometry.

Proteomic analysis was performed to construct a protein database for human laryngeal muscle. Thyroarytenoid (TA) muscle specimens were obtained from six post mortem cases within 24 h of death. Isoelectric focusing was performed by using immobilized pH gradient strips followed by 12% sodium dodecyl sulfate-polyacrylamide gel electrophoresis. Silver stained gels were then analyzed using PDQuest software to locate, quantify and match spots. Proteins were identified by matrix-assisted laser desorption/ionization-mass spectrometry on the basis of peptide mass fingerprinting following in-gel digestion with trypsin. Comparison of protein distribution between broad and narrow pH range gels demonstrated that 75% of all protein spots from human TA muscle were located within the pH range 5-8, and between mass 15-120 kDa. Based on peptide mass fingerprinting, 75 proteins were identified and classified into six functional groups. These include membrane proteins (8.5%), cytoskeletal and myofibrillar proteins (14.6%), energy production proteins (28%), proteins associated with stress responses (8.5%), and protein associated with transcription regulation (10.9%). Approximately one-third (29%) were categorized as "other proteins". This data provides an initial reference map for comparative studies of protein expression in human and laryngeal muscle. Further development of this database will provide a valuable resource for molecular analysis of normal and pathologic conditions affecting human striated muscle.

Databases as Topic↗

Gene-protein database of Escherichia coli K-12: edition 3.

The first two editions of the E. coli Gene-Protein Index were published to provide identifications of protein spots resolved by two-dimensional gel electrophoresis as the products of known genes. This third edition has been expanded to include information about genes and proteins gained directly from two-dimensional gel analysis--including information about protein spots not yet characterized genetically or biochemically--and is therefore more properly called a cellular protein database. An alpha-numeric designation has been uniquely assigned to each of the 616 polypeptide spots in the current database. To this, information is linked about the polypeptide's identification (protein name, gene name, Enzyme Commission--EC number), location on reference gels (x-y coordinates), genetics (Genbank code, DNA sequence reference), biochemistry (molecular weight, isoelectric point), and physiology (steady state level of the protein as a function of media and temperature, membership in various regulons and stimulons).

Bacterial Proteins↗

KDBI: Kinetic Data of Bio-molecular Interactions database.

Understanding of cellular processes and underlying molecular events requires knowledge about different aspects of molecular interactions, networks of molecules and pathways in addition to the sequence, structure and function of individual molecules involved. Databases of interacting molecules, pathways and related chemical reaction equations have been developed. The kinetic data for these interactions, which is important for mechanistic investigation, quantitative study and simulation of cellular processes and events, is not provided in the existing databases. We introduce a new database of Kinetic Data of Bio-molecular Interactions (KDBI) aimed at providing experimentally determined kinetic data of protein-protein, protein-RNA, protein-DNA, protein-ligand, RNA-ligand, DNA-ligand binding or reaction events described in the literature. KDBI contains information about binding or reaction event, participating molecules (name, synonyms, molecular formula, classification, SWISS-PROT AC or CAS number), binding or reaction equation, kinetic data and related references. The kinetic data is in terms of one or a combination of the following quantities as given in the literature of a particular event: association/dissociation or on/off rate constant, first/second/third/. order rate constant, equilibrium rate constant, catalytic rate constant, equilibrium association/dissociation constant, inhibition constant and binding affinity constant. Each entry can be retrieved through protein or nucleic acid or ligand name, SWISS-PROT AC number, ligand CAS number and full-text search of a binding or reaction event. KDBI currently contains 8273 entries of biomolecular binding or reaction events involving 1380 proteins, 143 nucleic acids and 1395 small molecules. Hyperlinks are provided for accessing references in Medline and available 3D structures in PDB and NDB. This database can be accessed at http://xin.cz3.nus.edu.sg/group/kdbi/kdbi.asp.

DNA↗

Protein restriction for diabetic renal disease.

OBJECTIVES: To determine whether protein restriction slows or prevents progression of diabetic nephropathy towards renal failure. SEARCH STRATEGY: Computerised databases MEDLINE (1976-1996) and EMBASE (1974-1996) were searched using keywords diabetes mellitus, diabetic nephropathy, dietary proteins, diet, protein restricted and uremia. Recent issues of selected journals (Diabetic Medicine, Diabetologia, Diabetes Care, Kidney International, Nephrology Dialysis and Transplantation) were handsearched for papers not yet in the computerised databases. Reference lists of papers were also checked. SELECTION CRITERIA: This review was not limited to randomised controlled trials. All trials involving people with insulin-dependent diabetes following a lower protein diet for at least 4 months were considered since the straight line nature of progression as reflected by GFR means that patients can act as their own controls in a before and after comparison. DATA COLLECTION AND ANALYSIS: Data were extracted for length of follow up, level of protein restriction, renal function and dietary compliance. No studies of the impact of protein restriction on outcomes such as the need for dialysis or transplantation were found. The trials reported only the effect on short-term indicators such as creatinine clearance. MAIN RESULTS: Overall a protein restricted diet (0.3-0. 8g/kg) does appear to slow the progression of diabetic nephropathy towards renal failure. REVIEWER'S CONCLUSIONS: The results show that reducing protein intake appears to slow progression to renal failure, but some questions remain unanswered. The first is what level of protein restriction we should be used? The trials aimed for a daily intake of between 0.3 to 0.8g/kg of protein. The second concerns compliance in routine care - what level would be acceptable to patients? The third concerns long term outcomes -the present trials use proxy indicators such as creatinine clearance rather than outcomes such as time to dialysis or prevention of ESRF. All trials were carried out in subjects with insulin-dependent diabetes. It remains to be seen if a lower protein intake would slow the progression of nephropathy affecting the non-insulin dependent diabetic population.

Diabetic Nephropathies↗

Proteome reference maps of Medicago truncatula embryogenic cell cultures generated from single protoplasts.

Using a combination of two-dimensional gel electrophoresis (2-DE) protein mapping and mass spectrometry (MS) analysis, we have established proteome reference maps of Medicago truncatula embryogenic tissue culture cells. The cultures were generated from single protoplasts, which provided a relatively homogeneous cell population. We used these to analyze protein expression at the globular stages of somatic embryogenesis, which is the earliest morphogenetic embryonic stage. Over 3000 proteins could reproducibly be resolved over a pI range of 4-11. Three hundred and twelve protein spots were extracted from colloidal Coomassie Blue-stained 2-DE gels and analyzed by matrix-assisted laser desorption/ionization-time of flight MS analysis and tandem MS sequencing. This enabled the identification of 169 protein spots representing 128 unique gene products using a publicly available expressed sequence tag database and the MASCOT search engine. These reference maps will be valuable for the investigation of the molecular events which occur during somatic embryogenesis in M. truncatula. The proteome reference maps and supplementary materials will be available and updated for public access at http://semele.anu.edu.au/.

Amino Acid Sequence↗

[Construction of standard human transcript dataset based on RefSeq and human genome sequence database].

The NCBI Reference Sequence (RefSeq) database aimed to provide a biologically non-redundant collection of DNA, RNA, and protein sequences and to promote the research on genes and proteins of human beings and other species. However, because of widely distributed polymorphisms and different quality control of experiments in individual laboratories, there are potential problems need to be identified in the RefSeq database. Regarding which, we herein define the concept, standard transcript, based on the Central Dogmas of Biology that each standard transcript should be perfectly mapped to the standard genomic DNA sequence at the exon level. A large scale analysis for mapping all of the RefSeq records of human being (2005-4-18) to the officially released human genome sequence database (2005-4-20) was further performed using BLAT, Sim4 and a homemade program, EIparser, which was especially designed for this purpose. The standard transcripts based on the RefSeq database were obtained according to the alignment with standard human genome database. There are 9,771 RefSeq records of human being labeled with "NM_" and "NR_" could be perfectly mapped to human genome sequences, while other 10,943 records could be considered as standard transcripts after reasonable revision by comparing with the genome sequences according to all of the three methods. Moreover, the left 203 unrevisable records and 2,676 inconsistent records reported by the above programs could not be considered as standard transcripts and should be checked critically before using because of potential errors in them. Our study has thus provided a reference standard dataset of human beings with high quality for further bioinformatic and experimental analysis such as polymorphism and mutation of human genes. The reference standard dataset based on above criteria could be retrieved from http://biocompute.bmi.ac.cn/transcriptome/index.htm.

Databases, Genetic↗

The EBI SRS server-new features.

MOTIVATION: Here we report on recent developments at the EBI SRS server (http://srs.ebi.ac.uk). SRS has become an integration system for both data retrieval and sequence analysis applications. The EBI SRS server is a primary gateway to major databases in the field of molecular biology produced and supported at EBI as well as European public access point to the MEDLINE database provided by US National Library of Medicine (NLM). It is a reference server for latest developments in data and application integration. The new additions include: concept of virtual databases, integration of XML databases like the Integrated Resource of Protein Domains and Functional Sites (InterPro), Gene Ontology (GO), MEDLINE, Metabolic pathways, etc., user friendly data representation in 'Nice views', SRSQuickSearch bookmarklets. AVAILABILITY: SRS6 is a licensed product of LION Bioscience AG freely available for academics. The EBI SRS server (http://srs.ebi.ac.uk) is a free central resource for molecular biology data as well as a reference server for the latest developments in data integration.

Computer Communication Networks↗

The RESID Database of protein structure modifications and the NRL-3D Sequence-Structure Database.

The RESID Database is a comprehensive collection of annotations and structures for protein post-translational modifications including N-terminal, C-terminal and peptide chain cross-link modifications. The RESID Database includes systematic and frequently observed alternate names, Chemical Abstracts Service registry numbers, atomic formulas and weights, enzyme activities, taxonomic range, keywords, literature citations with database cross-references, structural diagrams and molecular models. The NRL-3D Sequence-Structure Database is derived from the three-dimensional structure of proteins deposited with the Research Collaboratory for Structural Bioinformatics Protein Data Bank. The NRL-3D Database includes standardized and frequently observed alternate names, sources, keywords, literature citations, experimental conditions and searchable sequences from model coordinates. These databases are freely accessible through the National Cancer Institute-Frederick Advanced Biomedical Computing Center at these web sites: http://www. ncifcrf.gov/RESID, http://www.ncifcrf.gov/NRL-3D; or at these National Biomedical Research Foundation Protein Information Resource web sites: http://pir.georgetown.edu/pirwww/dbinfo/resid .html, http://pir.georgetown.edu/pirwww/dbinfo/nrl3d .html

Amino Acids↗

Codetection of a mixed population of candHPV62 containing wild-type and disrupted E1 open-reading frame in a 45-year-old woman with normal cytology.

We have cloned, sequenced, and characterized the complete genome of a novel human papillomavirus (HPV), candHPV62. During cloning, 2 candHPV62 viral isolates were recovered from a single cervical sample; 1 had all anticipated HPV open-reading frames (ORFs) intact, whereas the other exhibited an E1 frame-shift mutation. Further experiments indicated that the 2 strains were equivalent in abundance. It appears that an early mutation occurred within the E1 ORF, which was transcomplemented by an intact E1 protein. A search of the HPV database identified disruption of the E1 ORF in the cloned reference isolates of HPV16, HPV53, HPV56, and HPV72. These data suggest that disruption of the E1 ORF in genital HPVs is not uncommon.

Base Sequence↗

Inventory of high-abundance mRNAs in skeletal muscle of normal men.

G42875rial analysis of gene expression (SAGE) method was used to generate a catalog of 53,875 short (14 base) expressed sequence tags from polyadenylated RNA obtained from vastus lateralis muscle of healthy young men. Over 12,000 unique tags were detected. The frequency of occurrence of each tag reflects the relative abundance of the corresponding mRNA. The mRNA species that were detected 10 or more times, each comprising >/=0.02% of the mRNA population, accounted for 64% of the mRNA mass but <10% of the total number of mRNA species detected. Almost all of the abundant tags matched mRNA or EST sequences cataloged in GenBank. Mitochondrial transcripts accounted for approximately 20% of the polyadenylated RNA. Transcripts encoding proteins of the myofibrils were the most abundant nuclear-encoded mRNAs. Transcripts encoding ribosomal proteins, and those encoding proteins involved in energy metabolism, also were very abundant. The database can be used as a reference for investigations of alterations in gene expression associated with conditions that influence muscle function, such as muscular dystrophies, aging, and exercise.

Expressed Sequence Tags↗

FUGOID: functional genomics of organellar introns database.

FUGOID is a web-based, taxonomically broad organelle intron database that collects and integrates various functional and structural data on organellar (mitochondrial and chloroplast) introns. The main information provided by FUGOID includes intron sequence, subclass, resident ORF, self-splicing capability, host gene, protein factor(s) involved in splicing, mobility, insertion site, twintron, seminal references and taxonomic position of host organism. It is implemented in a relational database management system, allowing sophisticated, user-friendly searching, data entry and revision. Users can access the database by any common web browser using a variety of operating systems. The main page of the database is available at http://wnt.cc.utexas.edu/~ifmr530/introndata/main.htm.

Animals↗

Sequence search algorithm assessment and testing toolkit (SAT).

MOTIVATION: The Sequence Search Algorithm Assessment and Testing Toolkit (SAT) aims to be a complete package for the comparison of different protein homology search algorithms. The structural classification of proteins can provide us with a clear criterion for judgment in homology detection. There have been several assessments based on structural sequences with classifications but a good deal of similar work is now being repeated with locally developed procedures and programs. The SAT will provide developers with a complete package which will save time and produce more comparable performance assessments for search algorithms. The package is complete in the sense that it provides a non-redundant large sequence resource database, a well-characterized query database of proteins domains, all the parsers and some previous results from PSI-BLAST and a hidden markov model algorithm. RESULTS: An analysis on two different data sets was carried out using the SAT package. It compared the performance of a full protein sequence database (RSDB100) with a non-redundant representative sequence database derived from it (RSDB50). The performance measurement indicated that the full database is sub-optimal for a homology search. This result justifies the use of much smaller and faster RSDB50 than RSDB100 for the SAT. AVAILABILITY: A web site is up. The whole packa ge is accessible via www and ftp. ftp://ftp.ebi.ac.uk/pub/contrib/jong/SAT http://cyrah.ebi.ac.uk:1111/Proj/Bio/SAT http://www.mrc-lmb.cam.ac.uk/genomes/SAT In the package, some previous assessment results produced by the package can also be found for reference. CONTACT: jong@ebi.ac.uk

Algorithms↗

A database for G proteins and their interaction with GPCRs.

BACKGROUND: G protein-coupled receptors (GPCRs) transduce signals from extracellular space into the cell, through their interaction with G proteins, which act as switches forming hetero-trimers composed of different subunits (alpha,beta,gamma). The alpha subunit of the G protein is responsible for the recognition of a given GPCR. Whereas specialised resources for GPCRs, and other groups of receptors, are already available, currently, there is no publicly available database focusing on G Proteins and containing information about their coupling specificity with their respective receptors. DESCRIPTION: gpDB is a publicly accessible G proteins/GPCRs relational database. Including species homologs, the database contains detailed information for 418 G protein monomers (272 Galpha, 87 Gbeta and 59 Ggamma) and 2782 GPCRs sequences belonging to families with known coupling to G proteins. The GPCRs and the G proteins are classified according to a hierarchy of different classes, families and sub-families, based on extensive literature searchs. The main innovation besides the classification of both G proteins and GPCRs is the relational model of the database, describing the known coupling specificity of the GPCRs to their respective alpha subunit of G proteins, a unique feature not available in any other database. There is full sequence information with cross-references to publicly available databases, references to the literature concerning the coupling specificity and the dimerization of GPCRs and the user may submit advanced queries for text search. Furthermore, we provide a pattern search tool, an interface for running BLAST against the database and interconnectivity with PRED-TMR, PRED-GPCR and TMRPres2D. CONCLUSIONS: The database will be very useful, for both experimentalists and bioinformaticians, for the study of G protein/GPCR interactions and for future development of predictive algorithms. It is available for academics, via a web browser at the URL: http://bioinformatics.biol.uoa.gr/gpDB.

Databases, Protein↗

The KEGG database.

KEGG (http://www.genome.ad.jp/kegg/) is a suite of databases and associated software for understanding and simulating higher-order functional behaviours of the cell or the organism from its genome information. First, KEGG computerizes data and knowledge on protein interaction networks (PATHWAY database) and chemical reactions (LIGAND database) that are responsible for various cellular processes. Second, KEGG attempts to reconstruct protein interaction networks for all organisms whose genomes are completely sequenced (GENES and SSDB databases). Third, KEGG can be utilized as reference knowledge for functional genomics (EXPRESSION database) and proteomics (BRITE database) experiments. I will review the current status of KEGG and report on new developments in graph representation and graph computations.

Amino Acid Sequence↗

MycDB: an integrated mycobacterial database.

As part of ongoing efforts to investigate the molecular biology of the human pathogens in the genus Mycobacterium, a customized database was developed specifically for these organisms and implemented in ACEDB database manager software. The data loaded include the IMMYC Antigen List, details of reagents available from the CDC/WHO Antibody Bank, more than 1 Mb of sequences of mycobacterial genes and proteins from public databases, the physical maps of Mycobacterium leprae and Mycobacterium tuberculosis developed at the Institut Pasteur, as well as a subset of the references found in MedLine. The ACEDB software allows both quick and intuitive access to the data and to connections between facts by a simple mouse-driven interface, as well as by more powerful query mechanisms.

Amino Acid Sequence↗

The Database of Interacting Proteins: 2004 update.

The Database of Interacting Proteins (http://dip.doe-mbi.ucla.edu) aims to integrate the diverse body of experimental evidence on protein-protein interactions into a single, easily accessible online database. Because the reliability of experimental evidence varies widely, methods of quality assessment have been developed and utilized to identify the most reliable subset of the interactions. This CORE set can be used as a reference when evaluating the reliability of high-throughput protein-protein interaction data sets, for development of prediction methods, as well as in the studies of the properties of protein interaction networks.

Animals↗

The TIGRFAMs database of protein families.

TIGRFAMs is a collection of manually curated protein families consisting of hidden Markov models (HMMs), multiple sequence alignments, commentary, Gene Ontology (GO) assignments, literature references and pointers to related TIGRFAMs, Pfam and InterPro models. These models are designed to support both automated and manually curated annotation of genomes. TIGRFAMs contains models of full-length proteins and shorter regions at the levels of superfamilies, subfamilies and equivalogs, where equivalogs are sets of homologous proteins conserved with respect to function since their last common ancestor. The scope of each model is set by raising or lowering cutoff scores and choosing members of the seed alignment to group proteins sharing specific function (equivalog) or more general properties. The overall goal is to provide information with maximum utility for the annotation process. TIGRFAMs is thus complementary to Pfam, whose models typically achieve broad coverage across distant homologs but end at the boundaries of conserved structural domains. The database currently contains over 1600 protein families. TIGRFAMs is available for searching or downloading at www.tigr.org/TIGRFAMs.

Animals↗

Identification and kinetic analysis of a functional homolog of elongation factor 3, YEF3 in Saccharomyces cerevisiae.

Yeast and other fungi contain a soluble elongation factor 3 (EF-3) which is required for growth and protein synthesis. EF-3 contains two ABC cassettes, and binds and hydrolyses ATP. We identified a homolog of the YEF3 gene in the Saccharomyces cerevisiae genome database. This gene, designated YEF3B, is 84% identical in protein sequence to YEF3, which we will now refer to as YEF3A. YEF3B is not expressed during growth under laboratory conditions, and thus cannot rescue growth of YEF3A deletion strains. However, YEF3B can take the place of YEF3A in vivo when expressed from the YEF3A or ADH1 promoters. The products of the YEF3A and YEF3B genes, EF-3A and EF-3B, respectively, were expressed from the ADH1 promoter and purified. Both factors possessed basal and ribosomal-stimulated ATPase activity, and had similar affinity for yeast ribosomes (103 to 113 nM). K(m) values for ATP were similar, but the Kcat values differed significantly. Ribosome-dependent ATPase activity of EF-3A was more efficient than EF-3B, since the Kcat and Kcat/K(m) values for EF-3A were about two-fold higher; however, the difference in Kcat/K(m) values between the two factors was small for basal ATPase activity.

Base Sequence↗