Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Gene-protein database of Escherichia coli K-12: edition 3.

The first two editions of the E. coli Gene-Protein Index were published to provide identifications of protein spots resolved by two-dimensional gel electrophoresis as the products of known genes. This third edition has been expanded to include information about genes and proteins gained directly from two-dimensional gel analysis--including information about protein spots not yet characterized genetically or biochemically--and is therefore more properly called a cellular protein database. An alpha-numeric designation has been uniquely assigned to each of the 616 polypeptide spots in the current database. To this, information is linked about the polypeptide's identification (protein name, gene name, Enzyme Commission--EC number), location on reference gels (x-y coordinates), genetics (Genbank code, DNA sequence reference), biochemistry (molecular weight, isoelectric point), and physiology (steady state level of the protein as a function of media and temperature, membership in various regulons and stimulons).

Bacterial Proteins↗

Protein restriction for diabetic renal disease.

OBJECTIVES: To determine whether protein restriction slows or prevents progression of diabetic nephropathy towards renal failure. SEARCH STRATEGY: Computerised databases MEDLINE (1976-1996) and EMBASE (1974-1996) were searched using keywords diabetes mellitus, diabetic nephropathy, dietary proteins, diet, protein restricted and uremia. Recent issues of selected journals (Diabetic Medicine, Diabetologia, Diabetes Care, Kidney International, Nephrology Dialysis and Transplantation) were handsearched for papers not yet in the computerised databases. Reference lists of papers were also checked. SELECTION CRITERIA: This review was not limited to randomised controlled trials. All trials involving people with insulin-dependent diabetes following a lower protein diet for at least 4 months were considered since the straight line nature of progression as reflected by GFR means that patients can act as their own controls in a before and after comparison. DATA COLLECTION AND ANALYSIS: Data were extracted for length of follow up, level of protein restriction, renal function and dietary compliance. No studies of the impact of protein restriction on outcomes such as the need for dialysis or transplantation were found. The trials reported only the effect on short-term indicators such as creatinine clearance. MAIN RESULTS: Overall a protein restricted diet (0.3-0. 8g/kg) does appear to slow the progression of diabetic nephropathy towards renal failure. REVIEWER'S CONCLUSIONS: The results show that reducing protein intake appears to slow progression to renal failure, but some questions remain unanswered. The first is what level of protein restriction we should be used? The trials aimed for a daily intake of between 0.3 to 0.8g/kg of protein. The second concerns compliance in routine care - what level would be acceptable to patients? The third concerns long term outcomes -the present trials use proxy indicators such as creatinine clearance rather than outcomes such as time to dialysis or prevention of ESRF. All trials were carried out in subjects with insulin-dependent diabetes. It remains to be seen if a lower protein intake would slow the progression of nephropathy affecting the non-insulin dependent diabetic population.

Diabetic Nephropathies↗

The RESID Database of protein structure modifications and the NRL-3D Sequence-Structure Database.

The RESID Database is a comprehensive collection of annotations and structures for protein post-translational modifications including N-terminal, C-terminal and peptide chain cross-link modifications. The RESID Database includes systematic and frequently observed alternate names, Chemical Abstracts Service registry numbers, atomic formulas and weights, enzyme activities, taxonomic range, keywords, literature citations with database cross-references, structural diagrams and molecular models. The NRL-3D Sequence-Structure Database is derived from the three-dimensional structure of proteins deposited with the Research Collaboratory for Structural Bioinformatics Protein Data Bank. The NRL-3D Database includes standardized and frequently observed alternate names, sources, keywords, literature citations, experimental conditions and searchable sequences from model coordinates. These databases are freely accessible through the National Cancer Institute-Frederick Advanced Biomedical Computing Center at these web sites: http://www. ncifcrf.gov/RESID, http://www.ncifcrf.gov/NRL-3D; or at these National Biomedical Research Foundation Protein Information Resource web sites: http://pir.georgetown.edu/pirwww/dbinfo/resid .html, http://pir.georgetown.edu/pirwww/dbinfo/nrl3d .html

Amino Acids↗

Inventory of high-abundance mRNAs in skeletal muscle of normal men.

G42875rial analysis of gene expression (SAGE) method was used to generate a catalog of 53,875 short (14 base) expressed sequence tags from polyadenylated RNA obtained from vastus lateralis muscle of healthy young men. Over 12,000 unique tags were detected. The frequency of occurrence of each tag reflects the relative abundance of the corresponding mRNA. The mRNA species that were detected 10 or more times, each comprising >/=0.02% of the mRNA population, accounted for 64% of the mRNA mass but <10% of the total number of mRNA species detected. Almost all of the abundant tags matched mRNA or EST sequences cataloged in GenBank. Mitochondrial transcripts accounted for approximately 20% of the polyadenylated RNA. Transcripts encoding proteins of the myofibrils were the most abundant nuclear-encoded mRNAs. Transcripts encoding ribosomal proteins, and those encoding proteins involved in energy metabolism, also were very abundant. The database can be used as a reference for investigations of alterations in gene expression associated with conditions that influence muscle function, such as muscular dystrophies, aging, and exercise.

Expressed Sequence Tags↗

Sequence search algorithm assessment and testing toolkit (SAT).

MOTIVATION: The Sequence Search Algorithm Assessment and Testing Toolkit (SAT) aims to be a complete package for the comparison of different protein homology search algorithms. The structural classification of proteins can provide us with a clear criterion for judgment in homology detection. There have been several assessments based on structural sequences with classifications but a good deal of similar work is now being repeated with locally developed procedures and programs. The SAT will provide developers with a complete package which will save time and produce more comparable performance assessments for search algorithms. The package is complete in the sense that it provides a non-redundant large sequence resource database, a well-characterized query database of proteins domains, all the parsers and some previous results from PSI-BLAST and a hidden markov model algorithm. RESULTS: An analysis on two different data sets was carried out using the SAT package. It compared the performance of a full protein sequence database (RSDB100) with a non-redundant representative sequence database derived from it (RSDB50). The performance measurement indicated that the full database is sub-optimal for a homology search. This result justifies the use of much smaller and faster RSDB50 than RSDB100 for the SAT. AVAILABILITY: A web site is up. The whole packa ge is accessible via www and ftp. ftp://ftp.ebi.ac.uk/pub/contrib/jong/SAT http://cyrah.ebi.ac.uk:1111/Proj/Bio/SAT http://www.mrc-lmb.cam.ac.uk/genomes/SAT In the package, some previous assessment results produced by the package can also be found for reference. CONTACT: jong@ebi.ac.uk

Algorithms↗

MycDB: an integrated mycobacterial database.

As part of ongoing efforts to investigate the molecular biology of the human pathogens in the genus Mycobacterium, a customized database was developed specifically for these organisms and implemented in ACEDB database manager software. The data loaded include the IMMYC Antigen List, details of reagents available from the CDC/WHO Antibody Bank, more than 1 Mb of sequences of mycobacterial genes and proteins from public databases, the physical maps of Mycobacterium leprae and Mycobacterium tuberculosis developed at the Institut Pasteur, as well as a subset of the references found in MedLine. The ACEDB software allows both quick and intuitive access to the data and to connections between facts by a simple mouse-driven interface, as well as by more powerful query mechanisms.

Amino Acid Sequence↗

Identification and kinetic analysis of a functional homolog of elongation factor 3, YEF3 in Saccharomyces cerevisiae.

Yeast and other fungi contain a soluble elongation factor 3 (EF-3) which is required for growth and protein synthesis. EF-3 contains two ABC cassettes, and binds and hydrolyses ATP. We identified a homolog of the YEF3 gene in the Saccharomyces cerevisiae genome database. This gene, designated YEF3B, is 84% identical in protein sequence to YEF3, which we will now refer to as YEF3A. YEF3B is not expressed during growth under laboratory conditions, and thus cannot rescue growth of YEF3A deletion strains. However, YEF3B can take the place of YEF3A in vivo when expressed from the YEF3A or ADH1 promoters. The products of the YEF3A and YEF3B genes, EF-3A and EF-3B, respectively, were expressed from the ADH1 promoter and purified. Both factors possessed basal and ribosomal-stimulated ATPase activity, and had similar affinity for yeast ribosomes (103 to 113 nM). K(m) values for ATP were similar, but the Kcat values differed significantly. Ribosome-dependent ATPase activity of EF-3A was more efficient than EF-3B, since the Kcat and Kcat/K(m) values for EF-3A were about two-fold higher; however, the difference in Kcat/K(m) values between the two factors was small for basal ATPase activity.

Base Sequence↗

MITOP: database for mitochondria-related proteins, genes and diseases.

The MITOP database http://websvr.mips.biochem.mpg. de/proj/medgen/mitop/ consolidates information on both nuclear- and mitochondrial-encoded genes and their proteins. The five species files- Saccharomyces cerevisiae, Mus musculus, Caenorhabditis elegans, Neurospora crassa and Homo sapiens -include annotated data derived from a variety of online resources and the literature. A wide spectrum of search facilities is given in the interelated sections 'Gene catalogues', 'Protein catalogues', 'Homologies', 'Pathways and metabolism', and 'Human disease catalogue' including extensive references and hyperlinks for each entry. Precomputed FASTA searches using all the MITOP yeast protein entries and a list of the best EST hits with graphical cluster alignments related to the yeast reference sequence are presented. The MITOP orthologue tables with cross-listing to all the protein entries for each species in the database facilitate investigations into interspecies homology. A program (MITOPROT) is available to identify mitochondrial targeting sequences and graphical depictions of several important mitochondrial processes are included. The 'Human disease catalogue' lists a total of 101 disorders related to mitochondrial protein abnormalities, sorted by clinical criteria and age of onset.

Animals↗

Sequencing emm-specific PCR products for routine and accurate typing of group A streptococci.

Rapid sequence analysis of specific PCR products was used to accurately deduce emm types corresponding to the majority of the known group A streptococcal (GAS) M serotypes. The study involved 95 M type reference GAS strains and a survey of 74 recent clinical isolates. A high percentage of agreement between M type serology and the previously published 5' sequences of the emm genes of M type reference strains was noted. The 5' sequences for six established M protein genes--the emm-32, emm-34, emm-38, emm-40, emm-42, and emm-71 genes--were determined to supplement the existing emm sequence database. Rapid sequence analysis differentiated serologically M-nontypeable strains and was used to establish the probable.

Amino Acid Sequence↗

Two-dimensional gel electrophoresis of Escherichia coli homogenates: the Escherichia coli SWISS-2DPAGE database.

Numerous Escherichia coli proteins have already been characterized by two-dimensional gel electrophoresis (2-D PAGE), using carrier ampholytes in the first dimension (VanBogelen, R. A., Sankar, P., Clark, R. L., Bogan, J. A. and Neidhardt, F. C., Electrophoresis 1992, 13, 1014-1054). We present here a reference protein map of E. coli obtained with immobilized pH gradients (IPG) and available in a SWISS-2DPAGE format. Out of the protein spots identified in the E. coli gene protein database by Neidhardt's group, 153 have been identified in the E. coli gene protein database by Neihardt's group, 153 have been identified on the E. coli SWISS-2DPAGE database map by gel comparison and most of them were confirmed either by the analysis of amino acid composition (AAC) and/or N-terminal microsequencing. Additionally, five as yet unsequenced proteins were found. The E. coli SWISS-2DPAGE database is part of the ExPASy molecular biology server accessible through the Word Wide Web network.

Amino Acid Sequence↗

Database of mutations within the adenovirus 5 E1A oncogene.

The Ad5 E1A database is a listing of mutations affecting the early region 1A (E1A) proteins of human adenovirus type 5. The database contains the name of the mutation, the nucleic acid sequence changes, the resulting alterations in amino acid sequence and reference. Additional notes and references are provided on the effect of each mutation on E1A function. The database is contained within the Adenovirus 5 E1A page on the World Wide Web at: http://www.geocities.com/CapeCanaveral/Hangar /2541/

Adenovirus E1A Proteins↗

High-throughput screening of historic collections: observations on file size, biological targets, and file diversity.

At Pfizer Central Research, high-throughput screening has been an important source of new leads for drug discovery for a decade. Our experience with over 150 high-throughput screens can address questions about necessary file size, how well particular biological targets fare (with particular reference to protein-protein interactions), and what file diversity means in practice.

Databases, Factual↗

The human keratinocyte two-dimensional gel protein database (update 1992): towards an integrated approach to the study of cell proliferation, differentiation and skin diseases.

The master two-dimensional gel database of human keratinocytes currently lists 2980 cellular proteins (2098 isoelectric focusing, IEF; and 882 nonequilibrium pH gradient electrophoresis, NEPHGE) many of which correspond to posttranslational modifications. About 20% of all recorded proteins have been identified (protein name, organelle components, etc.) and they are listed in alphabetical order together with their M(r), pI, cellular localization and credit to the investigator(s) that aided in the identification. Also, we have listed 145 microsequenced proteins that are recorded in this database. As an aid in localizing the polypeptides we have included blow-ups of the master images (IEF, NEPHGE) displaying all the protein numbers. In the long run, the master keratinocyte database is expected to link protein and DNA sequencing and mapping information (Human Genome Program) and to provide an integrated picture of the expression levels and properties of the thousands of proteins that orchestrate various keratinocyte functions both in health and disease.

Cell Differentiation↗

Soy isoflavone analysis: quality control and a new internal standard.

Development of a database of the soy isoflavone content of foods requires accurate and precise evaluation of different food matrixes. To evaluate accuracy, we estimated recoveries of both internal and external standards in 5 different soyfoods weekly. Standards were evaluated daily for system quality assurance. To evaluate sample precision, we analyzed soybeans and soymilk bimonthly for within-day precision and over 4 d for day-to-day precision. CVs should be < or = 8%. We validated our methods for single and multiple recovery concentrations by using our new internal standard, 2,4,4'-trihydroxydeoxybenzoin, and the external standards daidzein, genistein, and genistin. Concentrations of 12 isoflavone isomers, 3 aglycones (daidzein, genistein, and glycitein), and 9 glucosides (daidzin, genistin, glycitin, acetyldaidzin, acetylgenistin, acetylglycitin, malonyldaidzin, malonylgenistin, and malonylglycitin) were measured in a variety of soybeans and soyfoods. The extraction methods used depended on soyfood type. The HPLC conditions for soy isoflavone analysis were improved, leading to good separation with a short analysis time (60 min/sample). A data bank of concentration and distribution of isoflavones in different soybean products was assembled. A wide range of isoflavone concentrations, from < 50 microg/g to > 20,000 microg/g, was found in different soy products. The glucoside forms are almost twice the molecular weight of the aglycones; reported isoflavone concentrations should be normalized to the aglycone mass (or an isoflavonoid equivalent) rather than a simple sum of all isomers.

Chromatography, High Pressure Liquid↗

Location on the human genetic linkage map of 26 genes involved in blood coagulation.

Several human genetic linkage maps have been constructed as part of the Human Genome Project. These maps show the positional order of closely linked, highly informative AC-repeat polymorphisms on each human chromosome, and are extremely useful in genetic linkage analysis of inheritable diseases. For a candidate gene approach the current linkage maps are less useful, since they consist mainly of anonymous markers rather than of specific genes. This situation also applies for inheritable disorders of blood coagulation. Numerous genes are involved in the blood coagulation cascade and its regulation, and can be considered as candidate genes for unexplained haemophilia and thrombophilia. We have selected 29 candidate genes that seem to be the ones most likely to be involved in thrombophilia. For 19 genes genotype data were already present in the CEPH database (version 7.0). We typed 7 additional genes in the CEPH reference families, i.e. the factor V, factor XII, protein C, protein S, prothrombin, thrombomodulin, and heparin cofactor II gene. The genotype data were used to integrate these 26 genes in the current genetic linkage map, and to identify closely linked AC-repeat polymorphisms. This information will benefit the investigation of inheritable disorders of blood coagulation, especially thrombophilia.

Base Sequence↗

Database of mutations that alter the large tumor antigen in simian virus 40.

The SV40 T antigen database is a listing of plasmids and/or viruses that express mutant forms of the virus-encoded large T antigen protein. The parental virus strain, nucleic acid sequence of the mutations, the effect of the mutation on the T antigen amino acid sequence, and key references are included in the listing. The database is available from the authors as a Macintosh FileMaker Pro file, and as a hard copy printout.

Antigens, Polyomavirus Transforming↗