Search PubMedSearch

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Gene-protein database of Escherichia coli K-12: edition 3.

The first two editions of the E. coli Gene-Protein Index were published to provide identifications of protein spots resolved by two-dimensional gel electrophoresis as the products of known genes. This third edition has been expanded to include information about genes and proteins gained directly from two-dimensional gel analysis--including information about protein spots not yet characterized genetically or biochemically--and is therefore more properly called a cellular protein database. An alpha-numeric designation has been uniquely assigned to each of the 616 polypeptide spots in the current database. To this, information is linked about the polypeptide's identification (protein name, gene name, Enzyme Commission--EC number), location on reference gels (x-y coordinates), genetics (Genbank code, DNA sequence reference), biochemistry (molecular weight, isoelectric point), and physiology (steady state level of the protein as a function of media and temperature, membership in various regulons and stimulons).

Bacterial Proteins

MycDB: an integrated mycobacterial database.

As part of ongoing efforts to investigate the molecular biology of the human pathogens in the genus Mycobacterium, a customized database was developed specifically for these organisms and implemented in ACEDB database manager software. The data loaded include the IMMYC Antigen List, details of reagents available from the CDC/WHO Antibody Bank, more than 1 Mb of sequences of mycobacterial genes and proteins from public databases, the physical maps of Mycobacterium leprae and Mycobacterium tuberculosis developed at the Institut Pasteur, as well as a subset of the references found in MedLine. The ACEDB software allows both quick and intuitive access to the data and to connections between facts by a simple mouse-driven interface, as well as by more powerful query mechanisms.

Amino Acid Sequence

Identification and kinetic analysis of a functional homolog of elongation factor 3, YEF3 in Saccharomyces cerevisiae.

Yeast and other fungi contain a soluble elongation factor 3 (EF-3) which is required for growth and protein synthesis. EF-3 contains two ABC cassettes, and binds and hydrolyses ATP. We identified a homolog of the YEF3 gene in the Saccharomyces cerevisiae genome database. This gene, designated YEF3B, is 84% identical in protein sequence to YEF3, which we will now refer to as YEF3A. YEF3B is not expressed during growth under laboratory conditions, and thus cannot rescue growth of YEF3A deletion strains. However, YEF3B can take the place of YEF3A in vivo when expressed from the YEF3A or ADH1 promoters. The products of the YEF3A and YEF3B genes, EF-3A and EF-3B, respectively, were expressed from the ADH1 promoter and purified. Both factors possessed basal and ribosomal-stimulated ATPase activity, and had similar affinity for yeast ribosomes (103 to 113 nM). K(m) values for ATP were similar, but the Kcat values differed significantly. Ribosome-dependent ATPase activity of EF-3A was more efficient than EF-3B, since the Kcat and Kcat/K(m) values for EF-3A were about two-fold higher; however, the difference in Kcat/K(m) values between the two factors was small for basal ATPase activity.

Base Sequence

MITOP: database for mitochondria-related proteins, genes and diseases.

The MITOP database http://websvr.mips.biochem.mpg. de/proj/medgen/mitop/ consolidates information on both nuclear- and mitochondrial-encoded genes and their proteins. The five species files- Saccharomyces cerevisiae, Mus musculus, Caenorhabditis elegans, Neurospora crassa and Homo sapiens -include annotated data derived from a variety of online resources and the literature. A wide spectrum of search facilities is given in the interelated sections 'Gene catalogues', 'Protein catalogues', 'Homologies', 'Pathways and metabolism', and 'Human disease catalogue' including extensive references and hyperlinks for each entry. Precomputed FASTA searches using all the MITOP yeast protein entries and a list of the best EST hits with graphical cluster alignments related to the yeast reference sequence are presented. The MITOP orthologue tables with cross-listing to all the protein entries for each species in the database facilitate investigations into interspecies homology. A program (MITOPROT) is available to identify mitochondrial targeting sequences and graphical depictions of several important mitochondrial processes are included. The 'Human disease catalogue' lists a total of 101 disorders related to mitochondrial protein abnormalities, sorted by clinical criteria and age of onset.

Animals

Sequencing emm-specific PCR products for routine and accurate typing of group A streptococci.

Rapid sequence analysis of specific PCR products was used to accurately deduce emm types corresponding to the majority of the known group A streptococcal (GAS) M serotypes. The study involved 95 M type reference GAS strains and a survey of 74 recent clinical isolates. A high percentage of agreement between M type serology and the previously published 5' sequences of the emm genes of M type reference strains was noted. The 5' sequences for six established M protein genes--the emm-32, emm-34, emm-38, emm-40, emm-42, and emm-71 genes--were determined to supplement the existing emm sequence database. Rapid sequence analysis differentiated serologically M-nontypeable strains and was used to establish the probable.

Amino Acid Sequence

Two-dimensional gel electrophoresis of Escherichia coli homogenates: the Escherichia coli SWISS-2DPAGE database.

Numerous Escherichia coli proteins have already been characterized by two-dimensional gel electrophoresis (2-D PAGE), using carrier ampholytes in the first dimension (VanBogelen, R. A., Sankar, P., Clark, R. L., Bogan, J. A. and Neidhardt, F. C., Electrophoresis 1992, 13, 1014-1054). We present here a reference protein map of E. coli obtained with immobilized pH gradients (IPG) and available in a SWISS-2DPAGE format. Out of the protein spots identified in the E. coli gene protein database by Neidhardt's group, 153 have been identified in the E. coli gene protein database by Neihardt's group, 153 have been identified on the E. coli SWISS-2DPAGE database map by gel comparison and most of them were confirmed either by the analysis of amino acid composition (AAC) and/or N-terminal microsequencing. Additionally, five as yet unsequenced proteins were found. The E. coli SWISS-2DPAGE database is part of the ExPASy molecular biology server accessible through the Word Wide Web network.

Amino Acid Sequence

Database of mutations within the adenovirus 5 E1A oncogene.

The Ad5 E1A database is a listing of mutations affecting the early region 1A (E1A) proteins of human adenovirus type 5. The database contains the name of the mutation, the nucleic acid sequence changes, the resulting alterations in amino acid sequence and reference. Additional notes and references are provided on the effect of each mutation on E1A function. The database is contained within the Adenovirus 5 E1A page on the World Wide Web at: http://www.geocities.com/CapeCanaveral/Hangar /2541/

Adenovirus E1A Proteins

High-throughput screening of historic collections: observations on file size, biological targets, and file diversity.

At Pfizer Central Research, high-throughput screening has been an important source of new leads for drug discovery for a decade. Our experience with over 150 high-throughput screens can address questions about necessary file size, how well particular biological targets fare (with particular reference to protein-protein interactions), and what file diversity means in practice.

Databases, Factual

The human keratinocyte two-dimensional gel protein database (update 1992): towards an integrated approach to the study of cell proliferation, differentiation and skin diseases.

The master two-dimensional gel database of human keratinocytes currently lists 2980 cellular proteins (2098 isoelectric focusing, IEF; and 882 nonequilibrium pH gradient electrophoresis, NEPHGE) many of which correspond to posttranslational modifications. About 20% of all recorded proteins have been identified (protein name, organelle components, etc.) and they are listed in alphabetical order together with their M(r), pI, cellular localization and credit to the investigator(s) that aided in the identification. Also, we have listed 145 microsequenced proteins that are recorded in this database. As an aid in localizing the polypeptides we have included blow-ups of the master images (IEF, NEPHGE) displaying all the protein numbers. In the long run, the master keratinocyte database is expected to link protein and DNA sequencing and mapping information (Human Genome Program) and to provide an integrated picture of the expression levels and properties of the thousands of proteins that orchestrate various keratinocyte functions both in health and disease.

Cell Differentiation

Soy isoflavone analysis: quality control and a new internal standard.

Development of a database of the soy isoflavone content of foods requires accurate and precise evaluation of different food matrixes. To evaluate accuracy, we estimated recoveries of both internal and external standards in 5 different soyfoods weekly. Standards were evaluated daily for system quality assurance. To evaluate sample precision, we analyzed soybeans and soymilk bimonthly for within-day precision and over 4 d for day-to-day precision. CVs should be < or = 8%. We validated our methods for single and multiple recovery concentrations by using our new internal standard, 2,4,4'-trihydroxydeoxybenzoin, and the external standards daidzein, genistein, and genistin. Concentrations of 12 isoflavone isomers, 3 aglycones (daidzein, genistein, and glycitein), and 9 glucosides (daidzin, genistin, glycitin, acetyldaidzin, acetylgenistin, acetylglycitin, malonyldaidzin, malonylgenistin, and malonylglycitin) were measured in a variety of soybeans and soyfoods. The extraction methods used depended on soyfood type. The HPLC conditions for soy isoflavone analysis were improved, leading to good separation with a short analysis time (60 min/sample). A data bank of concentration and distribution of isoflavones in different soybean products was assembled. A wide range of isoflavone concentrations, from < 50 microg/g to > 20,000 microg/g, was found in different soy products. The glucoside forms are almost twice the molecular weight of the aglycones; reported isoflavone concentrations should be normalized to the aglycone mass (or an isoflavonoid equivalent) rather than a simple sum of all isomers.

Chromatography, High Pressure Liquid

Location on the human genetic linkage map of 26 genes involved in blood coagulation.

Several human genetic linkage maps have been constructed as part of the Human Genome Project. These maps show the positional order of closely linked, highly informative AC-repeat polymorphisms on each human chromosome, and are extremely useful in genetic linkage analysis of inheritable diseases. For a candidate gene approach the current linkage maps are less useful, since they consist mainly of anonymous markers rather than of specific genes. This situation also applies for inheritable disorders of blood coagulation. Numerous genes are involved in the blood coagulation cascade and its regulation, and can be considered as candidate genes for unexplained haemophilia and thrombophilia. We have selected 29 candidate genes that seem to be the ones most likely to be involved in thrombophilia. For 19 genes genotype data were already present in the CEPH database (version 7.0). We typed 7 additional genes in the CEPH reference families, i.e. the factor V, factor XII, protein C, protein S, prothrombin, thrombomodulin, and heparin cofactor II gene. The genotype data were used to integrate these 26 genes in the current genetic linkage map, and to identify closely linked AC-repeat polymorphisms. This information will benefit the investigation of inheritable disorders of blood coagulation, especially thrombophilia.

Base Sequence

Database of mutations that alter the large tumor antigen in simian virus 40.

The SV40 T antigen database is a listing of plasmids and/or viruses that express mutant forms of the virus-encoded large T antigen protein. The parental virus strain, nucleic acid sequence of the mutations, the effect of the mutation on the T antigen amino acid sequence, and key references are included in the listing. The database is available from the authors as a Macintosh FileMaker Pro file, and as a hard copy printout.

Antigens, Polyomavirus Transforming

Structure prediction and modelling.

Protein structure prediction from sequence remains a major goal in molecular biology. The methods described in this review concentrate on deriving structural information through the detection of similarities between a test sequence and a database of known structures. Such methods are often referred to as knowledge-based strategies reflecting the use of a structural database in the analyses. The past year has seen considerable advances in both the development of automated procedures and their application to protein sequences of outstanding biological interest.

Algorithms

Metatranscriptomic analysis of viral sequences associated with Culex nigripalpus at an Alabama aquaculture site.

Mosquitoes associated with aquaculture habitats can harbor diverse viruses, yet the viromes of many locally abundant species remain poorly characterized. At an aquaculture-associated site in Auburn, Alabama, we surveyed mosquito populations and found Culex nigripalpus to be the dominant species collected. To characterize viruses associated with this mosquito, we performed RNA-seq on pooled female Cx. nigripalpus and compared complementary bioinformatic workflows for viral detection and genome recovery. One workflow removed host-associated reads by mapping to the closest available mosquito reference genome prior to assembly, whereas a second workflow used fully de novo assembly and viral database annotation. Additional protein-level filtering, cross-workflow comparison, and comparison of Trinity and rnaSPAdes assemblies were used to prioritize well-supported viral candidates. Across the original analyses, 16 submitted accessions corresponding to 12 collapsed virus/name groups were recovered, including Merida virus, Hubei mosquito virus 5, Zhejiang mosquito virus, Hubei virga-like virus 3, Rinkaby virus, Elemess virus, Qingnian mosquito virus, Serbia narna-like virus 2, XiangYun narna-levi-like virus 8, Ecclesville picorna-like virus, and baculovirus-like fragments. Several candidates were supported across multiple workflows, while others were recovered only under specific analytical conditions, indicating that candidate recovery was influenced by assembly and filtering choices. Selected viral contigs were independently supported by RT-PCR amplification. Overall, these results provide a first characterization of viral sequences associated with Cx. nigripalpus from an Alabama aquaculture-associated site and show that comparison across assembly and filtering strategies helped prioritize the most consistently supported viral candidates.

Animals

The Ribonuclease P database.

The Ribonuclease P Sequence database is a compilation of RNase P sequences, sequence alignments, secondary structures, three-dimensional models, and accessory information. In its initial form, the database contains information on RNase P RNA in bacteria and archaea, and RNase P protein in bacteria. The sequences themselves are presented phylogenetically ordered and aligned. The database also contains secondary structures of bacterial and archaeal RNAs, including specially annotated 'reference' secondary structures of Escherichia coli and Bacillus subtilis RNase P RNAs, a minimum phylogenetic consensus structure, and coordinates for models of three-dimensional structure.

Bacillus subtilis