Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Relationships between bacterial drug resistance pumps and other transport proteins.

We have used three reference sequences representative of bacterial drug resistance pumps and sugar transport proteins to collect the 91 most closely related sequences from a composite, nonredundant protein sequence database. Having eliminated certain very close relatives, the remainder were subjected to analysis and alignment by using two different similarity matrices: one of these was a matrix based on structural conservation of amino acid residues in proteins of known conformation and the other was based on the more familiar mutational matrix. Unrooted similarity trees for these proteins were constructed for each matrix and compared. A systematic analysis of the differences between these trees was undertaken and the sequences were analyzed for the presence or absence of certain sequence motifs. The results show that the clades created by the two methods are broadly comparable but that there are some clusters of sequences that are significantly different. Further analysis confirmed that (1) the sequences collected by this objective method are all known or putative 12-helix (in some cases reported as 14-helix) transmembrane proteins, (2) there is evidence for few cases of an origin based on gene duplication, (3) the bacterial drug resistance pumps are distributed in more than one clade and cannot be regarded as a definitive subset of these proteins, and that (4) the diversity is such that there is no evidence of a single ancestral protein. The possible extension of the methods to other cases of divergent protein sequences is discussed.

Amino Acid Sequence↗

Fission yeast tor1 functions in response to various stresses including nitrogen starvation, high osmolarity, and high temperature.

A target of rapamycin (TOR) protein is a protein kinase that exerts cellular signal transduction to regulate cell growth in response to extracellular nutrient conditions. In the Schizosaccharomyces pombe genome database, there are two genes encoding TOR-related proteins, but their functions have not been analyzed. Here we report that one of the genes, referred to as tor1+, is required for sexual development induced by nitrogen starvation. Ste11 is a key transcription factor for the initiation of sexual development. The expression of ste11+ is normally regulated in tor1- cells; and overexpression of ste11+ hardly rescues the defect in fertility in tor1-. Upon nitrogen starvation, tor1+ cells promote two rounds of the cell cycle to become arrested at the G1 phase before initiation of sexual development. The tor1- cells do not promote such a cell cycle, suggesting that Tor1 is necessary for the response to nitrogen starvation. The tor1- cells show no growth or very slow growth under various stress conditions, including external high pH, high concentrations of salts or sorbitol, and high temperature. These results suggest that Tor1 is necessary for any response to a wide range of stresses. The vegetative growth of tor1- cells is inhibited by rapamycin, although tor1+ cells are resistant to the drug. The tor1- cells are hypersensitive to fluphenazine and cyclosporin A, which specifically inhibit calmodulin and calcineurin, respectively.

Amino Acid Sequence↗

Proteomics reveals protein profile changes in doxorubicin--treated MCF-7 human breast cancer cells.

MCF-7 cells are extensively used as a cell model to investigate human breast tumors and the cellular mechanism of antitumor drugs such as doxorubicin (DOX), an anthracycline antitumor drug widely used in clinical chemotherapy. To understand the effects of DOX on the protein expression, we perform a comprehensive proteomics to survey global changes in proteins after DOX treatment in MCF-7 cells. Exposure of MCF-7 cells to 0.1 microM DOX for 2 days induced a differentiation-like phenotype with prominent perinuclear autocatalytic vacuoles, abundant filamentous material, and irregular microvilli at the cell surface. In this study, we also present a proteome reference map of MCF-7 cells with 21 identified protein spots via analysis of N-terminal sequencing, mass spectrometry, immunoblot and/or computer matching with protein database. Based on the proteome map, we found that DOX causes a markedly decrease in the levels of three isoforms of heat shock protein 27 (HSP27) whereas the levels of other stress associated proteins including HSP60, calreticulin, and protein disulfide isomerase were not significantly altered in DOX-treated MCF-7 cells. Taken together, we suggest that that action of DOX on breast tumor cells may be partly related to dysregulation of HSP27 expression. Modulation of HSP27 levels may be a clinically useful potential target for design of antitumor drugs and controlling breast tumor growth.

Antineoplastic Agents↗

A statistical analysis of N- and O-glycan linkage conformations from crystallographic data.

We have generated a database of 639 glycosidic linkage structures by an exhaustive survey of the available crystallographic data for isolated oligosaccharides, glycoproteins, and glycan-binding proteins. For isolated oligosaccharides there is relatively little crystallographic data available. A much larger number of glycoprotein and glycan-binding protein structures have now been solved in which two or more linked monosaccharides can be resolved. In the majority of these cases, only a few residues can be seen. Using the 639 glycosidic linkage structures, we have identified one or more distinct conformers for all the linkages. The O5-C1-O-C(x)' torsion angles for all these distinct conformers appear to be determined chiefly by the exo-anomeric effect. The Manalpha1-6Man linkage appears to be less restrained than the others, showing a wide degree of dispersion outside the ranges of the defined conformers. The identification of distinct conformers for glyco-sidic linkages allows "average" glycan structures to be modeled and also allows the easy identification of distorted glycosidic linkages. Such an analysis shows that the interactions between IgG Fc and its own N-linked glycan result in severe distortion of the terminal Galbeta1-4GlcNAc linkage only, indicating the strong interactions that must be present between the Gal residue and the protein surface. The applicability of this crystallographic based analysis to glycan structures in solution is discussed. This database of linkagestructures should be a very useful reference tool in three-dimensional structure determinations.

Carbohydrate Conformation↗

Functional coverage of the human genome by existing structures, structural genomics targets, and homology models.

The bias in protein structure and function space resulting from experimental limitations and targeting of particular functional classes of proteins by structural biologists has long been recognized, but never continuously quantified. Using the Enzyme Commission and the Gene Ontology classifications as a reference frame, and integrating structure data from the Protein Data Bank (PDB), target sequences from the structural genomics projects, structure homology derived from the SUPERFAMILY database, and genome annotations from Ensembl and NCBI, we provide a quantified view, both at the domain and whole-protein levels, of the current and projected coverage of protein structure and function space relative to the human genome. Protein structures currently provide at least one domain that covers 37% of the functional classes identified in the genome; whole structure coverage exists for 25% of the genome. If all the structural genomics targets were solved (twice the current number of structures in the PDB), it is estimated that structures of one domain would cover 69% of the functional classes identified and complete structure coverage would be 44%. Homology models from existing experimental structures extend the 37% coverage to 56% of the genome as single domains and 25% to 31% for complete structures. Coverage from homology models is not evenly distributed by protein family, reflecting differing degrees of sequence and structure divergence within families. While these data provide coverage, conversely, they also systematically highlight functional classes of proteins for which structures should be determined. Current key functional families without structure representation are highlighted here; updated information on the "most wanted list" that should be solved is available on a weekly basis from http://function.rcsb.org:8080/pdb/function_distribution/index.html.

Databases, Protein↗

Tertiary structure predictions on a comprehensive benchmark of medium to large size proteins.

We evaluate tertiary structure predictions on medium to large size proteins by TASSER, a new algorithm that assembles protein structures through rearranging the rigid fragments from threading templates guided by a reduced Calpha and side-chain based potential consistent with threading based tertiary restraints. Predictions were generated for 745 proteins 201-300 residues in length that cover the Protein Data Bank (PDB) at the level of 35% sequence identity. With homologous proteins excluded, in 365 cases, the templates identified by our threading program, PROSPECTOR_3, have a root-mean-square deviation (RMSD) to native < 6.5 angstroms, with >70% alignment coverage. After TASSER assembly, in 408 cases the best of the top five full-length models has a RMSD < 6.5 angstroms. Among the 745 targets are 18 membrane proteins, with one-third having a predicted RMSD < 5.5 A. For all representative proteins less than or equal to 300 residues that have corresponding multiple NMR structures in the Protein Data Bank, approximately 20% of the models generated by TASSER are closer to the NMR structure centroid than the farthest individual NMR model. These results suggest that reasonable structure predictions for nonhomologous large size proteins can be automatically generated on a proteomic scale, and the application of this approach to structural as well as functional genomics represent promising applications of TASSER.

Algorithms↗

Testing homology with Contact Accepted mutatiOn (CAO): a contact-based Markov model of protein evolution.

Point Accepted Mutation (PAM) is the Markov model of amino acid replacements in proteins introduced by Dayhoff and her co-workers (Dayhoff et al., 1978). The PAM matrices and other matrices based on the PAM model have been widely accepted as the standard scoring system of protein sequence similarity in protein sequence alignment tools. Here, we present Contact Accepted mutatiOn (CAO), a Markov model of protein residue contact mutations. The CAO model simulates the interchanging of structurally defined side-chain contacts, and introduces additional structural information into protein sequence alignments. Therefore, similarities between structurally conserved sequences can be detected even without apparent sequence similarity. CAO has been benchmarked on the HOMSTRAD database and a subset of the CATH database, by comparing sequence alignments with reference alignments derived from structural superposition. CAO yields scores that reflect coherently the structural quality of sequence alignments, which has implications particularly for homology modelling and threading techniques.

Amino Acid Sequence↗

Clinical laboratory differentiation of infectious versus non-infectious systemic inflammatory response syndrome.

OBJECTIVE: To evaluate the accuracy of C-reactive protein (CRP), procalcitonin (PCT), neopterin, and endotoxin in the differential diagnosis of sepsis and non-infectious systemic inflammatory response syndrome (SIRS). METHODS: A Medline database and references from identified articles were used to perform a literature search relating to the differential diagnosis of sepsis versus non-infectious SIRS. RESULTS: CRP, PCT, and neopterin are released both in sepsis and in non-infectious inflammatory disease. CRP and PCT are equally effective, although not perfect, in differentiating between sepsis and non-infectious SIRS. However, CRP and PCT have different kinetics and profiles. The kinetics of CRP is slower than that of PCT, and CRP levels may not further increase during more severe stages of sepsis. On the contrary, PCT rises in proportion to the severity of sepsis and reaches its highest levels in septic shock. PCT tends to be higher in nonsurvivor than in survivor. Therefore, PCT demonstrated a closer correlation with the severity of sepsis and outcome than CRP. Unlike CRP and PCT, neopterin is increased in viral infection as well as bacterial infection, and neopterin is also a useful indicator of sepsis. Endotoxemia was detected in no more than half of patients with Gram-negative bacteremia, and Gram-negative bacteremia was detected in half of patients with endotoxemia. CONCLUSIONS: The diagnostic capacity of PCT is superior to that of CRP due to the close correlation between PCT levels and the severity of sepsis and outcome. Neopterin is very useful in the diagnosis of viral infection. The endotoxin assay in combination with CRP, PCT, or neopterin may help as a diagnostic marker for Gram-negative bacterial infection.

Animals↗

Prediction of protein structural classes by support vector machines.

In this paper, we apply a new machine learning method which is called support vector machine to approach the prediction of protein structural class. The support vector machine method is performed based on the database derived from SCOP which is based upon domains of known structure and the evolutionary relationships and the principles that govern their 3D structure. As a result, high rates of both self-consistency and jackknife test are obtained. This indicates that the structural class of a protein inconsiderably correlated with its amino and composition, and the support vector machine can be referred as a powerful computational tool for predicting the structural classes of proteins.

Artificial Intelligence↗

AraC-XylS database: a family of positive transcriptional regulators in bacteria.

The AraC-XylS database contains information about a family of positive transcriptional regulators broadly distributed in bacteria. This specific database focuses on protein sequences and on the biological and functional features of each of the proteins that belong to this family. Each entry provides information on the protein itself, the annotated protein sequence and, when the crystal is available, a comprehensive representation of its three-dimensional structure. The organization of the database is based on an exhaustive analysis of the scientific literature. The data are interconnected and linked with other databases. Multiple alignments of the members of the family, an extensive collection of references and a tutorial about the family provide additional information. The AraC-XylS database is accessible on the World Wide Web at http://www.AraC-XylS.org.

Amino Acid Sequence↗

Towards the proteome of the marine bacterium Rhodopirellula baltica: mapping the soluble proteins.

The marine bacterium Rhodopirellula baltica, a member of the phylum Planctomycetes, has distinct morphological properties and contributes to remineralization of biomass in the natural environment. On the basis of its recently determined complete genome we investigated its proteome by 2-DE and established a reference 2-DE gel for the soluble protein fraction. Approximately 1000 protein spots were excised from a colloidal Coomassie-stained gel (pH 4-7), analyzed by MALDI-MS and identified by PMF. The non-redundant data set contained 626 distinct protein spots, corresponding to 558 different genes. The identified proteins were classified into role categories according to their predicted functions. The experimentally determined and the theoretically predicted proteomes were compared. Proteins, which were most abundant in 2-DE gels and the coding genes of which were also predicted to be highly expressed, could be linked mainly to housekeeping functions in glycolysis, tricarboxic acid cycle, amino acid biosynthesis, protein quality control and translation. Absence of predictable signal peptides indicated a localization of these proteins in the intracellular compartment, the pirellulosome. Among the identified proteins, 146 contained a predicted signal peptide suggesting their translocation. Some proteins were detected in more than one spot on the gel, indicating post-translational modification. In addition to identifying proteins present in the published sequence database for R. baltica, an alternative approach was used, in which the mass spectrometric data was searched against a maximal ORF set, allowing the identification of four previously unpredicted ORFs. The 2-DE reference map presented here will serve as framework for further experiments to study differential gene expression of R. baltica in response to external stimuli or cellular development and compartmentalization.

Bacteria↗

The ORFanage: an ORFan database.

As each newly sequenced genome contains a significant number of protein-coding ORFs that are species-, family- or lineage-specific, many interesting questions arise about the evolution and role of these ORFs and of the genomes they are part of. We refer to these poorly conserved ORFs as singleton or paralogous ORFans if they are unique to one genome, or as orthologous ORFans if they appear only in a family of closely related organisms and have no homolog in other genomes. In order to study and classify ORFans we have constructed the ORFanage, an ORFan database. This database consists of the predicted ORFs in fully sequenced microbial genomes, and enables searching for the three types of ORFans in any subset of the genomes chosen by the user. The ORFanage could help in choosing interesting targets for further genomic and evolutionary studies. The ORFanage is accessible via http://www.bioinformatics.buffalo. edu/ORFanage.

Computational Biology↗

Predicting co-complexed protein pairs using genomic and proteomic data integration.

BACKGROUND: Identifying all protein-protein interactions in an organism is a major objective of proteomics. A related goal is to know which protein pairs are present in the same protein complex. High-throughput methods such as yeast two-hybrid (Y2H) and affinity purification coupled with mass spectrometry (APMS) have been used to detect interacting proteins on a genomic scale. However, both Y2H and APMS methods have substantial false-positive rates. Aside from high-throughput interaction screens, other gene- or protein-pair characteristics may also be informative of physical interaction. Therefore it is desirable to integrate multiple datasets and utilize their different predictive value for more accurate prediction of co-complexed relationship. RESULTS: Using a supervised machine learning approach--probabilistic decision tree, we integrated high-throughput protein interaction datasets and other gene- and protein-pair characteristics to predict co-complexed pairs (CCP) of proteins. Our predictions proved more sensitive and specific than predictions based on Y2H or APMS methods alone or in combination. Among the top predictions not annotated as CCPs in our reference set (obtained from the MIPS complex catalogue), a significant fraction was found to physically interact according to a separate database (YPD, Yeast Proteome Database), and the remaining predictions may potentially represent unknown CCPs. CONCLUSIONS: We demonstrated that the probabilistic decision tree approach can be successfully used to predict co-complexed protein (CCP) pairs from other characteristics. Our top-scoring CCP predictions provide testable hypotheses for experimental validation.

Computational Biology↗

Extensive Analysis of Genetic Diversity in HLA-DMA, HLA-DMB, HLA-DOA and HLA-DOB: Characterisation of 236 Novel Alleles.

HLA-DMA, -DMB, -DOA and -DOB are non-classical HLA Class II genes that play a crucial role in the selection of highly stable HLA Class II/peptide complexes on antigen-presenting cells. Although the genes were initially thought to have a limited diversity with less than 13 alleles per gene documented in the IPD-IMGT/HLA Database in 2022, recent studies suggest a potential impact of certain alleles on the outcome of hematopoietic cell transplantation. To gain a deeper understanding of allelic diversity, we sequenced HLA-DMA, -DMB, -DOA and -DOB of 1880 potential stem cell donors from Germany, Poland, Great Britain and Chile, achieving full-gene resolution. Remarkably, we identified 3968 previously undescribed sequences, including 28 distinct novel proteins. The observed allele frequencies were consistent across all studied populations with one dominating protein for each gene: HLA-DMA*01:01 (>&#x2009;77%), HLA-DMB*01:01 (>&#x2009;63%), HLA-DOA*01:01 (>&#x2009;97%) and HLA-DOB*01:01 (>&#x2009;77%). Notably, a much higher diversity was observed in full-genomic resolution. Finally, we submitted 51 distinct novel sequences for HLA-DMA, 58 for HLA-DMB, 80 for HLA-DOA and 47 for HLA-DOB to the IPD-IMGT/HLA Database. This comprehensive reference database update will not only simplify future genotyping of HLA-DMA, -DMB, -DOA and -DOB but will hopefully also enhance our understanding of the complex process of peptide selection and loading to the HLA Class II proteins.

Humans↗

Analysis of mRNA expression and protein abundance data: an approach for the comparison of the enrichment of features in the cellular population of proteins and transcripts.

MOTIVATION: Protein abundance is related to mRNA expression through many different cellular processes. Up to now, there have been conflicting results on how correlated the levels of these two quantities are. Given that expression and abundance data are significantly more complex and noisy than the underlying genomic sequence information, it is reasonable to simplify and average them in terms of broad proteomic categories and features (e.g. functions or secondary structures), for understanding their relationship. Furthermore, it will be essential to integrate, within a common framework, the results of many varied experiments by different investigators. This will allow one to survey the characteristics of highly expressed genes and proteins. RESULTS: To this end, we outline a formalism for merging and scaling many different gene expression and protein abundance data sets into a comprehensive reference set, and we develop an approach for analyzing this in terms of broad categories, such as composition, function, structure and localization. As the various experiments are not always done using the same set of genes, sampling bias becomes a central issue, and our formalism is designed to explicitly show this and correct for it. We apply our formalism to the currently available gene expression and protein abundance data for yeast. Overall, we found substantial agreement between gene expression and protein abundance, in terms of the enrichment of structural and functional categories. This agreement, which was considerably greater than the simple correlation between these quantities for individual genes, reflects the way broad categories collect many individual measurements into simple, robust averages. In particular, we found that in comparison to the population of genes in the yeast genome, the cellular populations of transcripts and proteins (weighted by their respective abundances, the transcriptome and what we dub the translatome) were both enriched in: (i) the small amino acids Val, Gly, and Ala; (ii) low molecular weight proteins; (iii) helices and sheets relative to coils; (iv) cytoplasmic proteins relative to nuclear ones; and (v) proteins involved in 'protein synthesis,' 'cell structure,' and 'energy production.' SUPPLEMENTARY INFORMATION: http://genecensus.org/expression/translatome

Algorithms↗

Evolutionary relationships among G protein-coupled receptors using a clustered database approach.

Guanine nucleotide-binding protein-coupled receptors (GPCRs) comprise large and diverse gene families in fungi, plants, and the animal kingdom. GPCRs appear to share a common structure with 7 transmembrane segments, but sequence similarity is minimal among the most distant GPCRs. To reevaluate the question of evolutionary relationships among the disparate GPCR families, this study takes advantage of the dramatically increased number of cloned GPCRs. Sequences were selected from the National Center for Biotechnology Information (NCBI) nonredundant peptide database using iterative BLAST (Basic Local Alignment Search Tool) searches to yield a database of approximately 1700 GPCRs and unrelated membrane proteins as controls, divided into 34 distinct clusters. For each cluster, separate position-specific matrices were established to optimize sequence comparisons among GPCRs. This approach resulted in significant alignments between distant GPCR families, including receptors for the biogenic amine/peptide, VIP/secretin, cAMP, STE3/MAP3 fungal pheromones, latrophilin, developmental receptors frizzled and smoothened, as well as the more distant metabotrobic glutamate receptors, the STE2/MAM2 fungal pheromone receptors, and GPR1, a fungal glucose receptor. On the other hand, alignment scores between these recognized GPCR clades with p40 (putative GPCR) and pm1 (putative GPCR), as well as bacteriorhodopsins, failed to support a finding of homology. This study provides a refined view of GPCR ancestry and serves as a reference database with hyperlinks to other sources. Moreover, it may facilitate database annotation and the assignment of orphan receptors to GPCR families.

Amino Acid Sequence↗

An efficient disk based data structure for rapid searching of quantitative two-dimensional gel databases.

Fast access of two-dimensional (2-D) gel quantitative databases is important for rapid searching for protein differences between sets of 2-D gels from an experiment. The GELLAB-II system organizes corresponding spots from the gels in the database into reference or "Rspot" sets. These Rspot numeric names index fixed regions in the paged composite gel database file. This is adequate for an existing database, but has several problems. (i) Building the initial database requires guessing how much disk space to pre-allocate for each corresponding spot (i.e. spots from different gels). If it ever runs out of pre-allocated space during this process, it must expand the size of each corresponding set of spots copying the old database data into the new in-place on the disk. (ii) When adding new gels or editing the database, if a new spot is created, the system may also go into this expansion mode. The time spent and wasted disk space can be appreciable--depending on the size of the database (order of 100 gel database). (iii) Because each set of corresponding spots is the same size, we waste space in most spot sets since they do not require the additional space a few spot sets require which contain additional fragmented spots. We present a new low-level disk object-based structure and algorithm, paged indexed buckets (PIB), which optimizes disk space usage while having similar retrieval speed to the original method.

Algorithms↗

Spermatocytes and round spermatids of rat testis: protein patterns.

Spermatogenesis is a process in the testis that involves meiotic cell division and spermiogenesis. The mechanisms of regulation and its associated proteins are mostly unknown. This publication shows the two-dimensional (2-D) gel electrophoresis protein map obtained from rat testis using nonlinear 3.5-10 immobilized pH gradients for the first-dimensional separation. Eighteen proteins were successfully identified in the SWISS-PROT protein database using amino acid analysis of proteins recovered from polyvinylidene difluoride (PVDF) membranes and verified for one of them by comparison with Anderson's rat liver reference map. Fourteen new polypeptides were identified and four were previously known. Two of these new proteins were closely related to the spermatogenetic process. T-complex protein 1 is expressed in large amounts in germ cells. Androgen-dependent sperm-coating glycoprotein is secreted by epididymal cells. In order to detect changes in protein expression during meiosis and spermiogenesis, spermatocytes and round spermatid cell populations were purified by centrifugal elutriation and compared. In this way several proteins not found in the spermatocyte 2-D images could be high-lighted. The sperm-coating glycoprotein was thus shown to be present in large amounts in round spermatids.

Adolescent↗