Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Identification of garlic in old gildings by gas chromatography-mass spectrometry.

The proteinaceous content of garlic (Allium sativum) was characterised according to its amino acid composition by using a gas chromatography-mass spectrometry (GC-MS) analytical procedure. The procedure was tested on fresh and aged garlic samples as well as on reference gilding specimens prepared according to old recipes. The proteinaceous pattern showed a characteristic distribution of amino acids with glutamic acid being the major component. The average amino acidic composition was: glutamic acid (Glu; 29%), aspartic acid (Asp; 17%), serine (Ser; 11%), alanine, glycine, valine, leucine, lysine and phenylalanine (Ala, Gly, Val, Leu, Lys and Phe; 5-6%), isoleucine, proline and tyrosine (Ile, Pro and Tyr; 2-3%), methionine and hydroxyproline (Met and Hyp; 0.5%). In order to distinguish this material from animal glue and egg, which are the other proteinaceous media commonly used in gilding techniques, a database of amino acid percentages of the three proteins was built up and submitted to principal component analysis. Three separate clusters were obtained, allowing the protein identification. The application of the procedure on several gilding samples from Italian wall and easel paintings (13th-17th century) permitted to evidence the use of garlic as a gluing agent.

Garlic↗

E-MSD: an integrated data resource for bioinformatics.

The Macromolecular Structure Database (MSD) group (http://www.ebi.ac.uk/msd/) continues to enhance the quality and consistency of macromolecular structure data in the worldwide Protein Data Bank (wwPDB) and to work towards the integration of various bioinformatics data resources. One of the major obstacles to the improved integration of structural databases such as MSD and sequence databases like UniProt is the absence of up to date and well-maintained mapping between corresponding entries. We have worked closely with the UniProt group at the EBI to clean up the taxonomy and sequence cross-reference information in the MSD and UniProt databases. This information is vital for the reliable integration of the sequence family databases such as Pfam and Interpro with the structure-oriented databases of SCOP and CATH. This information has been made available to the eFamily group (http://www.efamily.org.uk/) and now forms the basis of the regular interchange of information between the member databases (MSD, UniProt, Pfam, Interpro, SCOP and CATH). This exchange of annotation information has enriched the structural information in the MSD database with annotation from wider sequence-oriented resources. This work was carried out under the 'Structure Integration with Function, Taxonomy and Sequences (SIFTS)' initiative (http://www.ebi.ac.uk/msd-srv/docs/sifts) in the MSD group.

Amino Acid Sequence↗

Cell wall proteomics of the green alga Haematococcus pluvialis (Chlorophyceae).

The green microalga Haematococcus pluvialis can synthesize and accumulate large amounts of the ketocarotenoid astaxanthin, and undergo profound changes in cell wall composition and architecture during the cell cycle and in response to environmental stresses. In this study, cell wall proteins (CWPs) of H. pluvialis were systematically analyzed by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) coupled with peptide mass fingerprinting (PMF) and sequence-database analysis. In total, 163 protein bands were analyzed, which resulted in positive identification of 81 protein orthologues. The highly complex and dynamic composition of CWPs is manifested by the fact that the majority of identified CWPs are differentially expressed at specific stages of the cell cycle along with a number of common wall-associated 'housekeeping' proteins. The detection of cellulose synthase orthologue in the vegetative cells suggested that the biosynthesis of cellulose occurred during primary wall formation, in contrast to earlier observations that cellulose was exclusively present in the secondary wall of the organism. A transient accumulation of a putative cytokinin oxidase at the early stage of encystment pointed to a possible role in cytokinin degradation while facilitating secondary wall formation and/or assisting in cell expansion. This work represents the first attempt to use a proteomic approach to investigate CWPs of microalgae. The reference protein map constructed and the specific protein markers obtained from this study provide a framework for future characterization of the expression and physiological functions of the proteins involved in the biogenesis and modifications in the cell wall of Haematococcus and related organisms.

Cell Cycle↗

Eight cDNA encoding putative aquaporins in Vitis hybrid Richter-110 and their differential expression.

The nucleotide sequences of eight cDNAs encoding putative aquaporins obtained from a leaf Vitis hybrid Richter-110 cDNA library are reported. They encode proteins ranging from 249 to 287 amino acids with characteristic sequences that clearly include them within the MIP family. According to available database sequence homologies, they can be classified into four groups belonging to two subfamilies: PIP (PIP1 and PIP2) and TIP (gamma-TIP and delta-TIP). In order to elucidate the expression patterns of these putative aquaporins in the plant, specific probes were developed and tissue specific differential expression was tested by reverse Northern and compared with two reference genes (malic enzyme and glutamate dehydrogenase). Clearly, most of the putative aquaporins had higher expression in roots, whereas expression in shoot and leaves was generally weaker than the reference genes.

Aquaporins↗

Column chromatographic prefractionation leads to the detection of 543 different gene products in human fetal brain.

In a previous publication a large series of proteins were identified in fetal human brain by the use of two-dimensional electrophoresis (2-DE) with subsequent matrix-assisted laser desorption/ionization-time of flight (MALDI-TOF) and MALDI-tandem time-of-flight (TOF/TOF) analysis. Further identification of many more different spots by traditional 2-DE without additional step such as narrow immobilized ph gradient (IPG) strips or prefractionation seems unlikely and we therefore decided to separate extracted brain proteins by ion-exchange chromatography using a TSK gel DEAE-5PW column followed by 2-DE of individual fractions and analysis by MALDI-TOF/TOF with LIFT technology in fetal brain of the early second trimester. About 1880 protein spots corresponding to 543 different gene products were identified. These proteins included housekeeping, signaling, cytoskeletal, metabolic, antioxidant, and neuron/synaptosomal specific proteins. Among these, 314 gene products (314/543, 57.8%), which have never been detected in traditional 2-DE of human fetal brain, were observed by this method. This updated map of fetal brain proteins may serve as data base and reference map for fetal brain proteins, and the methodology applied may be used as a valuable analytical tool for the basis of protein expressional studies in health and disease.

Brain↗

BRENDA, AMENDA and FRENDA: the enzyme information system in 2007.

The BRENDA (BRaunschweig ENzyme DAtabase) enzyme information system (http://www.brenda.uni-koeln.de) is the largest publicly available enzyme information system worldwide. The major parts of its contents are manually extracted from primary literature. It is not restricted to specific groups of enzymes, but includes information on all identified enzymes irrespective of the enzyme's source. The range of data encompasses functional, structural, sequence, localisation, disease-related, isolation, stability information on enzyme and ligand-related data. Each single entry is linked to the enzyme source and to a literature reference. Recently the data repository was complemented by text-mining data in AMENDA (Automatic Mining of ENzyme DAta) and FRENDA (Full Reference ENzyme DAta). A genome browser, membrane protein prediction and full-text search capacities were added. The newly implemented web service provides instant access to the data for programmers via a SOAP (Simple Object Access Protocol) interface. The BRENDA data can be downloaded in the form of a text file from the beginning of 2007.

Animals↗

An ontology for pharmaceutical ligands and its application for in silico screening and library design.

Annotation efforts in biosciences have focused in past years mainly on the annotation of genomic sequences. Only very limited effort has been put into annotation schemes for pharmaceutical ligands. Here we propose annotation schemes for the ligands of four major target classes, enzymes, G protein-coupled receptors (GPCRs), nuclear receptors (NRs), and ligand-gated ion channels (LGICs), and outline their usage for in silico screening and combinatorial library design. The proposed schemes cover ligand functionality and hierarchical levels of target classification. The classification schemes are based on those established by the EC, GPCRDB, NuclearDB, and LGICDB. The ligands of the MDL Drug Data Report (MDDR) database serve as a reference data set of known pharmacologically active compounds. All ligands were annotated according to the schemes when attribution was possible based on the activity classification provided by the reference database. The purpose of the ligand-target classification schemes is to allow annotation-based searching of the ligand database. In addition, the biological sequence information of the target is directly linkable to the ligand, hereby allowing sequence similarity-based identification of ligands of next homologous receptors. Ligands of specified levels can easily be retrieved to serve as comprehensive reference sets for cheminformatics-based similarity searches and for design of target class focused compound libraries. Retrospective in silico screening experiments within the MDDR01.1 database, searching for structures binding to dopamine D2, all dopamine receptors and all amine-binding class A GPCRs using known dopamine D2 binding compounds as a reference set, have shown that such reference sets are in particular useful for the identification of ligands binding to receptors closely related to the reference system. The potential for ligand identification drops with increasing phylogenetic distance. The analysis of the focus of a tertiary amine based combinatorial library compared to known amine binding class A GPCRs, peptide binding class A GPCRs, and LGIC ligands constitutes a second application scenario which illustrates how the focus of a combinatorial library can be treated quantitatively. The provided annotation schemes, which bridge chem- and bioinformatics by linking ligands to sequences, are expected to be of key utility for further systematic chemogenomics exploration of previously well explored target families.

Combinatorial Chemistry Techniques↗

A two-dimensional gel database of rat liver proteins useful in gene regulation and drug effects studies.

A standard two-dimensional (2-D) protein map of Fischer 344 rat liver (F344MST3) is presented, with a tabular listing of more than 1200 protein species. Sodium dodecyl sulfate (SDS) molecular mass and isoelectric point have been established, based on positions of numerous internal standards. This map has been used to connect and compare hundreds of 2-D gels of rat liver samples from a variety of studies, and forms the nucleus of an expanding database describing rat liver proteins and their regulation by various drugs and toxic agents. An example of such a study, involving regulation of cholesterol synthesis by cholesterol-lowering drugs and a high-cholesterol diet, is presented. Since the map has been obtained with a widely used and highly reproducible 2-D gel system (the Iso-Dalt system), it can be directly related to an expanding body of work in other laboratories.

Animals↗

A nanovirus-like DNA component associated with yellow vein disease of Ageratum conyzoides: evidence for interfamilial recombination between plant DNA viruses.

Yellow vein disease of Ageratum conyzoides, a weed species that is widely distributed throughout Asia, has been attributed to infection by the geminivirus Ageratum yellow vein virus (AYVV). In addition to a single AYVV genomic component (DNA A), we have previously demonstrated that infected plants contain chimeric defective viral components, comprising DNA A and nongeminiviral sequences, that act as defective interfering DNAs. A database search has revealed that the nongeminiviral sequences of one such defective component (def19) show significant homology with sequences of nanovirus components that encode replication-associated proteins (Reps). Primers designed to hybridise to the nongeminiviral DNA were used to PCR-amplify a full-length nanovirus-like component, referred to as DNA 1, from an extract of infected A. conyzoides. DNA 1 is unrelated to AYVV DNA A but resembles nanovirus components that encode Reps and is most closely related (73% identity) to a nanovirus-like DNA recently isolated from geminivirus-infected cotton. DNA 1 is dependent on AYVV DNA A for systemic infection of A. conyzoides and Nicotiana benthamiana and can systemically infect N. benthamiana in the presence of the bipartite geminivirus African cassava mosaic virus. A. conyzoides plants coinfected with AYVV DNA A and DNA 1 remain asymptomatic, indicating that additional factors are required to elicit yellow vein disease. Our results provide direct evidence for recombination between distinct families of plant single-stranded DNA viruses and suggest that coinfection by geminivirus and nanovirus-like pathogens may be a widespread phenomenon. The ability of plant DNA viruses to recombine in this way may greatly increase their scope for diversification.

Amino Acid Sequence↗

Survey of bacterial proteins released in cheese: a proteomic approach.

During the ripening of Emmental cheese, the bacterial ecosystem confers its organoleptic characteristics to the evolving curd both by the action of the living cells, and through the release of numerous proteins, including various types of enzymes into the cheese when the cells lyse. In Emmental cheese these proteins can be released from thermophilic lactic acid bacteria used as starters like Lactobacillus helveticus, Lb delbruecki subsp. lactis and Streptococcus salivarius subsp. thermophilus and ripening bacteria such as Propionibacterium freudenreichii. The aim of this study was to obtain a proteomic view of the different groups of proteins within the cheese using proteomic tools to create a reference map. A methodology was therefore developed to reduce the complexity of cheese matrix prior to 2D-PAGE analysis. The aqueous phase of cheese was prefractionated by size exclusion chromatography, bacterial and milk proteins were separated and subsequently characterised by mass spectrometry, prior to peptide mass fingerprint and sequence homology database search. Five functional groups of proteins were identified involved in: (i) proteolysis, (ii) glycolysis, (iii) stress response, (iv) DNA and RNA repair and (v) oxidoreduction. The results revealed stress responses triggered by thermophilic lactic acid bacteria and Propionibacterium strains at the end of ripening. Information was also obtained regarding the origin and nature of the peptidases released into the cheese, thus providing a greater understanding of casein degradation mechanisms during ripening. Different peptidases arose from St thermophilus and Lb helveticus, suggesting that streptococci are involved in peptide degradation in addition to the proteolytic activity of lactobacilli.

Autolysis↗

The effect of anemia treatment on selected health-related quality-of-life domains: a systematic review.

BACKGROUND: Anemia is a reduction in the oxygen-carrying capacity of red blood cells that results in a variety of symptoms, including dyspnea, headaches, light-headedness, and fatigue. Although anemia has been associated with reduced health-related quality of life (HRQoL), its treatment has not yet been consistently shown to improve HRQoL. OBJECTIVE: This systematic review of the literature was conducted to determine whether the treatment of anemia improves HRQoL domains, regardless of the type of underlying disease. METHODS: Data for this review were drawn from the clinical trial databases from 2 previous systematic literature reviews of erythropoiesis-stimulating protein treatment for renal insufficiency- and cancer-related anemia, both spanning the period January 1, 1980, through December 31, 2001. MEDLINE, Cancerlit, and Current Contents/Clinical Medicine were searched using the combined terms erythropoietin, kidney failure, neoplasms, and anemia. The reference lists of all identified articles were searched manually for additional relevant papers. The review included prospective studies that reported both HRQoL and hematocrit (Hct) in patients with cancer or renal insufficiency who received treatment for anemia with an erythropoiesis-stimulating protein. HRQoL was categorized by domain (overall, energy/fatigue, physical, activity); changes in HRQoL domains were expressed as effect sizes and meta-analyzed, as were correlation coefficients. The effects on HRQoL of dropout rate, study duration, baseline Hct, and change in Hct were examined in meta-regression analyses. RESULTS: Sixteen studies each were identified in patients with renal insufficiency (N = 2253) and patients with cancer (N = 10,695). The treated groups included 11,710 patients, and the control groups included 1238 patients. The baseline Hct in all treated groups averaged 26.0%: 28.3% in the group with cancer and 24.4% in the group with renal insufficiency. The mean improvement in Hct from baseline to the end of treatment was 8.3% (range, 1.0%-16.5%) in treated patients and 1.0% (range, 0.0%-3.3%) in controls. The Hct changes were similar in treated patients with cancer and treated patients with renal insufficiency, as was the HRQoL effect size (0.43). Dropout rate and study duration were not significant predictors of HRQoL changes, but change in Hct was a significant predictor in both conditions. Meta-analysis of the correlation coefficients, adjusting for HRQoL domains, showed a consistent and significant positive correlation between change in Hct and change in HRQoL (P < 0.001). CONCLUSION: The consistency in both direction and magnitude of effect across many studies and thousands of patients supports the hypothesis that treatment of anemia with erythropoiesis-stimulating protein improves selected HRQoL domains in patients with renal insufficiency- or cancer-related anemia.

Anemia↗

Allergen databases.

Allergies represent a significant medical and industrial problem. Molecular and clinical data on allergens are growing exponentially and in this article we have reviewed nine specialized allergen databases and identified data sources related to protein allergens contained in general purpose molecular databases. An analysis of allergens contained in public databases indicates a high level of redundancy of entries and a relatively low coverage of allergens by individual databases. From this analysis we identify current database needs for allergy research and, in particular, highlight the need for a centralized reference allergen database.

Allergens↗

Genome sequences of two closely related Vibrio parahaemolyticus phages, VP16T and VP16C.

Two bacteriophages of an environmental isolate of Vibrio parahaemolyticus were isolated and sequenced. The VP16T and VP16C phages were separated from a mixed lysate based on plaque morphology and exhibit 73 to 88% sequence identity over about 80% of their genomes. Only about 25% of their predicted open reading frames are similar to genes with known functions in the GenBank database. Both phages have cos sites and open reading frames encoding proteins closely related to coliphage lambda's terminase protein (the large subunit). Like in coliphage lambda and other siphophages, a large operon in each phage appears to encode proteins involved in DNA packaging and capsid assembly and presumably in host lysis; we refer to this as the structural operon. In addition, both phages have open reading frames closely related to genes encoding DNA polymerase and helicase proteins. Both phages also encode several putative transcription regulators, an apparent polypeptide deformylase, and a protein related to a virulence-associated protein, VapE, of Dichelobacter nodosus. Despite the similarity of the proteins and genome organization, each of the phages also encodes a few proteins not encoded by the other. We did not identify genes closely related to genes encoding integrase proteins belonging to either the tyrosine or serine recombinase family, and we have no evidence so far that these phages can lysogenize the V. parahaemolyticus strain 16 host. Surprisingly for active lytic viruses, the two phages have a codon usage that is very different than that of the host, suggesting the possibility that they may be relative newcomers to growth in V. parahaemolyticus. The DNA sequences should allow us to characterize the lifestyles of VP16T and VP16C and the interactions between these phages and their host at the molecular level, as well as their relationships to other marine and nonmarine phages.

Bacteriophages↗

Molecular biologist's guide to proteomics.

The emergence of proteomics, the large-scale analysis of proteins, has been inspired by the realization that the final product of a gene is inherently more complex and closer to function than the gene itself. Shortfalls in the ability of bioinformatics to predict both the existence and function of genes have also illustrated the need for protein analysis. Moreover, only through the study of proteins can posttranslational modifications be determined, which can profoundly affect protein function. Proteomics has been enabled by the accumulation of both DNA and protein sequence databases, improvements in mass spectrometry, and the development of computer algorithms for database searching. In this review, we describe why proteomics is important, how it is conducted, and how it can be applied to complement other existing technologies. We conclude that currently, the most practical application of proteomics is the analysis of target proteins as opposed to entire proteomes. This type of proteomics, referred to as functional proteomics, is always driven by a specific biological question. In this way, protein identification and characterization has a meaningful outcome. We discuss some of the advantages of a functional proteomics approach and provide examples of how different methodologies can be utilized to address a wide variety of biological problems.

Amino Acid Sequence↗

Key residues approach to the definition of protein families and analysis of sparse family signatures.

We extend the concept of the motif as a tool for characterizing protein families and explore the feasibility of a sparse "motif" that is the length of the protein sequence itself. The type of motif discussed is a sparse family signature consisting of a set of N key residue positions (A1, A2...AN) preceded by gaps (G) thus G1A1G2A2. ...GNAN. Both a residue and gap can be variable. A signature is matched to a protein sequence and scored using a dynamic programming algorithm which permits variability in gap distance and residue type. Generating a signature involves identifying residues associated with points of contact in interactions between secondary structure elements. A raw signature consists of a set of positions with potential key structural roles sampled from a sequence alignment constructed with reference to this contact data. Raw signatures are refined by sampling different gap-residue pairs until the specificity of a signature for the family cannot be further improved. We summarize signatures for nine families of protein of diverse fold and function and present results of scans against the OWL protein sequence database. The implications of such signatures are discussed.

Algorithms↗

Characterization of the gene expression profile of neuroblastoma cell line IMR-5 using serial analysis of gene expression.

The serial analysis of gene expression (SAGE) technique was used to generate a database of the most abundant transcripts of the MYCN-amplified neuroblastoma cell line IMR-5. A total of 8568 tags were sequenced and shown to represent 4034 unique tags, each of which corresponds to an individual transcript. Expression levels of genes are reflected by the frequency of occurrence of the respective tags. To validate fidelity of SAGE data, relative abundances of seven transcripts were evaluated by semiquantitative reverse transcriptase-polymerase chain reaction. Transcripts that were detected nine times or more (>0.1% of the total tag population) accounted for 36% of the total messenger RNA mass but only 3% of the total number of individual transcripts. A strong preponderance of genes involved in protein synthesis, in particular those encoding for ribosomal proteins, were observed among these high-abundance transcripts. Tags corresponding to the amplified gene DDX1 were conspicuously overrepresented in comparison to the other amplified genes MYCN, neuroblastoma amplified gene and MEIS1, which suggests an additional mechanism apart from genomic amplification contributing to the strong upregulation of this gene. This study provides a comprehensive gene expression profile of neuroblastoma cell line IMR-5 and may be used as a reference database for identification of candidate genes that are involved in etiology and pathogenesis of neuroblastoma.

Databases as Topic↗

Prediction of protein homo-oligomer types by pseudo amino acid composition: Approached with an improved feature extraction and Naive Bayes Feature Fusion.

The interaction of non-covalently bound monomeric protein subunits forms oligomers. The oligomeric proteins are superior to the monomers within the scope of functional evolution of biomacromolecules. Such complexes are involved in various biological processes, and play an important role. It is highly desirable to predict oligomer types automatically from their sequence. Here, based on the concept of pseudo amino acid composition, an improved feature extraction method of weighted auto-correlation function of amino acid residue index and Naive Bayes multi-feature fusion algorithm is proposed and applied to predict protein homo-oligomer types. We used the support vector machine (SVM) as base classifiers, in order to obtain better results. For example, the total accuracies of A, B, C, D and E sets based on this improved feature extraction method are 77.63, 77.16, 76.46, 76.70 and 75.06% respectively in the jackknife test, which are 6.39, 5.92, 5.22, 5.46 and 3.82% higher than that of G set based on conventional amino acid composition method with the same SVM. Comparing with Chou's feature extraction method of incorporating quasi-sequence-order effect, our method can increase the total accuracy at a level of 3.51 to 1.01%. The total accuracy improves from 79.66 to 80.83% by using the Naive Bayes Feature Fusion algorithm. These results show: 1) The improved feature extraction method is effective and feasible, and the feature vectors based on this method may contain more protein quaternary structure information and appear to capture essential information about the composition and hydrophobicity of residues in the surface patches that buried in the interfaces of associated subunits; 2) Naive Bayes Feature Fusion algorithm and SVM can be referred as a powerful computational tool for predicting protein homo-oligomer types.

Algorithms↗

Pair-preferences: a quantitative measure of regularities in protein sequences.

We present here the results obtained by applying several different methods to quantitatively measure regularities in protein sequences based on pair-preferences. We have studied the distribution of amino acid residues, singly as well as in pairs in a large data base and have attempted this task. We confirmed the existence of well-defined pair-preferences in proteins which were shown to be remarkably absent in simulated random sequences of similar amino acid distribution. The analysis of the sequences from the SWISS-PROT data base using simple statistical tests. Fourier analysis, fractal analysis and statistical thermodynamical tests were used to derive parameters to define a natural sequence. As a consequence of the existence of pair-preferences, parameters like fractal dimension (D), spectral exponent (beta), scaling parameter (H) and entropy (statistical) were found to be characteristic for natural sequences. For a reference state we chose a randomised state devoid of any pair-preference. The pair-preferences qualified well to be used as quantitative measures of regularities in protein sequences.

Algorithms↗