Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Expanding the protein catalogue in the proteome reference map of human breast cancer cells.

In this report we present a catalogue of 162 proteins (including isoforms and variants) identified in a prototype of proteomic map of breast cancer cells. This work represents the prosecution of previous studies describing the protein complement of breast cancer cells of the line 8701-BC, which has been well characterized for several parameters, providing to be a useful model for the study of breast cancer-associated candidate biomarkers. In particular, 110 spots were identified ex novo by PMF, or validated following previous gel matching identification method; 30 were identified by N-terminal microsequencing and the remaining by gel matching with maps available from our former work. As a consequence of the expanded number of proteins, we have updated our previous classification extending the number of protein groups from 4 to 13. In order to facilitate comparative proteome studies of different kinds of breast cancers, in this report we provide the whole complement of proteins so far identified and grouped into the new classification. A consistent number of them were not described before in other proteomic maps of breast cancer cells or tissues, and therefore they represent a valuable contribution for breast cancer protein databases and for future application in basic and clinical researches.

Biomarkers, Tumor↗

Characterization of Borrelia lusitaniae sp. nov. by 16S ribosomal DNA sequence analysis.

We determined the complete sequence of the rrs gene from five strains of genomic species PotiB2. Both distance and parsimony methods were used to infer the evolutionary relationships of the rrs gene sequence of this genomic species in comparison with the rrs gene sequence of Borrelia valaisiana and the rrs gene sequences of Borrelia burgdorferi sensu lato species obtained from sequence databases. The phylogenetic analysis revealed that the genomic species PotiB2 strains clustered in a separate lineage, which was consistent with data from previous DNA-DNA hybridization experiments (D. Postic, M. V. Assous, P. A. D. Grimont, and G. Baranton, Int. J. Syst. Bacteriol. 44:743-752, 1994). A PCR-restriction fragment length polymorphism analysis was used to identify genomic species PotiB2 and to differentiate it from B. burgdorferi sensu lato species. Moreover, signature nucleotide positions were identified for each B. burgdorferi sensu lato species. In accordance with DNA relatedness values, our findings suggest that genomic species PotiB2 can be more clearly defined and identified, and we propose that it should be referred to as a new species, Borrelia lusitaniae. The type strain is PotiB2.

Bacterial Proteins↗

Precision parameters of methods of analysis required for nutrition labeling. Part I. Major nutrients.

Major components of foods and feeds are fat, protein, and carbohydrates. Fat and protein are determined by direct measurements that are interpreted as the quantity of the constituent. Carbohydrates are usually calculated by difference. For this calculation, values for moisture/solids, ash, and "fiber" are also needed. The readily available collaborative studies for the determination of these major components are reviewed in an attempt to assign precision parameters to validated methods of analysis. When a number of studies for the same analyte, in the same food, by the same method are available, it is seen that the precision parameters among laboratories (standard deviations, SR; relative standard deviations, RSDR) and the ISO maximum tolerable difference functions (repeatability value, r; reproducibility value, R) are not characterized by any conventional distribution. The precision data are best summarized as a median or average parameter and the interval containing the centermost 90% of reported values. Typically, the precision of methods of analysis can be expressed as a function of concentration only, independent of analyte, matrix, and method. The average RSDR value from each collaborative data set can then be used as the numerator in a ratio containing, as the denominator, the value calculated from the Horwitz equation: RSDR = 2 exp (1 - 0.5 log C) where C is the concentration as a decimal fraction. A series of ratios consistently above 1, and especially above 2, probably indicates that a method is unacceptable with respect to precision. By this criterion, only the protein (Kjeldahl) determination is unqualifiedly acceptable with a 90% interval for RSDR of 1 to 3% at C values above about 0.01 (1 g/100 g). Fat, moisture/solids, and ash are acceptable down to limiting concentrations in the region of 1 to 5 g/100 g, if a test portion large enough to provide at least 50 mg of weighable residue or volatiles is specified. Measurements of individual carbohydrates and fiber-related analytes have unexpectedly poor precisions among laboratories. The variability, although high, may still be suitable for nutrition labeling. Reliability of analyses for the control of labeling of the primary nutrients must be achieved through quality assurance programs that require strict adherence to the directions of empirical methods and the use of suitable reference materials for absolute methods.

Databases, Bibliographic↗

Hierarchical protein folding pathways: a computational study of protein fragments.

We have previously presented a building block folding model. The model postulates that protein folding is a hierarchical top-down process. The basic unit from which a fold is constructed, referred to as a hydrophobic folding unit, is the outcome of combinatorial assembly of a set of "building blocks." Results obtained by the computational cutting procedure yield fragments that are in agreement with those obtained experimentally by limited proteolysis. Here we show that as expected, proteins from the same family give very similar building blocks. However, different proteins can also give building blocks that are similar in structure. In such cases the building blocks differ in sequence, stability, contacts with other building blocks, and in their 3D locations in the protein structure. This result, which we have repeatedly observed in many cases, leads us to conclude that while a building block is influenced by its environment, nevertheless, it can be viewed as a stand-alone unit. For small-sized building blocks existing in multiple conformations, interactions with sister building blocks in the protein will increase the population time of the native conformer. With this conclusion in hand, it is possible to develop an algorithm that predicts the building block assignment of a protein sequence whose structure is unknown. Toward this goal, we have created sequentially nonredundant databases of building block sequences. A protein sequence can be aligned against these, in order to be matched to a set of potential building blocks.

Algorithms↗

The MetaCyc Database.

MetaCyc is a metabolic-pathway database that describes 445 pathways and 1115 enzymes occurring in 158 organisms. MetaCyc is a review-level database in that a given entry in MetaCyc often integrates information from multiple literature sources. The pathways in MetaCyc were determined experimentally, and are labeled with the species in which they are known to occur based on literature references examined to date. MetaCyc contains extensive commentary and literature citations. Applications of MetaCyc include pathway analysis of genomes, metabolic engineering and biochemistry education. MetaCyc is queried using the Pathway Tools graphical user interface, which provides a wide variety of query operations and visualization tools. MetaCyc is available via the World Wide Web at http://ecocyc.org/ecocyc/metacyc.html, and is available for local installation as a binary program for the PC and the Sun workstation, and as a set of flatfiles. Contact metacyc-info@ai.sri.com for information on obtaining a local copy of MetaCyc.

Database Management Systems↗

Informatics issues in large-scale sequence analysis: elucidating the protein kinases of C. elegans.

With the availability of the nearly complete genomic sequence of C. elegans, the first multicellular organism to be sequenced, molecular biology has definitely entered the postgenomic era. Annotation of the genomic sequence, which refers to identifying the genes and other biologically relevant sections of the genome, is an important and nontrivial next step. A first-pass annotation will be necessarily incomplete but will drive further biological experiments, which in turn will help to annotate the genome better. Given the scale of the genome sequence analysis, it is clear that the annotation should be automated as much as possible without sacrificing the quality of analysis. In this work, we outline our approach to identifying the protein kinases of C. elegans from the genomic sequence. We describe new tools we have developed for analysis, management and visualization of genomic data. By developing modular and scalable solutions, this study has provided a framework for future analysis of the Drosophila and human genomes.

Animals↗

Gene expression profiling of the rat superior olivary complex using serial analysis of gene expression.

The superior olivary complex (SOC) is an auditory brainstem region that represents a favourable system to study rapid neurotransmission and the maturation of neuronal circuits. Here we performed serial analysis of gene expression (SAGE) on the SOC in 60-day-old Sprague-Dawley rats to identify genes specifically important for its function and to create a transcriptome reference for the subsequent identification of age-related or disease-related changes. Sequencing of 31 035 tags identified 10 473 different transcripts. Fifty-seven per cent of the unique tags with a count greater than four were statistically more highly represented in the SOC than in the hippocampus. Among them were genes encoding proteins involved in energy supply, the glutamate/glutamine shuttle, and myelination. Approximately 80 plasma membrane transporters, receptors, channels, and vesicular transporters were identified, and 25% of them displayed a significantly higher expression level in the SOC than in the hippocampus. Some of the plasma membrane proteins were not previously characterized in the SOC, e.g. the purinergic receptor subunit P2X(6) and the metabotropic GABA receptor Gpr51. Differential gene expression between SOC and hippocampus was confirmed using RNA in situ hybridization or immunohistochemistry. The extensive gene inventory presented here will alleviate the dissection of the molecular mechanisms underlying specific SOC functions and the comparison with other SAGE libraries from brain will ease the identification of promoters to generate region-specific transgenic animals. The analysis will be part of the publicly available database ID-GRAB.

Animals↗

Molecular cloning and characterization of iojap (ij), a pattern striping gene of maize.

Iojap (ij) is a recessive striped mutant of maize affecting the development of plastids in a local and position-dependent manner on the leaves. The ij-affected plastids are transmitted to some of the progeny even when the function of the nuclear gene is restored. Developmental defects during embryogenesis and leaf proliferation are other phenotypic characteristics of ij. The extent of striping and the degree of developmental arrest in ij depend upon genetic background. To understand the diverse and unique phenotypic expression of ij, a transposon tagging experiment has been conducted using Robertson's Mutator (Mu). A new ij mutant was obtained from crosses of the reference allele of (ij-ref) to Mu lines. Subsequent genetic and molecular studies showed that the mutant carried a new ij allele (ij-mum1) from the Mu lines and contained a Mu1 element that cosegregated with the iojap phenotype. A 6.0 kb EcoRI genomic DNA fragment containing the Mu1 element was cloned. ij-ref is unstable, and revertants (Ij-Rev) have been obtained. Using the flanking DNA from the genomic clone as a probe, DNA polymorphisms were detected between ij-ref and these revertants. Further, transcripts were restored to the normal level in Ij-Rev seedlings. Comparison of genomic DNA clones from ij-ref, ij-mum1 and Ij indicated that the ij-ref allele contained 1.5 kb of additional DNA related to a transposable element, Ds. Germinal and somatic revertant alleles were derived by excision of this 1.5 kb element from ij-ref. The structure of the Ij gene and the DNA sequence of its transcribed region were determined. The Ij gene encodes a 24.8 kDa protein that showed no significant sequence similarity with proteins listed in databases.

Alleles↗

Identification and functional characterization of variants in human concentrative nucleoside transporter 3, hCNT3 (SLC28A3), arising from single nucleotide polymorphisms in coding regions of the hCNT3 gene.

INTRODUCTION: Human concentrative nucleoside transporter 3, hCNT3 (SLC28A3), which mediates transport of purine and pyrimidine nucleosides and a variety of antiviral and anticancer nucleoside drugs, was investigated to determine if there are single nucleotide polymorphisms in the coding regions of the hCNT3 gene. METHODS AND RESULTS: Ninety-six DNA samples from Caucasians (Coriell Panel) were sequenced and sixteen variants in exons and flanking intronic regions were identified, of which five were coding variants; three of these were non-synonymous (S5N, L131F, Y513F) and were further investigated for functional alterations of the resulting recombinant proteins in Saccharomyces cerevisiae and Xenopus laevis oocytes. In yeast, immunostaining and fluorescence quantitation of the reference (wild-type) and variant CNT3 proteins showed similar levels of expression. Kinetic studies were undertaken in yeast with a high through-put semi-automated assay process; reference hCNT3 exhibited Km values of 1.7+/-0.3, 3.6+/-1.3, 2.2+/-0.7, and 2.1+/-0.6 muM and Vmax values of 1402+/-286, 1310+/-113, 1020+/-44, and 1740+/-114 pmol/mg/min, respectively, for uridine, cytidine, adenosine and inosine. Similar Km and Vmax values were obtained for the three variant proteins assayed in yeast under identical conditions. All of the characterized hCNT3 variants produced in oocytes retained sodium and proton dependence of uridine transport based on measurements of radioisotope flux and two-electrode voltage-clamp studies. CONCLUSION: These results suggested a high degree of conservation of function for hCNT3 in the Caucasian population.

Adenosine↗

Analysis of the Arabidopsis nuclear proteome and its response to cold stress.

The nucleus is the subcellular organelle that contains nearly all the genetic information required for the regulated expression of cellular proteins. In this study, we comprehensively characterized the Arabidopsis nuclear proteome. Nuclear proteins were isolated and analyzed using two-dimensional (2D) gel electrophoresis and matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS). Approximately 500-700 spots were detected in reference 2D gels of nuclear proteins. Proteomic analyses led to the identification of 184 spots corresponding to 158 different proteins implicated in a variety of cellular functions. We additionally analyzed the changes in the nuclear proteome in response to cold stress. Of the 184 identified proteins, 54 were up- or downregulated with a greater than twofold change in response to cold treatment. Among these, six proteins were selected for further characterization. Northern analysis data revealed that gene expression of these proteins was also altered by cold stress. Following transient expression in BY-2 protoplasts, two proteins were detected in both the cytoplasm and the nucleus and four others were detected exclusively in the nucleus, which correlates well with the nuclear localization patterns of the proteomic data. Our study provides an initial insight into the Arabidopsis nuclear proteome and its response to cold stress.

Arabidopsis↗

The key role of atom types, reference states, and interaction cutoff radii in the knowledge-based method: new variational approach.

We present a variational method to derive knowledge-based potentials. The method is based on an optimization procedure of objective variables: atom types, reference states, and interaction cutoff radii. We suggest and apply new unsymmetrical reference states. The cutoff radii and atom types are optimized to improve docking accuracy of the corresponding potentials. The atom types are varied along an atom type tree, with 6 root and 49 top atom types, and the set of 18 optimal atom types is obtained. We demonstrate strong dependence between the choice of atom types and the docking accuracy of the potentials derived with these atom types. The averaged root-mean square deviations (RMSDs) of the ligand docked positions relative to the experimentally determined positions decrease when the elements C, N, O are split into the optimal types.

Algorithms↗

Molecular evolution of serine protease and its inhibitor with special reference to domain evolution.

The evolution of serine protease and its inhibitor are discussed with special reference to domain evolution. It is now known that most proteins are composed of more than one functional domain. Because serine proteases such as urokinase and plasminogen are made of various functional domains, these proteins are typical examples of the so-called mosaic proteins. When Kringle domains in serine proteases and a Kunitz-type protease inhibitor domain in the amyloid beta precursor protein in Alzheimer's disease patients were examined by the molecular evolutionary analysis, the phylogenetic trees constructed showed that these functional domains had undergone dynamic changes in the evolutionary process. In particular, these domains are evolutionarily movable. Thus, it is concluded that various functional domains evolved independently of each other and that they have been shuffled to create the existent mosaic proteins. This conclusion leads us to the reasonable speculation that those functional domains must have been minigenes possibly at the time of primordial life or the origin of life. We call these minigenes 'ancestral minigenes'. Every effort should be made to answer the question about the minimum set of ancestral minigenes that must have existed and must have been needed for maintaining life forms. The DNA sequence database is useful for making attempts to answer such difficult but significant questions.

Alzheimer Disease↗

Two-dimensional gel electrophoresis maps of the proteome and phosphoproteome of primitively cultured rat mesangial cells.

Mesangial cells (MC) play an important role in maintaining the structure and function of the glomerulus. The proliferation of MC is a prominent feature of many kinds of glomerular disease. The first reference 2-DE maps of rat mesangial cells (RMC), stained with silver staining or Pro-Q Diamond dye, have been established here to describe the proteome and phosphoproteome of RMC, respectively. A total of 157 selected protein spots, corresponding to 118 unique proteins, have been identified by MALDI-TOF-MS or LC-ESI-IT-MS/MS, in which 37 protein spots representing 28 unique proteins have also been stained with Pro-Q Diamond, indicating that they are in phosphorylated forms. All the identified proteins were bioinformatically annotated in detail according to their physiochemical characteristics, subcellular location, and function. Most of the separated or identified protein spots are distributed in the area of mass 10-70 kDa and pI 5.0-8.0. The identified proteins include mainly cytoplasmic and nuclear proteins and some mitochondrial, endoplasmic reticulum, and membrane proteins. These proteins are classified into different functional groups such as structure and mobility proteins (21.2%), metabolic enzymes (16.9%), protein folding and metabolism proteins (13.6%), signaling proteins (14.4%), heat-shock proteins (7.6%), and other functional proteins (12.7%). While structure and mobility proteins are mostly represented by protein spots with high abundance, signaling proteins are mostly represented by protein spots with relatively low abundance. Such a 2-DE database for RMC, especially with many signaling proteins and phosphoproteins characterized, will provide a valuable resource for comparative proteomics analysis of normal and pathologic conditions affecting MC function or pathologic progress.

Animals↗

CFinder: locating cliques and overlapping modules in biological networks.

UNLABELLED: Most cellular tasks are performed not by individual proteins, but by groups of functionally associated proteins, often referred to as modules. In a protein association network modules appear as groups of densely interconnected nodes, also called communities or clusters. These modules often overlap with each other and form a network of their own, in which nodes (links) represent the modules (overlaps). We introduce CFinder, a fast program locating and visualizing overlapping, densely interconnected groups of nodes in undirected graphs, and allowing the user to easily navigate between the original graph and the web of these groups. We show that in gene (protein) association networks CFinder can be used to predict the function(s) of a single protein and to discover novel modules. CFinder is also very efficient for locating the cliques of large sparse graphs. AVAILABILITY: CFinder (for Windows, Linux and Macintosh) and its manual can be downloaded from http://angel.elte.hu/clustering. SUPPLEMENTARY INFORMATION: Supplementary data are available on Bioinformatics online.

Biology↗

HSP60 gene sequences as universal targets for microbial species identification: studies with coagulase-negative staphylococci.

A set of universal degenerate primers which amplified, by PCR, a 600-bp oligomer encoding a portion of the 60-kDa heat shock protein (HSP60) of both Staphylococcus aureus and Staphylococcus epidermidis were developed. However, when used as a DNA probe, the 600-bp PCR product generated from S. epidermidis failed to cross-hybridize under high-stringency conditions with the genomic DNA of S. aureus and vice versa. To investigate whether species-specific sequences might exist within the highly conserved HSP60 genes among different staphylococci, digoxigenin-labelled HSP60 probes generated by the degenerate HSP60 primers were prepared from the six most commonly isolated Staphylococcus species (S. aureus 8325-4, S. epidermidis 9759, S. haemolyticus ATCC 29970, S. schleiferi ATCC 43808, S. saprophyticus KL122, and S. lugdunensis CRSN 850412). These probes were used for dot blot hybridization with genomic DNA of 58 reference and clinical isolates of Staphylococcus and non-Staphylococcus species. These six Staphylococcus species HSP60 probes correctly identified the entire set of staphylococcal isolates. The species specificity of these HSP60 probes was further demonstrated by dot blot hybridization with PCR-amplified DNA from mixed cultures of different Staphylococcus species and by the partial DNA sequences of these probes. In addition, sequence homology searches of the NCBI BLAST databases with these partial HSP60 DNA sequences yielded the highest matching scores for both S. epidermidis and S. aureus with the corresponding species-specified probes. Finally, the HSP60 degenerate primers were shown to amplify an anticipated 600-bp PCR product from all 29 Staphylococcus species and from all but 2 of 30 other microbial species, including various gram-positive and gram-negative bacteria, mycobacteria, and fungi. These preliminary data suggest the presence of species-specific sequence variation within the highly conserved HSP60 genes of staphylococci. Further work is required to determine whether these degenerate HSP60 primers may be exploited for species-specific microbic identification and phylogenetic investigation of staphylococci and perhaps other microorganisms in general.

Base Sequence↗

Infectious pancreatic necrosis virus in Atlantic salmon, Salmo salar L., post-smolts in the Shetland Isles, Scotland: virus identification, histopathology, immunohistochemistry and genetic comparison with Scottish mainland isolates.

During mid-June 1999 peak mortalities of 11% of the total stock per week were seen at a sea cage site of Atlantic salmon, Salmo salar L., post-smolts in the Shetland Isles, Scotland. Virus was isolated on chinook salmon embryo (CHSE) cells in a standard diagnostic test and infectious pancreatic necrosis virus (IPNV) identified by enzyme-linked immunosorbent assay. IPNV was confirmed as serogroup A by a cell immunofluorescent antibody test using the cross-reactive monoclonal antibody AS-1. Four weeks after the main outbreak, virus titres in surviving moribund fish were assayed at >10(10) TCID50 g(-1) kidney. Histopathology of moribund fish was characterized by pancreatic acinar cell necrosis and a marked catarrhal enteritis of the intestinal mucosa. In the liver, necrosis, leucocytic infiltration and a generalized cell vacuolation were noted. IPNV-specific immunostaining was demonstrated in pancreas, liver, heart, gill and kidney tissue. The nucleotide sequence of the coding region of segment A was determined from the Shetland isolate. A 1180 bp fragment of the VP2 gene of this isolate was compared with a 1979 reference isolate from mainland Scottish Atlantic salmon, La/79 and another more recent mainland isolate, 432/00. Both A2 isolates were derived from carrier fish without signs of IPN and serotyped by a plaque neutralization test. The Shetland isolate shows a different nucleotide and amino acid sequence compared with the two isolates from carrier fish. These latter isolates showed identical amino acid sequences in the fragment examined, despite the 21 years separating the isolations. Sequence comparisons with other A2 (Sp) isolates on the database confirm all three Scottish isolates are A2 (Sp).

Animals↗

NCI Thesaurus: a semantic model integrating cancer-related clinical and molecular information.

Over the last 8 years, the National Cancer Institute (NCI) has launched a major effort to integrate molecular and clinical cancer-related information within a unified biomedical informatics framework, with controlled terminology as its foundational layer. The NCI Thesaurus is the reference terminology underpinning these efforts. It is designed to meet the growing need for accurate, comprehensive, and shared terminology, covering topics including: cancers, findings, drugs, therapies, anatomy, genes, pathways, cellular and subcellular processes, proteins, and experimental organisms. The NCI Thesaurus provides a partial model of how these things relate to each other, responding to actual user needs and implemented in a deductive logic framework that can help maintain the integrity and extend the informational power of what is provided. This paper presents the semantic model for cancer diseases and its uses in integrating clinical and molecular knowledge, more briefly examines the models and uses for drug, biochemical pathway, and mouse terminology, and discusses limits of the current approach and directions for future work.

Biomedical Research↗

Quantification of the variation in percentage identity for protein sequence alignments.

BACKGROUND: Percentage Identity (PID) is frequently quoted in discussion of sequence alignments since it appears simple and easy to understand. However, although there are several different ways to calculate percentage identity and each may yield a different result for the same alignment, the method of calculation is rarely reported. Accordingly, quantification of the variation in PID caused by the different calculations would help in interpreting PID values in the literature. In this study, the variation in PID was quantified systematically on a reference set of 1028 alignments generated by comparison of the protein three-dimensional structures. Since the alignment algorithm may also affect the range of PID, this study also considered the effect of algorithm, and the combination of algorithm and PID method. RESULTS: The maximum variation in PID due to the calculation method was 11.5% while the effect of alignment algorithm on PID was up to 14.6% across three popular alignment methods. The combined effect of alignment algorithm and PID calculation gave a variation of up to 22% on the test data, with an average of 5.3% +/- 2.8% for sequence pairs with < 30% identity. In order to see which PID method was most highly correlated with structural similarity, four different PID calculations were compared to similarity scores (Sc) from the comparison of the corresponding protein three-dimensional structures. The highest correlation coefficient for a PID calculation was 0.80. In contrast, the more sophisticated Z-score calculated by reference to randomized sequences gave a correlation coefficient of 0.84. CONCLUSION: Although it is well known amongst expert sequence analysts that PID is a poor score for discriminating between protein sequences, the apparent simplicity of the percentage identity score encourages its widespread use in establishing cutoffs for structural similarity. This paper illustrates that not only is PID a poor measure of sequence similarity when compared to the Z-score, but that there is also a large uncertainty in reported PID values. Since better alternatives to PID exist to quantify sequence similarity, these should be quoted where possible in preference to PID. The findings presented here should prove helpful to those new to sequence analysis, and in warning those who seek to interpret the value of a PID reported in the literature.

Algorithms↗