Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

PACSIN, a brain protein that is upregulated upon differentiation into neuronal cells.

To identify genes that are differentially expressed during self-repair processes in mouse brain, we screened a subtracted cDNA library enriched for brain-specific clones. One of these clones, H74, detected a 4.4-kb mRNA predominantly expressed in brain and dorsal root ganglia neurons. Expression increased continuously during the lifespan and the state of differentiation, but decreased after entorhinal-cortex lesion. A full-length cDNA clone was isolated from a cerebellum cDNA library and characterized. Sequence analysis and database search revealed high sequence similarity to FAP52, a protein expressed in focal-adhesion contacts, and uncharacterized Echinococcus and Caenorhabditis elegans gene products. Furthermore, peptide sequences derived from human cDNA fragments showed up to 65% sequence identity at the amino acid level. The presence of a C-terminal src homology 3 (SH3) domain and its phosphorylation by casein kinase 2 (CK2) and protein kinase C (PKC) imply a role in signaling. Here we demonstrate that the gene encodes a phosphoprotein, referred to as PACSIN, with a restricted spatial and temporal expression pattern.

Adaptor Proteins, Signal Transducing↗

The RITE assay: identifying effectors that target the transcription machinery using phage display technology.

We describe an approach using phage display to identify effectors (activators and repressors) of transcription based on the particular component of the general transcription machinery that they target. We refer to this approach as the reverse identification of transcriptional effectors (RITE) assay. A library of phages containing cDNA-encoded peptides displayed on their surfaces is screened using as the target a specific region of one of the general transcription factors (e.g., the C terminus of hTAFII135). The amino acid sequence encoded by the cDNA of an interacting phage is determined and analyzed in a database homology search to identify known or novel factors that may interact with the target protein. Candidate effectors from the homology search are synthesized from recombinant clones and tested for their abilities to bind to the target protein and to functionally modulate transcription in vivo when co-expressed with the transcriptional target protein. Because the RITE assay is a direct measure of the interactions between general transcription proteins and their effectors, it has an advantage over the well-known yeast two-hybrid system, which is not amenable to identifying transcription factor interactions.

Amino Acid Sequence↗

Identification and molecular characterization of a novel Bacillus strain capable of degrading Tween-80.

A Tween-80-degrading novel marine Bacillus strain, N10, has recently been isolated in Alexandria University, Egypt. The taxonomic position of this endospore forming bacterium was investigated on the basis of fatty acid analysis and 16S rRNA gene sequencing. Comparative computer database analyses revealed that the bacterium is a Bacillus subtilis strain. The gene encoding the small acid-soluble protein gamma-type (SASP-B), sspE, was successfully utilized in this study as a tool for discrimination between the two B. subtilis subspecies W23 and 168. Based on the alignment of 16S rRNA sequences and analysis of SASP-B relatedness, it has been demonstrated that the novel marine B. subtilis strain N10 is more closely related to the B. subtilis reference strain W23 than to 168. The strain, N10, has been deposited in the Bacillus Genetic Stock Center (BGSC) and assigned the accession number 3A17.

Amino Acid Sequence↗

Extending MapMan: application to legume genome arrays.

MOTIVATION: Based on a gene classification into hierarchical categories ('BINs'), MapMan was originally developed to display Arabidopsis thaliana gene expression in a functional context. We have created a bioinformatics system to extend MapMan to any organism by using a new BIN structure based on the KEGG database. Gene sequences are assigned to this ontology by homology relationships in four reference databases: KEGG, COG, Swiss-Prot and Gene Ontology. We applied this system to tailor MapMan to the GeneChips of two model legumes, Glycine max and Medicago truncatula. We also developed a module to identify the most relevant pathways involved. AVAILABILITY: All mapping files, pathway pictures and the analysis method are available at http://bioinfoserver.rsbs.anu.edu.au/

Algorithms↗

Protein folding via binding and vice versa.

The terms intermolecular and intramolecular recognition are often used when referring to binding and folding, highlighting the common ground between the two processes. Most studies, however, are aimed at either one process or the other. Here, we show how knowledge from binding can aid in understanding folding and vice versa.

Apolipoproteins↗

Comparative proteomics of human endothelial cell caveolae and rafts using two-dimensional gel electrophoresis and mass spectrometry.

The human endothelial cell plasma membrane harbors two subdomains of similar lipid composition, caveolae and rafts, both crucially involved in various essential cellular processes like transcytosis, signal transduction and cholesterol homeostasis. Caveolin-enriched membranes, isolated by either cationic silica or buoyant density methods, were explored by comparing large series of two-dimensional (2-D) maps and subsequent identification of over 100 protein spots by matrix-assisted laser desorption/ionization (MALDI) peptide mass fingerprinting. Improved representation and identification of membrane proteins and valuable information on various post-translational modifications was achieved by the presented optimized procedures for solubilization, destaining and database searching/computing. Whereas the cationic silica purification yielded predominantly known endoplasmic reticulum residents, the cold-detergent method yielded a large number of known caveolae residents, including caveolin-1. Thus, a large part of this subproteome was established, including known (trans-)membrane, signal transduction and glycosyl phosphatidylinositol (GPI)-anchored proteins. Several predicted proteins from the human genome were isolated for the first time from biological samples, including SGRP58, SLP-2, C8ORF2, and XRP-2. These findings and various optimized procedures can serve as a reference to study the differential composition of endothelial cell caveolae and rafts, known to be involved in pathologies like cancer and cardiovascular disease.

Blood Proteins↗

Evaluation of the sequence template method for protein structure prediction. Discrimination of the (beta/alpha)8-barrel fold.

A multiple alignment of five (beta/alpha)8-barrel enzymes has been derived from their structure. The eight beta-strands and eight alpha-helices of the (beta/alpha)8-barrel are correctly aligned and the equivalenced residues in these regions fulfil similar structural roles. Each beta-strand has a central core of usually four residues, two residues contribute side-chains to the barrel core and the other two residues are involved in beta-strand/alpha-helix contacts. However, the fold imposes no constraints on the volumes of the residues at either a local or global level: the volume of the beta-barrel core varies between 1088 A3 in glycolate oxidase and 1571 A3 in taka-amylase. Sequence motifs derived from the multiple alignment were scanned against a database of 124 protein sequences, including 17 (beta/alpha)8-barrel enzymes. The results were evaluated in terms of the discrimination of (beta/alpha)8-barrel sequences and the quality of the alignments obtained. One motif was able to identify the top 12% of high scoring sequences as forming (beta/alpha)8-barrels with 50% accuracy and the bottom 50% of sequences as not being (beta/alpha)8-barrel proteins with 100% accuracy. However, in most instances the alignments were poor. The reasons for this are discussed with reference to the (beta/alpha)8-barrel proteins and the sequence motif method in general.

Alcohol Oxidoreductases↗

The influence of gapped positions in multiple sequence alignments on secondary structure prediction methods.

All currently leading protein secondary structure prediction methods use a multiple protein sequence alignment to predict the secondary structure of the top sequence. In most of these methods, prior to prediction, alignment positions showing a gap in the top sequence are deleted, consequently leading to shrinking of the alignment and loss of position-specific information. In this paper we investigate the effect of this removal of information on secondary structure prediction accuracy. To this end, we have designed SymSSP, an algorithm that post-processes the predicted secondary structure of all sequences in a multiple sequence alignment by (i) making use of the alignment's evolutionary information and (ii) re-introducing most of the information that would otherwise be lost. The post-processed information is then given to a new dynamic programming routine that produces an optimally segmented consensus secondary structure for each of the multiple alignment sequences. We have tested our method on the state-of-the-art secondary structure prediction methods PHD, PROFsec, SSPro2 and JNET using the HOMSTRAD database of reference alignments. Our consensus-deriving dynamic programming strategy is consistently better at improving the segmentation quality of the predictions compared to the commonly used majority voting technique. In addition, we have applied several weighting schemes from the literature to our novel consensus-deriving dynamic programming routine. Finally, we have investigated the level of noise introduced by prediction errors into the consensus and show that predictions of edges of helices and strands are half the time wrong for all the four tested prediction methods.

Algorithms↗

Classification of spider neurotoxins using structural motifs by primary structure features. Single residue distribution analysis and pattern analysis techniques.

In recent years the data on the novel structures of spider toxins have been greatly increasing. The sequence data should be classified. We introduced two primary structure analysis techniques-single residue distribution analysis (SRDA) and pattern analysis for classifying spider polypeptide toxins with molecular weight less than 10kDa. For multiple sequence alignment, we also introduced three novel sequence representation formats named as a simple record, motif record and a pattern record, which can be useful for large-scale analysis of structures. About 300 sequences of spider toxins were analyzed and nine primary structure motifs were identified. New classification of spider toxins was proposed on the basis of previously described principal structural motif (PSM) and extra structural motif (ESM) [Kozlov, S.A., Malyavka, A.A., McCutchen, B., Lu, A., Schepers, E., Herrmann, R., Grishin, E.V., 2005. A novel strategy for the identification of toxin-like structures in spider venom. Proteins 59 (1), 131-140]. Five main structural classes were revealed, and for putative ion channel inhibitors from the most numerous classes 1, 2, and 3, five-digital personal ID numbers were introduced. A reference table with simple, motif and pattern representation sequence formats was created for all analyzed structures.

Amino Acid Motifs↗

Java editor for biological pathways.

SUMMARY: A visual Java-based tool for drawing and annotating biological pathways was developed. This tool integrates the possibilities of charting elements with different attributes (size, color, labels), drawing connections between elements in distinct characteristics (color, structure, width, arrows), as well as adding links to molecular biology databases, promoter sequences, information on the function of the genes or gene products, and references. It is easy to use and system independent. The result of the editing process is a PNG (portable network graphics) file for the images and XML (extended markup language) file for the appropriate links.

Documentation↗

Proteome analysis of B-cell maturation.

Proteins affected by anti-mIgM stimulation during B-cell maturation were identified using 2-DE-based proteomics. We investigated the proteome profiles of stimulated and nonstimulated Ramos B-cells at eight time points during 5 d and compared the obtained proteomic data to the corresponding data from DNA-microarray studies. Anti-mIgM stimulation of the cells resulted in significant differences (> or =twofold) in the protein abundance close to 100 proteins and differences in post-translational protein modifications. Forty-eight up- or down-regulated proteins were identified by mass spectrometric methods and database searches. The identities of a further nine proteins were revealed by comparing their positions to the known proteins in other lymphocyte 2-DE databases. Several of the proteins are directly related to the functional and morphological characteristics of B-cells, such as cytoskeleton rearrangement and intracellular signalling triggered by the crosslinking of B-cell receptors. In addition to proteins known to be involved in human B-cell maturation, we identified several proteins that were not previously linked to lymphocyte differentiation. The results provide deeper insights into the process of B-cell maturation and may lead to novel therapeutic strategies for immunodeficiencies. An interactive 2-DE reference map is available at http://bioinf.uta.fi/BcellProteome.

B-Lymphocytes↗

Cloning and characterization of human liver cytosolic beta-glycosidase.

Cytosolic beta-glucosidase (EC 3.2.1.21) from mammalian liver is a member of the family 1 glycoside hydrolases and is known for its ability to hydrolyse a range of beta-D-glycosides, including beta-D-glucoside and beta-D-galactoside. We therefore refer to this enzyme as cytosolic beta-glycosidase. We cloned the cDNA encoding the human cytosolic beta-glycosidase by performing PCR on cDNA prepared from total human liver RNA. Specific primers were based on human expressed sequence tags found in the expressed sequence tag database. The cloned cDNA contained 1407 nt with an open reading frame encoding 469 amino acid residues. Amino acid sequence analysis indicates that human cytosolic beta-glycosidase is most closely related to lactase phlorizin hydrolase and klotho protein. The enzyme was characterized by using cell lysates of COS-7 cells transfected with a eukaryotic expression vector containing the cDNA. The biochemical, kinetic and inhibition properties of the cloned enzyme were found to be identical with those reported for the enzyme purified from human liver.

Amino Acid Sequence↗

Microsatellite instability analysis in hereditary non-polyposis colon cancer using the Bethesda consensus panel of microsatellite markers in the absence of proband normal tissue.

BACKGROUND: Hereditary non-polyposis colon cancer (HNPCC) is an autosomal dominant syndrome predisposing to the early development of various cancers including those of colon, rectum, endometrium, ovarium, small bowel, stomach and urinary tract. HNPCC is caused by germline mutations in the DNA mismatch repair genes, mostly hMSH2 or hMLH1. In this study, we report the analysis for genetic counseling of three first-degree relatives (the mother and two sisters) of a male who died of colorectal adenocarcinoma at the age of 23. The family fulfilled strict Amsterdam-I criteria (AC-I) with the presence of extracolonic tumors in the extended pedigree. We overcame the difficulty of having a proband post-mortem non-tumor tissue sample for MSI testing by studying the alleles carried by his progenitors. METHODS: Tumor MSI testing is described as initial screening in both primary and metastasis tumor tissue blocks, using the reference panel of 5 microsatellite markers standardized by the National Cancer Institute (NCI) for the screening of HNPCC (BAT-25, BAT-26, D2S123, D5S346 and D17S250). Subsequent mutation analysis of the hMLH1 and hMSH2 genes was performed. RESULTS: Three of five microsatellite markers (BAT-25, BAT-26 and D5S346) presented different alleles in the proband's tumor as compared to those inherited from his parents. The tumor was classified as high frequency microsatellite instability (MSI-H). We identified in the HNPCC family a novel germline missense (c.1864C>A) mutation in exon 12 of hMSH2 gene, leading to a proline 622 to threonine (p.Pro622Thr) amino acid substitution. CONCLUSION: This approach allowed us to establish the tumor MSI status using the NCI recommended panel in the absence of proband's non-tumor tissue and before sequencing the obligate carrier. According to the Human Gene Mutation Database (HGMD) and the International Society for Gastrointestinal Hereditary Tumors (InSiGHT) Database this is the first report of this mutation.

Adult↗

Cloning and localization of a human diphthamide biosynthesis-like protein-2 gene, DPH2L2.

Sequence analysis of the candidate tumor suppressor OVCA1 revealed extensive sequence identity and similarity to proteins from a diverse number of species, including the yeast diphthamide biosynthesis protein-2, dph2, which suggested that OVCA1 may be the human homologue to this yeast gene. However, searches of the translated EST database for sequences in common with dph2 and OVCA1 uncovered an EST, h52976, with significant amino acid conservation with dph2. Isolation of a cDNA clone encompassing the EST by RACE methodologies and sequence analysis indicate the identification of a previously unidentified gene that is ubiquitously expressed and maps to chromosome 1p34. Based on amino acid sequence analysis, the 489-amino-acid protein encoded by this novel gene is distinct from OVCA1 and is more closely related to the yeast dph2 gene product. Therefore, we refer to this novel gene as DPH2L2, which constitutes one member of a novel gene family that may be involved in diphthamide biosynthesis in humans.

Amino Acid Sequence↗

Proteogenomic approaches for the molecular characterization of natural microbial communities.

At the present time we know little about how microbial communities function in their natural habitats. For example, how do microorganisms interact with each other and their physical and chemical surroundings and respond to environmental perturbations? We might begin to answer these questions if we could monitor the ways in which metabolic roles are partitioned amongst members as microbial communities assemble, determine how resources such as carbon, nitrogen, and energy are allocated into metabolic pathways, and understand the mechanisms by which organisms and communities respond to changes in their surroundings. Because many organisms cannot be cultivated, and given that the metabolisms of those growing in monoculture are likely to differ from those of organisms growing as part of consortia, it is vital to develop methods to study microbial communities in situ. Chemoautotrophic biofilms growing in mine tunnels hundreds of meters underground drive pyrite (FeS(2)) dissolution and acid and metal release, creating habitats that select for a small number of organism types. The geochemical and microbial simplicity of these systems, the significant biomass, and clearly defined biological-inorganic feedbacks make these ecosystem microcosms ideal for development of methods for the study of uncultivated microbial consortia. Our approach begins with the acquisition of genomic data from biofilms that are sampled over time and in different growth conditions. We have demonstrated that it is possible to assemble shotgun sequence data to reveal the gene complement of the dominant community members and to use these data to confidently identify a significant fraction of proteins from the dominant organisms by mass spectrometry (MS)-based proteomics. However, there are technical obstacles currently restricting this type of "proteogenomic" analysis. Composite genomic sequences assembled from environmental data from natural microbial communities do not capture the full range of genetic potential of the associated populations. Thus, it is necessary to develop bioinformatics approaches to generate relatively comprehensive gene inventories for each organism type. These inventories are critical for expression and functional analyses. In proteomic studies, for example, peptides that differ from those predicted from gene sequences can be measured, but they generally cannot be identified by database matching, even if the difference is only a single amino acid residue. Furthermore, many of the identified proteins have no known function. We propose that these challenges can be addressed by development of proteogenomic, biochemical, and geochemical methods that will be initially deployed in a simple, natural model ecosystem. The resulting approach should be broadly applicable and will enhance the utility and significance of genomic data from isolates and consortia for study of organisms in many habitats. Solutions draining pyrite-rich deposits are referred to as acid mine drainage (AMD). AMD is a very prevalent, international environmental problem associated with energy and metal resources. The biological-mineralogical interactions that define these systems can be harnessed for energy-efficient metal recovery and removal of sulfur from coal. The detailed understanding of microbial ecology and ecosystem dynamics resulting from the proposed work will provide a scientific foundation for dealing with the environmental challenges and technological opportunities, and yield new methods for analysis of more complex natural communities.

Ecosystem↗

Novel methods for secondary structure determination using low wavelength (VUV) circular dichroism spectroscopic data.

BACKGROUND: Circular Dichroism (CD) spectroscopy is a widely used method for studying protein structures in solution. Modern synchrotron radiation CD (SRCD) instruments have considerably higher photon fluxes than do conventional lab-based CD instruments, and hence have the ability to routinely measure CD data to much lower wavelengths. Recently a new reference dataset of SRCD spectra of proteins of known structure, designed to cover secondary structure and fold space, has been produced which includes low wavelength (vacuum ultraviolet - VUV) data. However, the existing algorithms used to calculate protein secondary structures from CD data have not been designed to take optimal advantage of the additional information in these low wavelength data. RESULTS: In this study, we have optimised secondary structure calculation methods based on the low wavelength CD data by examining existing algorithms and secondary structure assignment schemes, and then developing new methods which have produced clear improvements in prediction accuracy, especially for beta-sheet components. We have further shown that if precise measurements of protein concentrations, and therefore spectral magnitudes, are not available, the inclusion of the low wavelength data will significantly improve the analyses. However, we have also demonstrated that the new reference dataset, methods, and assignments can also improve the analyses of conventional circular dichroism data, even if the low wavelength data is not available. CONCLUSION: VUV CD data include important information on protein structure which can be exploited with the algorithms and methodologies described.

Algorithms↗

Towards the proteome of Burkholderia cenocepacia H111: setting up a 2-DE reference map.

Polyphasic-taxonomic studies of the past decade have shown that the Burkholderia cepacia complex (Bcc) comprises at least nine species, which share a high degree of 16S rDNA (98-100%) sequence similarity but only moderate levels of DNA-DNA hybridization. Members of the Bcc are well known as opportunistic pathogens of plants, animals and humans but also as biocontrol and bioremediation agents. In this study intra-, surface-associated and extracellular proteins of B. cenocepacia H111, which was isolated from a cystic fibrosis patient, were examined by 2-DE coupled to MALDI-TOF MS. MS and MS/MS data were searched against a database comprising all currently available annotated proteins of genetically closely related strains. In total 642 proteins spots were successfully identified corresponding to 390 different protein species, which were classified into functional categories. The majority of these proteins could be linked to housekeeping functions in energy production, amino acid metabolism, protein folding, post-translational modification and turnover, and translation. Noteworthy is the fact that a significant number of truly secreted and membrane proteins were identified in the extracellular and surface-associated sub-proteomes. This indicates that the pre-fractionation protocol used in this study is a highly valuable strategy for unravelling the cellular location of the identified proteins.

Bacterial Proteins↗