Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Characterization of the gene expression profile of neuroblastoma cell line IMR-5 using serial analysis of gene expression.

The serial analysis of gene expression (SAGE) technique was used to generate a database of the most abundant transcripts of the MYCN-amplified neuroblastoma cell line IMR-5. A total of 8568 tags were sequenced and shown to represent 4034 unique tags, each of which corresponds to an individual transcript. Expression levels of genes are reflected by the frequency of occurrence of the respective tags. To validate fidelity of SAGE data, relative abundances of seven transcripts were evaluated by semiquantitative reverse transcriptase-polymerase chain reaction. Transcripts that were detected nine times or more (>0.1% of the total tag population) accounted for 36% of the total messenger RNA mass but only 3% of the total number of individual transcripts. A strong preponderance of genes involved in protein synthesis, in particular those encoding for ribosomal proteins, were observed among these high-abundance transcripts. Tags corresponding to the amplified gene DDX1 were conspicuously overrepresented in comparison to the other amplified genes MYCN, neuroblastoma amplified gene and MEIS1, which suggests an additional mechanism apart from genomic amplification contributing to the strong upregulation of this gene. This study provides a comprehensive gene expression profile of neuroblastoma cell line IMR-5 and may be used as a reference database for identification of candidate genes that are involved in etiology and pathogenesis of neuroblastoma.

Databases as Topic↗

Prediction of protein homo-oligomer types by pseudo amino acid composition: Approached with an improved feature extraction and Naive Bayes Feature Fusion.

The interaction of non-covalently bound monomeric protein subunits forms oligomers. The oligomeric proteins are superior to the monomers within the scope of functional evolution of biomacromolecules. Such complexes are involved in various biological processes, and play an important role. It is highly desirable to predict oligomer types automatically from their sequence. Here, based on the concept of pseudo amino acid composition, an improved feature extraction method of weighted auto-correlation function of amino acid residue index and Naive Bayes multi-feature fusion algorithm is proposed and applied to predict protein homo-oligomer types. We used the support vector machine (SVM) as base classifiers, in order to obtain better results. For example, the total accuracies of A, B, C, D and E sets based on this improved feature extraction method are 77.63, 77.16, 76.46, 76.70 and 75.06% respectively in the jackknife test, which are 6.39, 5.92, 5.22, 5.46 and 3.82% higher than that of G set based on conventional amino acid composition method with the same SVM. Comparing with Chou's feature extraction method of incorporating quasi-sequence-order effect, our method can increase the total accuracy at a level of 3.51 to 1.01%. The total accuracy improves from 79.66 to 80.83% by using the Naive Bayes Feature Fusion algorithm. These results show: 1) The improved feature extraction method is effective and feasible, and the feature vectors based on this method may contain more protein quaternary structure information and appear to capture essential information about the composition and hydrophobicity of residues in the surface patches that buried in the interfaces of associated subunits; 2) Naive Bayes Feature Fusion algorithm and SVM can be referred as a powerful computational tool for predicting protein homo-oligomer types.

Algorithms↗

Pair-preferences: a quantitative measure of regularities in protein sequences.

We present here the results obtained by applying several different methods to quantitatively measure regularities in protein sequences based on pair-preferences. We have studied the distribution of amino acid residues, singly as well as in pairs in a large data base and have attempted this task. We confirmed the existence of well-defined pair-preferences in proteins which were shown to be remarkably absent in simulated random sequences of similar amino acid distribution. The analysis of the sequences from the SWISS-PROT data base using simple statistical tests. Fourier analysis, fractal analysis and statistical thermodynamical tests were used to derive parameters to define a natural sequence. As a consequence of the existence of pair-preferences, parameters like fractal dimension (D), spectral exponent (beta), scaling parameter (H) and entropy (statistical) were found to be characteristic for natural sequences. For a reference state we chose a randomised state devoid of any pair-preference. The pair-preferences qualified well to be used as quantitative measures of regularities in protein sequences.

Algorithms↗

Analysis of signaling pathways using functional proteomics.

Advances in analytical methods for protein analysis by mass spectrometry provide new tools for global analysis of the expressed protein profile of cells (referred to as proteomics). Currently, available methodology samples only part of the proteome. This is sufficient for analysis of signal transduction, because signaling pathways contain enzymes, which modify high-abundance proteins other than those of the pathway. Thus, modulation of the signaling through a pathway will produce a "footprint" in the proteome that is characteristic of a specific cell phenotype. Comparison of different samples to identify these differences in posttranslational modification or protein expression is referred to as functional proteomics. This review surveys the methods in widest use in functional proteomics, as well as a few promising new ones. Although proteomic analyses were first conducted 26 years ago, a renewed interest is fueled by several recent advances. Most important are the availability of public genome and protein databases and the development of high-sensitivity, easy-to-use mass spectrometers and database search engines capable of exploiting these databases. Other important advances include improved two-dimensional polyacrylamide gel electrophoresis (2D-PAGE), computer programs for analysis of the 2D-PAGE gel images, protocols for proteolytic digestion of proteins in excised gel pieces, and low-flow chromatography methods. Despite the limitations of these methods, they can distinguish subtle changes in the phenotype of cells, providing the basis for future studies in regulation of the phenotype.

Animals↗

The yeast proteome database (YPD) and Caenorhabditis elegans proteome database (WormPD): comprehensive resources for the organization and comparison of model organism protein information.

The Yeast Proteome Database (YPDtrade mark) has been for several years a resource for organized and accessible information about the proteins of Saccharomyces cerevisiae. We have now extended the YPD format to create a database containing complete proteome information about the model organism Caenorhabditis elegans (WormPDtrade mark). YPD and WormPD are designed for use not only by their respective research communities but also by the broader scientific community. In both databases, information gleaned from the literature is presented in a consistent, user-friendly Protein Report format: a single Web page presenting all available knowledge about a particular protein. Each Protein Report begins with a Title Line, a concise description of the function of that protein that is continually updated as curators review new literature. Properties and functions of the protein are presented in tabular form in the upper part of the Report, and free-text annotations organized by topic are presented in the lower part. Each Protein Report ends with a comprehensive reference list whose entries are linked to their MEDLINE s. YPD and WormPD are seamlessly integrated, with extensive links between the species. They are freely accessible to academic users on the WWW at http://www. proteome.com/databases/index.html, and are available by subscription to corporate users.

Animals↗

Identification and characterization of an ovary-selective isoform of epoxide hydrolase.

A novel ovary-selective gene was identified by suppression subtractive hybridization (SSH) that is expressed only during the mouse periovulatory phase of a stimulated estrous cycle. Analysis of the protein encoded by the full-length cDNA revealed that the majority of it, with the exception of the first 44 amino acids, matched soluble epoxide hydrolase (Ephx2, referred to as Ephx2A). By comparing the cDNA sequence of this newly identified variant of soluble epoxide hydrolase (referred to as Ephx2B) with the mouse genome database, an exon was identified that corresponds to its unique 5' cDNA sequence. Through the use of an Ephx2A-specific probe, Northern blot analysis revealed that this mRNA was also expressed in the ovary, with the highest level of expression occurring during the luteal phase of a stimulated estrous cycle. In situ hybridization revealed that Ephx2B mRNA expression was restricted to granulosa cells of preovulatory follicles. Ephx2A mRNA expression, however, was detectable in follicles at different stages of development, as well as in the corpus luteum. Total ovarian epoxide hydrolase activity increased following the induction of follicular development, and remained elevated through the periovulatory and postovulatory stages of a stimulated estrous cycle. The change in enzyme activity paralleled the combined mRNA expression profiles for both Ephx2A and Ephx2B, thus supporting a role for epoxide metabolism in ovarian function.

Amino Acid Sequence↗

In silico reconstruction of the metabolic pathways of Lactobacillus plantarum: comparing predictions of nutrient requirements with those from growth experiments.

On the basis of the annotated genome we reconstructed the metabolic pathways of the lactic acid bacterium Lactobacillus plantarum WCFS1. After automatic reconstruction by the Pathologic tool of Pathway Tools (http://bioinformatics.ai.sri.com/ptools/), the resulting pathway-genome database, LacplantCyc, was manually curated extensively. The current database contains refinements to existing routes and new gram-positive bacterium-specific reactions that were not present in the MetaCyc database. These reactions include, for example, reactions related to cell wall biosynthesis, molybdopterin biosynthesis, and transport. At present, LacplantCyc includes 129 pathways and 704 predicted reactions involving some 670 chemical species and 710 enzymes. We tested vitamin and amino acid requirements of L. plantarum experimentally and compared the results with the pathways present in LacplantCyc. In the majority of cases (32 of 37 cases) the experimental results agreed with the final reconstruction. LacplantCyc is the most extensively curated pathway-genome database for gram-positive bacteria and is open to the microbiology community via the World Wide Web (www.lacplantcyc.nl). It can be used as a reference pathway-genome database for gram-positive microbes in general and lactic acid bacteria in particular.

Amino Acids↗

Molecular chaperones in the Paracoccidioides brasiliensis transcriptome.

Paracoccidioides brasiliensis is a thermally dimorphic and a human pathogenic fungus. Our group has partially sequenced its transcriptome and generated a database of mycelial and yeast PbAESTs (P. brasiliensis assembled expressed sequence tags). In the present review we describe the identification of PbAESTs encoding molecular chaperones. These proteins, involved in protein folding and renaturation, are also implicated in several other biological processes, where the dimorphic transition is of particular interest. Another important issue concerning these proteins refers to their participation in the immunopathogenicity of infectious diseases. We have found 438 ESTs (184 in mycelium and 253 in yeast) encoding P. brasiliensis molecular chaperones and their co-chaperones, which were clustered in 48 genes. These genes were classified in families, corresponding to three small chaperones, nine HSP40s, 10 HSP60s, seven HSP70s, five HSP90s, four HSP100s, and 10 other chaperones. These results greatly increase the knowledge on P. brasiliensis molecular chaperones, since only eight of such proteins had been previously characterized.

DNA, Complementary↗

A protein class database organized with ProSite protein groups and PIR superfamilies.

A protein class (ProClass) database is developed as a "value-added" "second-generation" database organized according to family relationships. The database collects non-redundant protein sequence entries from SwissProt and PIR databases, and classifies them in families defined collectively by the ProSite protein groups and PIR superfamilies. The major objectives of the database are to maximize family information retrieval, to provide speedy family identification, and to help organizing existing protein sequence databases. The database has two sub-databases: PCFam (ProClass Family) to define protein families and provide links to ProSite patterns and PIR superfamilies, and PCSeq (ProClass Sequence) to describe sequence entries and provide links to PCFam, SwissProt, PIR, and ProSite databases. The current ProClass release has a total of 85,165 sequence entries, about half of which are classified in 3072 ProClass families; it also contains 10,431 newly established SwissProt-PIR links. The database can help reveal domain structures of related families, define new ProSite and PIR families, and provide family assignments for unclassified sequence entries. New ProSite and PIR family members are readily identified via database cross-reference, including 9437 SwissProt entries and 8522 PIR entries. False negative family members missed by both ProSite and PIR are detected using a neural network family identification system. The newly identified superfamily memberships are being incorporated into the current PIR database releases in a collaborative effort with the PIR. The ProClass database is accessible through anonymous FTP and on-line search on the World Wide Web.

Amino Acid Sequence↗

BacTregulators: a database of transcriptional regulators in bacteria and archaea.

MOTIVATION: The BacTregulators database is intended to collect and to integrate information on proteins belonging to defined families of transcriptional regulators in prokaryotes. RESULTS: The BacTregulators database currently contains data on two families of transcriptional regulators: AraC-XylS and TetR. The proteins included in the BacTregulators database have been identified by screening 123 genomes from archaea and bacteria and the SWISS-PROT and TrEMBL databases with profiles defining each family. As the result of an integration process, we have included 1326 different protein sequences from the AraC-XylS family and 1487 different protein sequences from the TetR family. The definition of an entry in BacTregulators is based on protein sequence, source organism, genome element and position in this genome element. The BacTregulators site allows the user to retrieve protein sequences, functional features and experimental evidence supporting the functions, references and the three-dimensional structure of the regulator when available. BacTregulators supplies an innovative tool that allows the researcher to obtain an integrated report that shows the data corresponding to other entries which are related by sequence similarity to the query entry. BacTregulators detects and classifies the regulators belonging to AraC-XylS and TetR families present in prokaryotic genomes, and thus contributes to a more accurate annotation of regulators in genomes. The information collected on each protein in the family can be useful to characterize a new regulator or compile information on the biological properties of a known regulator. AVAILABILITY: The BacTregulators is available at www.bactregulators.org

Archaea↗

KEGG as a glycome informatics resource.

Bioinformatics approaches to carbohydrate research have recently begun using large amounts of protein and carbohydrate data. In this field called glycome informatics, the foremost necessity is a comprehensive resource for genome-scale bioinformatics analysis of glycan data. Although the accumulation of experimental data may be useful as a reference of biological and biochemical information on carbohydrates, this is insufficient for bioinformatics analysis. Thus, we have developed a glycome informatics resource (http://www.genome.jp/kegg/glycan/) in KEGG (Kyoto Encyclopedia of Genes and Genomes), an integrated knowledge base of protein networks, genomic information, and chemical information. This review describes three noteworthy features: (1) GLYCAN, a database of carbohydrate structures; (2) glycan-related pathways; and (3) Composite Structure Map (CSM), a map illustrating all possible variations of carbohydrate structures within organisms. GLYCAN includes two useful tools: an intuitive drawing tool called KegDraw, and an efficient glycan search and alignment tool called KEGG Carbohydrate Matcher (KCaM). KEGG's glycan biosynthesis and metabolism pathways, integrating carbohydrate structures, proteins, and reactions, are also a pivotal resource. CSM is constructed as a bridge between carbohydrate functions and structures. CSM is able to display, for example, expression data of glycosyltransferases in a compact manner. In all the KEGG resources, various objects including KEGG pathways, chemical compounds, as well as carbohydrate structures are commonly represented as graphs, which are widely studied and utilized in the computer science field.

Carbohydrates↗

Wnt pathway curation using automated natural language processing: combining statistical methods with partial and full parse for knowledge extraction.

MOTIVATION: Wnt signaling is a very active area of research with highly relevant publications appearing at a rate of more than one per day. Building and maintaining databases describing signal transduction networks is a time-consuming and demanding task that requires careful literature analysis and extensive domain-specific knowledge. For instance, more than 50 factors involved in Wnt signal transduction have been identified as of late 2003. In this work we describe a natural language processing (NLP) system that is able to identify references to biological interaction networks in free text and automatically assembles a protein association and interaction map. RESULTS: A 'gold standard' set of names and assertions was derived by manual scanning of the Wnt genes website (http://www.stanford.edu/~rnusse/wntwindow.html) including 53 interactions involved in Wnt signaling. This system was used to analyze a corpus of peer-reviewed articles related to Wnt signaling including 3369 Pubmed and 1230 full text papers. Names for key Wnt-pathway associated proteins and biological entities are identified using a chi-squared analysis of noun phrases over-represented in the Wnt literature as compared to the general signal transduction literature. Interestingly, we identified several instances where generic terms were used on the website when more specific terms occur in the literature, and one typographic error on the Wnt canonical pathway. Using the named entity list and performing an exhaustive assertion extraction of the corpus, 34 of the 53 interactions in the 'gold standard' Wnt signaling set were successfully identified (64% recall). In addition, the automated extraction found several interactions involving key Wnt-related molecules which were missing or different from those in the canonical diagram, and these were confirmed by manual review of the text. These results suggest that a combination of NLP techniques for information extraction can form a useful first-pass tool for assisting human annotation and maintenance of signal pathway databases. AVAILABILITY: The pipeline software components are freely available on request to the authors. CONTACT: dstates@umich.edu SUPPLEMENTARY INFORMATION: http://stateslab.bioinformatics.med.umich.edu/software.html.

Animals↗

Kazusa mammalian cDNA resources: towards functional characterization of KIAA gene products.

The Kazusa cDNA project pioneered an extensive sequencing project of human cDNAs in their entirety and focused sequencing efforts particularly on large cDNAs encoding large proteins. More than 2000 human genes, referred to as 'KIAA' genes, were initially identified through this cDNA project. Since many KIAA genes still remain functionally uncharacterized, our current focus is to determine their biological functions in vivo. In this review, we describe the current status of the Kazusa mammalian cDNA resources and the future direction of the functional characterization of KIAA genes.

Animals↗

A large family of endosome-localized proteins related to sorting nexin 1.

Sorting nexin 1 (SNX1), a peripheral membrane protein, has previously been shown to regulate the cell-surface expression of the human epidermal growth factor receptor [Kurten, Cadena and Gill (1996) Science 272, 1008-1010]. Searches of human expressed sequence tag databases with SNX1 revealed eleven related human cDNA sequences, termed SNX2 to SNX12, eight of them novel. Analysis of SNX1-related sequences in the Saccharomyces cerevisiae genome clearly shows a greatly expanded SNX family in humans in comparison with yeast. On the basis of the predicted protein sequences, all members of this family of hydrophilic molecules contain a conserved 70-110-residue Phox homology (PX) domain, referred to as the SNX-PX domain. Within the SNX family, subgroups were identified on the basis of the sequence similarities of the SNX-PX domain and the overall domain structure of each protein. The members of one subgroup, which includes human SNX1, SNX2, SNX4, SNX5 and SNX6 and the yeast Vps5p and YJL036W, all contain coiled-coil regions within their large C-terminal domains and are found distributed in both membrane and cytosolic fractions, typical of hydrophilic peripheral membrane proteins. Localization of the human SNX1 subgroup members in HeLa cells transfected with the full-length cDNA species revealed a similar intracellular distribution that in all cases overlapped substantially with the early endosome marker, early endosome autoantigen 1. The intracellular localization of deletion mutants and fusions with green fluorescent protein showed that the C-terminal regions of SNX1 and SNX5 are responsible for their endosomal localization. On the basis of these results, the functions of these SNX molecules are likely to be unique to endosomes, mediated in part by interactions with SNX-specific C-terminal sequences and membrane-associated determinants.

Amino Acid Sequence↗

A systematic review of the influence of different titanium surfaces on proliferation, differentiation and protein synthesis of osteoblast-like MG63 cells.

OBJECTIVES: Titanium is the standard material for dental and orthopaedical implants. The good biocompatibility has been proven in many experimental and clinical investigations. Different titanium topographies were tested in vitro using different cell culture models. The aim of this systematic review was to evaluate and summarize the medical/dental literature to assess on which kind of titanium surface structure the osteoblast-like osteosarcoma cells MG63 show the best proliferation and differentiation rate, and the best protein synthesis. METHODS: A systematic search was carried out using different on-line databases (PubMed, Web of Science, Cochrane Library, International Poster Journal), supplemented by handsearch in selected journals and by examination of the bibliographies of the identified articles. Inclusion and exclusion criterias were applied when considering relevant articles. Studies which met the inclusion criteria were included and data extraction was undertaken by one reviewer. RESULTS: The search yielded 348 references. Nine articles referring to nine different studies were relevant to our question. Additionally 8 less relevant articles were identified. It was found that regularly textured surfaces of pure titanium with R(a) values (average roughness) of around 4 mum are well-accepted by MG63 cells. CONCLUSIONS: The surfaces and culture conditions vary widely. Therefore it is still difficult to recommend one particular surface. It seems that there are no differences in cell proliferation and differentiation on surfaces treated by blasting and etching. Standardization in fabrication and size of the different test surfaces as well as homogeneity in culture times and plating densities should be aspects for future research.

Algorithms↗

Comparative proteome analysis of breast cancer and normal breast.

Breast cancer is a leading cause of death for women. The underlying molecular mechanism is still not well understood. In this study, two-dimensional gel electrophoresis combined with mass spectrometry was used to analyze changes in the proteome of infiltrating ductal carcinoma compared to normal breast tissue. Ten sets of two-dimensional gels per experimental condition were analyzed and more than 500 spots each were detected. This revealed 39 spots for which expression in breast cancer cells were reproducibly altered more than twofold compared to normal controls (p < 0.01). These spots represented 25 different proteins after identification using the database search after mass spectrometry, comprising cell defense proteins, enzymes involved in glycolytic energy metabolism and homeostasis, protein folding and structural proteins, proteins involved in cytoskeleton and cell motility, and proteins involved in other functions. In addition, 28 nondifferentially expressed proteins with different functions were also mapped and identified, which might help to establish a two-dimensional gel electrophoresis reference map of human breast cancer. Our study shows that proteomics offers a powerful methodology to detect the proteins that show different expression patterns in breast cancer tissue and may provide an accurate molecular classification. The differentially expressed proteins may be used as potential candidate markers for diagnostic purposes or for determination of tumor sensitivity to therapy. The functional implications of the identified proteins are discussed.

Biomarkers, Tumor↗

Comparative proteomic analysis of mammalian animal tissues and body fluids: bovine proteome database.

Characterizing the complete proteome of multicellular organisms is a challenging task using the currently available technologies. With the increasing degree of genetic complexity, animals acquire a broader repertoire of options to meet environmental challenges. Mammalian cells from different tissues/body fluids express different thousands of proteins with a predicted dynamic range of up to five to six orders of magnitude, thus necessitating the whole arsenal of dedicated analytical strategies for a detailed proteome characterization. Nevertheless, 2D-E analysis of whole cellular lysates still remains the most used initial approach for the proteomic description of specialized cells. It enables to obtain an overview of the main soluble protein components of a specific tissue/body fluid, allowing comparison between different cellular types and molecular description of organ specialization. Massive proteomic investigations have been reported mainly in the case of human, mouse and rat, allowing comparative analysis. For this reason, a research project focused on the 2D-E characterization of tissues and biological fluids from other domestic mammals has been undertaken in our laboratory. A number of high-resolution reference electrophoretic maps have been established for liver, kidney, muscle, plasma and red blood cells samples from Holstein Friesian bovine female individuals. Among the 1863 distinct protein features detected, 534 species were identified and associated to 209 different genes by a combination of MALDI-TOF mass fingerprint, capillary LC-ESI-IT-MS-MS and image gel matching procedures. Identified polypeptide species and differences in expression profiles between various tissues/fluids clearly reflected organ biochemical specialization. This experimental output allowed establishing a 2D-E bovine database accessible at the URL address for image comparison.

Animals↗

Gene annotation from scientific literature using mappings between keyword systems.

MOTIVATION: The description of genes in databases by keywords helps the non-specialist to quickly grasp the properties of a gene and increases the efficiency of computational tools that are applied to gene data (e.g. searching a gene database for sequences related to a particular biological process). However, the association of keywords to genes or protein sequences is a difficult process that ultimately implies examination of the literature related to a gene. RESULTS: To support this task, we present a procedure to derive keywords from the set of scientific abstracts related to a gene. Our system is based on the automated extraction of mappings between related terms from different databases using a model of fuzzy associations that can be applied with all generality to any pair of linked databases. We tested the system by annotating genes of the SWISS-PROT database with keywords derived from the abstracts linked to their entries (stored in the MEDLINE database of scientific references). The performance of the annotation procedure was much better for SWISS-PROT keywords (recall of 47%, precision of 68%) than for Gene Ontology terms (recall of 8%, precision of 67%). AVAILABILITY: The algorithm can be publicly accessed and used for the annotation of sequences through a web server at http://www.bork.embl.de/kat

Abstracting and Indexing↗