Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Annotation matters: the effect of structural gene annotation on orthology inference.

MOTIVATION: In silico gene annotation, the process of identifying the genes present in a genome, remains a challenging task. As genome assemblies rapidly increase, the corresponding gene models and repertoires often fall short in quality. Despite advances in annotation methods, a lack of community standards means that most published gene annotations result from ad hoc pipelines. As a result, only a few species have nearly complete and accurate gene models. This annotation quality is thought to affect downstream analyses, including orthology inference, often the first step of comparative genomics studies. RESULTS: We show that different annotation methods yield markedly distinct orthology inferences. We compared orthology assignments of gene models obtained by four prominent protein-coding gene model sources: the NCBI Eukaryotic Genome Annotation Pipeline, the Ensembl Gene Annotation System, the UniProt Reference Proteomes, and Augustus 3.4 (an ab initio pipeline). We observe significant discrepancies between sources, namely in the proportion of orthologous genes per genome, the completeness of Hierarchical Orthologous Groups, and the accuracy and recall of the predicted orthologs on a standard orthology benchmark.

Molecular Sequence Annotation↗

A statistical framework for combining and interpreting proteomic datasets.

MOTIVATION: To identify accurately protein function on a proteome-wide scale requires integrating data within and between high-throughput experiments. High-throughput proteomic datasets often have high rates of errors and thus yield incomplete and contradictory information. In this study, we develop a simple statistical framework using Bayes' law to interpret such data and combine information from different high-throughput experiments. In order to illustrate our approach we apply it to two protein complex purification datasets. RESULTS: Our approach shows how to use high-throughput data to calculate accurately the probability that two proteins are part of the same complex. Importantly, our approach does not need a reference set of verified protein interactions to determine false positive and false negative error rates of protein association. We also demonstrate how to combine information from two separate protein purification datasets into a combined dataset that has greater coverage and accuracy than either dataset alone. In addition, we also provide a technique for estimating the total number of proteins which can be detected using a particular experimental technique. AVAILABILITY: A suite of simple programs to accomplish some of the above tasks is available at www.unm.edu/~compbio/software/DatasetAssess

Algorithms↗

Proteome analysis of Spiroplasma melliferum (A56) and protein characterisation across species boundaries.

Spiroplasma melliferum (Class: Mollicutes) is a wall-less, helical bacterium with a genome of approximately 1460 kbp encoding 800-1000 gene-products. A two-dimensional electrophoresis gel reference map of S. melliferum was produced by Phoretix 2-D gel software analysis of eight high quality gels. The reference map showed 456 silver-stained and replicated protein spots. 156 proteins (34% of visible protein spots) from S. melliferum were further characterised by one, or a combination, of the following: amino acid analysis, peptide-mass fingerprinting via matrix assisted laser desorption ionisation-time of flight (MALDI-TOF) mass spectrometry, and N-terminal protein microsequencing. Proteins with close relationship to those previously determined from other species were identified across species barriers. Thus, this study represents the first larger-scale analysis of a proteome based upon the attribution of predominantly 'unique numerical parameters' for protein characterisation across species boundaries, as opposed to a sequence-based approach. This approach allowed all database entries to be screened for homology, as is currently the case for studies based on nucleic acid or protein sequence information. Several proteins studied from this organism were identified as hypothetical, or having no close homolog already present in the databases. Gene-products from major families such as glycolysis, translation, transcription, cellular processes, energy metabolism and protein synthesis were identified. Several gene-products characterised in S. melliferum were not previously found in studies of the entire Mycoplasma genitalium and Mycoplasma pneumoniae (both closely related Mollicutes) genomes. The presence of such gene-products in S. melliferum is discussed in terms of genome size as compared with the smallest known free-living organisms. Finally, the levels of expression of S. melliferum gene-products were determined with respect to total optical intensity associated with all visible proteins expressed in exponentially grown cells.

Amino Acid Sequence↗

Predicting protein functions with message passing algorithms.

MOTIVATION: In the last few years, a growing interest in biology has been shifting toward the problem of optimal information extraction from the huge amount of data generated via large-scale and high-throughput techniques. One of the most relevant issues has recently emerged that of correctly and reliably predicting the functions of a given protein with that of functions exploiting information coming from the whole network of proteins physically interacting with the functionally undetermined one. In the present work, we will refer to an 'observed' protein as the one present in the protein-protein interaction networks published in the literature. METHODS: The method proposed in this paper is based on a message passing algorithm known as Belief Propagation, which accepts the network of protein's physical interactions and a catalog of known protein's functions as input, and returns the probabilities for each unclassified protein of having one chosen function. The implementation of the algorithm allows for fast online analysis, and can easily be generalized into more complex graph topologies taking into account hypergraphs, i.e. complexes of more than two interacting proteins. RESULTS: Benchmarks of our method are the two Saccharomyces cerevisiae protein-protein interaction networks and the Database of Interacting Proteins. The validity of our approach is successfully tested against other available techniques. CONTACT: leone@isiosf.isi.it SUPPLEMENTARY INFORMATION: http://isiosf.isi.it/~pagnani

Algorithms↗

Genes conserved in yeast and humans.

Evolutionary conservation of homologous gene products from distantly related organisms provides an information resource of great value for elucidating protein structure and function. Sequence similarities also serve as molecular cross-references between diverse organisms that offer different, or complementary, experimental approaches for analyzing gene expression and biochemistry in normal and abnormal states. There are now countless examples of information about a protein from one species contributing to the understanding of biological phenomena or disease in another species. Such connections are often unanticipated and surprising, but there is an opportunity to make them more systematically as concerted genome sequencing projects progress. In the present review we focus on connections between yeast and human proteins and their functional implications. We present several 'case studies' as well as survey results derived from comprehensive sequence comparisons among all yeast and human proteins currently present in the public databases.

Conserved Sequence↗

Kinesin superfamily proteins (KIFs) in the mouse transcriptome.

In the post genomic era where virtually all the genes and the proteins are known, an important task is to provide a comprehensive analysis of the expression of important classes of genes, such as those that are required for intracellular transport. We report the comprehensive analysis of the Kinesin Superfamily, which is the first and only large protein family whose constituents have been completely identified and confirmed in silico and at the cDNA, mRNA level. In FANTOM2, we have found 90 clones from 33 Kinesin Superfamily Protein (KIF) gene loci. The clones were analyzed in reference to sequence state, library of origin, detection methods, and alternative splicing. More than half of the representative transcriptional units (TU) were full length. The FANTOM2 library also contains novel splice variants previously unreported. We have compared and evaluated various protein classification tools and protein search methods using this data set. This report provides a foundation for future research of the intracellular transport along microtubules and proves the significance of intracellular transport protein transcripts as part of the transcriptome.

Alternative Splicing↗

In vitro model system for the identification and characterization of proteins involved in inflammatory processes.

An in vitro model featuring important inflammatory cellular states was established, based on the murine monocyte/macrophage cell line RAW 264.7. Macrophages are key players in chronic inflammation, and major parts of the biochemical reactions taking place in vivo, e.g., the production of proinflammatory cytokines, can be triggered in vitro by stimulation of the cells with bacterial lipopolysaccharide (LPS). A mastergel, representing a synthetic image of the expressed basic set of cellular proteins, was designed by a computer-assisted overlay of a statistically significant number of two-dimensional electrophoresis (2-DE) gels of unstimulated RAW 264.7 cells. This image served as a reference for qualitative and quantitative changes in the protein pattern induced by stimulation of the macrophages with LPS. The optimal conditions for LPS stimulation were evaluated by monitoring the expression and secretion of the proinflammatory cytokine tumor necrosis factor-alpha(TNF-alpha). The comparison of the mastergel with the 2-DE gels of LPS-stimulated cells revealed several changes in the protein pattern. In order to prove the relevance of the presented model system, we focused on two low molecular weight proteins, which showed significant changes in the apparent concentration in a 2-DE pattern. These proteins were further characterized by microsequencing of internal peptides. A comparison of the obtained sequences with protein databases identified them as cofilin and keratinocyte lipid-binding protein.

Amino Acid Sequence↗

Analysis of flanking sequences from dissociation insertion lines: a database for reverse genetics in Arabidopsis.

We have generated Dissociation (Ds) element insertions throughout the Arabidopsis genome as a means of random mutagenesis. Here, we present the molecular analysis of genomic sequences that flank the Ds insertions of 931 independent transposant lines. Flanking sequences from 511 lines proved to be identical or homologous to DNA or protein sequences in public databases, and disruptions within known or putative genes were indicated for 354 lines. Because a significant portion (45%) of the insertions occurred within sequences defined by GenBank BAC and P1 clones, we were able to assess the distribution of Ds insertions throughout the genome. We discovered a significant preference for Ds transposition to the regions adjacent to nucleolus organizer regions on chromosomes 2 and 4. Otherwise, the mapped insertions appeared to be evenly dispersed throughout the genome. For any given gene, insertions preferentially occurred at the 5' end, although disruption was clearly possible at any intragenic position. The insertion sites of >500 lines that could be characterized by reference to public databases are presented in a tabular format at http://www.plantcell. org/cgi/content/full/11/12/2263/DC1. This database should be of value to researchers using reverse genetics approaches to determine gene function.

Arabidopsis↗

Novel six-nucleotide deletion in the hepatic fructose-1,6-bisphosphate aldolase gene in a patient with hereditary fructose intolerance and enzyme structure-function implications.

Hereditary fructose intolerance (HFI) is an autosomal recessive human disease that results from the deficiency of the hepatic aldolase isoenzyme. Affected individuals will succumb to the disease unless it is readily diagnosed and fructose eliminated from the diet. Simple and non-invasive diagnosis is now possible by direct DNA analysis that scans for known and unknown mutations. Using a combination of several PCR-based methods (restriction enzyme digestion, allele specific oligonucleotide hybridisation, single strand conformation analysis and direct sequencing) we identified a novel six-nucleotide deletion in exon 6 of the aldolase B gene (delta 6ex6) that leads to the elimination of two amino acid residues (Leu182 and Val183) leaving the message inframe. The three-dimensional structural alterations induced in the enzyme by delta 6ex6 have been elucidated by molecular graphics analysis using the crystal structure of the rabbit muscle aldolase as reference model. These studies showed that the elimination of Leu182 and Val183 perturbs the correct orientation of adjacent catalytic residues such as Lys146 and Glu187.

Amino Acid Sequence↗

A comprehensive web resource on RNA helicases from the baker's yeast Saccharomyces cerevisiae.

Members of the RNA helicase protein family are defined by several motifs that have been widely conserved during evolution. They are found in all organisms-from bacteria to humans-and many viruses. The minimum number of RNA helicases present within a eukaryotic cell can be predicted from the complete sequence of the Saccharomyces cerevisiae genome. Recent progress in the functional analysis of various family members has confirmed the significance of RNA helicases for most cellular RNA metabolic processes. We have assembled a web resource that focuses on RNA helicases from the budding yeast Saccharomyces cerevisiae. It includes descriptions of RNA helicases and their functions, links to sequence- and yeast-specific databases, an extensive list of references, and links to non-yeast helicase web resources.

Databases, Factual↗

Proteome profiling of human epithelial ovarian cancer cell line TOV-112D.

A proteome profiling of the epithelial ovarian cancer cell line TOV-112D was initiated as a protein expression reference in the study of ovarian cancer. Two complementary proteomic approaches were used in order to maximise protein identification: two-dimensional gel electrophoresis (2DE) protein separation coupled to matrix assisted laser desorption/ionisation time-of-flight mass spectrometry (MALDI-TOF MS) and one-dimensional gel electrophoresis (1DE) coupled to liquid-chromatography tandem mass spectrometry (LC MS/MS). One hundred and seventy-two proteins have been identified among 288 spots selected on two-dimensional gels and a total of 579 proteins were identified with the 1DE LC MS/MS approach. This proteome profiling covers a wide range of protein expression and identifies several proteins known for their oncogenic properties. Bioinformatics tools were used to mine databases in order to determine whether the identified proteins have previously been implicated in pathways associated with carcinogenesis or cell proliferation. Indeed, several of the proteins have been reported to be specific ovarian cancer markers while others are common to many tumorigenic tissues or proliferating cells. The diversity of proteins found and their association with known oncogenic pathways validate this proteomic approach. The proteome 2D map of the TOV-112D cell line will provide a valuable resource in studies on differential protein expression of human ovarian carcinomas while the 1DE LC MS/MS approach gives a picture of the actual protein profile of the TOV-112D cell line. This work represents one of the most complete ovarian protein expression analysis reports to date and the first comparative study of gene expression profiling and proteomic patterns in ovarian cancer.

Cell Line, Transformed↗

The Chemical Shift Index method applied to resin-bound peptides.

The Chemical Shift Index (CSI) method proposed by Wishart et al. [Biochemistry (1992) 31, 1647-1651] to evaluate the secondary structure of peptides in aqueous solution uses as its reference the chemical shift values of each of the 20 natural amino acids (X) in a typical nonstructured sequence GGXAGG (17-20). In order to apply the CSI method to protected resin-bound peptides, we established a new database of chemical shift values for the same GGXAGG sequences in their protected form and anchored to a polystyrene resin swollen in DMF-d7. The predictive value of this new reference set in the CSI protocol was tested on different resin-bound peptides that were previously characterized by a full NOE analysis.

Amino Acid Sequence↗

REBASE--restriction enzymes and methylases.

REBASE contains comprehensive information about restriction enzymes, DNA methylases and related proteins such as nicking enzymes, specificity subunits and control proteins. It contains published and unpublished references, recognition and cleavage sites, isoschizomers, commercial availability, methylation sensitivity, crystal data and sequence data. Homing endonucleases are also included. Most recently, extensive information about the methylation sensitivity of restriction enzymes has been added and a new feature contains complete analyses of the putative restriction systems in the sequenced bacterial and archaeal genomes. The data is distributed via email, ftp (ftp.neb.com) and the Web (http://rebase. neb.com).

Base Sequence↗

[A two-dimensional reference map of mouse ovary proteins].

OBJECTIVE: To perform a preliminary proteomic analysis of mouse ovaries and to study the protein's function in mouse ovary. METHODS: The two-dimensional gel electrophoresis (2-DE) and matrix assisted laser desorption/ionization-time of flight mass spectrometry (MALDI-TOF MS) were used to analyze mouse ovarian proteome. A 12.5% sodium dodecyl sulfate (SDS) reference gel was generated by immobilized pH gradient isoelectric focusing of mouse ovary proteins in a non-linear gradient (pH 3-10). And GRP78 was selected to perform with immunohistochemistry within mouse ovaries. RESULTS: Based on peptide mass fingerprinting, 52 proteins were identified and classified into seven functional groups: Cell/organism defense and antioxidant, cell signaling/communications proteins, cell structure/motility proteins, metabolism proteins, RNA synthesis processing, protein synthesis and processing, and unclassified proteins. The immunoreactivity of GRP78 was detected in GCs in the follicular, and with during GCs Luteinizing in the menstrual cycle, the protein expression (brown) increased continually and came to a head when ovulation happened. CONCLUSION: This work provides a first step toward the establishment of a systematic ovary protein database and stands as a valuable resource for molecular analyses of normal and pathologic conditions affecting mouse ovaries.

Animals↗

Carbohydrate supplementation of human milk to promote growth in preterm infants.

BACKGROUND: This section is under preparation and will be included in the next issue. OBJECTIVES: The main objective was to determine if addition of carbohydrate supplements to human milk leads to improved growth and neurodevelopmental outcomes without significant adverse effects in preterm infants. SEARCH STRATEGY: The standard search strategy of the Neonatal Review Group was used. This includes searches of the Oxford Database of Perinatal Trials, MEDLINE, previous reviews including cross references, abstracts, conferences and symposia proceedings, expert informants, journal handsearching mainly in the English language. SELECTION CRITERIA: All trials utilising random or quasi-random allocation evaluating the supplementation of human milk with carbohydrate in preterm infants within a nursery setting were eligible. DATA COLLECTION AND ANALYSIS: Not applicable. MAIN RESULTS: No eligible trials were found. REVIEWER'S CONCLUSIONS: There are no studies which have specifically evaluated the addition of carbohydrate alone for the purpose of improving growth and neurodevelopmental outcomes. No recommendations for practice can be made. Research should be directed towards comparison of different quantities and types of carbohydrate in multicomponent fortifiers containing protein and minerals, specifically evaluating short-term growth and long-term growth and neurodevelopmental outcomes.

Dietary Carbohydrates↗

A nifH-based oligonucleotide microarray for functional diagnostics of nitrogen-fixing microorganisms.

Nitrogen fixation is an important process in biogeochemical cycles exclusively carried out by prokaryotes, mostly by an evolutionarily conserved nitrogenase protein complex, of which one of the structural genes (nifH) is highly valuable for phylogenetic and diversity analyses. We developed a nifH-based short oligonucleotide microarray (nifH diagnostic microarray) as a rapid tool to effectively monitor nitrogen-fixing diazotrophic populations in a wide range of environments. Taking account of the overwhelming predominance of environmental nifH fragments from uncultivated microorganisms in public databases, our nifH microarray is mainly based on nifH sequences from as yet unidentified prokaryotes. Standard conditions for microarray performance were determined, and criteria for the design of specific oligonucleotides were defined. A primary set of 56 oligonucleotides was validated with fluorescence-labeled single-stranded nifH targets from five reference strains, 26 environmental clones, and artificial mixtures of reference strains. The nifH microarray was applied to analyze the diversity (based on DNA) and activity (based on mRNA) of diazotrophs in roots of wild rice samples from Namibia. Results demonstrated that only a small subset of diazotrophs being present in the sample were actually fixing nitrogen actively. Our data suggest that the developed nifH microarray is a highly reproducible and semiquantitative method for mapping the variability of diazotrophic diversity, allowing rapid comparisons of the relative abundance and activity of diazotrophic prokaryotes in the environment. A further refined nifH microarray comprising of 194 oligonucleotide probes now covers more than 90% of sequences in our nifH database.

Azoarcus↗

A novel type of conserved DNA-binding domain in the transcriptional regulators of the AlgR/AgrA/LytR family.

Sequence analysis of bacterial genomes revealed a novel DNA-binding domain. This domain is found in several response regulators of the two-component signal transduction system, such as Pseudomonas aeruginosa AlgR, involved in the regulation of alginate biosynthesis and in the pathogenesis of cystic fibrosis; Clostridium perfringens VirR, a regulator of virulence factors, and in several regulators of bacteriocin biosynthesis, previously unified in the AgrA/ComE family. Most of the transcriptional regulators that contain this DNA-binding domain are involved in biosynthesis of extracellular polysaccharides, fimbriation, expression of exoproteins, including toxins, and quorum sensing. We refer to it as the LytTR ('litter') domain, after Bacillus subtilis LytT and Staphylococcus aureus LytR response regulators, involved in regulation of cell autolysis. In addition to response regulators, the LytTR domain is found in combination with MHYT, PAS and other sensor domains.

Amino Acid Sequence↗

Reproducibility of polypeptide spot positions in two-dimensional gels run using carrier ampholytes in the isoelectric focusing dimension.

The reproducibility of complex protein patterns in two-dimensional (2-D) gels run with carrier ampholytes in the first dimension has been investigated. Two different laboratories collaborated in the study and 18 or 19 gels were run in each laboratory for comparison. The electrophoresis chemicals, running devices, and samples were standardized in both labs. The resulting 37 gels were scanned with a charge-coupled device (CCD) camera and spots were located, counted, quantified, and matched using a commercially available image analysis system. Subsequently, the reproducibility of spot position was determined. To perform the statistical analysis, the test gels were initially each matched to a master reference gel. Next, three sets of 12 gels (the image analysis software database could analyze only 12 gels at a time) were analyzed and the isoelectric point (pI) and molecular weight (M(r)) positional variation of all the spots that matched across the gels in each set was determined. The resulting statistical analysis indicates very high reproducibility of the carrier ampholyte technique.

Buffers↗