Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Bioinformatic insights from metagenomics through visualization.

Cutting-edge biological and bioinformatics research seeks a systems perspective through the analysis of multiple types of high-throughput and other experimental data for the same sample. Systems-level analysis requires the integration and fusion of such data, typically through advanced statistics and mathematics. Visualization is a complementary computational approach that supports integration and analysis of complex data or its derivatives. We present a bioinformatics visualization prototype, Juxter, which depicts categorical information derived from or assigned to these diverse data for the purpose of comparing patterns across categorizations. The visualization allows users to easily discern correlated and anomalous patterns in the data. These patterns, which might not be detected automatically by algorithms, may reveal valuable information leading to insight and discovery. We describe the visualization and interaction capabilities and demonstrate its utility in a new field, metagenomics, which combines molecular biology and genetics to identify and characterize genetic material from multi-species microbial samples.

Algorithms↗

A comprehensive whole genome bacterial phylogeny using correlated peptide motifs defined in a high dimensional vector space.

As whole genome sequences continue to expand in number and complexity, effective methods for comparing and categorizing both genes and species represented within extremely large datasets are required. Methods introduced to date have generally utilized incomplete and likely insufficient subsets of the available data. We have developed an accurate and efficient method for producing robust gene and species phylogenies using very large whole genome protein datasets. This method relies on multidimensional protein vector definitions supplied by the singular value decomposition (SVD) of a large sparse data matrix in which each protein is uniquely represented as a vector of overlapping tetrapeptide frequencies. Quantitative pairwise estimates of species similarity were obtained by summing the protein vectors to form species vectors, then determining the cosines of the angles between species vectors. Evolutionary trees produced using this method confirmed many accepted prokaryotic relationships. However, several unconventional relationships were also noted. In addition, we demonstrate that many of the SVD-derived right basis vectors represent particular conserved protein families, while many of the corresponding left basis vectors describe conserved motifs within these families as sets of correlated peptides (copeps). This analysis represents the most detailed simultaneous comparison of prokaryotic genes and species available to date.

Amino Acid Motifs↗

An SVD-based comparison of nine whole eukaryotic genomes supports a coelomate rather than ecdysozoan lineage.

BACKGROUND: Eukaryotic whole genome sequences are accumulating at an impressive rate. Effective methods for comparing multiple whole eukaryotic genomes on a large scale are needed. Most attempted solutions involve the production of large scale alignments, and many of these require a high stringency pre-screen for putative orthologs in order to reduce the effective size of the dataset and provide a reasonably high but unknown fraction of correctly aligned homologous sites for comparison. As an alternative, highly efficient methods that do not require the pre-alignment of operationally defined orthologs are also being explored. RESULTS: A non-alignment method based on the Singular Value Decomposition (SVD) was used to compare the predicted protein complement of nine whole eukaryotic genomes ranging from yeast to man. This analysis resulted in the simultaneous identification and definition of a large number of well conserved motifs and gene families, and produced a species tree supporting one of two conflicting hypotheses of metazoan relationships. CONCLUSIONS: Our SVD-based analysis of the entire protein complement of nine whole eukaryotic genomes suggests that highly conserved motifs and gene families can be identified and effectively compared in a single coherent definition space for the easy extraction of gene and species trees. While this occurs without the explicit definition of orthologs or homologous sites, the analysis can provide a basis for these definitions.

Amino Acid Motifs↗

Automatic detection of false annotations via binary property clustering.

BACKGROUND: Computational protein annotation methods occasionally introduce errors. False-positive (FP) errors are annotations that are mistakenly associated with a protein. Such false annotations introduce errors that may spread into databases through similarity with other proteins. Generally, methods used to minimize the chance for FPs result in decreased sensitivity or low throughput. We present a novel protein-clustering method that enables automatic separation of FP from true hits. The method quantifies the biological similarity between pairs of proteins by examining each protein's annotations, and then proceeds by clustering sets of proteins that received similar annotation into biological groups. RESULTS: Using a test set of all PROSITE signatures that are marked as FPs, we show that the method successfully separates FPs in 69% of the 327 test cases supplied by PROSITE. Furthermore, we constructed an extensive random FP simulation test and show a high degree of success in detecting FP, indicating that the method is not specifically tuned for PROSITE and performs well on larger scales. We also suggest some means of predicting in which cases this approach would be successful. CONCLUSION: Automatic detection of FPs may greatly facilitate the manual validation process and increase annotation sensitivity. With the increasing number of automatic annotations, the tendency of biological properties to be clustered, once a biological similarity measure is introduced, may become exceedingly helpful in the development of such automatic methods.

Algorithms↗

No simple dependence between protein evolution rate and the number of protein-protein interactions: only the most prolific interactors tend to evolve slowly.

BACKGROUND: It has been suggested that rates of protein evolution are influenced, to a great extent, by the proportion of amino acid residues that are directly involved in protein function. In agreement with this hypothesis, recent work has shown a negative correlation between evolutionary rates and the number of protein-protein interactions. However, the extent to which the number of protein-protein interactions influences evolutionary rates remains unclear. Here, we address this question at several different levels of evolutionary relatedness. RESULTS: Manually curated data on the number of protein-protein interactions among Saccharomyces cerevisiae proteins was examined for possible correlation with evolutionary rates between S. cerevisiae and Schizosaccharomyces pombe orthologs. Only a very weak negative correlation between the number of interactions and evolutionary rate of a protein was observed. Furthermore, no relationship was found between a more general measure of the evolutionary conservation of S. cerevisiae proteins, based on the taxonomic distribution of their homologs, and the number of protein-protein interactions. However, when the proteins from yeast were assorted into discrete bins according to the number of interactions, it turned out that 6.5% of the proteins with the greatest number of interactions evolved, on average, significantly slower than the rest of the proteins. Comparisons were also performed using protein-protein interaction data obtained with high-throughput analysis of Helicobacter pylori proteins. No convincing relationship between the number of protein-protein interactions and evolutionary rates was detected, either for comparisons of orthologs from two completely sequenced H. pylori strains or for comparisons of H. pylori and Campylobacter jejuni orthologs, even when the proteins were classified into bins by the number of interactions. CONCLUSION: The currently available comparative-genomic data do not support the hypothesis that the evolutionary rates of the majority of proteins substantially depend on the number of protein-protein interactions they are involved in. However, a small fraction of yeast proteins with the largest number of interactions (the hubs of the interaction network) tend to evolve slower than the bulk of the proteins.

Bacterial Proteins↗

NovelFam3000--uncharacterized human protein domains conserved across model organisms.

BACKGROUND: Despite significant efforts from the research community, an extensive portion of the proteins encoded by human genes lack an assigned cellular function. Most metazoan proteins are composed of structural and/or functional domains, of which many appear in multiple proteins. Once a domain is characterized in one protein, the presence of a similar sequence in an uncharacterized protein serves as a basis for inference of function. Thus knowledge of a domain's function, or the protein within which it arises, can facilitate the analysis of an entire set of proteins. DESCRIPTION: From the Pfam domain database, we extracted uncharacterized protein domains represented in proteins from humans, worms, and flies. A data centre was created to facilitate the analysis of the uncharacterized domain-containing proteins. The centre both provides researchers with links to dispersed internet resources containing gene-specific experimental data and enables them to post relevant experimental results or comments. For each human gene in the system, a characterization score is posted, allowing users to track the progress of characterization over time or to identify for study uncharacterized domains in well-characterized genes. As a test of the system, a subset of 39 domains was selected for analysis and the experimental results posted to the NovelFam3000 system. For 25 human protein members of these 39 domain families, detailed sub-cellular localizations were determined. Specific observations are presented based on the analysis of the integrated information provided through the online NovelFam3000 system. CONCLUSION: Consistent experimental results between multiple members of a domain family allow for inferences of the domain's functional role. We unite bioinformatics resources and experimental data in order to accelerate the functional characterization of scarcely annotated domain families.

Animals↗

Identification of components of protein complexes.

Protocols are given for a variety of techniques used in protein identification of complexes, including identification of in-gel separated proteins and LC-MS/MS. Gels, staining procedures, and peptide extraction protocols that are compatible with mass spectrometry are described. The detection limits of the various staining procedures and their compatibility with mass spectrometry are discussed. The various mass spectrometric techniques used (MALDI-MS, MALDI-MS/MS, nanospray, and ESI/LC-MS/MS) are also described, along with an indication of the advantages and disadvantages of each, and when they would most appropriately be used. Common pitfalls associated with database searching are also discussed.

Animals↗

Feature selection and the class imbalance problem in predicting protein function from sequence.

When the standard approach to predict protein function by sequence homology fails, other alternative methods can be used that require only the amino acid sequence for predicting function. One such approach uses machine learning to predict protein function directly from amino acid sequence features. However, there are two issues to consider before successful functional prediction can take place: identifying discriminatory features, and overcoming the challenge of a large imbalance in the training data. We show that by applying feature subset selection followed by undersampling of the majority class, significantly better support vector machine (SVM) classifiers are generated compared with standard machine learning approaches. As well as revealing that the features selected could have the potential to advance our understanding of the relationship between sequence and function, we also show that undersampling to produce fully balanced data significantly improves performance. The best discriminating ability is achieved using SVMs together with feature selection and full undersampling; this approach strongly outperforms other competitive learning algorithms. We conclude that this combined approach can generate powerful machine learning classifiers for predicting protein function directly from sequence.

Algorithms↗

Integrating medical and genomic data: a successful example for rare diseases.

The recent advances on genomics and proteomics research bring up a significant grow on the information that is publicly available. However, navigating through genetic and bioinformatics databases can be a too complex and unproductive task for a primary care physician. In this paper we present diseasecard, a web portal for rare disease that provides transparently to the user a virtually integration of distributed and heterogeneous information.

Computational Biology↗

Pathogenomics analysis of Leishmania spp.: flagellar gene families of putative virulence factors.

The trypanosomatid flagellar apparatus contains conventional and unique features, whose roles in infectivity are still enigmatic. Although the flagellum and the flagellar pocket are critical organelles responsible for all vesicular trafficking between the cytoplasm and cell surface, still very little is known about their roles in pathogenesis and how molecules get to and from the flagellar pocket. The ongoing analysis of the genome sequences and proteome profiles of Leishmania major and L infantum, Trypanosoma cruzi, T. brucei, and T. gambiensi ( www.genedb.org ), coupled with our own work on L. chagasi (as part of the Brazilian Northeast Genome Program- www.progene.ufpe.br ), prompted us to scrutinize flagellar genes and proteins of Leishmania spp. promastigotes that could be virulence factors in leishmaniasis. We have identified some overlooked parasite factors such as the MNUDC-1 (a protein involved in nuclear development and genomic fusion) and SQS (an enzyme of sterol biosynthesis), among the described flagellar gene families. A database concerning the results of this work, as well as of other studies of Leishmania and its organelles, is available at http://nugen.lcc.uece.br/LPGate . It will serve as a convenient bioinformatics resource on genomics and pathology of the etiological agents of leishmaniasis.

Amino Acid Sequence↗

Selective analysis of phosphopeptides within a protein mixture by chemical modification, reversible biotinylation and mass spectrometry.

A new method combining chemical modification and affinity purification is described for the characterization of serine and threonine phosphopeptides in proteins. The method is based on the conversion of phosphoserine and phosphothreonine residues to S-(2-mercaptoethyl)cysteinyl or beta-methyl-S-(2-mercaptoethyl)cysteinyl residues by beta-elimination/1,2-ethanedithiol addition, followed by reversible biotinylation of the modified proteins. After trypsin digestion, the biotinylated peptides were affinity-isolated and enriched, and subsequently subjected to structural characterization by liquid chromatography/tandem mass spectrometry (LC/MS/MS). Database searching allowed for automated identification of modified residues that were originally phosphorylated. The applicability of the method is demonstrated by the identification of all known phosphorylation sites in a mixture of alpha-casein, beta-casein, and ovalbumin. The technique has potential for adaptations to proteome-wide analysis of protein phosphorylation.

Amino Acid Sequence↗

Integration with the human genome of peptide sequences obtained by high-throughput mass spectrometry.

A crucial aim upon the completion of the human genome is the verification and functional annotation of all predicted genes and their protein products. Here we describe the mapping of peptides derived from accurate interpretations of protein tandem mass spectrometry (MS) data to eukaryotic genomes and the generation of an expandable resource for integration of data from many diverse proteomics experiments. Furthermore, we demonstrate that peptide identifications obtained from high-throughput proteomics can be integrated on a large scale with the human genome. This resource could serve as an expandable repository for MS-derived proteome information.

Amino Acid Sequence↗

Discriminant models for high-throughput proteomics mass spectrometer data.

We use several different multivariate analysis methods to discriminate between diseased and healthy patients using protein mass spectrometer data provided by Duke University. Two problems were presented by the university; one in which the responses (diseased or healthy) of the patients were not known and second, when the responses were known. In the latter case, the data can be used as a 'training' set. We attempted both problems. In particular, we use principle component analysis along with clustering methods to discriminate for the first problem set and partial least squares coupled with logistic and discriminant methods when the responses were known. In addition, we were able to detect regions of interest in the spectrum where there were differences in the protein patterns between healthy and diseased patients. There was considerable effort involved in the preprocessing of the data. We used a binning approach to reduce the number of variables rather than peak heights or peak areas. We performed a square root transformation on the data to help stabilize the variance; this in turn made a significant improvement in clustering results.

Computational Biology↗

PROTICdb: a web-based application to store, track, query, and compare plant proteome data.

PROTICdb is a web-based application, mainly designed to store and analyze plant proteome data obtained by two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) and mass spectrometry (MS). The purposes of PROTICdb are (i) to store, track, and query information related to proteomic experiments, i.e., from tissue sampling to protein identification and quantitative measurements, and (ii) to integrate information from the user's own expertise and other sources into a knowledge base, used to support data interpretation (e.g., for the determination of allelic variants or products of post-translational modifications). Data insertion into the relational database of PROTICdb is achieved either by uploading outputs of image analysis and MS identification software, or by filling web forms. 2-D PAGE annotated maps can be displayed, queried, and compared through a graphical interface. Links to external databases are also available. Quantitative data can be easily exported in a tabulated format for statistical analyses. PROTICdb is based on the Oracle or the PostgreSQL Database Management System and is freely available upon request at the following URL: http://moulon.inra.fr/ bioinfo/PROTICdb.

Computer Graphics↗

Proteomic analysis of cellular responses to low concentration N-methyl-N'-nitro-N-nitrosoguanidine in human amnion FL cells.

We have shown previously that exposure to a low concentration of N-methyl-N'-nitro-N-nitrosoguanidine (MNNG) induces comprehensive changes in the protein expression profile of human amnion FL cells, including the induction, suppression, upregulation, and downregulation of various proteins. In addition, by proteomic analysis combining two-dimensional gel electrophoresis (2-DE) and mass spectrometry, some of the induced and suppressed proteins were identified. In this study, we identified an additional 18 proteins among those that were either up- or downregulated by MNNG treatment. The proteins identified were a heterogeneous group that included several zinc finger proteins, proteins involved in signal transduction, cytoskeletal proteins, cell-cycle regulation proteins, and proteins with unknown functions. The involvement of these proteins in the cellular responses to alkylating agents has not been reported before and their physiological relevance is not clear. Therefore, our findings may help better understand the global cellular stress responses to chemical carcinogens, and may lead to new studies on the functions of these MNNG-responsive proteins. Furthermore, some of these proteins may serve as biomarkers for detecting exposure of human populations to environmental carcinogens.

Amnion↗