Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa↗

InSilicoSpectro: an open-source proteomics library.

We present a new proteomics open-source project, InSilicoSpectro, aimed at implementing recurrent computations that are necessary for proteomics data analysis. Illustrative examples are mass list file format conversions, protein sequence digestion, theoretical peptide and fragment mass computations, graphical display, matching with experimental data, isoelectric point estimation, and peptide retention time prediction. The project library is written in Perl, a widely used scripting language in bioinformatics, and it offers a unique framework of integrated objects to implement complex proteomics data analyses. For instance, only a few lines of code are required to digest a protein with fixed and variable modifications, label peptides with 18O, compute the fragmentation spectra and display their match with experimental spectra. We believe that InSilicoSpectro will be of great help to bioinformaticians, without detailed knowledge of proteomics specifics, and to mass spectrometrists with computer programming interest as well.

Amino Acid Sequence↗

A comprehensive protein resource for the study of bladder cancer: http://biobase.dk/cgi-bin/celis.

In our laboratories we are exploring the possibility of using proteome expression profiles of fresh bladder tumors (transitional cell carcinomas, TCCs; squamous cell carcinomas, SCCs) and random biopsies as fingerprints to subclassify histopathological types and as a starting point to search for protein markers that may form the basis for diagnosis, prognosis, and treatment. Ultimately, the goal of these studies is to identify signaling pathways and components that are affected at various stages of bladder cancer progression and that may provide novel leads in drug discovery. Here we present our ongoing efforts to establish comprehensive two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) databases of TCCs and SCCs which are being constructed based on the proteomic and immunohistochemical analysis of hundreds of fresh tumors, random biopsies and cystectomies received shortly after operation (http://biobase.dk/cgi-bin/celis).

Animals↗

NLSdb: database of nuclear localization signals.

NLSdb is a database of nuclear localization signals (NLSs) and of nuclear proteins. NLSs are short stretches of residues mediating transport of nuclear proteins into the nucleus. The database contains 114 experimentally determined NLSs that were obtained through an extensive literature search. Using 'in silico mutagenesis' this set was extended to 308 experimental and potential NLSs. This final set matched over 43% of all known nuclear proteins and matches no currently known non-nuclear protein. NLSdb contains over 6000 predicted nuclear proteins and their targeting signals from the PDB and SWISS-PROT/TrEMBL databases. The database also contains over 12 500 predicted nuclear proteins from six entirely sequenced eukaryotic proteomes (Homo sapiens, Mus musculus, Drosophila melanogaster, Caenorhabditis elegans, Arabidopsis thaliana and Saccharomyces cerevisiae). NLS motifs often co-localize with DNA-binding regions. This observation was used to also annotate over 1500 DNA-binding proteins. NLSdb can be accessed via the web site: http://cubic.bioc.columbia.edu/db/NLSdb/.

Active Transport, Cell Nucleus↗

GenAge: a genomic and proteomic network map of human ageing.

The aim of this work was to provide an overview of the genetics of human ageing to gain novel insights about the mechanisms involved. By incorporating findings from model organisms to humans, such as mutations that either delay or accelerate ageing in mice, we constructed the gene networks previously related to ageing: namely, the network related to DNA metabolism and the network involving the GH/IGF-1 axis. Gathering data about the interacting partners of these proteins allowed us to suggest the involvement in ageing of a number of proteins through a "guilt-by-association" methodology. To organize our data, we developed the first curated database of genes related to human ageing: GenAge. With over 200 entries, GenAge may serve as a reference database of genes related to human ageing. Moreover, we rendered the first proteomic network map of human ageing, which suggests a relationship between the genetics of development and the genetics of ageing. Our work serves as a framework upon which a systems-biology understanding of ageing can be developed. GenAge is freely available for academic purposes at: http://genomics.senescence.info/genes/.

Aging↗

Large-scale analysis of 73 329 physcomitrella plants transformed with different gene disruption libraries: production parameters and mutant phenotypes.

Gene targeting in the moss Physcomitrella patens has created a new platform for plant functional genomics. We produced a mutant collection of 73 329 Physcomitrella plants and evaluated the phenotype of each transformant in comparison to wild type Physcomitrella. Production parameters and morphological changes in 16 categories, such as plant structure, colour, coverage with gametophores, cell shape, etc., were listed and all data were compiled in a database (mossDB). Our mutant collection consists of at least 1804 auxotrophic mutants which showed growth defects on minimal Knop medium but were rescued on supplemented medium. 8129 haploid and 11 068 polyploid transformants had morphological alterations. 9 % of the haploid transformants had deviations in the leaf shape, 7 % developed less gametophores or had a different leaf cell shape. Other morphological deviations in plant structure, colour, and uniformity of leaves on a moss colony were less frequently observed. Preculture conditions of the plant material and the cDNA library (representing genes from either protonema, gametophore or sporophyte tissue) used to transform Physcomitrella had an effect on the number of transformants per transformation. We found correlations between ploidy level and plant morphology and growth rate on Knop medium. In haploid transformants correlations between the percentage of plants with specific phenotypes and the cDNA library used for transformation were detected. The number of different cDNAs present during transformation had no effect on the number of transformants per transformation, but it had an effect on the overall percentage of plants with phenotypic deviations. We conclude that by linking incoming molecular, proteome, and metabolome data of the transformants in the future, the database mossDB will be a valuable biological resource for systems biology.

Bryopsida↗

A database of protein expression in lung cancer.

We have developed a comprehensive approach to identifying molecular changes in lung cancer that includes both genomic and proteomic analyses. The related effort has produced a large amount of data pertaining to gene expression at the RNA and protein levels. As a result, we have constructed a database that contains protein expression data on lung cancer as well as other relevant data including DNA microarray derived data. A large number of proteins that are expressed in different types of lung cancer have been identified and have been correlated with the expression measures for their corresponding genes at the RNA level. The database is intended to facilitate our effort at developing novel classification schemes for lung cancer and the identification of novel markers for early diagnosis.

Biomarkers, Tumor↗

Fast approximate motif statistics.

We present in this article a fast approximate method for computing the statistics of a number of non-self-overlapping matches of motifs in a random text in the nonuniform Bernoulli model. This method is well suited for protein motifs where the probability of self-overlap of motifs is small. For 96% of the PROSITE motifs, the expectations of occurrences of the motifs in a 7-million-amino-acids random database are computed by the approximate method with less than 1% error when compared with the exact method. Processing of the whole PROSITE takes about 30 seconds with the approximate method. We apply this new method to a comparison of the C. elegans and S. cerevisiae proteomes.

Amino Acid Motifs↗

A combination of chemical derivatisation and improved bioinformatic tools optimises protein identification for proteomics.

The identification of individual protein species within an organism's proteome has been optimised by increasing the information produced from mass spectral analysis through the chemical derivatisation of tryptic peptides and the development of new software tools. Peptide fragments are subjected to two forms of derivatisation. First, lysine residues are converted to homoarginine moieties by guanidination. This procedure has two advantages, first, it usually identifies the C-terminal amino acid of the tryptic peptide and also greatly increases the total information content of the mass spectrum by improving the signal response of C-terminal lysine fragments. Second, an Edman-type phenylthiocarbamoyl (PTC) modification is carried out on the N-terminal amino acid. The renders the first peptide bond highly susceptible to cleavage during mass spectrometry (MS) analysis and consequently allows the ready identification of the N-terminal residue. The utility of the procedure has been demonstrated by developing novel bioinformatic tools to exploit the additional mass spectral data in the identification of proteome proteins from the yeast Saccharomyces cerevisiae. With this combination of novel chemistry and bioinformatics, it should be possible to identify unambiguously any yeast protein spot or band from either two-dimensional or one-dimensional electropheretograms.

Databases, Factual↗

Microorganism identification by matrix-assisted laser/desorption ionization mass spectrometry and model-derived ribosomal protein biomarkers.

An improved data analysis method is described for rapid identification of intact microorganisms from MALDI-TOF-MS data. The method makes no use of mass spectral fingerprints. Instead, a microorganism database is automatically generated that contains biomarker masses derived from ribosomal protein sequences and a model of N-terminal Met loss. We quantitatively validate the method via a blind study that seeks to identify microorganisms with known ribosomal protein sequences. We also include in the database microorganisms with incompletely known sets of ribosomal proteins to test the specificity of the method. With an optimal MALDI protocol, and at the 95% confidence level, microorganisms represented in the database with 20 or more biomarkers (i.e., those with complete or nearly completely sequenced genomes) are correctly identified from their spectra 100% of the time, with no incorrect identifications. Microorganisms with seven or less biomarkers (i.e., incompletely sequenced genomes) are either not identified or misidentified. Robustness with respect to variations in sample preparation protocol and mass analysis protocol is demonstrated by collecting data with two different matrixes and under two different ion-mode configurations. Statistical analysis suggests that, even without further improvement, the method described here would successfully scale up to microorganism databases with roughly 1000 microorganisms. The results demonstrate that microorganism identification based on proteome data and modeling can perform as well as methods based on mass spectral fingerprinting.

Bacteria↗

The power and the limitations of cross-species protein identification by mass spectrometry-driven sequence similarity searches.

Mass spectrometry-driven BLAST (MS BLAST) is a database search protocol for identifying unknown proteins by sequence similarity to homologous proteins available in a database. MS BLAST utilizes redundant, degenerate, and partially inaccurate peptide sequence data obtained by de novo interpretation of tandem mass spectra and has become a powerful tool in functional proteomic research. Using computational modeling, we evaluated the potential of MS BLAST for proteome-wide identification of unknown proteins. We determined how the success rate of protein identification depends on the full-length sequence identity between the queried protein and its closest homologue in a database. We also estimated phylogenetic distances between organisms under study and related reference organisms with completely sequenced genomes that allow substantial coverage of unknown proteomes.

Animals↗

CHOP proteins into structural domain-like fragments.

We developed a method CHOP dissecting proteins into domain-like fragments. The basic idea was to cut proteins beginning from very reliable experimental information (PDB), proceeding to expert annotations of domain-like regions (Pfam-A), and completing through cuts based on termini of known proteins. In this way, CHOP dissected more than two thirds of all proteins from 62 proteomes. Analysis of our structural domain-like fragments revealed four surprising results. First, >70% of all dissected proteins contained more than one fragment. Second, most domains spanned on average over approximately 100 residues. This average was similar for eukaryotic and prokaryotic proteins, and it is also valid-although previously not described-for all proteins in the PDB. Third, single-domain proteins were significant longer than most domains in multidomain proteins. Fourth, three fourths of all domains appeared shorter than 210 residues. We believe that our CHOP fragments constituted an important resource for functional and structural genomics. Nevertheless, our main motivation to develop CHOP was that the single-linkage clustering method failed to adequately group full-length proteins. In contrast, CLUP-the simple clustering scheme CLUP introduced here-succeeded largely to group the CHOP fragments from 62 proteomes such that all members of one cluster shared a basic structural core. CLUP found >63,000 multi- and >118,000 single-member clusters. Although most fragments were restricted to a particular cluster, approximately 24% of the fragments were duplicated in at least two clusters. Our thresholds for grouping two fragments into the same cluster were rather conservative. Nevertheless, our results suggested that structural genomics initiatives have to target >30,000 fragments to at least cover the multimember clusters in 62 proteomes.

Amino Acids↗

Advances in mass spectrometry for proteome analysis.

The most demanding problems in proteomics continue to challenge modern mass spectrometry. Recent developments in instrument design have led to lower limits of detection, while new ion activation techniques and improved understanding of gas-phase ion chemistry have enhanced the capabilities of tandem mass spectrometry for peptide and protein structure elucidation. Future developments must address the., understanding of protein-protein interactions and the characterisation of the dynamic proteome.

Databases, Factual↗

Using ontologies in PROTEUS for modeling proteomics data mining applications.

Bioinformatics applications are often characterized by a combination of (pre) processing of raw data representing biological elements, (e.g. sequence alignment, structure prediction), and an high level data mining analysis. Developing such applications needs knowledge of both data mining and bioinformatics domains, that can be effectively achieved by combining ontology about the application domain and ontology about the approaches and processes to solve the given problem. In this paper we talk about using ontologies to model proteomics in silico experiments. In particular data mining of mass spectrometry proteomics data is considered.

Computational Biology↗

Reliable automatic protein identification from matrix-assisted laser desorption/ionization mass spectrometric peptide fingerprints.

Matrix-assisted laser desorption/ionization (MALDI) mass spectrometry of protein samples from two-dimensional (2-D) gels in conjunction with protein sequence database searches is frequently used to identify proteins. Moreover, the automatic analysis of complete 2-D gels with hundreds and even thousands of protein spots ("proteome analysis") is possible, without human intervention, with the availability of highly accurate mass spectrometry instruments, and high-throughput facilities for preparation and handling of protein samples from 2-D gels. However, the lack of software for precise automatic analysis and annotation of mass spectra, as well as software for in-batch sequence database queries, is increasingly becoming a significant bottleneck for the proteomics work flow. In the present paper we outline an algorithm for reliable, accurate, and automatic evaluation of mass spectrometric data and database searches. We show here that simply selecting from the sequence database the protein that has the most matching fragment masses often leads to false-positive results. Reliable protein identification is dependent on several parameters: the accuracy of fragment mass determination, the number of masses submitted for query, the mass distribution of query masses, the number of masses matching between sample and database protein, the size of the sequence database, and the kind and number of modifications considered. Using these parameters, we derive a simple statistical estimation that can be used to calculate the probability of true-positive protein identification.

Automation↗

The Human Proteome Organization Plasma Proteome Project pilot phase: reference specimens, technology platform comparisons, and standardized data submissions and analyses.

A comprehensive, systematic characterization of cirolating proteins in health and disease will greatly facilitate development of biomarkers for prevention, diagnosis, and therapy of cancers and other diseases. The Human Proteome Organization Plasma Proteome Project pilot phase aims to (1) compare the advantages and limitations of many technology platforms; (2) contrast reference specimens of human plasma (ethylenediaminetetra acetic acid, heparin, citrate-anticoagulated) and serum, in terms of numbers of proteins identified and any interferences with various technology platforms; and (3) create a global knowledge base/data repository.

Biomarkers↗