Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

In silico discovery of human natural antisense transcripts.

BACKGROUND: Several high-throughput searches for potential natural antisense transcripts (NATs) have been performed recently, but most of the reports were focused on cis type. A thorough in silico analysis of human transcripts will help expand our knowledge of NATs. RESULTS: We have identified 568 NATs from human RefSeq RNA sequences. Among them, 403 NATs are reported for the first time, and at least 157 novel NATs are trans type. According to the pairing region of a sense and antisense RNA pair, hNATs are divided into 6 classes, of which about 87% involve 5' or 3' UTR sequences, supporting the regulatory role of UTRs. Among a total of 535 NAT pairs related with splice variants, 77.4% (414/535) have their pairing regions affected or completely eliminated by alternative splicing, suggesting significant relationship of alternative splicing and antisense-directed regulation. The extensive occurrence of splice variants in hNATs and other multiple pairing patterns results in a one-to-many relationship, allowing the formation of complex regulation networks. Based on microarray data from Stanford Microarray Database, two hNAT pairs were found to display significant inverse expression patterns before and after insulin injection. CONCLUSION: NATs might carry out more extensive and complex functions than previously thought. Combined with endogenous micro RNAs, hNATs could be regarded as a special group of transcripts contributing to the complex regulation networks.

Algorithms↗

Plant metabolomics: potential for practical operation.

In the postgenomic era, metabolomics is expected to be the newest useful omics science for functional genomics. However, in plant science, the present metabolomics technology cannot be considered a universal tool to perfectly elucidate perturbations imposed on sample plants although this is desired by plant physiologists. Despite it being an immature technology, metabolomics has already been used as a powerful tool for precise phenotyping, particularly for industrial application. Metabolomics is the best technology for the analysis of large mutant or transgenic libraries of model experimental plants, such as Arabidopsis, rice, etc. Here, we review the applications and technical problems of metabolomics. We also suggest the potential of metabolomics for plant post-genomic science.

Computational Biology↗

Intrinsic protein disorder in complete genomes.

Intrinsic protein disorder refers to segments or to whole proteins that fail to fold completely on their own. Here we predicted disorder on protein sequences from 34 genomes, including 22 bacteria, 7 archaea, and 5 eucaryotes. Predicted disordered segments > or = 50, > or = 40, and > or = 30 in length were determined as well as proteins estimated to be wholly disordered. The five eucaryotes were separated from bacteria and archaea by having the highest percentages of sequences predicted to have disordered segments > or = 50 in length: from 25% for Plasmodium to 41% for Drosophila. Estimates of wholly disordered proteins in the bacteria ranged from 1% to 8%, averaging to 3 +/- 2%, estimates in various archaea ranged from 2 to 11%, plus an apparently anomalous 18%, averaging to 7 +/- 5% that drops to 5 +/- 3% if the high value is discarded. Estimates in the 5 eucarya ranged from 3 to 17%. The putative wholly disordered proteins were often ribosomal proteins, but in addition about equal numbers were of known and unknown function. Overall, intrinsic disorder appears to be a common, with eucaryotes perhaps having a higher percentage of native disorder than archaea or bacteria.

Animals↗

Evaluation of sequest result filter-Xcorr and Unified Score.

OBJECTIVE: To estimate the effect of two simple filters, two or more positive peptide filter and Unified Score filter on the true positive rate of protein and peptide. METHODS: Twenty-two LC-MS/MS datasets were from 18 known protein mixture. Two or more positive peptide filter and Unified Score filter were applied to the 22 datasets. The filters effect was evaluated according to the true positive rate of protein and peptide for each filter. RESULTS: The positive rates of protein and peptide from two or more peptide filter raised from 56.49% to 92.86%-99.12% (for protein) and from 90.67% to 97.74%-99.62% (for peptide), but many positive proteins were filtered out. The positive rates of protein and peptide from Unified Score (ThermoFinnigan value 2400) were only about 35.51% and 82.99%, but after adjusted the value (3900) according to the number of false positive peptide, those positive rate raised to 63.61% (for protein) and 91.97% (for peptide). CONCLUSIONS: Two or more peptides requirement could significantly decrease false positive rate, but it also may filter out many true positive proteins especially low molecular weight and less abundant proteins. Unified Score may be a better filter than Xcorr and DeltaCn combination and the value of 3900 is found to be more suitable for this particular datasets.

Algorithms↗

Improving computational predictions of cis-regulatory binding sites.

The location of cis-regulatory binding sites determine the connectivity of genetic regulatory networks and therefore constitute a natural focal point for research into the many biological systems controlled by such regulatory networks. Accurate computational prediction of these binding sites would facilitate research into a multitude of key areas, including embryonic development, evolution, pharmacogenemics, cancer and many other transcriptional diseases, and is likely to be an important precursor for the reverse engineering of genome wide, genetic regulatory networks. Many algorithmic strategies have been developed for the computational prediction of cis-regulatory binding sites but currently all approaches are prone to high rates of false positive predictions, and many are highly dependent on additional information, limiting their usefulness as research tools. In this paper we present an approach for improving the accuracy of a selection of established prediction algorithms. Firstly, it is shown that species specific optimization of algorithmic parameters can, in some cases, significantly improve the accuracy of algorithmic predictions. Secondly, it is demonstrated that the use of non-linear classification algorithms to integrate predictions from multiple sources can result in more accurate predictions. Finally, it is shown that further improvements in prediction accuracy can be gained with the use of biologically inspired post-processing of predictions.

Algorithms↗

Protein identification from two-dimensional gel electrophoresis analysis of Klebsiella pneumoniae by combined use of mass spectrometry data and raw genome sequences.

Separation of proteins by two-dimensional gel electrophoresis (2-DE) coupled with identification of proteins through peptide mass fingerprinting (PMF) by matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS) is the widely used technique for proteomic analysis. This approach relies, however, on the presence of the proteins studied in public-accessible protein databases or the availability of annotated genome sequences of an organism. In this work, we investigated the reliability of using raw genome sequences for identifying proteins by PMF without the need of additional information such as amino acid sequences. The method is demonstrated for proteomic analysis of Klebsiella pneumoniae grown anaerobically on glycerol. For 197 spots excised from 2-DE gels and submitted for mass spectrometric analysis 164 spots were clearly identified as 122 individual proteins. 95% of the 164 spots can be successfully identified merely by using peptide mass fingerprints and a strain-specific protein database (ProtKpn) constructed from the raw genome sequences of K. pneumoniae. Cross-species protein searching in the public databases mainly resulted in the identification of 57% of the 66 high expressed protein spots in comparison to 97% by using the ProtKpn database. 10 dha regulon related proteins that are essential for the initial enzymatic steps of anaerobic glycerol metabolism were successfully identified using the ProtKpn database, whereas none of them could be identified by cross-species searching. In conclusion, the use of strain-specific protein database constructed from raw genome sequences makes it possible to reliably identify most of the proteins from 2-DE analysis simply through peptide mass fingerprinting.

Journal Article↗

Lipid mediator informatics and proteomics in inflammation resolution.

Lipid mediator informatics is an emerging area denoted to the identification of bioactive lipid mediators (LMs) and their biosynthetic profiles and pathways. LM informatics and proteomics applied to inflammation, systems tissues research provides a powerful means of uncovering key biomarkers for novel processes in health and disease. By incorporating them with system biology analysis, we review here our initial steps toward elucidating relationships among a range of bimolecular classes and provide an appreciation of their roles and activities in the pathophysiology of disease. LM informatics employing liquid chromatography-ultraviolet-tandem mass spectrometry (LC-UV-MS/MS), gas chromatography-mass spectrometry (GC-MS), computer-based automated systems equipped with databases and novel searching algorithms, and enzyme-linked immunosorbent assay (ELISA) to evaluate and profile temporal and spatial production of mediators combined with proteomics at defined points during experimental inflammation and its resolution enable us to identify novel mediators in resolution. The automated system including databases and searching algorithms is crucial for prompt and accurate analysis of these lipid mediators biosynthesized from precursor polyunsaturated fatty acids such as eicosanoids, resolvins, and neuroprotectins, which play key roles in human physiology and many prevalent diseases, especially those related to inflammation. This review presents detailed protocols used in our lab for LM informatics and proteomics using LC-UV-MS/MS, GC-MS, ELISA, novel databases and searching algorithms, and 2-dimensional gel electrophoresis and LC-nanospray-MS/MS peptide mapping.

Algorithms↗

Quality classification of tandem mass spectrometry data.

UNLABELLED: Peptide identification by tandem mass spectrometry is an important tool in proteomic research. Powerful identification programs exist, such as SEQUEST, ProICAT and Mascot, which can relate experimental spectra to the theoretical ones derived from protein databases, thus removing much of the manual input needed in the identification process. However, the time-consuming validation of the peptide identifications is still the bottleneck of many proteomic studies. One way to further streamline this process is to remove those spectra that are unlikely to provide a confident or valid peptide identification, and in this way to reduce the labour from the validation phase. RESULTS: We propose a prefiltering scheme for evaluating the quality of spectra before the database search. The spectra are classified into two classes: spectra which contain valuable information for peptide identification and spectra that are not derived from peptides or contain insufficient information for interpretation. The different spectral features developed for the classification are tested on a real-life material originating from human lymphoblast samples and on a standard mixture of 9 proteins, both labelled with the ICAT-reagent. The results show that the prefiltering scheme efficiently separates the two spectra classes.

Algorithms↗

Publishing large proteome datasets: scientific policy meets emerging technologies.

Currently, there are various approaches to proteomic analyses based on either 2D gel or HPLC separation platforms, generating data of different formats, structures and types. Identification of these separated proteins or peptide fragments is typically achieved by mass spectrometry (MS) measurements that use either accurate mass measurements or fragmentation (MS-MS) information. Integrating the information generated from these different platforms is essential if proteomics is to succeed. A further challenge lies in generating standards that can accept the hundreds-of-thousands of mass spectra produced per analysis based on threshold or probability measurements. Finally, peer review and electronic publication processes will be crucial to the dissemination and use of proteomic information. Merging the policy requirements of data-intensive research with information technology will enable scientists to gain real value from global proteomics information.

Chromatography, Liquid↗

Information management for the study of allergies.

Microarrays and other large-scale screening technologies produce quantities of increasingly complex allergy data. These data link molecular and clinical measurements and observations and provide fertile ground for improving our understanding of the processes involved in allergic reactions. Information technology is employed in gathering, storage, retrieval and analysis of these data. The increasing proportion of allergy data are generated from genomics and proteomics approaches. The major activity focuses on characterization of allergens including IgE reactivity, structural properties, and mapping of IgE and T-cell epitopes. Because of the complexity of allergy data, their utilization requires bioinformatics approaches. Allergen data are stored in the general and specialist databases. At least a dozen of important allergen databases and data repositories have been developed to date. These data are analysed using general and specialist bioinformatics tools. The major applications of bioinformatics include support for allergen characterization, assessment of allergenicity, and identification of allergic cross-reactivity. These applications in turn support the development of vaccines and therapies for allergic disease. In this article we review allergen databases and tools for the analysis of allergens, and discuss the new directions in the field supported by large scale screening involving genomics, proteomics, and bioinformatics support.

Allergens↗

AMPDB: the Arabidopsis Mitochondrial Protein Database.

The Arabidopsis Mitochondrial Protein Database is an Internet-accessible relational database containing information on the predicted and experimentally confirmed protein complement of mitochondria from the model plant Arabidopsis thaliana (http://www.ampdb.bcs.uwa.edu.au/). The database was formed using the total non-redundant nuclear and organelle encoded sets of protein sequences and allows relational searching of published proteomic analyses of Arabidopsis mitochondrial samples, a set of predictions from six independent subcellular-targeting prediction programs, and orthology predictions based on pairwise comparison of the Arabidopsis protein set with known yeast and human mitochondrial proteins and with the proteome of Rickettsia. A variety of precomputed physical-biochemical parameters are also searchable as well as a more detailed breakdown of mass spectral data produced from our proteomic analysis of Arabidopsis mitochondria. It contains hyperlinks to other Arabidopsis genomic resources (MIPS, TIGR and TAIR), which provide rapid access to changing gene models as well as hyperlinks to T-DNA insertion resources, Massively Parallel Signature Sequencing (MPSS) and Genome Tiling Array data and a variety of other Arabidopsis online resources. It also incorporates basic analysis tools built into the query structure such as a BLAST facility and tools for protein sequence alignments for convenient analysis of queried results.

Arabidopsis Proteins↗

Unveiling the Molecular Secrets of Seaweeds: A Comprehensive Review of Bioinformatics Applications in Algal Research.

Recent advances in high-throughput sequencing, bioinformatics, and multi-omics technologies have transformed seaweed research by overcoming long-standing challenges associated with complex genomes, diverse life cycles, and limited genomic resources. This review provides a comprehensive overview of bioinformatics approaches used to investigate seaweed genomics, transcriptomics, proteomics, metabolomics, microbiomes, and functional genomics, with emphasis on the computational tools and databases that support these analyses. Applications of bioinformatics in phylogenetics, drug discovery, microbiome characterization, and the development of biofuels, nutraceuticals, pharmaceuticals, and sustainable agriculture are also discussed. Particular attention is given to emerging strategies involving multi-omics integration, genome editing, artificial intelligence, machine learning, and synthetic biology that are reshaping seaweed research. The review further examines current challenges, including incomplete genomic resources, data standardization, and the need for experimental validation of computational predictions. Collectively, these advances highlight the growing role of bioinformatics in enabling systems-level understanding of seaweed biology and accelerating their translation into sustainable biotechnological and marine bioeconomy applications.

macroalgal genomics↗

Comparative proteomic analysis of high cell density cultivations with two recombinant Bacillus megaterium strains for the production of a heterologous dextransucrase.

High cell density cultivations were performed under identical conditions for two Bacillus megaterium strains (MS941 and WH320), both carrying a heterologous dextransucrase (dsrS) gene under the control of the xylA promoter. At characteristic points of the cultivations (end of batch, initial feeding, before and after induction) the proteome was analyzed based on two dimensional gel electrophoresis and mass spectrometric protein identification using the protein database "bmegMEC.v2" recently made available. High expression but no secretion of DsrS was found for the chemical mutant WH320 whereas for MS 941, a defined protease deficient mutant of the same parent strain (DSM319), not even expression of DsrS could be detected. The proteomic analysis resulted in the identification of proteins involved in different cellular pathways such as in central carbon and overflow metabolism, in protein synthesis, protein secretion and degradation, in cell wall metabolism, in cell division and sporulation, in membrane transport and in stress responses. The two strains exhibited considerable variations in expression levels of specific proteins during the different phases of the cultivation process, whereas induction of DsrS production had, in general, little effect. The largely differing behaviour of the two strains with regard to DsrS expression can be attributed, at least in part, to changes observed in the proteome which predominantly concern biosynthetic enzymes and proteins belonging to the membrane translocation system, which were strongly down-regulated at high cell densities in MS941 compared with WH320. At the same time a cell envelope-associated quality control protease and two peptidoglycan-binding proteins related to cell wall turnover were strongly expressed in MS941 but not found in WH320. However, to further explain the very different physiological responses of the two strains to the same cultivation conditions, it is necessary to identify the mutated genes in WH320 in addition to the known lacZ. In view of the results of this proteomic study it seems that at high cell density conditions and hence low growth rates MS941, in contrast to WH320, does not maintain a vegetative growth which is essential for the expression of the foreign dsrS gene by using the xylA promoter. It is conceivable that applications of a promoter which is highly active under nutrient-limited cultivation conditions is necessary, at least for MS941, for the overexpression of recombinant genes in such B. megaterium fed-batch cultivation process. However to obtain a heterologous protein in secreted and properly folded form stills remains a big challenge.

Journal Article↗

Shotgun proteomics: tools for the analysis of complex biological systems.

Recent interest in proteomics has been fueled by the completion of multiple genome projects and ignited by the common need of biologists to rapidly and comprehensively evaluate complex samples of proteins on a global level. 'Shotgun proteomics' refers to the direct analysis of complex protein mixtures to rapidly generate a global profile of the protein complement within the mixture. This approach has been facilitated by the use of multidimensional protein identification technology (MudPIT), which incorporates multidimensional high-pressure liquid chromatography (LC/LC), tandem mass spectrometry (MS/MS) and database-searching algorithms. This review will focus on the most recent advances in methodologies for shotgun proteomics and address the limitations of the application of each to real biological samples.

Algorithms↗

'Harvester': a fast meta search engine of human protein resources.

SUMMARY: We have developed a Web-based tool named 'Harvester' that bulk-collects bioinformatic data on human proteins from various databases and prediction servers. The information on every single protein is assembled on a single HTML page as a combination of database screen-shots and plain text. A full text meta search engine, similar to Google trade mark, allows screening of the whole genome proteome for current protein functions and predictions in a few seconds. With Harvester it is now possible to compare and check the quality of different database entries and prediction algorithms on a single page. A feedback forum allows users to comment on Harvester and to report database inconsistencies. AVAILABILITY: The service is freely available to the academic community at http://harvester.embl.de.

Database Management Systems↗

Processing complex mixtures of intact proteins for direct analysis by mass spectrometry.

For analysis of intact proteins by mass spectrometry (MS), a new twist to a two-dimensional approach to proteome fractionation employs an acid-labile detergent instead of sodium dodecyl sulfate during continuous-elution gel electrophoresis. Use of this acid-labile surfactant (ALS) facilitates subsequent reversed-phase liquid chromatography (RPLC) for a net two-dimensional fractionation illustrated by transforming thousands of intact proteins from Saccharomyces cerevisiae to mixtures of 5-20 components (all within approximately 5 kDa of one another) for presentation via electrospray ionization (ESI) to a Fourier transform MS (FTMS). Between 3 and 13 proteins have been detected directly using ESI-FTMS (or MALDI-TOF), and the fractionation showed a peak capacity of approximately 400 between 0 and 70 kDa. A probability-based identification was made automatically from raw MS/MS data (obtained using a quadrupole-FTMS hybrid instrument) for one protein that differed from that predicted in a yeast database of approximately 19,000 protein forms. This ALS-PAGE/RPLC approach to proteome processing ameliorates the "front end" problem that accompanies direct analysis of whole proteins and assists the future realization of protein identification with 100% sequence coverage in a high-throughput format.

Algorithms↗