Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “PROTEOMICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

A streamlined approach to high-throughput proteomics.

Proteomics has rapidly become an important tool for life science research, allowing the integrated analysis of global protein expression from a single experiment. To accommodate the complexity and dynamic nature of any proteome, researchers must use a combination of disparate protein biochemistry techniques, often a highly involved and time-consuming process. Whilst highly sophisticated, individual technologies for each step in studying a proteome are available, true high-throughput proteomics that provides a high degree of reproducibility and sensitivity has been difficult to achieve. The development of high-throughput proteomic platforms, encompassing all aspects of proteome analysis and integrated with genomics and bioinformatics technology, therefore represents a crucial step for the advancement of proteomics research. ProteomIQ (Proteome Systems) is the first fully integrated, start-to-finish proteomics platform to enter the market. Sample preparation and tracking, centralized data acquisition and instrument control, and direct interfacing with genomics and bioinformatics databases are combined into a single suite of integrated hardware and software tools, facilitating high reproducibility and rapid turnaround times. This review will highlight some features of ProteomIQ, with particular emphasis on the analysis of proteins separated by 2D polyacrylamide gel electrophoresis.

Automation↗

Nuclear Proteome Map of Mouse Heart Chambers.

Heart specialization involves nuclear programs; however, chamber-specific regulation of the nuclear proteome landscape remains unknown. In this study, we isolated the nucleus from four major anatomical regions of healthy mouse heart (fresh) and employed quantitative mass spectrometry-based proteomics to construct a comprehensive nuclear proteome landscape of left ventricle (LV, 2403 proteins), right ventricle (RV, 2242 proteins), left atrium (LA, 2368 proteins), and right atrium (RA, 1816 proteins). This led to the discovery of nuclear regional proteome signatures (ventricular signature, 297 proteins; atrial signature, 183 proteins) associated with oxidative metabolism and redox regulation, ferroptosis, extracellular-matrix remodeling, SUMO- and stress-responsive control and transcriptional regulation. Chamber-level analyses further identify distinct nuclear features in LV (120 proteins), LA (188 proteins), and RA (72 proteins). In addition, we defined conserved core nuclear proteome (230 proteins) shared across all anatomical regions, enriched for transcription-regulator complexes, nucleolar/ribosome-associated, RNA-processing, and chromatin-organization components. Within this core network, we report 78 transcription factors/co-factors and select nuclear, chromatin and RNA export-associated proteins, including 29 specific factors (e.g., Alpk3, Rbm14, Arglu1, Hmgb1, Myef2, Sf1) associated with the heart. Regionally, we verified spatial localization in heart of H2ac21 and Sun2 in LA and Ptbp2 in LV by immunofluorescence. This study provides insights into the chamber-resolved view of the nuclear proteome in the heart, establishes a framework for linking nuclear proteomic signatures to atrial and ventricular biology, unique features of the heart nuclear proteome landscape relative to other organs, and a baseline for studying nuclear remodeling in cardiac pathophysiology.

Animals↗

A comparative proteome analysis of hippocampal tissue from schizophrenic and Alzheimer's disease individuals.

The proteins expressed by a genome have been termed the proteome. Comparative proteome analysis of brain tissue offers a novel means to identify biologically significant gene products that underlie psychopathology. In this study we collected post mortem hippocampal tissue from the brains of seven schizophrenic, seven Alzheimer's disease (AD) and seven control individuals. Hippocampal proteomes were visualised by two-dimensional gel electrophoresis of homogenised tissue. A mean of 549 (s.d. 35) proteins were successfully matched between each disease group and the control group. In comparison with the control hippocampal proteome, eight proteins in the schizophrenic hippocampal proteome were found to be decreased and eight increased in concentration, whereas, in the AD hippocampal proteome, 35 proteins were decreased and 73 were increased in concentration (P<0.05). One protein, which was decreased in concentration in both diseases, was characterised as diazepam binding inhibitor (DBI) by N-terminal sequence analysis. DBI can regulate the action of the GABA(A) receptor. Protein changes involved 6% of the assessed AD hippocampal proteome, whereas, in schizophrenia protein changes involved less than 1% of the assessed hippocampal proteome. We conclude that schizophrenia has a subtle neuropathological presentation and comparative proteome analysis is a viable means by which to investigate diseases of the brain at the molecular level.

Adult↗

Target identification and validation in drug discovery: the role of proteomics.

Proteomics, the study of cellular protein expression, is an evolving technology platform that has the potential to identify novel proteins involved in key biological processes in the cell that may serve as potential drug targets. While proteomics has considerable theoretical promise, individual cells/tissues have the potential to generate many millions of proteins while the current analytical technologies that involve the use of time-consuming two dimensional gel electrophoresis (2DIGE) and various mass spectrometry (MS) techniques are unable to handle complex biological samples without multiple high-resolution purification steps to reduce their complexity. This can significantly limit the speed of data generation and replication and requires the use of bioinformatic algorithms to reconstitute the parent proteome, a process that does not always result in a reproducible outcome. In addition, membrane bound proteins, e.g., receptors and ion channels, that are the targets of many existing drugs, are not amenable to study due, in part, to limitations in current proteomic techniques and also to these being present in low abundance and thus disproportionally represented in proteome profiles. Subproteomes with reduced complexity have been used to generate data related to specific, hypothesis-driven questions regarding target identification, protein-interaction networks and signaling pathways. However progress to date, with the exception of diagnostic proteomics in the field of cancer, has been exceedingly slow with an inability to put such studies in the context of a larger proteome, limiting the value of the information. Additionally the pathway for target validation (which can be more accurately described at the preclinical level as target confidence building) remains unclear. It is important that the ability to measure and interrogate proteomes matches expectations, avoiding a repetition of the disappointment and subsequent skepticism that accompanied what proved to be unrealistic expectations for the rapid contribution of data based on the genome maps, to biomedical research.

Computational Biology↗

Transpulmonary proteomic gradient analysis in women with pulmonary arterial hypertension associated with systemic sclerosis.

This study investigated proteomic alterations in the pulmonary circulation of patients with pulmonary arterial hypertension associated with systemic sclerosis (PAH-SSc) by analyzing the transpulmonary protein gradient and comparing the proteomic profiles with systemic sclerosis (SSc) without PAH. Twenty women were included (10 PAH-SSc, 64.6&#xa0;&#xb1;&#xa0;10.8&#xa0;years; 10 SSc, 62.8&#xa0;&#xb1;&#xa0;11.5&#xa0;years). The transpulmonary gradient was defined as the difference in biomarker concentrations between wedge-position and pulmonary artery blood samples. Peptides were analysed using liquid chromatography-mass spectrometry, and differentially abundant proteins were identified with Proteome Discoverer. Protein-protein interaction networks were generated with STRING and visualized in Cytoscape. A total of 270 proteins were detected, with no significant transpulmonary gradient alterations. However, patients with PAH-SSc showed distinct proteomic profiles compared to SSc. Multivariate analysis identified 48 differentially abundant proteins in pulmonary artery plasma, with 15 overrepresented and 33 downregulated in PAH-SSc. Among these, the downregulation of transforming growth factor-beta-induced protein ig-h3 (TGF&#x3b2;I/ig-h3) points to a potential involvement of the TGF-&#x3b2;-related extracellular matrix remodelling pathway in PAH-SSc. However, further validation in larger and independent cohorts is required before its relevance as a biomarker or therapeutic target can be established. In conclusion, while no transpulmonary proteomic gradient was observed, the proteomic profiles of PAH-SSc and SSc were different. The profile in PAH-SSc was characterized by differences in immune response, lipid metabolism, and hemostatic proteins. SIGNIFICANCE: This study offers the first proteomic characterization of the transpulmonary gradient in PAH-SSc and SSc. Although no differences in the gradient were found, the pulmonary artery plasma proteome of PAH-SSc patients showed a distinct pattern compared to SSc. Several proteins associated with immune function, haemostasis, and cellular processes were altered, which may indicate specific pathophysiological features of PAH-SSc or suggest how lung dysfunction develops in SSc. Targeting dysregulated proteins like TGF&#x3b2;I/ig-h3 or addressing immune-coagulation imbalances may support future research studies. Overall, these findings refine the molecular profile of PAH-SSc and provide a basis for future large-scale studies aimed at clarifying disease mechanisms and identifying clinically relevant molecular signatures.

Humans↗

Application of proteomics to the study of molecular mechanisms in neurotoxicology.

The proteome is the protein compliment of the genome and is the result of genetic expression, ribosomal synthesis and proteolytic degradation. Proteins participate in most major cell processes and their function is highly regulated by post-translational modifications such as phosphorylation and glycosylation. As a result, neurotoxicant-induced changes in protein levels, function or regulation could have a negative impact on neuronal viability. At the molecular level, direct oxidative or covalent modifications of individual proteins by various chemicals or drugs is likely to lead to perturbation of tertiary structure and a loss of function. The proteome and the functional determinants of its individual protein components are, therefore, likely targets of neurotoxicant action and resulting characteristic disruptions could be critically involved in corresponding mechanisms of neurotoxicity. Clearly, investigating changes in the proteome can provide important clues for deciphering mechanisms of toxicant action and, therefore, proteomics, the study of the proteome, is currently, and will likely remain, a significant experimental approach for mechanistic research in neurotoxicology. The purpose of this review is to discuss proteomics as a tool for neurotoxicological investigations. A variety of classic proteomic techniques (e.g. liquid chromatography (LC)/tandem mass spectroscopy, two-dimensional gel image analysis) as well as more recently developed approaches (e.g. two-hybrid systems, antibody arrays, protein chips, isotope-coded affinity tags, ICAT) are available to determine protein levels, identify components of multiprotein complexes and to detect post-translational changes. Proteomics, therefore, offers a comprehensive overview of cell proteins, and in the case of neurotoxicant exposure, can provide quantitative data regarding changes in corresponding expression levels and/or post-translational modifications that might be associated with neuron injury.

Animals↗

Comparison of indirect and direct approaches using ion-trap and Fourier transform ion cyclotron resonance mass spectrometry for exploring viperid venom proteomes.

In a sense, the field of snake venom proteomics has been under investigation since the very earliest biochemical studies where it was soon recognized that venoms are comprised of complex mixtures of bioactive molecules, most of which are proteins. Only with the re-emergence of 2D polyacrylamide gel electrophoresis (2D PAGE) and the recent developments in mass spectrometry for the identification/characterization of proteins coupled with venom gland transcriptomes has the field of snake venom proteomics began to flourish and provide exciting insights into the protein composition of venoms and subsequently their pathological activities. In this manuscript we will briefly discuss the state of snake venom proteomics followed by the presentation of several straightforward experiments designed to explore approaches to investigating venom proteomics. The first set of experiments used 1D gel electrophoresis (1D PAGE) of Crotalus atrox venom followed by slice-by-slice analysis of the proteins using liquid chromatography/mass spectrometry/mass spectrometry (LC/MS/MS). In the second set of experiments, C. atrox and Bothrops jararaca venoms were subjected to in-solution digestion followed by Fourier transform ion cyclotron resonance (FTICR) LC/MS/MS. The peptide ion-maps of these venoms were compared along with the proteins identified. In addition, the results were compared to the results observed from the 1D PAGE approach. From these studies it is clear that sample de-complexation/fractionation before mass spectrometry is still the best approach for maximum proteome coverage. Furthermore, comparison of venom proteomes based on tryptic peptide identities between the proteomes is not particularly effective since there does not appear to be a sufficient number of such identical peptides, derived from related proteins, present in venoms. Finally, as has previously been recognized without either better databases of venom protein sequences or facile and rapid de novo sequencing technologies for mass spectrometry, snake venom proteome investigation will remain a laborious task.

Animals↗

Synergistic computational and experimental proteomics approaches for more accurate detection of active serine hydrolases in yeast.

An analysis of the structurally and catalytically diverse serine hydrolase protein family in the Saccharomyces cerevisiae proteome was undertaken using two independent but complementary, large-scale approaches. The first approach is based on computational analysis of serine hydrolase active site structures; the second utilizes the chemical reactivity of the serine hydrolase active site in complex mixtures. These proteomics approaches share the ability to fractionate the complex proteome into functional subsets. Each method identified a significant number of sequences, but 15 proteins were identified by both methods. Eight of these were unannotated in the Saccharomyces Genome Database at the time of this study and are thus novel serine hydrolase identifications. Three of the previously uncharacterized proteins are members of a eukaryotic serine hydrolase family, designated as Fsh (family of serine hydrolase), identified here for the first time. OVCA2, a potential human tumor suppressor, and DYR-SCHPO, a dihydrofolate reductase from Schizosaccharomyces pombe, are members of this family. Comparing the combined results to results of other proteomic methods showed that only four of the 15 proteins were identified in a recent large-scale, "shotgun" proteomic analysis and eight were identified using a related, but similar, approach (neither identifies function). Only 10 of the 15 were annotated using alternate motif-based computational tools. The results demonstrate the precision derived from combining complementary, function-based approaches to extract biological information from complex proteomes. The chemical proteomics technology indicates that a functional protein is being expressed in the cell, while the computational proteomics technology adds details about the specific type of function and residue that is likely being labeled. The combination of synergistic methods facilitates analysis, enriches true positive results, and increases confidence in novel identifications. This work also highlights the risks inherent in annotation transfer and the use of scoring functions for determination of correct annotations.

Amino Acid Sequence↗

HUPO initiatives relevant to clinical proteomics.

The past few years have seen a tremendous interest in the potential of proteomics to address unmet needs in biomedicine. Such unmet needs include more effective strategies for early disease detection and monitoring and more effective therapies, in addition to developing a better understanding of disease pathogenesis. Proteomics is particularly suited for investigating biological fluids to identify disease-related alterations and to develop molecular signatures for disease processes. However, much of the effort undertaken in clinical proteomics to date represents either demonstrations of principles or relatively small-scale studies when compared with genomics effort and accomplishments or more pertinently when contrasted with the tremendous untapped potential of clinical proteomics. Clearly, we are in the early stages. What seems to be urgently needed is an organized effort to build a solid foundation for proteomics that includes developing a much needed infrastructure with adequate resources. The Human Proteome Organization (HUPO) is fostering an organized international effort in proteomics that includes initiatives around organ systems and biological fluids that have disease relevance as well as development of proteomics resources.

Humans↗

Shotgun proteomics using the iTRAQ isobaric tags.

Shotgun proteomic methods involving isobaric tagging of peptides enable high-throughput proteomic analysis. iTRAQ reagents allow simultaneous identification and quantitation of proteins in four different samples using tandem mass spectrometry (MS). In this article, we provide a brief description of proteome analysis using iTRAQ reagents and review the current applications of these reagents in proteomic studies. We also compare different aspects of protein identification including protein sequence coverage and proteome coverage obtained using iTRAQ reagents with those using other shotgun proteomic techniques. We briefly discuss the issue of isotope purity correction in measured peak areas during protein quantitation using iTRAQ reagents. Finally, we conclude with some of the current challenges in MS-based proteomic analysis that are limiting protein identifications obtained by different shotgun proteomic methods.

Animals↗

ECLIPSE: exploring the dark proteome of ESKAPE pathogens through the sequence similarity network of the Protein Universe Atlas.

MOTIVATION: The accelerating crisis of antimicrobial resistance among the critical so-called ESKAPE pathogens demands the urgent identification of novel molecular targets. However, a substantial fraction of ESKAPE proteomes remains functionally uncharacterized, with many genes annotated as encoding hypothetical proteins. These protein sequences often lack significant similarity to known protein families when conventional homology-based annotation methods are used and thus remain "dark". This limits our ability to explore their roles in pathogenicity, and it is thus crucial to bridge this substantial gap in pathogen biology by developing new strategies to illuminate these "dark" regions of the ESKAPE pan-proteome. RESULTS: We introduce ECLIPSE (ESKAPE Connectome Linkage and Inference for Proteome Sequence Exploration), a network-based computational framework that systematically identifies and prioritizes functionally dark protein families in ESKAPE pan-proteomes. ECLIPSE embeds target ESKAPE pathogen proteomes within the global sequence similarity network of the Protein Universe Atlas. It detects connected components composed entirely of unannotated proteins, called the "dark proteome." As a case study, we applied ECLIPSE to a pan-proteome of 3&#x2006;460&#x2006;657 protein sequences from 635 strains of Pseudomonas aeruginosa (PA). ECLIPSE identified 120&#x2006;985 proteins (4%) residing in completely dark connected components. Furthermore, we have performed a taxonomic diversity analysis using normalized Shannon indices to characterize each dark component by its enrichment in ESKAPE pathogens. The analysis utilized the evenness (E) value (see Methods 2.1), which distinguishes Pseudomonas-specific (target-specific) from ESKAPE-enriched dark components. We then developed the Dark Proteome Prioritization Score (DPPS), a composite multidimensional scoring framework (see Methods 2.5). It ranks these dark components by biological relevance across four orthogonal axes: (i) functional darkness, (ii) P. aeruginosa proportion in the Atlas, (iii) AMR-clade taxonomic restriction, and (iv) conservation across the 635 P. aeruginosa strains. This framework outputs a robust four-tier scoring system; the prioritized Tier I components were validated by weight sensitivity analysis and remained stable across 500 Monte Carlo weight perturbations. Structural characterization of one of the top-ranked ESKAPE-enriched dark components revealed that it belongs to the beta-barrel fold DUF1302 (PF06980) family, for which no experimentally solved three-dimensional structure exists in the PDB. The genomic context analysis indicates that it is co-localized with a LuxR-type transcriptional regulator. Collectively, ECLIPSE identifies evolutionarily conserved, structurally defined, and functionally dark proteins enriched across ESKAPE pathogens; these dark proteins can further be utilized as alternative antimicrobial targets for experimental characterization. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available for free at: Github: https://github.com/surabhilata/ECLIPSE.git, Zenodo: DOI: 10.5281/zenodo.21064323.

Proteome↗

The Escherichia coli proteome: past, present, and future prospects.

Proteomics has emerged as an indispensable methodology for large-scale protein analysis in functional genomics. The Escherichia coli proteome has been extensively studied and is well defined in terms of biochemical, biological, and biotechnological data. Even before the entire E. coli proteome was fully elucidated, the largest available data set had been integrated to decipher regulatory circuits and metabolic pathways, providing valuable insights into global cellular physiology and the development of metabolic and cellular engineering strategies. With the recent advent of advanced proteomic technologies, the E. coli proteome has been used for the validation of new technologies and methodologies such as sample prefractionation, protein enrichment, two-dimensional gel electrophoresis, protein detection, mass spectrometry (MS), combinatorial assays with n-dimensional chromatographies and MS, and image analysis software. These important technologies will not only provide a great amount of additional information on the E. coli proteome but also synergistically contribute to other proteomic studies. Here, we review the past development and current status of E. coli proteome research in terms of its biological, biotechnological, and methodological significance and suggest future prospects.

Bacterial Proteins↗

Proteomic Signatures Related to Physical Activity Are Associated with Risks of Future Disease.

PURPOSE: Physical activity (PA) can lower the risk of developing chronic diseases. However, few studies have examined the proteomic signatures linked to PA, and the role of these signatures in the connection between PA levels and future disease risk remains unclear. This study aimed to investigate whether proteomic signatures indicative of PA are associated with the risk of developing common chronic diseases and to explore their role as statistical links in the relationship between PA levels and disease development. METHODS: We used data from a subcohort of UK Biobank participants. PA intensity data were collected from accelerometers worn by each participant. Plasma proteomics results were obtained through Olink analysis. The risks of developing each primary chronic disease were evaluated for types of PA and their associated proteomic signatures, adjusting for age, sex, ethnicity, socioeconomic status, lifestyle factors, and key measurement time-lag covariates. RESULTS: Based on the UK Biobank, we identified significant differences among the proteomic signatures of accelerometer-measured light PA, moderate-to-vigorous PA, and total PA. The main enriched pathways of these proteomic signatures included cell adhesion, cell migration, and immune response. Higher levels of accelerometer-measured PA and their associated proteomic signatures correlated with a lower risk of developing cardiometabolic disorders, cancers, psychological or neurological disorders, and respiratory diseases. CONCLUSIONS: Our findings show that PA and PA-related proteomic signatures are statistically associated with lower risks of chronic diseases. Further analyses identified proteins that were correlated with both PA and disease risk. These results need to be confirmed through longitudinal studies involving diverse populations.

Humans↗

[Introduction of proteomic approach to environmental medicine].

Recent progress in life science technology and the availability of much information on genes obtained by genome analysis has enabled us to analyze the changes of proteins on a large scale. Sets of proteins are called proteomes, and proteomics is the scientific field of proteome analysis including differential, post translational modification and interaction analyses. Various proteomic techniques, particularly two-dimensional gel electrophoresis (2-DE), mass spectrometry, protein chip methods, and surface plasmon resonance (SPR), are very useful for acquiring proteomes in cells, tissues and body fluid, and for analyzing interactions between a protein and other biofactors including proteins. A proteomic approach is also useful for determining biomarkers of diseases and key proteins involved in various stages of metabolism such as differentiation, cell cycle and apoptosis. Environmental pollutants including endocrine disruptors inhibit activities of various organs in wild animals and humans. Proteomic approaches could be very useful tools for elucidating the mechanisms of damage caused by environmental pollutants. In this review, we describe the application of a proteomic approach to the field of environmental medicine.

Animals↗

Proteomic studies in plants.

Proteomics is a leading technology for the high-throughput analysis of proteins on a genome-wide scale. With the completion of genome sequencing projects and the development of analytical methods for protein characterization, proteomics has become a major field of functional genomics. The initial objective of proteomics was the large-scale identification of all protein species in a cell or tissue. The applications are currently being extended to analyze various functional aspects of proteins such as post-translational modifications, protein-protein interactions, activities and structures. Whereas the proteomics research is quite advanced in animals and yeast as well as Escherichia coli, plant proteomics is only at the initial phase. Major studies of plant proteomics have been reported on subcellular proteomes and protein complexes (e.g. proteins in the plasma membranes, chloroplasts, mitochondria and nuclei). Here several plant proteomics studies will be presented, followed by a recent work using multidimensional protein identification technology (MudPIT).

Cell Membrane↗

Complexity analysis of yeast proteome network.

Topological and compositional complexity of protein-protein networks is assessed in a variety of ways making use of graph theory and information theory. The methodology used is borrowed from mathematical chemistry and includes complexity descriptors such as substructure count, overall connectivity, walk count, and information on various vertex distributions. The approach is applied to the (incomplete) proteome of Saccharomyces cerevisiae containing 232 protein complexes of a total of 1,440 proteins. The proteome network and each of its nine functional subsets of protein complexes are disconnected graphs, containing a number of noninteracting species and a major component. A weighted edge between two vertices in these graphs stands for the number of shared proteins between the respective complexes. The major component is a highly connected, 'small-world' network, in which the average vertex distance between protein complexes does not exceed 2.2 (2.4 for the entire proteome), whereas the maximum distance does not exceed 4 (or 5 for the proteome). The vertex degree distribution in the major proteome component with 199 complexes follows the power law P(k) approximately k(-gamma), with gamma approximately = 1.7. The analysis of the functional organization of the yeast proteome has shown that, for any pair of biological functions, there always exist many proteins that can perform both functions. The potential application of the quantitative proteome descriptors discussed includes quantitative relationships between the structure and biological action of dynamic protein complexes in changing environment, identification of targets for markers/drugs, as well as system analysis and comparative studies of proteomes.

Fungal Proteins↗

Proteomics in biomarker discovery and drug development.

Proteomics is a research field aiming to characterize molecular and cellular dynamics in protein expression and function on a global level. The introduction of proteomics has been greatly broadening our view and accelerating our path in various medical researches. The most significant advantage of proteomics is its ability to examine a whole proteome or sub-proteome in a single experiment so that the protein alterations corresponding to a pathological or biochemical condition at a given time can be considered in an integrated way. Proteomic technology has been extensively used to tackle a wide variety of medical subjects including biomarker discovery and drug development. By complement with other new technique advances in genomics and bioinformatics, proteomics has a great potential to make considerable contribution to biomarker identification and to revolutionize drug development process. This article provides a brief overview of the proteomic technologies and their application in biomarker discovery and drug development.

Animals↗

Data mining techniques for cancer detection using serum proteomic profiling.

OBJECTIVE: Pathological changes in an organ or tissue may be reflected in proteomic patterns in serum. It is possible that unique serum proteomic patterns could be used to discriminate cancer samples from non-cancer ones. Due to the complexity of proteomic profiling, a higher order analysis such as data mining is needed to uncover the differences in complex proteomic patterns. The objectives of this paper are (1) to briefly review the application of data mining techniques in proteomics for cancer detection/diagnosis; (2) to explore a novel analytic method with different feature selection methods; (3) to compare the results obtained on different datasets and that reported by Petricoin et al. in terms of detection performance and selected proteomic patterns. METHODS AND MATERIAL: Three serum SELDI MS data sets were used in this research to identify serum proteomic patterns that distinguish the serum of ovarian cancer cases from non-cancer controls. A support vector machine-based method is applied in this study, in which statistical testing and genetic algorithm-based methods are used for feature selection respectively. Leave-one-out cross validation with receiver operating characteristic (ROC) curve is used for evaluation and comparison of cancer detection performance. RESULTS AND CONCLUSIONS: The results showed that (1) data mining techniques can be successfully applied to ovarian cancer detection with a reasonably high performance; (2) the classification using features selected by the genetic algorithm consistently outperformed those selected by statistical testing in terms of accuracy and robustness; (3) the discriminatory features (proteomic patterns) can be very different from one selection method to another. In other words, the pattern selection and its classification efficiency are highly classifier dependent. Therefore, when using data mining techniques, the discrimination of cancer from normal does not depend solely upon the identity and origination of cancer-related proteins.

Biomarkers, Tumor↗