Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus↗

Analysis of peptide MS/MS spectra from large-scale proteomics experiments using spectrum libraries.

A widespread proteomics procedure for characterizing a complex mixture of proteins combines tandem mass spectrometry and database search software to yield mass spectra with identified peptide sequences. The same peptides are often detected in multiple experiments, and once they have been identified, the respective spectra can be used for future identifications. We present a method for collecting previously identified tandem mass spectra into a reference library that is used to identify new spectra. Query spectra are compared to references in the library to find the ones that are most similar. A dot product metric is used to measure the degree of similarity. With our largest library, the search of a query set finds 91% of the spectrum identifications and 93.7% of the protein identifications that could be made with a SEQUEST database search. A second experiment demonstrates that queries acquired on an LCQ ion trap mass spectrometer can be identified with a library of references acquired on an LTQ ion trap mass spectrometer. The dot product similarity score provides good separation of correct and incorrect identifications.

Amino Acid Sequence↗

Analysis of the Saccharomyces cerevisiae proteome with PeptideAtlas.

We present the Saccharomyces cerevisiae PeptideAtlas composed from 47 diverse experiments and 4.9 million tandem mass spectra. The observed peptides align to 61% of Saccharomyces Genome Database (SGD) open reading frames (ORFs), 49% of the uncharacterized SGD ORFs, 54% of S. cerevisiae ORFs with a Gene Ontology annotation of 'molecular function unknown', and 76% of ORFs with Gene names. We highlight the use of this resource for data mining, construction of high quality lists for targeted proteomics, validation of proteins, and software development.

Codon↗

[A evolutionary approach to identification of orthologous relationship across proteomes].

How to identify the true orthologous and paralogous relationships among protein families is still a key problem in genome annotation and comparative protemics. Here, a evolutionary approach to ascertainment of the orthologous relationships across the genomes is developed. Forty-four cases of protein families are used in the test for the evolutionary approach. Compared with the method of COG (cluster of orthologous groups of proteins), this approach can generally identify the orthologous relationships and accurately predict the genome function.

Animals↗

Thermo-search: lifestyle and thermostability analysis.

Thermo-search is an online web tool for the analysis of proteomes and individual proteins according to the ratio of two couplets of preferred and avoided amino acids in hyperthermophiles, thermophiles and mesophiles. It displays the ratio between glutamic acid plus lysine (E+K) and glutamine plus histidine (Q+H), which is higher in thermophilic proteomes and thermostable proteins than in mesophilic proteomes and thermo labile proteins. Thermo-search allows a rapid screen of the CRM database for thermostable proteins in their functional categories and a visualization of the (E+K)/(Q+H) average ratio between organisms, allowing a comparison of their lifestyles.

Amino Acids↗

The clinical promise of mass spectrometry-based single-cell proteomics: from bedside to bench.

INTRODUCTION: Single-cell proteomics (SCP) is entering into a transformative phase, moving beyond technically demanding benchmarking studies toward robust and reproducible workflows capable of quantifying thousands of proteins per cell. These advances highlight SCP's potential to address clinically relevant questions by resolving cellular and pathological heterogeneity that remains obscured in bulk proteomics. AREAS COVERED: This review discusses current advances, challenges, and clinical applications of SCP based on literature identified through searches in major scientific databases. Many clinically relevant samples remain underexplored in SCP studies, in part because their application requires careful evaluation of pre-analytical variables that can strongly influence proteomic readouts. Current SCP methodologies vary according to sample type, experimental conditions, and available resources. Compared with single-cell RNA sequencing, SCP remains limited in cellular throughput, making it challenging to define optimal sample sizes and to reliably detect both abundant and rare cell populations. These limitations also make dataset integration difficult, as reduced cellular coverage and sampling depth increase data sparsity. Moreover, implementing quality control strategies across sequential SCP experiments is essential to ensure data robustness, comparability, and accurate biological interpretation. EXPERT OPINION: Applying SCP to clinical samples advances our understanding of biological complexity and holds potential to drive progress in translational and precision medicine.

Humans↗

The role of mass spectrometry in proteome studies.

Mass spectrometry (MS) is an important tool in modern protein chemistry. In proteome analyses the expression of hundreds or thousands of proteins can be monitored at the same time. First, complex protein mixtures are separated by two-dimensional gel electrophoresis (2-DE) and then individual proteins are identified by using MS followed by database searches. Recent developments in this field have made it possible to do automated, high-throughput protein identification that is needed in proteome analyses. MS can also be used to characterize post-translational modifications in proteins and to study protein complexes. This review will introduce the current MS methods used in proteome studies, and discuss their advantages and disadvantages. New instrumental MS developments are also presented that are useful in these analyses.

Databases, Protein↗

Expanding the subproteome of the inner mitochondria using protein separation technologies: one- and two-dimensional liquid chromatography and two-dimensional gel electrophoresis.

Currently no single proteomics technology has sufficient analytical power to allow for the detection of an entire proteome of an organelle, cell, or tissue. One approach that can be used to expand proteome coverage is the use of multiple separation technologies especially if there is minimal overlap in the proteins observed by the different methods. Using the inner mitochondrial membrane subproteome as a model proteome, we compared for the first time the ability of three protein separation methods (two-dimensional liquid chromatography using the ProteomeLab PF 2D Protein Fractionation System from Beckman Coulter, one-dimensional reversed phase high performance liquid chromatography, and two-dimensional gel electrophoresis) to determine the relative overlap in protein separation for these technologies. Data from these different methods indicated that a strikingly low number of proteins overlapped with less than 24% of proteins common between any two technologies and only 7% common among all three methods. Utilizing the three technologies allowed the creation of a composite database totaling 348 non-redundant proteins. 82% of these proteins had not been observed previously in proteomics studies of this subproteome, whereas 44% had not been identified in proteomics studies of intact mitochondria. Each protein separation method was found to successfully resolve a unique subset of proteins with the liquid chromatography methods being more suited for the analysis of transmembrane domain proteins and novel protein discovery. We also demonstrated that both the one- and two-dimensional LC allowed for the separation of the alpha-subunit of F1F0 ATP synthase that differed due to a change in pI or hydrophobicity.

Animals↗

The proteogenomic landscape of the human kidney and implications for cardio-kidney-metabolic health.

Nearly one-third of the global population is affected by cardio-kidney-metabolic (CKM) diseases; however, the molecular mechanisms underlying CKM diseases are poorly understood. Here we show that tissue proteomics provide critical insights not captured by tissue gene expression or blood proteomics information by performing whole-genome and RNA sequencing and proteomics analysis of human kidney samples (n = 337), and we generated a publicly available database. Via Bayesian co-localization and Mendelian randomization analyses of kidney protein quantitative trait loci and 36 CKM genome-wide association studies, we prioritized 89 proteins for CKM traits. We prioritized relationships that could underlie the interconnectedness of CKM traits and discovered multiple and targetable mechanisms for CKM diseases, including the potential role of kidney angiopoietin-like protein 3 (ANGPTL3) in serum lipid levels and kidney function as well as the role of charged multivesicular body protein 1A in kidney function and hypertension. Notably, we identify pathways with confluence of evidence from genetic loci, tissue gene expression and protein levels for CKM traits. In summary, our large-scale kidney proteomics study uncovers proteins and targetable mechanisms prioritized for CKM diseases.

Humans↗

Large-scale identification of Caenorhabditis elegans proteins by multidimensional liquid chromatography-tandem mass spectrometry.

A proteome of a model organism, Caenorhabditis elegans, was analyzed by an integrated liquid chromatography (LC)-based protein identification system, which was constructed by microscale two-dimensional liquid chromatography (2DLC) coupled with electrospray ionization (ESI) tandem mass spectrometry (MS/MS) on a high-resolution hybrid mass spectrometer with an automated data analysis system. Soluble and insoluble protein fractions were prepared from a mixed growth phase culture of the worm C. elegans, digested with trypsin, and fractionated separately on the 2DLC system. The separated peptides were directly analyzed by on-line ESI-MS/MS in a data-dependent mode, and the resultant spectral data were automatically processed to search a genome sequence database, wormpep 66, for protein identification. The total number of proteins of the composite proteome identified in this method was 1,616, including 110 secreted/targeted proteins and 242 transmembrane proteins. The codon adaptation indices of the identified proteins suggested that the system could identify proteins of relatively low abundance, which are difficult to identify by conventional 2D-gel electrophoresis (GE) followed by an offline mass spectrometric analysis such as peptide mass fingerprinting. Among the approximately 5,400 peptides assigned in this study, many peptides with post-translational modifications, such as N-terminal acetylation and phosphorylation, were detected. This expression profile of C. elegans, containing 571 hypothetical gene products, will serve as the basic data of a major proteome set expressed in the worm.

Animals↗

An integrated map of the murine hippocampal proteome based upon five mouse strains.

With the advent of proteomics technologies it is possible to simultaneously demonstrate the expression of hundreds of proteins. The information offered by proteomics provides context-based understanding of cellular protein networks and has been proven to be a valuable approach in neuroscience studies. The mouse hippocampus has been a major target of analysis in the search for molecular correlates to neuronal information storage. Although human and rat hippocampal samples have been successfully subjected to proteomic profiling, no elaborate analysis providing the fundamental experimental basis for protein-expression studies in the mouse hippocampus has been carried out as yet. This led us to construct a master map generated from the individual hippocampal proteomes of five different mouse strains. A proteomic approach, based upon 2-DE coupled to MS (MALDI-TOF/TOF) has been chosen in an attempt to establish a comprehensive reference database of proteins expressed in the mouse hippocampus. 469 individual proteins, represented by 1156 spots displaying various functional states of the respective gene products were identified. Proteomic profiling of the hippocampus, a brain region with a pivotal role for neuronal information processing and storage may provide insight into the characteristics of proteins serving this highly sophisticated function.

Animals↗

Biological spectra analysis: Linking biological activity profiles to molecular structure.

Establishing quantitative relationships between molecular structure and broad biological effects has been a longstanding challenge in science. Currently, no method exists for forecasting broad biological activity profiles of medicinal agents even within narrow boundaries of structurally similar molecules. Starting from the premise that biological activity results from the capacity of small organic molecules to modulate the activity of the proteome, we set out to investigate whether descriptor sets could be developed for measuring and quantifying this molecular property. Using a 1,567-compound database, we show that percent inhibition values, determined at single high drug concentration in a battery of in vitro assays representing a cross section of the proteome, provide precise molecular property descriptors that identify the structure of molecules. When broad biological activity of molecules is represented in spectra form, organic molecules can be sorted by quantifying differences between biological spectra. Unlike traditional structure-activity relationship methods, sorting of molecules by using biospectra comparisons does not require knowledge of a molecule's putative drug targets. To illustrate this finding, we selected as starting point the biological activity spectra of clotrimazole and tioconazole because their putative target, lanosterol demethylase (CYP51), was not included in the bioassay array. Spectra similarity obtained through profile similarity measurements and hierarchical clustering provided an unbiased means for establishing quantitative relationships between chemical structures and biological activity spectra. This methodology, which we have termed biological spectra analysis, provides the capability not only of sorting molecules on the basis of biospectra similarity but also of predicting simultaneous interactions of new molecules with multiple proteins.

Molecular Structure↗

Bacterial alpha2-macroglobulins: colonization factors acquired by horizontal gene transfer from the metazoan genome?

BACKGROUND: Invasive bacteria are known to have captured and adapted eukaryotic host genes. They also readily acquire colonizing genes from other bacteria by horizontal gene transfer. Closely related species such as Helicobacter pylori and Helicobacter hepaticus, which exploit different host tissues, share almost none of their colonization genes. The protease inhibitor alpha2-macroglobulin provides a major metazoan defense against invasive bacteria, trapping attacking proteases required by parasites for successful invasion. RESULTS: Database searches with metazoan alpha2-macroglobulin sequences revealed homologous sequences in bacterial proteomes. The bacterial alpha2-macroglobulin phylogenetic distribution is patchy and violates the vertical descent model. Bacterial alpha2-macroglobulin genes are found in diverse clades, including purple bacteria (proteobacteria), fusobacteria, spirochetes, bacteroidetes, deinococcids, cyanobacteria, planctomycetes and thermotogae. Most bacterial species with bacterial alpha2-macroglobulin genes exploit higher eukaryotes (multicellular plants and animals) as hosts. Both pathogenically invasive and saprophytically colonizing species possess bacterial alpha2-macroglobulins, indicating that bacterial alpha2-macroglobulin is a colonization rather than a virulence factor. CONCLUSIONS: Metazoan alpha2-macroglobulins inhibit proteases of pathogens. The bacterial homologs may function in reverse to block host antimicrobial defenses. Alpha2-macroglobulin was probably acquired one or more times from metazoan hosts and has then spread widely through other colonizing bacterial species by more than 10 independent horizontal gene transfers. yfhM-like bacterial alpha2-macroglobulin genes are often found tightly linked with pbpC, encoding an atypical peptidoglycan transglycosylase, PBP1C, that does not function in vegetative peptidoglycan synthesis. We suggest that YfhM and PBP1C are coupled together as a periplasmic defense and repair system. Bacterial alpha2-macroglobulins might provide useful targets for enhancing vaccine efficacy in combating infections.

Alphaproteobacteria↗

Ontologies for proteomics: towards a systematic definition of structure and function that scales to the genome level.

A principal aim of post-genomic biology is elucidating the structures, functions and biochemical properties of all gene products in a genome. However, to adequately comprehend such a large amount of information we need new descriptions of proteins that scale to the genomic level. In short, we need a unified ontology for proteomics. Much progress has been made towards this end, including a variety of approaches to systematic structural and functional classification and initial work towards developing standardized, unified descriptions for protein properties. In relation to function, there is a particularly great diversity of approaches, involving placing a protein in structured hierarchies or more-generalized networks and a recent approach based on circumscribing a protein's function through systematic enumeration of molecular interactions.

Computational Biology↗

Identification of new intrinsic proteins in Arabidopsis plasma membrane proteome.

Identification and characterization of anion channel genes in plants represent a goal for a better understanding of their central role in cell signaling, osmoregulation, nutrition, and metabolism. Though channel activities have been well characterized in plasma membrane by electrophysiology, the corresponding molecular entities are little documented. Indeed, the hydrophobic protein equipment of plant plasma membrane still remains largely unknown, though several proteomic approaches have been reported. To identify new putative transport systems, we developed a new proteomic strategy based on mass spectrometry analyses of a plasma membrane fraction enriched in hydrophobic proteins. We produced from Arabidopsis cell suspensions a highly purified plasma membrane fraction and characterized it in detail by immunological and enzymatic tests. Using complementary methods for the extraction of hydrophobic proteins and mass spectrometry analyses on mono-dimensional gels, about 100 proteins have been identified, 95% of which had never been found in previous proteomic studies. The inventory of the plasma membrane proteome generated by this approach contains numerous plasma membrane integral proteins, one-third displaying at least four transmembrane segments. The plasma membrane localization was confirmed for several proteins, therefore validating such proteomic strategy. An in silico analysis shows a correlation between the putative functions of the identified proteins and the expected roles for plasma membrane in transport, signaling, cellular traffic, and metabolism. This analysis also reveals 10 proteins that display structural properties compatible with transport functions and will constitute interesting targets for further functional studies.

Arabidopsis↗

Mass Spectrometry-Based Profiling of Personalized Immunopeptidomes in Thai Renal Cell Carcinoma.

This study profiles the personalized immunopeptidomes of 13 Thai patients with renal cell carcinoma (RCC), addressing a critical knowledge gap in Southeast Asian populations characterized by distinct HLA allele distributions. We combined whole-exome sequencing (WES)-based personalized proteome construction with liquid chromatography-tandem mass spectrometry (LC-MS/MS), using both database-driven searches and de novo peptide sequencing. HLA typing identified several class I allotypes that are underrepresented in publicly available immunopeptidome resources, including seven alleles not previously represented in the databases examined; HLA-A*11:01 was the most frequent allele in this cohort. Database-based analysis identified a single tumor-specific neoantigen derived from a mutant JADE2 peptide in the patient with the highest tumor mutational burden, which was validated by a mutant-specific ELISPOT response. In contrast, de novo sequencing revealed numerous noncanonical peptides, a subset of which were supported by proteogenomic validation using PepQuery and detected exclusively in cancer proteomes but not in normal tissue data sets, indicating their potential as tumor-associated antigen candidates. Together, these results establish an integrated and scalable framework for identifying HLA-presented tumor-derived peptides and provide a foundational immunopeptidome resource to support personalized cancer immunotherapy development in Southeast Asia.

Humans↗

Intimate evolution of proteins. Proteome atomic content correlates with genome base composition.

Discerning the significant relations that exist within and among genome sequences is a major step toward the modeling of biopolymer evolution. Here we report the systematic analysis of the atomic composition of proteins encoded by organisms representative of each kingdoms. Protein atomic contents are shown to vary largely among species, the larger variations being observed for the main architectural component of proteins, the carbon atom. These variations apply to the bulk proteins as well as to subsets of ortholog proteins. A pronounced correlation between proteome carbon content and genome base composition is further evidenced, with high G+C genome content being related to low protein carbon content. The generation of random proteomes and the examination of the canonical genetic code provide arguments for the hypothesis that natural selection might have driven genome base composition.

Animals↗