Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

GenAge: a genomic and proteomic network map of human ageing.

The aim of this work was to provide an overview of the genetics of human ageing to gain novel insights about the mechanisms involved. By incorporating findings from model organisms to humans, such as mutations that either delay or accelerate ageing in mice, we constructed the gene networks previously related to ageing: namely, the network related to DNA metabolism and the network involving the GH/IGF-1 axis. Gathering data about the interacting partners of these proteins allowed us to suggest the involvement in ageing of a number of proteins through a "guilt-by-association" methodology. To organize our data, we developed the first curated database of genes related to human ageing: GenAge. With over 200 entries, GenAge may serve as a reference database of genes related to human ageing. Moreover, we rendered the first proteomic network map of human ageing, which suggests a relationship between the genetics of development and the genetics of ageing. Our work serves as a framework upon which a systems-biology understanding of ageing can be developed. GenAge is freely available for academic purposes at: http://genomics.senescence.info/genes/.

Aging↗

Vaccination against infections in chronic lymphocytic leukemia.

Chronic lymphocytic leukemia (CLL) is a well-defined mature B-cell neoplasm associated with increased susceptibility to infections. Two major options in prevention of infections in CLL, intravenous gammaglobulin treatment and antimicrobial chemoprophylaxis, have not resulted in satisfactory outcome. A third strategy, antimicrobial vaccination, is the topic of this minireview. We collected articles and their references concerning CLL vaccination from the Medline database starting from 1966 and thirteen relevant studies were found. Plain bacterial polysaccharide vaccines would seem to be ineffective in antibody formation in patients with CLL. However, protein and conjugate vaccines appear to be more immunogenic and their responses may be further enhanced with ranitidine adjuvant treatment. New well-designed investigations are needed to develop appropriate vaccination strategies and evaluate vaccination efficacy in infection morbidity and mortality in CLL.

Bacterial Vaccines↗

HI-FEVER: a Nextflow pipeline for the high-throughput discovery and annotation of endogenous viral elements.

SUMMARY: Endogenous viral elements (EVEs) offer valuable insights into virus and host evolution, but their detection remains computationally and biologically challenging. We present HI-FEVER, a user-friendly Nextflow pipeline for the discovery of EVEs in eukaryotic host genomes. HI-FEVER is highly parallelizable and customizable, ensuring computational efficiency while allowing researchers to fine-tune parameters to their specific needs. Its output provides a comprehensive analysis of discovered EVEs, including detailed annotations which can provide evolutionary insights. HI-FEVER scales seamlessly to handle millions of viral protein queries across multiple host genomes on both laptops and high-performance computing nodes. AVAILABILITY AND IMPLEMENTATION: The HI-FEVER source code is available on GitHub at https://github.com/PaleovirologyLab/hi-fever. Minimal reference databases, test datasets and benchmarking results are hosted on the Open Science Framework at https://osf.io/y357r. A detailed wiki is available at https://github.com/PaleovirologyLab/hi-fever/wiki, including usage instructions, parameter descriptions, and guidance on interpreting outputs. The pipeline includes a Pixi environment compatible with Conda and Apptainer containerization, and Docker images. HI-FEVER has been tested on Linux, Windows (via WSL2), and macOS (Intel and ARM64).

Software↗

IMGT, the international ImMunoGeneTics information system, http://imgt.cines.fr: the reference in immunoinformatics.

IMGT, the international ImMunoGeneTics information system (http://imgt.cines.fr), is a high quality integrated information system specializing in immunoglobulins (IG), T cell receptors (TR), major histocompatibility complex (MHC) and related proteins of the immune system of human and other vertebrates, created in 1989, by the Laboratoire d'ImmunoGénétique Moléculaire (LIGM), at the Université Montpellier II, CNRS, Montpellier, France. IMGT is the global reference in immunogenetics and immunoinformatics and provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT includes three sequence databases (IMGT/LIGM-DB, IMGT/MHC-DB hosted at EBI, IMGT/PRIMER-DB), one genome database (IMGT/GENE-DB), one 3D structure database (IMGT/3Dstructure-DB), Web resources comprising 8000 HTML pages ("IMGT Marie-Paule page") and interactive tools for sequence (IMGT/V-QUEST, IMGT/JunctionAnalysis, IMGT/Allele-Align, IMGT/PhyloGene) and genome (IMGT/GeneSearch, IMGT/GeneView, IMGT/LocusView) analysis. IMGT data are expertly annotated according to the rules of the IMGT Scientific chart, based on the IMGT-ONTOLOGY concepts. IMGT tools are particularly useful for the analysis of the IG and TR repertoires in physiological normal and pathological situations. IMGT has important applications in medical research (repertoire analysis in autoimmune diseases, AIDS, leukemias, lymphomas, myelomas), biotechnology related to antibody engineering (phage displays, combinatorial libraries) and therapeutic approaches (graft, immunotherapy). IMGT is freely available at http://imgt.cines.fr.

Animals↗

Using standard positions and image fusion to create proteome maps from collections of two-dimensional gel electrophoresis images.

Databases for two-dimensional protein gels pose new challenges in extracting meaningful information from large numbers of experiments. In order to create expression profiles, positions of corresponding protein spots across all gel images have to be established. In larger gel sets errors may accumulate rapidly during this spot matching process, effectively limiting the number of samples available for data mining. Here we present a novel approach for organizing spot data based on the concept of a standard position for a protein species. Standard positions are meaningful average positions that are determined using all occurrences of a protein species. They can be extended to spots that are not annotated via interpolation. The standard position of a spot can serve as a unifying index across all gels in a database, thus allowing creation and analysis of expression profiles that span the whole collection. The standard position gives a much more accurate estimation of a spot's position on a gel than can be obtained using theoretical isoelectric point and molecular weight. Positional indexing is a complement to a priori identifications (e.g. by mass spectrometry or Edman degradation). Moreover it can be used in advance to select spots that are worth identifying because they show relevant expression profiles. Furthermore, we show how to combine all spots that occur on any of the gels into one synthetic but nevertheless realistic-looking image. This composite image is produced such that all spots have their standard positions. It can serve as a proteome reference map for an organism. As an application, we have computed a reference map from 23 gel images of Bacillus subtilis, using an enhanced prerelease version of the gel analysis software Delta2D (DECODON, Greifswald, Germany).

Bacterial Proteins↗

TCOF1 mutation database: novel mutation in the alternatively spliced exon 6A and update in mutation nomenclature.

Recently, a novel exon was described in TCOF1 that, although alternatively spliced, is included in the major protein isoform. In addition, most published mutations in this gene do not conform to current mutation nomenclature guidelines. Given these observations, we developed an online database of TCOF1 mutations in which all the reported mutations are renamed according to standard recommendations and in reference to the genomic and novel cDNA reference sequences (www.genoma.ib.usp.br/TCOF1_database). We also report in this work: 1) results of the first screening for large deletions in TCOF1 by Southern blot in patients without mutation detected by direct sequencing; 2) the identification of the first pathogenic mutation in the newly described exon 6A; and 3) statistical analysis of pathogenic mutations and polymorphism distribution throughout the gene.

Alternative Splicing↗

Multiple SVM-RFE for gene selection in cancer classification with expression data.

This paper proposes a new feature selection method that uses a backward elimination procedure similar to that implemented in support vector machine recursive feature elimination (SVM-RFE). Unlike the SVM-RFE method, at each step, the proposed approach computes the feature ranking score from a statistical analysis of weight vectors of multiple linear SVMs trained on subsamples of the original training data. We tested the proposed method on four gene expression datasets for cancer classification. The results show that the proposed feature selection method selects better gene subsets than the original SVM-RFE and improves the classification accuracy. A Gene Ontology-based similarity assessment indicates that the selected subsets are functionally diverse, further validating our gene selection method. This investigation also suggests that, for gene expression-based cancer classification, average test error from multiple partitions of training and test sets can be recommended as a reference of performance quality.

Algorithms↗

A World Wide Web-service to aid the development of AMBER parameters using analogy to standard parameters.

A World-Wide Web service has been constructed to assist the development of force field parameters extending the AMBER force fields. This service extracts parameters from the standard AMBER force field parameter databases. From a Web-based interface the user can choose between bond, angle and torsional parameters with certain constraints on element and/or atom hybridization. The software constructed for the purpose of finding appropriate parameters will locate standard AMBER force field parameters matching the user specification. For bond and angle parameters a scatter plot of the reference values against force constants is provided. This service has been produced to assist in extraction and evaluation of parameters that may be useful for molecules other than proteins and nucleic acids.

Biochemical Phenomena↗

LISTA, LISTA-HOP and LISTA-HON: a comprehensive compilation of protein encoding sequences and its associated homology databases from the yeast Saccharomyces.

We continued our effort to make a comprehensive database (LISTA) for the yeast Saccharomyces cerevisiae. In this database each sequence has been attributed a single genetic name. In the case of duplicated sequences a simple method has been applied to distinguish between sequences of one and the same gene from non-allelic sequences of duplicated genes. If necessary, synonyms are given in the case of allelic duplicated sequences. Thus sequences can be found either by the name or by synonyms given in LISTA. Each entry contains the genetic name, the mnemonic from the EMBL data bank, the codon bias, reference of the publication of the sequence, Chromosomal location as far as known, Swissprot and EMBL accession numbers. To obtain more information on the included sequences, each entry has been screened against non-redundant nucleotide and protein data bank collections resulting in LISTA-HON and LISTA-HOP. The LISTA data base can be linked to the associated data sets or to nucleotide and protein banks by the Sequence Retrieval System (SRS).

Base Sequence↗

Open source system for analyzing, validating, and storing protein identification data.

This paper describes an open-source system for analyzing, storing, and validating proteomics information derived from tandem mass spectrometry. It is based on a combination of data analysis servers, a user interface, and a relational database. The database was designed to store the minimum amount of information necessary to search and retrieve data obtained from the publicly available data analysis servers. Collectively, this system was referred to as the Global Proteome Machine (GPM). The components of the system have been made available as open source development projects. A publicly available system has been established, comprised of a group of data analysis servers and one main database server.

Computational Biology↗

Separation and characterization of rice proteins.

Rice proteins from nine tissues and one organelle (leaf, chloroplast, stem, root, germ, dark germinated seedling, seed, bran, chaff and callus) were isolated and then separated by two-dimensional gel electrophoresis (2-DE). The protein spots were characterized according to molecular weight, isoelectric point and partial amino-terminal sequence. Electrophoresis was carried out by isoelectric focusing (IEF), nonequilibrium pH gradient electrophoresis (NEPHGE) and immobilized pH gradient (IPG) in the first dimension, and by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) in the second dimension. With the aid of nine marker proteins, the patterns of IEF, NEPHGE and IPG 2-DE gels were graphically combined by computer into a single synthetic image for each tissue, respectively, and these images for the nine tissues and one organelle were again combined into a single 2-DE image for the integrated rice protein spots. The rice 2-DE gel image resolved 4892 proteins. About 3% of the spots are characterized by amino-terminal sequencing.

Amino Acid Sequence↗

Two-dimensional gel protein database of Saccharomyces cerevisiae.

With the systematic sequencing of the yeast genome, yeast biology has entered a new era where novel challenges have to be faced. One challenge is the identification of the function of the several hundred novel genes discovered by genome sequencing. Another is to understand how all yeast genes act in concert to ensure and maintain cell organization. Two-dimensional (2-D) gel electrophoresis is the technique of choice to take up these challenges because it provides the opportunity of obtaining an overall view of genome expression. In prospect of these studies we have undertaken the construction of a yeast 2-D gel protein database that contains information on polypeptides of the yeast protein map. In this paper we report the information presently contained in this database. The reported information includes the identification of 250 protein spots and the characterization of polypeptides corresponding to N-terminal acetylated proteins, mitochondrial proteins, glucose-repressed proteins, heat shock induced proteins and proteins encoded by intron-containing genes. In all, 600 spots are annotated. These data can be accessed on the Yeast Protein Map server through the World Wide Web network.

Computer Communication Networks↗

Derangement of hypothetical proteins in fetal Down's syndrome brain.

The success of the Human Genome Project (HGP) enables prediction of proteins by computer programs from nucleic acid sequences and for which there is no experimental evidence. Clues for function of hypothetical proteins are provided by sequence similarity with proteins of known function in model organisms. The availability of this bulk of new data is of immediate importance to Down's syndrome (DS) research. DS is the most common human chromosomal abnormality caused by an extra copy of chromosome 21 and is characterized by somatic anomalies and mental retardation. In addition, overexpression of chromosome 21 genes is directly or indirectly responsible for mental retardation and other phenotypic abnormalities of DS. To allow insight into how trisomy 21 represents the phenotype of DS, we constructed a two-dimensional protein map and investigated expression of 8 hypothetical proteins in fetal DS (n = 7) and control (n = 7) brains (cortex). Two-dimensional electrophoresis (2-DE) with subsequent in-gel digestion of spots and matrix-assisted laser desorption/ionization (MALDI) spectroscopic identification followed by quantification of spots with specific software was applied. Quantitative analysis of hypothetical protein FLJ10849, hypothetical protein FLJ20113, and activator of hsp90 ATPase homologue 1 (AHA1) revealed levels comparable between DS and controls. By contrast, expression levels of hypothetical protein KIAA1185, hypothetical protein 55.2 kDa, hypothetical protein 58.8 kDa, actin-related protein 3beta (ARP3beta), and putative GTP-binding protein PTD004 were significantly decreased (P < 0.05) in fetal DS brain, and domain analysis suggests involvement in cytoskeleton, signaling, and chaperone system abnormalities.

Abortion, Induced↗

DynaProt 2D: an advanced proteomic database for dynamic online access to proteomes and two-dimensional electrophoresis gels.

DynaProt 2D presents an advanced online database for dynamic access to proteomes and two-dimensional (2D) gels. The database was designed to administer complete in silico proteomes and links them with experimental proteomic data in the manner of 2D electrophoresis gels (IPG-Dalt). The 2D gels serve as reference maps in 2D gel analysis as well as tools for navigation of the database to switch between experimental and predicted data. Therefore, all identified spots in the gels are clickable and linked with summarized protein information. The protein information tables contain calculated characteristics, which are often used in proteomics, such as the molecular weight, isoelectric point, codon adaptation index, grand average of hydropathicity, etc. The design of the database permits online extension of gel data and protein attributes without knowledge of any software language. Besides navigation via 2D gels, the clear graphical user interface permits quick and intuitive searching throughout complete proteomes and supports, e.g. the search for proteins with isoelectric points within pH ranges of interest or protein classes (e.g. ribosomal proteins or transporters). The first organism implemented in the database is Lactococcus lactis. The database is available at www.wzw.tum.de/proteomik/lactis.

Bacterial Proteins↗

Unique NS5b hepatitis C virus gene sequence consensus database is essential for standardization of genotype determinations in multicenter epidemiological studies.

A multicenter study of NS5b hepatitis C virus (HCV) genotype determination involving 12 laboratories demonstrates that any laboratory with expertise in sequencing techniques would be able to provide a reliable HCV genotype for clinical and epidemiological purposes as long as they are provided a consensus reference sequence database.

Consensus Sequence↗

A database application for pre-processing, storage and comparison of mass spectra derived from patients and controls.

BACKGROUND: Statistical comparison of peptide profiles in biomarker discovery requires fast, user-friendly software for high throughput data analysis. Important features are flexibility in changing input variables and statistical analysis of peptides that are differentially expressed between patient and control groups. In addition, integration the mass spectrometry data with the results of other experiments, such as microarray analysis, and information from other databases requires a central storage of the profile matrix, where protein id's can be added to peptide masses of interest. RESULTS: A new database application is presented, to detect and identify significantly differentially expressed peptides in peptide profiles obtained from body fluids of patient and control groups. The presented modular software is capable of central storage of mass spectra and results in fast analysis. The software architecture consists of 4 pillars, 1) a Graphical User Interface written in Java, 2) a MySQL database, which contains all metadata, such as experiment numbers and sample codes, 3) a FTP (File Transport Protocol) server to store all raw mass spectrometry files and processed data, and 4) the software package R, which is used for modular statistical calculations, such as the Wilcoxon-Mann-Whitney rank sum test. Statistic analysis by the Wilcoxon-Mann-Whitney test in R demonstrates that peptide-profiles of two patient groups 1) breast cancer patients with leptomeningeal metastases and 2) prostate cancer patients in end stage disease can be distinguished from those of control groups. CONCLUSION: The database application is capable to distinguish patient Matrix Assisted Laser Desorption Ionization (MALDI-TOF) peptide profiles from control groups using large size datasets. The modular architecture of the application makes it possible to adapt the application to handle also large sized data from MS/MS- and Fourier Transform Ion Cyclotron Resonance (FT-ICR) mass spectrometry experiments. It is expected that the higher resolution and mass accuracy of the FT-ICR mass spectrometry prevents the clustering of peaks of different peptides and allows the identification of differentially expressed proteins from the peptide profiles.

Algorithms↗

SUBA: the Arabidopsis Subcellular Database.

Knowledge of protein localisation contributes towards our understanding of protein function and of biological inter-relationships. A variety of experimental methods are currently being used to produce localisation data that need to be made accessible in an integrated manner. Chimeric fluorescent fusion proteins have been used to define subcellular localisations with at least 1100 related experiments completed in Arabidopsis. More recently, many studies have employed mass spectrometry to undertake proteomic surveys of subcellular components in Arabidopsis yielding localisation information for approximately 2600 proteins. Further protein localisation information may be obtained from other literature references to analysis of locations (AmiGO: approximately 900 proteins), location information from Swiss-Prot annotations (approximately 2000 proteins); and location inferred from gene descriptions (approximately 2700 proteins). Additionally, an increasing volume of available software provides location prediction information for proteins based on amino acid sequence. We have undertaken to bring these various data sources together to build SUBA, a SUBcellular location database for Arabidopsis proteins. The localisation data in SUBA encompasses 10 distinct subcellular locations, >6743 non-redundant proteins and represents the proteins encoded in the transcripts responsible for 51% of Arabidopsis expressed sequence tags. The SUBA database provides a powerful means by which to assess protein subcellular localisation in Arabidopsis (http://www.suba.bcs.uwa.edu.au).

Arabidopsis Proteins↗

Proteomic analysis of protein components in periodontal ligament fibroblasts.

BACKGROUND: Characterization of periodontal ligament (PDL) fibroblast proteome is an important tool for understanding PDL physiology and regulation and for identifying disease-related protein markers. PDL fibroblast protein expression has been studied using immunological methods, although limited to previously identified proteins for which specific antibodies are available. METHODS: We applied proteomic analysis coupled with mass spectrometry and database knowledge to human PDL fibroblasts. RESULTS: We detected 900 spots and identified 117 protein spots originating in 74 different genes. In addition to scaffold cytoskeletal proteins, e.g., actin, tubulin, and vimentin, we identified proteins implicated with cellular motility and membrane trafficking, chaparonine, stress and folding proteins, metabolic enzymes, proteins associated with detoxification and membrane activity, biodegradative metabolism, translation and transduction, extracellular proteins, and cell cycle regulation proteins. CONCLUSIONS: Most of these identified proteins are closely related to the extensive PDL fibroblasts' functions and homeostasis. Our PDL fibroblast proteome map can serve as a reference map for future clinical studies as well as basic research.

Adolescent↗