Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

SUBA: the Arabidopsis Subcellular Database.

Knowledge of protein localisation contributes towards our understanding of protein function and of biological inter-relationships. A variety of experimental methods are currently being used to produce localisation data that need to be made accessible in an integrated manner. Chimeric fluorescent fusion proteins have been used to define subcellular localisations with at least 1100 related experiments completed in Arabidopsis. More recently, many studies have employed mass spectrometry to undertake proteomic surveys of subcellular components in Arabidopsis yielding localisation information for approximately 2600 proteins. Further protein localisation information may be obtained from other literature references to analysis of locations (AmiGO: approximately 900 proteins), location information from Swiss-Prot annotations (approximately 2000 proteins); and location inferred from gene descriptions (approximately 2700 proteins). Additionally, an increasing volume of available software provides location prediction information for proteins based on amino acid sequence. We have undertaken to bring these various data sources together to build SUBA, a SUBcellular location database for Arabidopsis proteins. The localisation data in SUBA encompasses 10 distinct subcellular locations, >6743 non-redundant proteins and represents the proteins encoded in the transcripts responsible for 51% of Arabidopsis expressed sequence tags. The SUBA database provides a powerful means by which to assess protein subcellular localisation in Arabidopsis (http://www.suba.bcs.uwa.edu.au).

Arabidopsis Proteins↗

Proteomic analysis of protein components in periodontal ligament fibroblasts.

BACKGROUND: Characterization of periodontal ligament (PDL) fibroblast proteome is an important tool for understanding PDL physiology and regulation and for identifying disease-related protein markers. PDL fibroblast protein expression has been studied using immunological methods, although limited to previously identified proteins for which specific antibodies are available. METHODS: We applied proteomic analysis coupled with mass spectrometry and database knowledge to human PDL fibroblasts. RESULTS: We detected 900 spots and identified 117 protein spots originating in 74 different genes. In addition to scaffold cytoskeletal proteins, e.g., actin, tubulin, and vimentin, we identified proteins implicated with cellular motility and membrane trafficking, chaparonine, stress and folding proteins, metabolic enzymes, proteins associated with detoxification and membrane activity, biodegradative metabolism, translation and transduction, extracellular proteins, and cell cycle regulation proteins. CONCLUSIONS: Most of these identified proteins are closely related to the extensive PDL fibroblasts' functions and homeostasis. Our PDL fibroblast proteome map can serve as a reference map for future clinical studies as well as basic research.

Adolescent↗

The PROSITE database.

The PROSITE database consists of a large collection of biologically meaningful signatures that are described as patterns or profiles. Each signature is linked to a documentation that provides useful biological information on the protein family, domain or functional site identified by the signature. The PROSITE database is now complemented by a series of rules that can give more precise information about specific residues. During the last 2 years, the documentation and the ScanProsite web pages were redesigned to add more functionalities. The latest version of PROSITE (release 19.11 of September 27, 2005) contains 1329 patterns and 552 profile entries. Over the past 2 years more than 200 domains have been added, and now 52% of UniProtKB/Swiss-Prot entries (release 48.1 of September 27, 2005) have a cross-reference to a PROSITE entry. The database is accessible at http://www.expasy.org/prosite/.

Amino Acids↗

Descriptive analytical data and consequences for calculation of common reference intervals in the Nordic Reference Interval Project 2000.

In the Nordic Reference Interval Project (NORIP), data from 102 Nordic clinical chemical laboratories were obtained. Each laboratory reported analytical data on up to 25 of the most commonly used clinical biochemical properties, including results from each of a minimum of 25 reference individuals. A reference material consisting of a liquid frozen pool of serum with values traceable to reference methods (used as the project "calibrator" for non-enzymes to correct reference values) was measured together with other serum pool controls in each laboratory in the same analytical series as the project samples. The data on the controls were used to evaluate the analytical quality of the routine methods. For reference interval calculations, only such reference values on enzymes were accepted that were obtained by applying the International Federation of Clinical Chemistry (IFCC) compatible methods (37 degrees C), while "calibrator"-corrected reference values were used in the cases of non-enzymes. For each property, gender- and age-specific reference intervals were estimated, based on simple non-parametric calculations and using objective criteria to perform partitioning into subgroups. It is concluded that the same reference intervals are applicable in all five Nordic countries. The following descriptive data for the considered properties are presented in the tables: number of measurement values from each country and measurement system, certified/indicative target values for controls, differences between methods and measurement systems together with coefficients of variation, effects of control correction on the measurement values, differences between subgroups as determined by age, gender, country and material, and comparison of the new reference intervals with those presented in standard textbooks. The 25 components involved in this project were (listed in alphabetical order): Alanine transaminase, albumin, alkaline phosphatase, amylase, amylase pancreatic type, aspartate transaminase, bilirubin, calcium, carbamide, cholesterol, creatine kinase, creatininium, gamma-glutamyltransferase, glucose, HDL-cholesterol, iron, iron-binding capacity, lactate dehydrogenase, magnesium, phosphate, potassium, protein, sodium, triglyceride and urate.

Blood Chemical Analysis↗

Characterisation of rice anther proteins expressed at the young microspore stage.

In combination with two-dimensional polyacrylamide gel electrophoresis (2-DE) protein mapping and mass spectrometry analysis, the pattern of gene expression in specific tissues at a specific stage can be displayed and characterised. We used this approach for rice (Oryza sativa L. cultivar Doongara) to display and assign identity to proteins in the anthers at the young microspore stage. Over 4000 anther proteins in the pI range of 4-11 and molecular mass range of 6-122 kDa were reproducibly resolved after silver staining, representing about 10% of the estimated total genomic output of rice. Two hundred and seventy-three protein spots have been extracted either from polyninylidene diffluoride membrane blots or from colloidal Coomassie blue stained 2-DE gels and analysed by N-terminal sequencing, Matrix-assisted laser desorption/ionization-time of flight mass spectrometry (MS) analysis or tandem MS sequencing. This enabled identification of 53 anther protein spots representing 43 different proteins. Using the publicly available rice expressed sequence tag (EST) database at the National Centre for Biotechnology Information, a further 37 protein spots were matched to ESTs. After BLAST searching with these ESTs, we were able to predict the identity of 22 of these protein spots. Proteome reference maps of rice anthers have been constructed according to the SWISS-2DPAGE standards and are available for public access at http://semele.anu.edu.au/2d/2d.html.

Electrophoresis, Gel, Two-Dimensional↗

EcoGene: a genome sequence database for Escherichia coli K-12.

The EcoGene database provides a set of gene and protein sequences derived from the genome sequence of Escherichia coli K-12. EcoGene is a source of re-annotated sequences for the SWISS-PROT and Colibri databases. EcoGene is used for genetic and physical map compilations in collaboration with the Coli Genetic Stock Center. The EcoGene12 release includes 4293 genes. EcoGene12 differs from the GenBank annotation of the complete genome sequence in several ways, including (i) the revision of 706 predicted or confirmed gene start sites, (ii) the correction or hypothetical reconstruction of 61 frame-shifts caused by either sequence error or mutation, (iii) the reconstruction of 14 protein sequences interrupted by the insertion of IS elements, and (iv) pre-dictions that 92 genes are partially deleted gene fragments. A literature survey identified 717 proteins whose N-terminal amino acids have been verified by sequencing. 12 446 cross-references to 6835 literature citations and s are provided. EcoGene is accessible at a new website: http://bmb.med.miami.edu/EcoGene/EcoWeb. Users can search and retrieve individual EcoGene GenePages or they can download large datasets for incorporation into database management systems, facilitating various genome-scale computational and functional analyses.

Databases, Factual↗

MAPU: Max-Planck Unified database of organellar, cellular, tissue and body fluid proteomes.

Mass spectrometry (MS)-based proteomics has become a powerful technology to map the protein composition of organelles, cell types and tissues. In our department, a large-scale effort to map these proteomes is complemented by the Max-Planck Unified (MAPU) proteome database. MAPU contains several body fluid proteomes; including plasma, urine, and cerebrospinal fluid. Cell lines have been mapped to a depth of several thousand proteins and the red blood cell proteome has also been analyzed in depth. The liver proteome is represented with 3200 proteins. By employing high resolution MS and stringent validation criteria, false positive identification rates in MAPU are lower than 1:1000. Thus MAPU datasets can serve as reference proteomes in biomarker discovery. MAPU contains the peptides identifying each protein, measured masses, scores and intensities and is freely available at http://www.mapuproteome.com using a clickable interface of cell or body parts. Proteome data can be queried across proteomes by protein name, accession number, sequence similarity, peptide sequence and annotation information. More than 4500 mouse and 2500 human proteins have already been identified in at least one proteome. Basic annotation information and links to other public databases are provided in MAPU and we plan to add further analysis tools.

Animals↗

GeneViTo: visualizing gene-product functional and structural features in genomic datasets.

BACKGROUND: The availability of increasing amounts of sequence data from completely sequenced genomes boosts the development of new computational methods for automated genome annotation and comparative genomics. Therefore, there is a need for tools that facilitate the visualization of raw data and results produced by bioinformatics analysis, providing new means for interactive genome exploration. Visual inspection can be used as a basis to assess the quality of various analysis algorithms and to aid in-depth genomic studies. RESULTS: GeneViTo is a JAVA-based computer application that serves as a workbench for genome-wide analysis through visual interaction. The application deals with various experimental information concerning both DNA and protein sequences (derived from public sequence databases or proprietary data sources) and meta-data obtained by various prediction algorithms, classification schemes or user-defined features. Interaction with a Graphical User Interface (GUI) allows easy extraction of genomic and proteomic data referring to the sequence itself, sequence features, or general structural and functional features. Emphasis is laid on the potential comparison between annotation and prediction data in order to offer a supplement to the provided information, especially in cases of "poor" annotation, or an evaluation of available predictions. Moreover, desired information can be output in high quality JPEG image files for further elaboration and scientific use. A compilation of properly formatted GeneViTo input data for demonstration is available to interested readers for two completely sequenced prokaryotes, Chlamydia trachomatis and Methanococcus jannaschii. CONCLUSIONS: GeneViTo offers an inspectional view of genomic functional elements, concerning data stemming both from database annotation and analysis tools for an overall analysis of existing genomes. The application is compatible with Linux or Windows ME-2000-XP operating systems, provided that the appropriate Java Runtime Environment is already installed in the system.

Bacterial Proton-Translocating ATPases↗

Biological Macromolecule Crystallization Database, Version 3.0: new features, data and the NASA archive for protein crystal growth data.

Version 3.0 of the NIST/NASA/CARB Biological Macromolecule Crystallization Database (BMCD) includes crystal and crystallization data on all forms of biological macromolecules which have produced crystals suitable for X-ray diffraction studies. The data include summary information on each of the macromolecules, crystal data, crystallization conditions and comments about the crystallization procedure if it varies from the traditional methods employed for crystal growth. The database-management software maintains continuity with previous versions providing similar search procedures and displays. Version 3.0 of the BMCD includes protocols and results of crystallization experiments undertaken in space. These new data are comprised of both the NASA Protein Crystal Growth Archive, which includes information on all NASA-sponsored protein crystal growth experiments, and data describing other internationally sponsored microgravity macromolecule crystallization studies. The entries for the space growth crystallization experiments contain the crystallization protocols, apparatus descriptions, flight summary data, indication of success or failure of the experiments, references, etc. Other new features of the BMCD include the addition of crystallization procedures for small peptides and cross references to other structural biology databases.

Journal Article↗

The Role of Small Segmental Duplications in Generating Identical Isoforms Through Alternative Splicing Sites.

Alternative splicing plays a crucial role in expanding proteomic diversity but can also generate identical isoforms under certain conditions. While mutually exclusive splicing of tandem exons has occasionally been reported to produce identical isoforms, the extent to which other splicing events contribute to this phenomenon remains unclear. In this study, we demonstrate that alternative 5' and 3' splice site selection can also lead to the formation of identical isoforms, providing an additional type of splicing event for functional redundancy in transcriptomes. To address this, we analyzed reference genome annotations from 15 plant species, including Arabidopsis thaliana and wheat (Triticum aestivum), obtained from the RefSeq database. Identical isoforms were computationally defined as transcripts with distinct exon-intron structures but identical coding sequences. Our analysis reveals that the majority of alternative 5' and 3' fragments originate from small segmental duplications, suggesting that sequence repetition within gene regions facilitates the emergence of such splicing patterns. We also observed differences in the annotated 5' UTRs of some identical isoforms. However, since the alternative splicing sites themselves were not located within UTRs, these differences may reflect annotation uncertainty rather than genuine AS-derived variation. Given that UTR predictions in reference databases are not always precise, such observations should be interpreted cautiously. Expression analysis using an isoform-specific k-mer approach confirmed that identical isoforms can be differentially regulated. These findings suggest that, beyond expanding protein diversity, alternative splicing can also generate redundant isoforms that are differentially expressed at the RNA level, indicating potential regulatory roles. By elucidating the structural and regulatory factors contributing to the formation and retention of identical isoforms, our study provides new insights into the evolutionary and functional significance of alternative splicing in plants.

Alternative Splicing↗

A comprehensive proteomic analysis of the accessory sex gland fluid from mature Holstein bulls.

The expression of proteins in accessory sex gland fluid (AGF) of proven, high use mature Holstein bulls was evaluated. Thirty-seven bulls with documented fertility based on their non-return rates were studied. AGF was obtained by artificial vagina after bulls were surgically equipped with cannulae in the vasa deferentia. Samples of AGF were evaluated by two-dimensional SDS-PAGE, gels stained with Coomassie blue and polypeptide maps analyzed by PDQuest software. A master gel generated by the software representing the best pattern of spots in the AGF polypeptide maps was used as a reference for protein identification. Proteins were identified by Western blots and capillary liquid chromatography-nanoelectrospray ionization tandem-mass spectrometry (CapLC-MS/MS). The product ion spectra were processed using Protein Lynx Global Server 2.1 prior to database search with both PLGS and MASCOT (Matrix Science) software. The entire NCBI database was considered for mass fingerprint matching. An average of 52+/-5 spots was detected in the AGF 2D gels, which corresponded to proteins potentially involved in capacitation (bovine seminal plasma protein-BSP-A1/A2 and A3, BSP 30 kDa, albumin); sperm membrane protection, prevention of oxidative stress, complement-mediated sperm destruction and anti-microbial activity (albumin, clusterin, acidic seminal fluid protein--aSFP, 5'-nucleotidase--5'-NT, phospholipase A2--PLA2); acrosome reaction and sperm-oocyte interaction (PLA2, osteopontin); interaction with the extracellular matrix (tissue inhibitor of metalloproteinase 2, clusterin) and sperm motility (aSFP, spermadhesin Z13, 5'-NT). The 20 spots distinguished in all gels were matched to proteins associated with these functions. Proteins identified by tandem mass spectrometry as ecto-ADP-ribosyltransferase 5 and nucleobindin, never described before in the accessory sex gland secretions, were also detected. In summary, we identified a diverse range of components in the accessory sex gland fluid of a select group of Holstein bulls with documented fertility. Known characteristics of these proteins suggest that they play important roles in sperm physiology after ejaculation.

Animals↗

The Peptaibol Database: a database for sequences and structures of naturally occurring peptaibols.

The Peptaibol Database is a sequence and structure resource for the unusual class of peptides known as peptaibols. These peptides exhibit antibiotic and membrane channel-forming activities. The database includes sequence, biological source and bibliographical data for the naturally occurring peptaibols. Information is also collated for the growing number of peptaibol 3D structures determined by either crystallography or NMR spectroscopy. The database can be obtained as a whole or can be queried by name, group, sequence motif, biological origin and/or literature reference. The Peptaibol Database can be freely accessed at http://www.cryst.bbk.ac.uk/peptaibol.

Anti-Bacterial Agents↗

The Retinome - defining a reference transcriptome of the adult mammalian retina/retinal pigment epithelium.

BACKGROUND: The mammalian retina is a valuable model system to study neuronal biology in health and disease. To obtain insight into intrinsic processes of the retina, great efforts are directed towards the identification and characterization of transcripts with functional relevance to this tissue. RESULTS: With the goal to assemble a first genome-wide reference transcriptome of the adult mammalian retina, referred to as the retinome, we have extracted 13,037 non-redundant annotated genes from nearly 500,000 published datasets on redundant retina/retinal pigment epithelium (RPE) transcripts. The data were generated from 27 independent studies employing a wide range of molecular and biocomputational approaches. Comparison to known retina-/RPE-specific pathways and established retinal gene networks suggest that the reference retinome may represent up to 90% of the retinal transcripts. We show that the distribution of retinal genes along the chromosomes is not random but exhibits a higher order organization closely following the previously observed clustering of genes with increased expression. CONCLUSION: The genome wide retinome map offers a rational basis for selecting suggestive candidate genes for hereditary as well as complex retinal diseases facilitating elaborate studies into normal and pathological pathways. To make this unique resource freely available we have built a database providing a query interface to the reference retinome 1.

Adult↗

Evaluation of the antihyperlipidemic properties of dietary supplements.

We reviewed the published literature regarding the antihyperlipidemic effects of dietary supplements. A search of MEDLINE database, EMBASE Drugs and Pharmacology database, and the Internet was performed, and pertinent studies were identified and evaluated. References from published articles and tertiary references were used to gather additional data. Published trials indicate that red yeast rice, tocotrienols, gugulipid, garlic, and soy protein all have antihypercholesterolemic effects. These supplements, as well as omega-3 fatty acids, also have antihypertriglyceridemic effects. In clinical trials none of the agents led to a reduction in low-density lipoproteins greater than 25%, suggesting modest efficacy. When recommending these supplements, clinicians should keep in mind that their long-term safety is not established and patients should be monitored closely.

Adult↗

In silico prediction of the impact of genomic variations in the small conductance calcium activated potassium channel SK3 structure and function.

The small-conductance calcium-activated potassium channel SK3, encoded by the KCNN3 gene, plays a critical role in regulating dopaminergic neuron (DN) firing patterns by modulating after hyperpolarization currents. SK3 dysfunction has been implicated in neuropsychiatric and neurodegenerative disorders. We analyzed structural and functional consequences of KCNN3 splicing and genetic variation. Alternative splicing variants of the KCNN3 gene were retrieved from the Ensembl database and aligned using T-Coffee, manually inspected and curated. Protein domains were identified with Pfam 35.0, SMART 9.0, and InterPro 98.0, and visualized. An AlphaFold2 model of SK3 full-length protein (UniProt: Q9UGI6) used as reference and structural models of its splicing variants were predicted with ColabFold. Functional domains (S1-S6 transmembrane helices, H5 pore loop, and calmodulin-binding) were defined and superimposed onto the AlphaFold2 reference. Domain integrity was assessed based on completeness of all expected residue indices within each functional region. SNPs and CNVs across all coding KCNN3 splicing variants were analyzed, classified, and filtered to isolate pathogenic variants prioritizing non-synonymous amino acid substitutions. Differential variant impacts across splicing isoforms were assessed by mapping variant positions to individual transcript protein sequences and used to predict functional consequences. Two long and two short splicing variants are known. Short variants lack the motif required for potassium channels. Pathogenic variants result from missense mutations resulting in amino acid substitutions. In all cases, the consequential effects depend on the specific location and role of the amino acid being changed.

SK3 channels↗

Structural interpretation of mutations and SNPs using STRAP-NT.

Visualization of residue positions in protein alignments and mapping onto suitable structural models is an important first step in the interpretation of mutations or polymorphisms in terms of protein function, interaction, and thermodynamic stability. Selecting and highlighting large numbers of residue positions in a protein structure can be time-consuming and tedious with currently available software. Previously, a series of tasks and analyses had to be performed one-by-one to map mutations onto 3D protein structures; STRAP-NT is an extension of STRAP that automates these tasks so that users can quickly and conveniently map mutations onto 3D protein structures. When the structure of the protein of interest is not yet available, a related protein can frequently be found in the structure databases. In this case the alignment of both proteins becomes the crucial part of the analysis. Therefore we embedded these program modules into the Java-based multiple sequence alignment program STRAP-NT. STRAP-NT can simultaneously map an arbitrary number of mutations denoted using either the nucleotide or amino acid sequence. When the designations of the mutations refer to genomic sites, STRAP-NT translates them into the corresponding amino acid positions, taking intron-exon boundaries into account. STRAP-NT tightly integrates a number of current protein structure viewers (currently PYMOL, RASMOL, JMOL, and VMD) with which mutations and polymorphisms can be directly displayed on the 3D protein structure model. STRAP-NT is available at the PDB site and at http://www.charite.de/bioinf/strap/ or http://strapjava.de.

DNA Mutational Analysis↗

Serial analysis of gene expression in the hippocampus of patients with mesial temporal lobe epilepsy.

Hippocampal sclerosis constitutes the most frequent neuropathological finding in patients with medically intractable mesial temporal lobe epilepsy. Serial analysis of gene expression was used to get a global view of the gene profile in human hippocampus in control condition and in epileptic condition associated with hippocampal sclerosis. Libraries were generated from control hippocampus, obtained by rapid autopsy, and from hippocampal surgical specimens of patients with mesial temporal lobe epilepsy and the classical pattern of hippocampal sclerosis. More than 50,000 tags were analyzed (28,282, control hippocampus; 25,953, hippocampal sclerosis) resulting in 9206 (control hippocampus) and 9599 (hippocampal sclerosis) unique tags (genes), each representing a specific mRNA transcript. Comparison of the two libraries resulted in the identification of 143 transcripts that were differentially expressed. These genes belong to a variety of functional classes, including basic metabolism, transcription regulation, protein synthesis and degradation, signal transduction, structural proteins, regeneration and synaptic plasticity and genes of unknown identity of function. The database generated by this study provides an extensive inventory of genes expressed in human control hippocampus, identifies new high-abundant genes associated with altered hippocampal morphology in patients with mesial temporal lobe epilepsy and serves as a reference for future studies aimed at detecting hippocampal transcriptional responses under various pathological conditions.

Base Sequence↗

Automatic extraction of mutations from Medline and cross-validation with OMIM.

Mutations help us to understand the molecular origins of diseases. Researchers, therefore, both publish and seek disease-relevant mutations in public databases and in scientific literature, e.g. Medline. The retrieval tends to be time-consuming and incomplete. Automated screening of the literature is more efficient. We developed extraction methods (called MEMA) that scan Medline abstracts for mutations. MEMA identified 24,351 singleton mutations in conjunction with a HUGO gene name out of 16,728 abstracts. From a sample of 100 abstracts we estimated the recall for the identification of mutation-gene pairs to 35% at a precision of 93%. Recall for the mutation detection alone was >67% with a precision rate of >96%. This shows that our system produces reliable data. The subset consisting of protein sequence mutations (PSMs) from MEMA was compared to the entries in OMIM (20,503 entries versus 6699, respectively). We found 1826 PSM-gene pairs to be in common to both datasets (cross-validated). This is 27% of all PSM-gene pairs in OMIM and 91% of those pairs from OMIM which co-occur in at least one Medline abstract. We conclude that Medline covers a large portion of the mutations known to OMIM. Another large portion could be artificially produced mutations from mutagenesis experiments. Access to the database of extracted mutation-gene pairs is available through the web pages of the EBI (refer to http://www.ebi. ac.uk/rebholz/index.html).

Animals↗