Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Soy isoflavone analysis: quality control and a new internal standard.

Development of a database of the soy isoflavone content of foods requires accurate and precise evaluation of different food matrixes. To evaluate accuracy, we estimated recoveries of both internal and external standards in 5 different soyfoods weekly. Standards were evaluated daily for system quality assurance. To evaluate sample precision, we analyzed soybeans and soymilk bimonthly for within-day precision and over 4 d for day-to-day precision. CVs should be < or = 8%. We validated our methods for single and multiple recovery concentrations by using our new internal standard, 2,4,4'-trihydroxydeoxybenzoin, and the external standards daidzein, genistein, and genistin. Concentrations of 12 isoflavone isomers, 3 aglycones (daidzein, genistein, and glycitein), and 9 glucosides (daidzin, genistin, glycitin, acetyldaidzin, acetylgenistin, acetylglycitin, malonyldaidzin, malonylgenistin, and malonylglycitin) were measured in a variety of soybeans and soyfoods. The extraction methods used depended on soyfood type. The HPLC conditions for soy isoflavone analysis were improved, leading to good separation with a short analysis time (60 min/sample). A data bank of concentration and distribution of isoflavones in different soybean products was assembled. A wide range of isoflavone concentrations, from < 50 microg/g to > 20,000 microg/g, was found in different soy products. The glucoside forms are almost twice the molecular weight of the aglycones; reported isoflavone concentrations should be normalized to the aglycone mass (or an isoflavonoid equivalent) rather than a simple sum of all isomers.

Chromatography, High Pressure Liquid↗

Location on the human genetic linkage map of 26 genes involved in blood coagulation.

Several human genetic linkage maps have been constructed as part of the Human Genome Project. These maps show the positional order of closely linked, highly informative AC-repeat polymorphisms on each human chromosome, and are extremely useful in genetic linkage analysis of inheritable diseases. For a candidate gene approach the current linkage maps are less useful, since they consist mainly of anonymous markers rather than of specific genes. This situation also applies for inheritable disorders of blood coagulation. Numerous genes are involved in the blood coagulation cascade and its regulation, and can be considered as candidate genes for unexplained haemophilia and thrombophilia. We have selected 29 candidate genes that seem to be the ones most likely to be involved in thrombophilia. For 19 genes genotype data were already present in the CEPH database (version 7.0). We typed 7 additional genes in the CEPH reference families, i.e. the factor V, factor XII, protein C, protein S, prothrombin, thrombomodulin, and heparin cofactor II gene. The genotype data were used to integrate these 26 genes in the current genetic linkage map, and to identify closely linked AC-repeat polymorphisms. This information will benefit the investigation of inheritable disorders of blood coagulation, especially thrombophilia.

Base Sequence↗

SubtiList: the reference database for the Bacillus subtilis genome.

SubtiList is the reference database dedicated to the genome of Bacillus subtilis 168, the paradigm of Gram-positive endospore-forming bacteria. Developed in the framework of the B.subtilis genome project, SubtiList provides a curated dataset of DNA and protein sequences, combined with the relevant annotations and functional assignments. Information about gene functions and products is continuously updated by linking relevant bibliographic references. Recently, sequence corrections arising from both systematic verifications and submissions by individual scientists were included in the reference genome sequence. SubtiList is based on a generic relational data schema and a World Wide Web interface developed for the handling of bacterial genomes, called GenoList. The World Wide Web interface was designed to allow users to easily browse through genome data and retrieve information according to common biological queries. SubtiList also provides more elaborate tools, such as pattern searching, which are tightly connected to the overall browsing system. SubtiList is accessible at http://genolist.pasteur.fr/SubtiList/. Similar bacterial databases are accessible at http://genolist.pasteur.fr/.

Bacillus subtilis↗

GOLD.db: genomics of lipid-associated disorders database.

BACKGROUND: The GOLD.db (Genomics of Lipid-Associated Disorders Database) was developed to address the need for integrating disparate information on the function and properties of genes and their products that are particularly relevant to the biology, diagnosis management, treatment, and prevention of lipid-associated disorders. DESCRIPTION: The GOLD.db http://gold.tugraz.at provides a reference for pathways and information about the relevant genes and proteins in an efficiently organized way. The main focus was to provide biological pathways with image maps and visual pathway information for lipid metabolism and obesity-related research. This database provides also the possibility to map gene expression data individually to each pathway. Gene expression at different experimental conditions can be viewed sequentially in context of the pathway. Related large scale gene expression data sets were provided and can be searched for specific genes to integrate information regarding their expression levels in different studies and conditions. Analytic and data mining tools, reagents, protocols, references, and links to relevant genomic resources were included in the database. Finally, the usability of the database was demonstrated using an example about the regulation of Pten mRNA during adipocyte differentiation in the context of relevant pathways. CONCLUSIONS: The GOLD.db will be a valuable tool that allow researchers to efficiently analyze patterns of gene expression and to display them in a variety of useful and informative ways, allowing outside researchers to perform queries pertaining to gene expression results in the context of biological processes and pathways.

Adipocytes↗

Nh3D: a reference dataset of non-homologous protein structures.

BACKGROUND: The statistical analysis of protein structures requires datasets in which structural features can be considered independently distributed, i.e. not related through common ancestry, and that fulfil minimal requirements regarding the experimental quality of the structures it contains. However, non-redundant datasets based on sequence similarity invariably contain distantly related homologues. Here we provide a reference dataset of non-homologous protein domains, assuming that structural dissimilarity at the topology level is incompatible with recognizable common ancestry. The dataset is based on domains at the Topology level of the CATH database which hierarchically classifies all protein structures. It contains the best refined representatives of each Topology level, validates structural dissimilarity and removes internally duplicated fragments. The compilation of Nh3D is fully scripted. RESULTS: The current Nh3D list contains 570 domains with a total of 90780 residues. It covers more than 70% of folds at the Topology level of the CATH database and represents more than 90% of the structures in the PDB that have been classified by CATH. We observe that even though all protein pairs are structurally dissimilar, some pairwise sequence identities after global alignment are greater than 30%. CONCLUSION: Nh3D is freely available as a reference dataset for the statistical analysis of sequence and structure features of proteins in the PDB. Regularly updated versions of Nh3D and the corresponding PDB-formatted coordinate sets are accessible from our Web site http://www.schematikon.org.

Algorithms↗

Database of mutations that alter the large tumor antigen in simian virus 40.

The SV40 T antigen database is a listing of plasmids and/or viruses that express mutant forms of the virus-encoded large T antigen protein. The parental virus strain, nucleic acid sequence of the mutations, the effect of the mutation on the T antigen amino acid sequence, and key references are included in the listing. The database is available from the authors as a Macintosh FileMaker Pro file, and as a hard copy printout.

Antigens, Polyomavirus Transforming↗

Rationally selected basis proteins: a new approach to selecting proteins for spectroscopic secondary structure analysis.

Protein basis sets have been extensively used as reference data for the determination of protein structure with optical methods such as circular dichroism and infrared spectroscopies. We have taken a new approach to basis protein selection by utilizing three crystal structure classification databases: CATH, SCOP, and PDB_SELECT. Through the use of the information available in these and other online resources, we identified 115 commercially available proteins as potential basis set candidates. By carefully screening the quality of the crystal structures and commercial protein preparations, we obtained a final set of 50 rationally selected proteins (RaSP50) that has been optimized for use in spectroscopic protein structure determination studies. These proteins span the full range of known protein folds as well as alpha-helix and beta-sheet contents, and they represent a more comprehensive variety of fold types than any previous reference set. This report includes a detailed presentation of the reasoning behind the rational protein selection process, a description of the properties of the RaSP50 set, and a discussion of the types of structural and spectral variations that are represented in the set.

Circular Dichroism↗

MitoP2, an integrated database on mitochondrial proteins in yeast and man.

The aim of the MitoP2 database (http://ihg.gsf.de/mitop2) is to provide a comprehensive list of mitochondrial proteins of yeast and man. Based on the current literature we created an annotated reference set of yeast and human proteins. In addition, data sets relevant to the study of the mitochondrial proteome are integrated and accessible via search tools and links. They include computational predictions of signalling sequences, and summarize results from proteome mapping, mutant screening, expression profiling, protein-protein interaction and cellular sublocalization studies. For each individual approach, specificity and sensitivity for allocating mitochondrial proteins was calculated. By providing the evidence for mitochondrial candidate proteins the MitoP2 database lends itself to the genetic characterization of human mitochondriopathies.

Computational Biology↗

Structure prediction and modelling.

Protein structure prediction from sequence remains a major goal in molecular biology. The methods described in this review concentrate on deriving structural information through the detection of similarities between a test sequence and a database of known structures. Such methods are often referred to as knowledge-based strategies reflecting the use of a structural database in the analyses. The past year has seen considerable advances in both the development of automated procedures and their application to protein sequences of outstanding biological interest.

Algorithms↗

AffinDB: a freely accessible database of affinities for protein-ligand complexes from the PDB.

AffinDB is a database of affinity data for structurally resolved protein-ligand complexes from the Protein Data Bank (PDB). It is freely accessible at http://www.agklebe.de/affinity. Affinity data are collected from the scientific literature, both from primary sources describing the original experimental work of affinity determination and from secondary references which report affinity values determined by others. AffinDB currently contains over 730 affinity entries covering more than 450 different protein-ligand complexes. Besides the affinity value, PDB summary information and additional data are provided, including the experimental conditions of the affinity measurement (if available in the corresponding reference); 2D drawing, SMILES code and molecular weight of the ligand; links to other databases, and bibliographic information. AffinDB can be queried by PDB code or by any combination of affinity range, temperature and pH value of the measurement, ligand molecular weight, and publication data (author, journal and year). Search results can be saved as tabular reports in text files. The database is supposed to be a valuable resource for researchers interested in biomolecular recognition and the development of tools for correlating structural data with affinities, as needed, for example, in structure-based drug design.

Databases, Protein↗

Proteins from bovine tissues and biological fluids: defining a reference electrophoresis map for liver, kidney, muscle, plasma and red blood cells.

A number of high resolution two-dimensional electrophoresis (2-DE) reference maps for bovine tissues and biological fluids have been determined for animals in basal state. Among the 1863 distinct protein features detected in samples of liver, kidney, muscle, plasma and red blood cells, 509 species were identified and associated to 209 different genes. Difficulties in the identification were related to the poorly characterized Bos taurus genome and were solved by a combined matrix-assisted laser desorption/ionisation-mass spectrometry and liquid chromatography-electrospray ionization tandem mass spectrometry approach. The experimental output allowed us to establish a 2-DE database accessible through the World Wide Web network at the URL address (http://www.iabbam.na.cnr.it/Biochem). These reference maps may serve as a tool in future veterinary medical studies aimed at the evaluation of changes in protein repertoire for altered animal physiological conditions and infectious diseases, to the definition of molecular markers for novel diagnostic kits and vaccines, as well as the characterization of protein modifications in bovine materials following technological processes used in the food industry.

Animals↗

Gene expression analysis in the hippocampal formation of tree shrews chronically treated with cortisol.

Adrenal corticosteroids influence the function of the hippocampus, the brain structure in which the highest expression of glucocorticoid receptors is found. Chronic high levels of cortisol elicited by stress or through exogenous administration can cause irreversible damage and cognitive deficits. In this study, we searched for genes expressed in the hippocampal formation after chronic cortisol treatment in male tree shrews. Animals were treated orally with cortisol for 28 days. At the end of the experiments, we generated two subtractive hippocampal hybridization libraries from which we sequenced 2,246 expressed sequenced tags (ESTs) potentially regulated by cortisol. To validate this approach further, we selected some of the candidate clones to measure mRNA expression levels in hippocampus using real-time PCR. We found that 66% of the sequences tested (10 of 15) were differentially represented between cortisol-treated and control animals. The complete set of clones was subjected to a bioinformatic analysis, which allowed classification of the ESTs into four different main categories: 1) known proteins or genes (approximately 28%), 2) ESTs previously published in the database (approximately 16%), 3) novel ESTs matching only the reference human or mouse genome (approximately 5%), and 4) sequences that do not match any public database (50%). Interestingly, the last category was the most abundant. Hybridization assays revealed that several of these clones are indeed expressed in hippocampal tissue from tree shrew, human, and/or rat. Therefore, we discovered an extensive inventory of new molecular targets in the hippocampus that serves as a reference for hippocampal transcriptional responses under various conditions. Finally, a detailed analysis of the genomic localization in human and mouse genomes revealed a survey of putative novel splicing variants for several genes of the nervous system.

Animals↗

Metatranscriptomic analysis of viral sequences associated with Culex nigripalpus at an Alabama aquaculture site.

Mosquitoes associated with aquaculture habitats can harbor diverse viruses, yet the viromes of many locally abundant species remain poorly characterized. At an aquaculture-associated site in Auburn, Alabama, we surveyed mosquito populations and found Culex nigripalpus to be the dominant species collected. To characterize viruses associated with this mosquito, we performed RNA-seq on pooled female Cx. nigripalpus and compared complementary bioinformatic workflows for viral detection and genome recovery. One workflow removed host-associated reads by mapping to the closest available mosquito reference genome prior to assembly, whereas a second workflow used fully de novo assembly and viral database annotation. Additional protein-level filtering, cross-workflow comparison, and comparison of Trinity and rnaSPAdes assemblies were used to prioritize well-supported viral candidates. Across the original analyses, 16 submitted accessions corresponding to 12 collapsed virus/name groups were recovered, including Merida virus, Hubei mosquito virus 5, Zhejiang mosquito virus, Hubei virga-like virus 3, Rinkaby virus, Elemess virus, Qingnian mosquito virus, Serbia narna-like virus 2, XiangYun narna-levi-like virus 8, Ecclesville picorna-like virus, and baculovirus-like fragments. Several candidates were supported across multiple workflows, while others were recovered only under specific analytical conditions, indicating that candidate recovery was influenced by assembly and filtering choices. Selected viral contigs were independently supported by RT-PCR amplification. Overall, these results provide a first characterization of viral sequences associated with Cx. nigripalpus from an Alabama aquaculture-associated site and show that comparison across assembly and filtering strategies helped prioritize the most consistently supported viral candidates.

Animals↗

Screening of Bifidobacterium strains isolated from human faeces for antagonistic activities against potentially bacterial pathogens.

As probiotic bacteria, strains belonging to the genus Bifidobacterium colonise the gastro-intestinal tract of humans and animals at the time of birth, and they are found in young as well as in adult individuals in great numbers. Moreover, they can interact with the development of enteric infections by the production of antimicrobial metabolites. In this work 281 strains of bifidobacteria were anaerobically isolated from human faecal samples, supplied by volunteers of different ages (youngs, adults, elders), and preliminarly described by microscopic observation. All strains were screened by the fructose 6-phosphate phosphoketolase (F6PPK) test in order to confirm their classification within the genus Bifidobacterium. Selected strains were used to evaluate their antagonistic activities against Escherichia coli, Salmonella thyphimurium, Staphylococcus lentus, Enterococcus faecalis, Acinetobacter calcoaceticus, Sphingomonas paucimobilis, Listeria monocytogenes, Yersinia enterocolitica, Bacillus cereus, Clostridium sporogenes. Experiments were performed in vitro by different methods based on the observation of growth inhibition in Petri dishes. The strains that showed the highest inhibiting activities were compared by SDS-PAGE for total cell proteins, using type strains of human origin as references. Representative isolates were metabolically characterised by the BIOLOG system; a specific database was created with strains obtained from our collection and a statistical evaluation for metabolic patterns was carried out.

Adolescent↗

The human cornea proteome: bioinformatic analyses indicate import of plasma proteins into the cornea.

Increased biochemical knowledge of normal and diseased corneas is essential for the understanding of corneal homeostasis and pathophysiology. In a recent study, we characterized the proteome of the normal human cornea and identified 141 distinct proteins. This dataset represents the most comprehensive protein study of the cornea to date and provides a useful reference for further studies of normal and diseased human corneas. The list of identified proteins is available at the Cornea Protein Database. In the present paper, we review the utilized procedures for extraction and fractionation of corneal proteins and discuss the potential roles of the identified proteins in relation to homeostasis, diseases, and wound-healing of the cornea. In addition, we compare the list of identified proteins with high quality gene expression libraries (cDNA libraries) and Serial Analysis of Gene Expression (SAGE) data. Of the 141 proteins, 86 (61%) were recognized in cDNA libraries from the corneas of dogs and rabbits, or humans with keratoconus, and 98 (69.5%) were recognized in SAGE data of mouse and human corneas. However, the percentages of identified genes in each of the protein functional groups differed markedly. Thus, exceptionally few of the traditional blood/plasma proteins and immune defense proteins that were identified in the human cornea were recognized in the gene expression libraries of the cornea. This observation strongly indicates that these abundant corneal proteins are not expressed in the cornea but originate from the surrounding pericorneal tissue.

Animals↗

Generalization of a targeted library design protocol: application to 5-HT7 receptor ligands.

Herein a general concept for the design of targeted libraries for proteins with binding sites that are divided into subsites is laid out, including several practical aspects and their solutions. The design is based on a chemogenomic classification of the subsites followed by collection of bioactive molecular fragments and virtual library generation. The general process is outlined and applied to the assembly of a library of 500 molecules targeting the serotonin type 7 (5-HT7) receptor, a class A G-Protein Coupled Receptor (GPCR). Utilizing commercially available building blocks of similar size and composition, a reference library was created. Control sets of known ligands for the 5-HT7 receptor, other GPCRs, and nuclear receptors were collected from literature sources. Principal component analysis of molecular descriptors for the two libraries and the literature sets, displayed a focusing of the targeted library to the region in the chemical space defined by the literature actives, suggesting a denser coverage of the bioactive region than for the more diverse reference library. Additional computational validations, including PCA class predictions, 3D pharmacophore modeling, and docking calculations all indicated an enrichment factor of 5-HT7 ligand-like molecules in the range of 2-4 for the targeted library compared to the reference library.

Binding Sites↗

The Ribonuclease P database.

The Ribonuclease P Sequence database is a compilation of RNase P sequences, sequence alignments, secondary structures, three-dimensional models, and accessory information. In its initial form, the database contains information on RNase P RNA in bacteria and archaea, and RNase P protein in bacteria. The sequences themselves are presented phylogenetically ordered and aligned. The database also contains secondary structures of bacterial and archaeal RNAs, including specially annotated 'reference' secondary structures of Escherichia coli and Bacillus subtilis RNase P RNAs, a minimum phylogenetic consensus structure, and coordinates for models of three-dimensional structure.

Bacillus subtilis↗

Twenty thousand ORFan microbial protein families for the biologist?

The genomes of most newly sequenced organisms contain a significant fraction of ORFs (open reading frames) that match no other sequence in the databases. We refer to these singleton ORFs as sequence ORFans. Because little can be learned about ORFans by homology, the origin and functions of ORFans remain a mystery. However, in this era of full genome sequencing, it seems that ORFans have been underemphasized. In this minireview, we draw attention to the increasing number of ORFans and to the consequences of this growth to biological research in the postgenomic era.

Animals↗