Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

New and simple equations to estimate the energy and fat contents and energy density of humans in sickness and health.

Two formulas were derived to estimate the energy content of the human body which use only body mass, total body water by 3H2O dilution space and body minerals assessed by anthropometry. The formulas were tested in a body composition database of 561 patients and 151 normal volunteers using established metabolizable energy values for protein, fat and glycogen. Total body protein was determined by in vivo neutron activation analysis (IVNAA), body water by dilution of tritium and body minerals from skeletal frame size. Body glycogen was assumed to be 14.6% of the mineral component. Body fat was obtained by difference, body mass less the sum of water, protein, minerals and glycogen. The standard deviation in the estimate of body energy content was 30 MJ or 4.1% of the energy content of reference man. Two formulas for body energy content were derived by regression with body mass, total body water and body minerals or height. Two formulas for energy density and formulas for percentage body fat were similarly derived.

Adolescent↗

Classification of protein quaternary structure by functional domain composition.

BACKGROUND: The number and the arrangement of subunits that form a protein are referred to as quaternary structure. Quaternary structure is an important protein attribute that is closely related to its function. Proteins with quaternary structure are called oligomeric proteins. Oligomeric proteins are involved in various biological processes, such as metabolism, signal transduction, and chromosome replication. Thus, it is highly desirable to develop some computational methods to automatically classify the quaternary structure of proteins from their sequences. RESULTS: To explore this problem, we adopted an approach based on the functional domain composition of proteins. Every protein was represented by a vector calculated from the domains in the PFAM database. The nearest neighbor algorithm (NNA) was used for classifying the quaternary structure of proteins from this information. The jackknife cross-validation test was performed on the non-redundant protein dataset in which the sequence identity was less than 25%. The overall success rate obtained is 75.17%. Additionally, to demonstrate the effectiveness of this method, we predicted the proteins in an independent dataset and achieved an overall success rate of 84.11% CONCLUSION: Compared with the amino acid composition method and Blast, the results indicate that the domain composition approach may be a more effective and promising high-throughput method in dealing with this complicated problem in bioinformatics.

Algorithms↗

CAFTAN: a tool for fast mapping, and quality assessment of cDNAs.

BACKGROUND: The German cDNA Consortium has been cloning full length cDNAs and continued with their exploitation in protein localization experiments and cellular assays. However, the efficient use of large cDNA resources requires the development of strategies that are capable of a speedy selection of truly useful cDNAs from biological and experimental noise. To this end we have developed a new high-throughput analysis tool, CAFTAN, which simplifies these efforts and thus fills the gap between large-scale cDNA collections and their systematic annotation and application in functional genomics. RESULTS: CAFTAN is built around the mapping of cDNAs to the genome assembly, and the subsequent analysis of their genomic context. It uses sequence features like the presence and type of PolyA signals, inner and flanking repeats, the GC-content, splice site types, etc. All these features are evaluated in individual tests and classify cDNAs according to their sequence quality and likelihood to have been generated from fully processed mRNAs. Additionally, CAFTAN compares the coordinates of mapped cDNAs with the genomic coordinates of reference sets from public available resources (e.g., VEGA, ENSEMBL). This provides detailed information about overlapping exons and the structural classification of cDNAs with respect to the reference set of splice variants. The evaluation of CAFTAN showed that is able to correctly classify more than 85% of 5950 selected "known protein-coding" VEGA cDNAs as high quality multi- or single-exon. It identified as good 80.6 % of the single exon cDNAs and 85 % of the multiple exon cDNAs. The program is written in Perl and in a modular way, allowing the adoption of this strategy to other tasks like EST-annotation, or to extend it by adding new classification rules and new organism databases as they become available. We think that it is a very useful program for the annotation and research of unfinished genomes. CONCLUSION: CAFTAN is a high-throughput sequence analysis tool, which performs a fast and reliable quality prediction of cDNAs. Several thousands of cDNAs can be analyzed in a short time, giving the curator/scientist a first quick overview about the quality and the already existing annotation of a set of cDNAs. It supports the rejection of low quality cDNAs and helps in the selection of likely novel splice variants, and/or completely novel transcripts for new experiments.

Chromosome Mapping↗

Comparison of in-gel and on-membrane digestion methods at low to sub-pmol level for subsequent peptide and fragment-ion mass analysis using matrix-assisted laser-desorption/ionization mass spectrometry.

The success of the mass spectrometric-based approaches for the identification of gel-separated proteins relies upon recovery of peptides, without high levels of ionization-suppressing contaminants, in solvents compatible with the mass spectrometer being employed. We sought to determine whether in-gel or on-membrane digestion provided a significant advantage when low to sub-pmol quantities of gel-separated proteins were analyzed by matrix-assisted laser-desorption/ionization mass spectrometry (MALDI-MS) with respect to the number and size of released peptides. Serial dilutions of five standard proteins of M(r) 17,000 to 97,000 (from 16 pmol to 125 fmol) were electrophoresed and subjected to in-gel digestion (using a microcolumn clean-up protocol, Courchesne, P.L. and Patterson, S. D., BioTechniques, 1997, in press) or on-membrane digestion following blotting to the PVDF-based membranes, Immobilon-P and Immobilon-CD. Peptide maps were able to be obtained for all proteins at the detection limit of each method (Immobilon-P and Immobilon-CD, 0.5 pmol; and in-gel, 125 fmol), and searches of Swiss-Prot or a non-redundant database (> 193000 entries) successfully identified all of the proteins, except beta-casein. Fragment-ion spectra using a curved-field reflector MALDI-MS were obtained from more than one peptide per protein at loads down to 250 fmol (except beta-casein). Using the uninterpreted data, a search of the nonredundant database and a six-way translation of GenBank dbEST (> 2,208,000 entries total) was able to identify myoglobin, carbonic anhydrase II, and phosphorylase b.

Acrylic Resins↗

The bovine fatty acid binding protein 4 gene is significantly associated with marbling and subcutaneous fat depth in Wagyu x Limousin F2 crosses.

Fatty acid binding protein 4 (FABP4), which is expressed in adipose tissue, interacts with peroxisome proliferator-activated receptors and binds to hormone-sensitive lipase and therefore, plays an important role in lipid metabolism and homeostasis in adipocytes. The objective of this study was to investigate associations of the bovine FABP4 gene with fat deposition. Both cDNA and genomic DNA sequences of the bovine gene were retrieved from the public databases and aligned to determine its genomic organization. Primers targeting two regions of the FABP4 gene were designed: from nucleotides 5433-6106 and from nucleotides 7417-7868 (AAFC01136716). Direct sequencing of polymerase chain reaction (PCR) products on two DNA pools from high- and low-marbling animals revealed two single nucleotide polymorphisms (SNPs): AAFC01136716.1:g.7516G>C and g.7713G>C. The former SNP, detected by PCR-restriction fragment length polymorphism using restriction enzyme MspA1I, was genotyped on 246 F2 animals in a Waygu x Limousin F2 reference population. Statistical analysis showed that the FABP4 genotype significantly affected marbling score (P = 0.0398) and subcutaneous fat depth (P = 0.0246). The FABP4 gene falls into a suggestive/significant quantitative trait loci interval for beef marbling that was previously reported on bovine chromosome 14 in three other populations.

Animals↗

Molecular characterization of a novel pattern recognition protein from nonspecific cytotoxic cells: sequence analysis, phylogenetic comparisons and anti-microbial activity of a recombinant homologue.

Nonspecific cytotoxic cells (NCC) are the first identified and most extensively studied killer cell population in teleosts. NCC kill a wide variety of target cells including tumor cells, virally transformed cells and protozoan parasites. The present study identified a novel evolutionarily conserved oligodeoxynucleotide (ODN) binding membrane protein expressed by channel catfish (Ictalurus punctatus) NCC. Peptide fingerprinting analysis of the ODN binding protein (referred to as NCC cationic anti-microbial protein-1/ncamp-1) identified a peptide that was used to design degenerate primers. A catfish NCC cDNA library was used as template with these primers and the PCR-amplified product was sequenced. The translated sequence contained 203 amino acids (molecular mass of 22,064.63 Da) with characteristic lysine rich regions and a pI=pH 10.75. Sequence comparisons of this protein indicated similarity to zebrafish (51.2%) histone family member 1-X and (to a lesser extent) to trout H1. A search of EST databases confirmed that ncamp-1 is also expressed in various tissues of channel catfish as well as zebrafish. Inspection for signature repeats in ncamp-1 and comparisons with histone-like peptides from different species indicated the presence of multiple lysine based motifs composed of AKKA or PKK repeats. The novel protein was cloned, expressed in E. coli and the recombinant was used to generate rabbit anti-serum. The recombinant ncamp-1 bound GpC and CpG ODNs and was detected with homologous anti-ncamp-1 polyclonal antibodies. Western blots of NCC membranes using anti-ncamp-1 serum detected a 29 kDa protein. Binding competition experiments demonstrated that anti-ncamp-1 antibodies and GpC bound to the same protein on NCC. Two different truncated forms of ncamp-1 as well as the full-length recombinant protein exhibited anti-microbial activity. The present study demonstrated the expression by NCC of a new membrane protein that may participate in the recognition of bacterial DNA and as such participate in innate anti-microbial immune responses in teleosts.

Amino Acid Motifs↗

Cloning and characterization of the Caenorhabditis elegans CeCRMP/DHP-1 and -2; common ancestors of CRMP and dihydropyrimidinase?

The vertebrate CRMP (collapsin-response-mediator protein) gene family comprises at least four members. These CRMPs exhibit about 60% amino acid identity with vertebrate dihydropyrimidinase (DHP), an amidohydrolase involved in the pyrimidine degradation pathway. CRMP is also referred to as DRP (DHP-related protein), TOAD-64 (turned on after division, 64 kDa) and Ulip (Unc-33-like phosphoprotein). These vertebrate CRMPs are expressed mainly in early neuronal differentiation, which suggests that they play a role in neuronal development. In this study we isolated two cDNA clones from nematode C. elegans based on their sequence homology to vertebrate CRMPs and DHP. These two molecules, termed CeCRMP/DHP-1 and -2, turned out to be Ulip-B and -A, respectively, which were previously identified in the C. elegans genomic database by Byk et al. (1998). These newly isolated molecules were believed to represent a common ancestral state before the gene duplication between CRMPs and DHP. CeCRMP/DHP-1 and -2 protein retained all putative zinc-binding residues thought to be essential for the amidohydrolase activity of DHP and exhibited a weak amidohydrolase activity when 5-bromo-dihydrouracil was used as a substrate. Whole-mount in situ hybridization and expression analysis using GFP fusions revealed that CeCRMP/DHP-1 was transiently expressed in the hypodermis of C. elegans during the early larva stage. CeCRMP/DHP-1 was also expressed in a single nerve cell between the pharynx and ring neuropil. On the other hand, expression of CeCRMP/DHP-2 was observed in the body wall muscle throughout the lifespan of C. elegans. These results indicate that a major site of CeCRMP/DHP-1 and -2 expression is non-neuronal. Targeted gene disruption of CeCRMP/DHP-2 caused no particular difference in appearance or movement phenotype.

Amidohydrolases↗

Norwalk virus nonstructural protein p48 forms a complex with the SNARE regulator VAP-A and prevents cell surface expression of vesicular stomatitis virus G protein.

Norwalk virus (NV), a reference strain of human calicivirus in the Norovirus genus of the family Caliciviridae, contains a positive-strand RNA genome with three open reading frames. ORF1 encodes a 1,789-amino-acid polyprotein that is processed into nonstructural proteins that include an NTPase, VPg, protease, and RNA-dependent RNA polymerase. The N-terminal protein p48 of ORF1 shows no significant sequence similarity to viral or cellular proteins, and its function in the human calicivirus replication cycle is not known. The lack of sequence similarity to any protein in the public databases suggested that p48 may have a unique function in the NV replication cycle or, alternatively, may perform a characterized function in replication by a unique mechanism. In this report, it is shown that p48 displays a vesicular localization pattern in transfected cells when fused to the fluorescent reporter EYFP. A predicted transmembrane domain at the C terminus of p48 was not necessary for the observed localization pattern, but this domain was sufficient to redirect localization of EYFP to a fluorescent pattern consistent with the Golgi apparatus. A yeast two-hybrid screen identified the SNARE regulator vesicle-associated membrane protein-associated protein A (VAP-A) as a binding partner of p48. Biochemical assays confirmed that p48 and VAP-A interact and form a stable complex in mammalian cells. Furthermore, expression of the vesicular stomatitis virus G glcyoprotein on the cell surface was inhibited when cells coexpressed p48, suggesting that p48 disrupts intracellular protein trafficking.

Animals↗

System, trends and perspectives of proteomics in dicot plants Part II: Proteomes of the complex developmental stages.

This review is devoted to the proteomes of the complex developmental stages of dicotyledoneous (dicot) plant materials. The two core technologies, two-dimensional gel electrophoresis (2-DGE) and mass spectrometry (MS), independently or in combination with each other, are propelling dicot plant proteomics to new discoveries and functions, with the establishment of tissue-specific and organelle proteomes, mostly in Arabidopsis thaliana and Medicago truncatula, revealing their complexity and specificity. These experimental proteomes have provided a good start towards the establishment of high-density 2-DGE reference maps and peptide mass fingerprint databases, for not only the model dicot plants, A. thaliana and M. truncatula, but also other important dicot plants, which will serve as a basis for proteomes of many other dicot plants and plant materials.

Arabidopsis↗

Hierarchical clustering algorithm for comprehensive orthologous-domain classification in multiple genomes.

Ortholog identification is a crucial first step in comparative genomics. Here, we present a rapid method of ortholog grouping which is effective enough to allow the comparison of many genomes simultaneously. The method takes as input all-against-all similarity data and classifies genes based on the traditional hierarchical clustering algorithm UPGMA. In the course of clustering, the method detects domain fusion or fission events, and splits clusters into domains if required. The subsequent procedure splits the resulting trees such that intra-species paralogous genes are divided into different groups so as to create plausible orthologous groups. As a result, the procedure can split genes into the domains minimally required for ortholog grouping. The procedure, named DomClust, was tested using the COG database as a reference. When comparing several clustering algorithms combined with the conventional bidirectional best-hit (BBH) criterion, we found that our method generally showed better agreement with the COG classification. By comparing the clustering results generated from datasets of different releases, we also found that our method showed relatively good stability in comparison to the BBH-based methods.

Algorithms↗

PALS db: Putative Alternative Splicing database.

PALS db is a collection of Putative Alternative Splicing information from 19 936 human UniGene clusters and 16 615 mouse UniGene clusters. Alternative splicing (AS) sites were predicted by using the longest messenger RNA (mRNA) sequence in each UniGene cluster as the reference sequence. This sequence was aligned with related sequences in UniGene and dbEST to reveal the AS. This information was presented with six features: (i) literature aliases were used to improve the result of a gene name search; (ii) the quality of a prediction can be easily judged from the color-coded similarity and the scaled length of an alignment; (iii) we have clustered those EST sequences that support the same AS site together to enhance the users' confidence on a prediction; (iv) the users can also set up the alignment criteria interactively to recover false negatives; (v) tissue distribution can be displayed by placing the mouse cursor over an alignment; (vi) gene features will be analyzed at foreign sites by submitting the selected mRNA or its encoded protein as a query. Using these features, the users cannot only discover putative AS sites in silico, but also make new observations by combining AS information with tissue distributions or with gene features. PALS db is available at http://palsdb.ym.edu.tw/.

Alternative Splicing↗

In silico-initiated cloning and molecular characterization of a novel human member of the L1 gene family of neural cell adhesion molecules.

To discover genes contributing to mental retardation in 3p- syndrome patients we have used in silico searches for neural genes in NCBI databases (dbEST and Uni-Gene). An EST with strong homology to the rat CAM L1 gene subsequently mapped to 3p26 was used to isolate a full-length cDNA. Molecular analysis of this cDNA, referred to as CALL (cell adhesion L1-like), showed that it is encoded by a chromosome 3p26 locus and is a novel member of the L1 gene family of neural cell adhesion molecules. Multiple lines of evidence suggest CALL is likely the human ortholog of the murine gene CHL1: it is 84% identical on the protein level, has the same domain structure, same membrane topology, and a similar expression pattern. The orthology of CALL and CHL1 was confirmed by phylogenetic analysis. By in situ hybridization, CALL is shown to be expressed regionally in a timely fashion in the central nervous system, spinal cord, and peripheral nervous system during rat development. Northern analysis and EST representation reveal that it is expressed in the brain and also outside the nervous system in some adult human tissues and tumor cell lines. The cytoplasmic domain of CALL is conserved among other members of the L1 subfamily and features sequence motifs that may involve CALL in signal transduction pathways.

Amino Acid Sequence↗

Analysis of relative positions of ribonucleotide bases in a crystal structure of ribosome.

Relative positions of bases to bases in a crystal structure of ribosome were analyzed extensively. It was found that there is no clear relation between bases apart more than 15 A and, thus, the relative location of bases can be analyzed within 15 A of the reference bases. As for base pairing, major positioning was found to be due to the Watson-Crick type base pairs. Some other positions corresponding to non-Watson-Crick type base pairs were also found in some extents. As for base-base stacking, it was observed that the bases stacked to adenine base are dispersive. It was found that less non-Watson-Crick base pairs was found close to the protein binding site, suggesting that the protein components have a tendency to bind to the regular stem structures. The database of relative location of bases must be useful for improvement of structural determination and structural modeling systems.

Algorithms↗

Coxiella burnetii whole cell lysate protein identification by mass spectrometry and tandem mass spectrometry.

The whole cell lysate of Coxiella burnetii strain RSA 493 was separated by two-dimensional electrophoresis and more than 500 protein spots were found on silver-stained reference map. Spots from the gels were subjected to identification based on peptide mass fingerprinting (PMF). In order to identify additional proteins, tandem mass spectrometry (MS/MS) using electrospray and matrix-assisted laser desorption/ionization techniques was applied. The three independent approaches resulted in the identification of 197 open reading frames (ORFs). Fifty-two proteins were identified by PMF and at least with one of the MS/MS methods, 37 proteins with both MS/MS instruments, and 19 proteins with all three techniques applied. All predicted C. burnetii ORFs were compared with the Clusters of Orthologous Groups database. The data related to identified proteins were stored and indexed in a file that can be read and searched using Microsoft Access.

Bacterial Proteins↗

A current genotoxicity database for heterocyclic thermic food mutagens. I. Genetically relevant endpoints.

Cooking, heat processing, or pyrolysis of protein-rich foods induce the formation of a series of structurally related heterocyclic aromatic bases that have been found to be mutagens. The primary genetic assay utilized to detect and isolate these mutagens has been the his reversion assay in Salmonella typhimurium. The classification and nomenclature of these chemicals is revised to reflect recent advances. The findings of short-term tests for genetic injury that have been applied to these agents are presented in a systematic way. Cell-free, bacterial, mammalian cell culture, and in vivo systems are included. Major results, the mutagens tested, and key references are presented in tabular form, with text commentary. Integrated conclusions on the state of current knowledge of the genetic toxicity of thermic food mutagens are presented. Areas in need of further research are defined. Finally, an outline is presented of a suggested path leading to the determination whether normal methods of food preparation and processing constitute a human health hazard.

Animals↗

Applied proteomics: mitochondrial proteins and effect on function.

The identification of a majority of the polypeptides in mitochondria would be invaluable because they play crucial and diverse roles in many cellular processes and diseases. The endogenous production of reactive oxygen species (ROS) is a major limiter of life as illustrated by studies in which the transgenic overexpression in invertebrates of catalytic antioxidant enzymes results in increased lifespans. Mitochondria have received considerable attention as a principal source---and target---of ROS. Mitochondrial oxidative stress has been implicated in heart disease including myocardial preconditioning, ischemia/reperfusion, and other pathologies. In addition, oxidative stress in the mitochondria is associated with the pathogenesis of Alzheimer's disease, Parkinson's disease, prion diseases, and amyotrophic lateral sclerosis (ALS) as well as aging itself. The rapidly emerging field of proteomics can provide powerful strategies for the characterization of mitochondrial proteins. Current approaches to mitochondrial proteomics include the creation of detailed catalogues of the protein components in a single sample or the identification of differentially expressed proteins in diseased or physiologically altered samples versus a reference control. It is clear that for any proteomics approach prefractionation of complex protein mixtures is essential to facilitate the identification of low-abundance proteins because the dynamic range of protein abundance within cells has been estimated to be as high as 10(7). The opportunities for identification of proteins directly involved in diseases associated with or caused by mitochondrial dysfunction are compelling. Future efforts will focus on linking genomic array information to actual protein levels in mitochondria.

Animals↗

New insights into the rat spermatogonial proteome: identification of 156 additional proteins.

Despite the essential role played by spermatogonia in testicular function, little is known about these cells. To improve our understanding of their biology, our group recently identified a set of 53 spermatogonial proteins using two-dimensional (2-D) gel electrophoresis and mass spectrometry. To continue this work, we investigated a subset of the spermatogonial proteome using narrow range immobilized pH gradients to favor the detection of less abundant proteins. A 2-D reference map of spermatogonia in the pH range 4-9 was created, and protein entities fractionated in a pH 5-6 2-D gel were further processed for protein identification. A new set of 156 polypeptides was identified by peptide mass fingerprinting and tandem mass spectrometry. These polypeptides corresponded to 102 different proteins, which reflect the complexity of post-translational modifications. Seventy-nine of these proteins were identified for the first time in spermatogonia. All identified proteins were classified into functional groups. This work represents a first step toward the establishment of a systematic spermatogonia protein database.

Animals↗

Comparison of knowledge-based and distance geometry approaches for generation of molecular conformations.

A knowledge-based approach for generating conformations of molecules has been developed. The method described here provides a good sampling of the molecule's conformational space by restricting the generated conformations to those consistent with the reference database. The present approach, internally named et for enumerate torsions, differs from previous database-mining approaches by employing a library of much larger substructures while treating open chains, rings, and combinations of chains and rings in the same manner. In addition to knowledge in the form of observed torsion angles, some knowledge from the medicinal chemist is captured in the form of which substructures are identified. The knowledge-based approach is compared to Blaney et al.'s distance geometry (DG) algorithm for sampling the conformational space of molecules. The structures of 113 protein-bound molecules, determined by X-ray crystallography, were used to compare the methods. The present knowledge-based approach (i) generates conformations closer to the experimentally determined conformation, (ii) generates them sooner, and (iii) is significantly faster than the DG method.

Algorithms↗