Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

A large family of endosome-localized proteins related to sorting nexin 1.

Sorting nexin 1 (SNX1), a peripheral membrane protein, has previously been shown to regulate the cell-surface expression of the human epidermal growth factor receptor [Kurten, Cadena and Gill (1996) Science 272, 1008-1010]. Searches of human expressed sequence tag databases with SNX1 revealed eleven related human cDNA sequences, termed SNX2 to SNX12, eight of them novel. Analysis of SNX1-related sequences in the Saccharomyces cerevisiae genome clearly shows a greatly expanded SNX family in humans in comparison with yeast. On the basis of the predicted protein sequences, all members of this family of hydrophilic molecules contain a conserved 70-110-residue Phox homology (PX) domain, referred to as the SNX-PX domain. Within the SNX family, subgroups were identified on the basis of the sequence similarities of the SNX-PX domain and the overall domain structure of each protein. The members of one subgroup, which includes human SNX1, SNX2, SNX4, SNX5 and SNX6 and the yeast Vps5p and YJL036W, all contain coiled-coil regions within their large C-terminal domains and are found distributed in both membrane and cytosolic fractions, typical of hydrophilic peripheral membrane proteins. Localization of the human SNX1 subgroup members in HeLa cells transfected with the full-length cDNA species revealed a similar intracellular distribution that in all cases overlapped substantially with the early endosome marker, early endosome autoantigen 1. The intracellular localization of deletion mutants and fusions with green fluorescent protein showed that the C-terminal regions of SNX1 and SNX5 are responsible for their endosomal localization. On the basis of these results, the functions of these SNX molecules are likely to be unique to endosomes, mediated in part by interactions with SNX-specific C-terminal sequences and membrane-associated determinants.

Amino Acid Sequence↗

A systematic review of the influence of different titanium surfaces on proliferation, differentiation and protein synthesis of osteoblast-like MG63 cells.

OBJECTIVES: Titanium is the standard material for dental and orthopaedical implants. The good biocompatibility has been proven in many experimental and clinical investigations. Different titanium topographies were tested in vitro using different cell culture models. The aim of this systematic review was to evaluate and summarize the medical/dental literature to assess on which kind of titanium surface structure the osteoblast-like osteosarcoma cells MG63 show the best proliferation and differentiation rate, and the best protein synthesis. METHODS: A systematic search was carried out using different on-line databases (PubMed, Web of Science, Cochrane Library, International Poster Journal), supplemented by handsearch in selected journals and by examination of the bibliographies of the identified articles. Inclusion and exclusion criterias were applied when considering relevant articles. Studies which met the inclusion criteria were included and data extraction was undertaken by one reviewer. RESULTS: The search yielded 348 references. Nine articles referring to nine different studies were relevant to our question. Additionally 8 less relevant articles were identified. It was found that regularly textured surfaces of pure titanium with R(a) values (average roughness) of around 4 mum are well-accepted by MG63 cells. CONCLUSIONS: The surfaces and culture conditions vary widely. Therefore it is still difficult to recommend one particular surface. It seems that there are no differences in cell proliferation and differentiation on surfaces treated by blasting and etching. Standardization in fabrication and size of the different test surfaces as well as homogeneity in culture times and plating densities should be aspects for future research.

Algorithms↗

Comparative proteome analysis of breast cancer and normal breast.

Breast cancer is a leading cause of death for women. The underlying molecular mechanism is still not well understood. In this study, two-dimensional gel electrophoresis combined with mass spectrometry was used to analyze changes in the proteome of infiltrating ductal carcinoma compared to normal breast tissue. Ten sets of two-dimensional gels per experimental condition were analyzed and more than 500 spots each were detected. This revealed 39 spots for which expression in breast cancer cells were reproducibly altered more than twofold compared to normal controls (p < 0.01). These spots represented 25 different proteins after identification using the database search after mass spectrometry, comprising cell defense proteins, enzymes involved in glycolytic energy metabolism and homeostasis, protein folding and structural proteins, proteins involved in cytoskeleton and cell motility, and proteins involved in other functions. In addition, 28 nondifferentially expressed proteins with different functions were also mapped and identified, which might help to establish a two-dimensional gel electrophoresis reference map of human breast cancer. Our study shows that proteomics offers a powerful methodology to detect the proteins that show different expression patterns in breast cancer tissue and may provide an accurate molecular classification. The differentially expressed proteins may be used as potential candidate markers for diagnostic purposes or for determination of tumor sensitivity to therapy. The functional implications of the identified proteins are discussed.

Biomarkers, Tumor↗

Comparative proteomic analysis of mammalian animal tissues and body fluids: bovine proteome database.

Characterizing the complete proteome of multicellular organisms is a challenging task using the currently available technologies. With the increasing degree of genetic complexity, animals acquire a broader repertoire of options to meet environmental challenges. Mammalian cells from different tissues/body fluids express different thousands of proteins with a predicted dynamic range of up to five to six orders of magnitude, thus necessitating the whole arsenal of dedicated analytical strategies for a detailed proteome characterization. Nevertheless, 2D-E analysis of whole cellular lysates still remains the most used initial approach for the proteomic description of specialized cells. It enables to obtain an overview of the main soluble protein components of a specific tissue/body fluid, allowing comparison between different cellular types and molecular description of organ specialization. Massive proteomic investigations have been reported mainly in the case of human, mouse and rat, allowing comparative analysis. For this reason, a research project focused on the 2D-E characterization of tissues and biological fluids from other domestic mammals has been undertaken in our laboratory. A number of high-resolution reference electrophoretic maps have been established for liver, kidney, muscle, plasma and red blood cells samples from Holstein Friesian bovine female individuals. Among the 1863 distinct protein features detected, 534 species were identified and associated to 209 different genes by a combination of MALDI-TOF mass fingerprint, capillary LC-ESI-IT-MS-MS and image gel matching procedures. Identified polypeptide species and differences in expression profiles between various tissues/fluids clearly reflected organ biochemical specialization. This experimental output allowed establishing a 2D-E bovine database accessible at the URL address for image comparison.

Animals↗

Gene annotation from scientific literature using mappings between keyword systems.

MOTIVATION: The description of genes in databases by keywords helps the non-specialist to quickly grasp the properties of a gene and increases the efficiency of computational tools that are applied to gene data (e.g. searching a gene database for sequences related to a particular biological process). However, the association of keywords to genes or protein sequences is a difficult process that ultimately implies examination of the literature related to a gene. RESULTS: To support this task, we present a procedure to derive keywords from the set of scientific abstracts related to a gene. Our system is based on the automated extraction of mappings between related terms from different databases using a model of fuzzy associations that can be applied with all generality to any pair of linked databases. We tested the system by annotating genes of the SWISS-PROT database with keywords derived from the abstracts linked to their entries (stored in the MEDLINE database of scientific references). The performance of the annotation procedure was much better for SWISS-PROT keywords (recall of 47%, precision of 68%) than for Gene Ontology terms (recall of 8%, precision of 67%). AVAILABILITY: The algorithm can be publicly accessed and used for the annotation of sequences through a web server at http://www.bork.embl.de/kat

Abstracting and Indexing↗

COACH: profile-profile alignment of protein families using hidden Markov models.

MOTIVATION: Alignments of two multiple-sequence alignments, or statistical models of such alignments (profiles), have important applications in computational biology. The increased amount of information in a profile versus a single sequence can lead to more accurate alignments and more sensitive homolog detection in database searches. Several profile-profile alignment methods have been proposed and have been shown to improve sensitivity and alignment quality compared with sequence-sequence methods (such as BLAST) and profile-sequence methods (e.g. PSI-BLAST). Here we present a new approach to profile-profile alignment we call Comparison of Alignments by Constructing Hidden Markov Models (HMMs) (COACH). COACH aligns two multiple sequence alignments by constructing a profile HMM from one alignment and aligning the other to that HMM. RESULTS: We compare the alignment accuracy of COACH with two recently published methods: Yona and Levitt's prof_sim and Sadreyev and Grishin's COMPASS. On two sets of reference alignments selected from the FSSP database, we find that COACH is able, on average, to produce alignments giving the best coverage or the fewest errors, depending on the chosen parameter settings. AVAILABILITY: COACH is freely available from www.drive5.com/lobster

Algorithms↗

Comparison of six microcomputer dietary analysis systems with the USDA Nutrient Data Base for Standard Reference.

We compared the general operating features and nutrient databases of six microcomputer dietary analysis systems. A 3-day food record with 73 food items was entered into each program; nutrient averages were compared with the US Department of Agriculture Nutrient Data Base for Standard Reference (USDA NDB), full version, release 9, for microcomputers. The six programs were found to vary widely in cost, number of foods and nutrients in the database, use of non-USDA data and imputation of data for missing values, number of print/export options, time to analyze the 3-day food record, and overall ease of use. Although all of the microcomputer dietary analysis systems were within 7% of the USDA NDB for energy, protein, total fat, and total carbohydrates, the proportion of other nutrients varying more than 15% from the USDA NDB varied considerably between programs. Variance among programs for 3-day food record nutrient values occurred because of differences in the number of food items included in the database (leading to varying degrees of substitution), the recency of the nutrient data (whether or not the most recent USDA releases had been incorporated), and the number of missing values (the degree to which non-USDA sources or estimated calculations were used to fill in the blanks from the USDA standard). Our results demonstrate that it is important for each dietitian to carefully choose a microcomputer dietary analysis system that is suitable to specific and predetermined needs.

Databases, Factual↗

Comparison of eight microcomputer dietary analysis programs with the USDA Nutrient Data Base for Standard Reference.

OBJECTIVE: To compare the general operating features and nutrient databases of eight microcomputer dietary analysis programs. DESIGN: A 3-day food record with 73 food items was entered into each program by the authors. The general operating features of the program were summarized and evaluated. The nutrient database was evaluated by comparing the nutrient analysis output with the 1993 US Department of Agriculture (USDA) Nutrient Data Base for Standard Reference (NDB), full version, release 10, for microcomputers. RESULTS: The programs varied in cost, number of foods and nutrients in the database, use of non-USDA data, and inputting of data for missing values. We also found differences in the quality of user manuals and help screens, ease of food entry and averaging of 3-day nutrient intake, speed of analyzing and printing results, quality and number of print/export options, and overall ease of learning and using the program. All but one of the programs were within 15% of the USDA NDB for energy, protein, total fat, and total carbohydrates. However, there was some difference in the number of other nutrients and food components varying more than 15% from the USDA NDB. These differences occurred because of variations in the number of food items included in each programs' database and the number of missing nutrient values in the database. APPLICATIONS: Our results demonstrate the importance of carefully choosing a microcomputer dietary analysis program that is suitable to the user's specific and predetermined needs.

Databases, Factual↗

Toward the quantitative prediction of T-cell epitopes: QSAR studies on peptides having affinity with the class I MHC molecular HLA-A*0201.

It would be useful for vaccine development to develop a method of rapidly identifying peptide epitopes. In this paper, the empirical three-dimensional quantitative structure-affinity relationship (3D-QSAR) methods were used to study the relationship between the three dimensional structural parameters (the isotropic surface area, ISA, and the electronic charge index, ECI) of the HLA-A*0201 binding peptide and the HLA-A*0201/peptide binding affinities. A set of 102 peptides having affinity with the class I MHC HLA-A*0201 molecule was used as training set. A test set of 40 peptides was used to determine the predictive value of the models. The 3D-QSAR models yielded a q2 = 0.5724 and a high rpred2 = 0.6955. The standard regression coefficients indicated that the hydrophobic interactions played an important role in peptide-MHC molecule binding and predicted the specific amino acid residue essential at a certain position of the peptide. The approach tested in the current paper is highly complementary to many of the methods described in references and possesses good predictability. It is a rapid and convenient method to detect high affinity peptide epitopes.

Amino Acid Sequence↗

Plasma protein map: an update by microsequencing.

The reference plasma protein map, obtained with immobilized pH gradients in the first dimension of two-dimensional electrophoresis, is presented. By microsequencing, more than 40 polypeptide chains were identified. The new polypeptides and previously known proteins are listed in a table and labeled on the protein map, thus providing an update of the human plasma two-dimensional gel database.

Amino Acid Sequence↗

A robust approach for the analysis of peptides in the low femtomole range by capillary electrophoresis-tandem mass spectrometry.

A capillary electrophoresis-tandem mass spectrometry (CE-MS/MS) approach has been developed for routine application in proteomic studies. Robustness of the coupling is achieved by using a standard coaxial sheath-flow sprayer. Thereby, greater stability than nanoelectrospray ionization-mass spectrometry coupling of sheathless capillary electrophoresis or nanoliquid chromatography (nano-LC) is achieved, resulting in stable operation for several weeks and unattended overnight sequences. The applied sheath flow is reduced to 1-2 microL/min in order to increase sensitivity. Standard peptides and those of digests of standard proteins and gel-separated proteins can be detected in the low femtomole range (full scan and MS/MS). Detection limits are found to be as low as 500 amol. Low femtomole amounts are required for unequivocal identification by MS/MS experiments in the ion trap and subsequent database search. By applying a simple pH-mediated stacking the concentration sensitivity can be lowered to some tens of fmol/microL (nM), depending on capillary size. This sensitivity is close to published values for sheathless CE-MS and nano-LC-MS, respectively (a comparison to reference values is presented). Moreover, with capillaries of about 50 cm in length separations in less than 10 min are possible resulting in a throughput of up to four analyses per hour. This is a factor of 4-12 times faster than nano-LC separation, being the state-of-the-art techniques for proteomic studies.

Animals↗

NCBI Reference Sequence project: update and current status.

The goal of the NCBI Reference Sequence (RefSeq) project is to provide the single best non-redundant and comprehensive collection of naturally occurring biological molecules, representing the central dogma. Nucleotide and protein sequences are explicitly linked on a residue-by-residue basis in this collection. Ideally all molecule types will be available for each well-studied organism, but the initial database collection pragmatically includes only those molecules and organisms that are most readily identified. Thus different amounts of information are available for different organisms at any given time. Furthermore, for some organisms additional intermediate records are provided when the genome sequence is not yet finished. The collection is supplied by NCBI through three distinct pipelines in addition to collaborations with community groups. The collection is curated on an ongoing basis. Additional information about the NCBI RefSeq project is available at http://www.ncbi.nih.gov/RefSeq/.

Alternative Splicing↗

High throughput processing of the structural information in the protein data bank.

The protein data bank (PDB) is the largest, most comprehensive, freely available depository of protein structural information, containing more than 37,500 deposited structures. On one hand, the form and the organization of the PDB seems to be perfectly adequate for gathering information from specific protein structures, by using the bibliographic references and the informative remark fields. On the other hand, however, it seems to be impossible to automatically review remark fields and journal references for processing hundreds or thousands of PDB files. We present here a family of combinatorial algorithms to solve some of these problems. Our algorithms are capable to automatically analyze PDB structural information, identify missing atoms, repair chain ID information, and most importantly, the algorithms are capable of identifying ligands with their respective binding sites.

Algorithms↗

Flexsim-R: a virtual affinity fingerprint descriptor to calculate similarities of functional groups.

Methods to describe the similarity of fragments occurring in drug-like molecules are of fundamental importance in computational drug design. In the early phase of lead discovery, they can help to select diverse building blocks for combinatorial compound libraries intended for broad screening. In lead optimization, such methods can guide bioisosteric replacements of one functional group by another or serve as descriptors for QSAR calculations. In this paper, we outline the development of a novel 3D descriptor, termed Flexsim-R, which is a further extension of our virtual affinity fingerprint idea. Descriptors are calculated based on docking of small fragments such as building blocks for combinatorial chemistry or functional groups of drug-like molecules into a reference panel of protein binding sites. The method is validated by examining the neighborhood behavior of the affinity fingerprints and by deriving predictive QSAR models for a couple of literature peptide data sets.

Algorithms↗

Electrostatics of ligand binding: parametrization of the generalized Born model and comparison with the Poisson-Boltzmann approach.

An accurate and fast evaluation of the electrostatics in ligand-protein interactions is crucial for computer-aided drug design. The pairwise generalized Born (GB) model, a fast analytical method originally developed for studying the solvation of organic molecules, has been widely applied to macromolecular systems, including ligand-protein complexes. However, this model involves several empirical scaling parameters, which have been optimized for the solvation of organic molecules, peptides, and nucleic acids but not for energetics of ligand binding. Studies have shown that a good solvation energy does not guarantee a correct model of solvent-mediated interactions. Thus, in this study, we have used the Poisson-Boltzmann (PB) approach as a reference to optimize the GB model for studies of ligand-protein interactions. Specifically, we have employed the pairwise descreening approximation proposed by Hawkins et al.(1) for GB calculations and DelPhi for PB calculations. The AMBER all-atom force field parameters have been used in this work. Seventeen protein-ligand complexes have been used as a training database, and a set of atomic descreening parameters has been selected with which the pairwise GB model and the PB model yield comparable results on atomic Born radii, the electrostatic component of free energies of ligand binding, and desolvation energies of the ligands and proteins. The energetics of the 15 test complexes calculated with the GB model using this set of parameters also agrees well with the energetics calculated with the PB method. This is the first time that the GB model has been parametrized and thoroughly compared with the PB model for the electrostatics of ligand binding.

Ligands↗

Human Plasma PeptideAtlas.

Peptide identifications of high probability from 28 LC-MS/MS human serum and plasma experiments from eight different laboratories, carried out in the context of the HUPO Plasma Proteome Project, were combined and mapped to the EnsEMBL human genome. The 6929 distinct observed peptides were mapped to approximately 960 different proteins. The resulting compendium of peptides and their associated samples, proteins, and genes is made publicly available as a reference for future research on human plasma.

Blood Proteins↗

Evaluation of protein fold comparison servers.

When a new protein structure has been determined, comparison with the database of known structures enables classification of its fold as new or belonging to a known class of proteins. This in turn may provide clues about the function of the protein. A large number of fold comparison programs have been developed, but they have never been subjected to a comprehensive and critical comparative analysis. Here we describe an evaluation of 11 publicly available, Web-based servers for automatic fold comparison. Both their functionality (e.g., user interface, presentation, and annotation of results) and their performance (i.e., how well established structural similarities are recognized) were assessed. The servers were subjected to a battery of performance tests covering a broad spectrum of folds as well as special cases, such as multidomain proteins, Calpha-only models, new folds, and NMR-based models. The CATH structural classification system was used as a reference. These tests revealed the strong and weak sides of each server. On the whole, CE, DALI, MATRAS, and VAST showed the best performance, but none of the servers achieved a 100% success rate. Where no structurally similar proteins are found by any individual server, it is recommended to try one or two other servers before any conclusions concerning the novelty of a fold are put on paper.

Computational Biology↗