Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Knowledge-based chemoinformatic approaches to drug discovery.

The modern drug discovery process is steadily becoming more information driven. Structural, physicochemical and ADME-Tox property profiles of reference (successful) ligands, along with structural information of their target proteins, have been extremely useful for early-stage drug discovery. Recently, databases of known biologically active ligands (knowledge bases) have become more focused toward different protein-target classes. The number of new chemoinformatics tools used to analyze structures and properties of successful molecules has also increased enormously. Scientists in this area are exploring new physicochemical properties and appropriate drug sets to understand druglike properties. In this review, the various uses of the ligand knowledge bases in the drug discovery process have been critically reviewed.

Drug Design↗

[Cloning and sequence analysis of a gene encoding amastin from Leishmania major].

OBJECTIVE: To clone a gene encoding surface protein from Leishmania major. METHODS: Using T. cruzi amastin DNA sequence as a reference, computer search was done on GenBank and dbEST databases by using BLAST path. A Leishmania major DNA library has been constructed and screened by in situ colony hybridization. RESULTS: A 309nt DNA fragment from Leishmania major was found in dbEST. Leishmania major DNA library was screened using specific primers synthesized according to 309 nt DNA sequence, and a full-length coding sequence for Leishmania major amastin was cloned. The coding sequence consisted of 552 nt, and translated into 183 amino acid residues. The homology is 23.5% at amino acid sequence level between Leishmania major and T. cruzi amastins. CONCLUSION: A full length amastin coding gene for Leishmania major has been cloned.

Amino Acid Sequence↗

Confident protein identification using the average peptide score method coupled with search-specific, ab initio thresholds.

Perhaps the greatest difficulty in interpreting large sets of protein identifications derived from mass spectrometric methods is whether or not to trust the results. For such experiments, the level of confidence in each protein identification made needs to be far greater than the often used 95% significance threshold to avoid the identification of many false-positives. To provide higher confidence results, we have developed an innovative scoring strategy coupling the recently published Average Peptide Score (APS) method with pre-filtering of peptide identifications, using a simple peptide quality filter. Iterative generation of these filters in conjunction with reversed database searching is used to determine the correct levels at which the APS and peptide quality thresholds should be set to return virtually zero false-positive reports. This proceeds without the need to reference a known dataset.

Computational Biology↗

Reference distributions for apolipoproteins AI and B and the apolipoprotein B/AI ratios: a practical and clinically relevant approach in a large cohort.

The two serum apolipoproteins in the highest concentrations, apolipoprotein (apo) AI and apolipoprotein B, and the apolipoprotein B/AI ratio are measured to assess clinical risk for atherosclerotic heart and peripheral vascular diseases. The study is based on a cohort of over 37,000 Caucasian individuals from northern New England measured in one laboratory by immunonephelometry using standardized reference materials. All samples received for protein analyses were accepted provided adequate identifying information was available. Laboratory and demographic information was entered into a single database for subsequent study. Our results show that for males without evidence of inflammation, values of apo AI change little through life. For females, however, values gradually increase until about 60 years of age then fall somewhat thereafter. Among adults, females have higher apo AI values on average, than males. Apo B values change significantly through life, increasing after the end of the second decade to a peak during the sixth decade, then falling thereafter. In the past, concern has been expressed that apo AI is an acute phase reactant (APR), thus complicating cardiovascular risk assessment. The effects of an APR (C-reactive protein >or=10 mg/L) on apo AI, but not on apo B, are measurable for both sexes, most noticeably beyond the age of 60 years in males and females. When values were expressed as age- and gender-specific multiples of the median (MoMs), the resulting distributions fit a log-Gaussian distribution well over a broad range. The size of the relatively homogenous cohort, by a standardized approach, provides a firm basis for comparison to preexisting reference intervals and for establishing a clinically useful and current reference interval for the three main apolipoprotein values.

Age Distribution↗

Protein supplementation of human milk for promoting growth in preterm infants.

BACKGROUND: For term infants, human milk provides adequate nutrition to facilitate growth, as well as potential beneficial effects on immunity and the maternal-infant emotional state. However, the role of human milk in preterm infants is less well defined as it contains insufficient quantities of some nutrients to meet the estimated needs of the infant. Preterm infants require higher protein intakes than term infants to attain adequate growth rates, and have relatively higher protein turnover rates. Inadequate protein intakes may be partly responsible for low serum albumin and blood urea concentrations in preterm infants. OBJECTIVES: The main objective was to determine if addition of protein to human milk leads to improved growth and neurodevelopmental outcomes without significant adverse effects in preterm infants. SEARCH STRATEGY: The standard search strategy of the Cochrane Neonatal Review Group was used. This includes searches of the Oxford Database of Perinatal Trials, MEDLINE, previous reviews including cross references, abstracts, conferences and symposia proceedings, expert informants, and journal handsearching mainly in the English language. SELECTION CRITERIA: All trials utilizing random or quasi-random allocation to supplementation of human milk with protein or no supplementation in preterm infants who remained in hospital were eligible. DATA COLLECTION AND ANALYSIS: Data were extracting using the standard methods of the Cochrane Neonatal Review Group, with separate evaluation of trial quality and data extraction by each author and synthesis of data using relative risk and weighted mean difference. MAIN RESULTS: Protein supplementation of human milk results in increases in short term weight gain (WMD 3.6 g/kg/day, 95% CI 2.4 to 4.8 g/kg/day), linear growth (WMD 0.28 cm/week, 95% CI 0.18 to 0.38 cm/week) and head growth (WMD 0.15 cm/week, 95% CI 0.06 to 0.23 cm/week). There are insufficient data to evaluate long term neurodevelopmental and growth outcomes. There are too few infants studied to be certain that adverse effects of protein supplementation are not increased. Blood urea levels are increased (WMD 1.0 mmol/l, 95% CI 0.8 to 1.2 mmol/l). REVIEWER'S CONCLUSIONS: Protein supplementation of human milk in relatively well preterm infants results in increases in short term weight gain, linear and head growth. Urea levels are increased, which may reflect adequate rather than excessive dietary protein intake. Further research should be directed towards the evaluation of specific levels of protein intake in preterm infants and the clinical effects of supplementation with protein, including long term growth and neurodevelopmental outcomes. This may best be done in the context of refinement of available multicomponent fortifier preparations.

Dietary Proteins↗

The PDBbind database: methodologies and updates.

We have developed the PDBbind database to provide a comprehensive collection of binding affinities for the protein-ligand complexes in the Protein Data Bank (PDB). This paper gives a full description of the latest version, i.e., version 2003, which is an update to our recently reported work. Out of 23 790 entries in the PDB release No.107 (January 2004), 5897 entries were identified as protein-ligand complexes that meet our definition. Experimentally determined binding affinities (K(d), K(i), and IC(50)) for 1622 of these were retrieved from the references associated with these complexes. A total of 900 complexes were selected to form a "refined set", which is of particular value as a standard data set for docking and scoring studies. All of the final data, including binding affinity data, reference citations, and processed structural files, have been incorporated into the PDBbind database accessible on-line at http:// www.pdbbind.org/.

Databases, Protein↗

Ensembl 2007.

The Ensembl (http://www.ensembl.org/) project provides a comprehensive and integrated source of annotation of chordate genome sequences. Over the past year the number of genomes available from Ensembl has increased from 15 to 33, with the addition of sites for the mammalian genomes of elephant, rabbit, armadillo, tenrec, platypus, pig, cat, bush baby, common shrew, microbat and european hedgehog; the fish genomes of stickleback and medaka and the second example of the genomes of the sea squirt (Ciona savignyi) and the mosquito (Aedes aegypti). Some of the major features added during the year include the first complete gene sets for genomes with low-sequence coverage, the introduction of new strain variation data and the introduction of new orthology/paralog annotations based on gene trees.

Animals↗

Protein palmitoylation by a family of DHHC protein S-acyltransferases.

Protein palmitoylation refers to the posttranslational addition of a 16 carbon fatty acid to the side chain of cysteine, forming a thioester linkage. This acyl modification is readily reversible, providing a potential regulatory mechanism to mediate protein-membrane interactions and subcellular trafficking of proteins. The mechanism that underlies the transfer of palmitate or other long-chain fatty acids to protein was uncovered through genetic screens in yeast. Two related S-palmitoyltransferases were discovered. Erf2 palmitoylates yeast Ras proteins, whereas Akr1 modifies the yeast casein kinase, Yck2. Erf2 and Akr1 share a common sequence referred to as a DHHC (aspartate-histidine-histidine-cysteine) domain. Numerous genes encoding DHHC domain proteins are found in all eukaryotic genome databases. Mounting evidence is consistent with this signature motif playing a direct role in protein acyltransferase (PAT) reactions, although many questions remain. This review presents the genetic and biochemical evidence for the PAT activity of DHHC proteins and discusses the mechanism of protein-mediated palmitoylation.

Acetyltransferases↗

ACNUC--a portable retrieval system for nucleic acid sequence databases: logical and physical designs and usage.

ACNUC is a database structure and retrieval software for use with either the GenBank or EMBL nucleic acid sequence data collections. The nucleotide and textual data furnished by both collections are each restructured into a database that allows sequence retrieval on a multi-criterion basis. The main selection criteria are: species (or higher order taxon), keyword, reference, journal, author, and organelle; all logical combinations of these criteria can be used. Direct access to sequence regions that code for a specific product (protein, tRNA or rRNA) is provided. A versatile extraction procedure copies selected sequences, or fragments of them, from the database to user files suitable to be analysed by user-supplied application programs. A detailed help mechanism is provided to aid the user at any time during the retrieval session. All software has been written in FORTRAN 77 which guarantees a high degree of transportability to minicomputers or mainframes.

Base Sequence↗

YPL.db: the Yeast Protein Localization database.

The Yeast Protein Localization database (YPL.db) contains information about the localization patterns of yeast proteins resulting from microscopic analyses. The data and parameters of the experiments to obtain the localization information, together with images from confocal or video microscopy, are stored in a relational database, building an archive of, and the documentation for, all experiments. The database can be queried based on gene name, protein localization, growth conditions and a number of additional parameters. All experiment parameters are selectable from predefined lists to ensure database integrity and conformity across different investigators. The database provides a structure reference resource to allow for better characterization of unknown or ambiguous localization patterns. Links to MIPS, YPD and SGD databases are provided to allow fast access to further information not contained in the localization database itself. YPL.db is available at http://ypl.tugraz.at.

Computer Graphics↗

Continuous and discontinuous domains: an algorithm for the automatic generation of reliable protein domain definitions.

An algorithm is presented for the fast and accurate definition of protein structural domains from coordinate data without prior knowledge of the number or type of domains. The algorithm explicitly locates domains that comprise one or two continuous segments of protein chain. Domains that include more than two segments are also located. The algorithm was applied to a nonredundant database of 230 protein structures and the results compared to domain definitions obtained from the literature, or by inspection of the coordinates on molecular graphics. For 70% of the proteins, the derived domains agree with the reference definitions, 18% show minor differences and only 12% (28 proteins) show very different definitions. Three screens were applied to identify the derived domains least likely to agree with the subjective definition set. These screens revealed a set of 173 proteins, 97% of which agree well with the subjective definitions. The algorithm represents a practical domain identification tool that can be run routinely on the entire structural database. Adjustment of parameters also allows smaller compact units to be identified in proteins.

Actins↗

Protein analysis by mass spectrometry and sequence database searching: a proteomic approach to identify human lymphoblastoid cell line proteins.

Lymphoblastoid cell lines correspond to in vitro EBV-immortalized lymphocyte B-cells. These cells display a suitable model for experiments dealing with changes in protein expression occurring upon B-cell differentiation, after drug treatment, or after inhibition of some transcription factors. For all these reasons we have undertaken an effort aimed at developing a hematopoietic cell line protein two-dimensional electrophoresis (2-DE) database, containing B-lymphoblastoid 2-DE maps. In this work, matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF-MS) peptide mass fingerprinting analysis was adopted for protein identification. The peptide mass fingerprinting identification and the sequence coverage obtained on colloidal Coomassie blue (CBB) stained gel was close to that obtained using zinc-imidazole staining. Everything considered, CBB being more comfortable for subsequent spot manipulations, CBB staining was chosen for identification of a larger number of polypeptides. The results suggest that reticulation of the gel can interfere preventing the uptake of the enzyme during the in-gel digestion step. Consequently, low molecular mass proteins appear more difficult to identify by mass fingerprinting. Finally, the information provided in this study allows the construction of a new annoted reference map of human lymphoblastoid cell proteins. Among the identified proteins 60% were not yet positioned on 2-DE maps in three of the most important well-documented databases. The annoted map will be accessible via Internet on the LBPP server at URL:http:// www-smbh.univ-paris13.fr/lbtp/index.htm.

Acrylic Resins↗

CyanoBase, the genome database for Synechocystis sp. strain PCC6803: status for the year 2000.

CyanoBase provides an online resource for access to data on genomic information about the cyanobacterium Synechocystis sp. strain PCC6803. The database contains annotations for each protein-coding gene deduced from the entire nucleotide sequence of the genome, gene classification lists, and keyword and similarity search engines. Core portions of CyanoBase consist of annotations for each of the 3168 protein genes deduced from the entire nucleotide sequence of this genome. The contents of each gene were improved by updating with the results of similarity searches and by introducing references for analysis in bioinformatics. The database now contains repository facilities that store and provide experimental information, in addition to providing proposals for the function of each gene. This information should help to avoid unnecessary, overlapping experiments and should assist communication between scientists who wish to elucidate the function of putative genes on the cyanobacteria genome. The current URL of CyanoBase is http://www.kazusa.or.jp:8080/cyano/

Cyanobacteria↗

The genomic heterogeneity among Mycobacterium terrae complex displayed by sequencing of 16S rRNA and hsp 65 genes.

The species identification within Mycobacterium terrae complex has been known to be very difficult. In this study, the genomic diversity of M. terrae complex with eighteen clinical isolates, which were initially identified as M. terrae complex by phenotypic method, was investigated, including that of three type strains (M. terrae, M. nonchromogenicum, and M. triviale ). 16S rRNA and 65-kDa heat shock protein (hsp 65) gene sequences of mycobacteria were determined and aligned with eleven other references for the comparison using similarity search against the GenBank and Ribosomal Database Project II (RDP) databases. 16S rRNA and hsp 65 genes of M. terrae complex showed genomic heterogeneity. Amongst the eighteen clinical isolates, nine were identified as M. nonchromogenicum, eight as M. terrae, one as M. mucogenicum with the molecular characteristic of rapid growth. M. nonchromogenicum could be subdivided into three subgroups, while M. terrae could be subdivided into two subgroups using a 5 bp criterion (>1% difference). Seven isolates in two subgroups of M. nonchromogenicum were Mycobacterium sp. strain MCRO 6, which was closely related to M. nonchromogenicum. The hsp 65 gene could not differentiate one M. nonchromogenicum from M. avium or one M. terrae from M. intracellulare. The nucleotide sequence analysis of 16S rRNA and hsp 65 genes was shown to be useful in identifying the M. terrae complex, but hsp 65 was less discriminating than 16S rRNA.

Bacterial Proteins↗

The ICOH and IUPAC international programme for establishing reference values of metals.

In cooperation with the ICOH Scientific Committee on the Toxicology of metals and IUPAC Commission on Toxicology, we have developed evaluation criteria for derivation of reference values for metal concentrations in human tissues and fluids. In a first attempt to illustrate how these criteria may be used, tentative reference values for mercury in human blood were derived. For persons who do not eat fish, a mean value of 10 mumol/1 (2 micrograms/1) was suggested. It was pointed out, however, that this value was based on information that did not meet the desired quality requirements, which, unfortunately were not met by any of the published reports.

Animals↗

Separation of human erythrocyte membrane associated proteins with one-dimensional and two-dimensional gel electrophoresis followed by identification with matrix-assisted laser desorption/ionization-time of flight mass spectrometry.

A classical proteomic analysis was used to establish a reference map of proteins associated with healthy human erythrocyte ghosts. Following osmotic lysis and differential centrifugation, ghost proteins were separated by either one-dimensional gel electrophoresis (1-DE) or two-dimensional gel electrophoresis (2-DE). Selected protein bands or spots were excised and trypsinized before mass spectrometric analyses and data mining was performed using the SWISS-PROT and NCBI nonredundant databases. A total of 102 protein spots from a 2-D gel were successfully identified. These corresponded to 59 distinct polypeptides with the remaining 43 being isoforms. As for the 1-D gel, 44 polypeptides were identified, of which 19 were also found on the 2-D gel. Most of the 19 common polypeptides were membrane cytoskeletal proteins that are often referred to as the "band" proteins. The remaining 25 polypeptides that were found exclusively on 1-D gels were proteins with high hydrophobicity (e.g., sorbitol dehydrogenase and glucose transporter) and high molecular mass (e.g., Kell blood group glycoprotein and Janus-kinase 2). A higher number of signaling proteins was also identified on 1-D gels compared to 2-D gels. These included Ras, cAMP dependent protein kinase and TGF-beta receptor type 1 precursor.

Centrifugation↗

Sequence database search using jumping alignments.

We describe a new algorithm for amino acid sequence classification and the detection of remote homologues. The rationale is to exploit both vertical and horizontal information of a multiple alignment in a well balanced manner. This is in contrast to established methods like profiles and hidden Markov models which focus on vertical information as they model the columns of the alignment independently. In our setting, we want to select from a given database of "candidate sequences" those proteins that belong to a given superfamily. In order to do so, each candidate sequence is separately tested against a multiple alignment of the known members of the superfamily by means of a new jumping alignment algorithm. This algorithm is an extension of the Smith-Waterman algorithm and computes a local alignment of a single sequence and a multiple alignment. In contrast to traditional methods, however, this alignment is not based on a summary of the individual columns of the multiple alignment. Rather, the candidate sequence at each position is aligned to one sequence of the multiple alignment, called the "reference sequence". In addition, the reference sequence may change within the alignment, while each such jump is penalized. To evaluate the discriminative quality of the jumping alignment algorithm, we compared it to hidden Markov models on a subset of the SCOP database of protein domains. The discriminative quality was assessed by counting the number of false positives that ranked higher than the first true positive (FP-count). For moderate FP-counts above five, the number of successful searches with our method was considerably higher than with hidden Markov models.

Animals↗

CCR9A and CCR9B: two receptors for the chemokine CCL25/TECK/Ck beta-15 that differ in their sensitivities to ligand.

We isolated cDNAs for a chemokine receptor-related protein having the database designation GPR-9-6. Two classes of cDNAs were identified from mRNAs that arose by alternative splicing and that encode receptors that we refer to as CCR9A and CCR9B. CCR9A is predicted to contain 12 additional amino acids at its N terminus as compared with CCR9B. Cells transfected with cDNAs for CCR9A and CCR9B responded to the chemokine CC chemokine ligand 25 (CCL25)/thymus-expressed chemokine (TECK)/chemokine beta-15 (CK beta-15) in assays for both calcium flux and chemotaxis. No other chemokines tested produced responses specific for the cDNA-transfected cells. mRNA for CCR9A/B is expressed predominantly in the thymus, coincident with the expression of CCL25, and highest expression for CCR9A/B among thymocyte subsets was found in CD4+CD8+ cells. mRNAs encoding the A and B forms of the receptor were expressed at a ratio of approximately 10:1 in immortalized T cell lines, in PBMC, and in diverse populations of thymocytes. The EC50 of CCL25 for CCR9A was lower than that for CCR9B, and CCR9A was desensitized by doses of CCL25 that failed to silence CCR9B. CCR9 is the first example of a chemokine receptor in which alternative mRNA splicing leads to proteins of differing activities, providing a mechanism for extending the range of concentrations over which a cell can respond to increments in the concentration of ligand. The study of CCR9A and CCR9B should enhance our understanding of the role of the chemokine system in T cell biology, particularly during the stages of thymocyte development.

Alternative Splicing↗