Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

The genomic heterogeneity among Mycobacterium terrae complex displayed by sequencing of 16S rRNA and hsp 65 genes.

The species identification within Mycobacterium terrae complex has been known to be very difficult. In this study, the genomic diversity of M. terrae complex with eighteen clinical isolates, which were initially identified as M. terrae complex by phenotypic method, was investigated, including that of three type strains (M. terrae, M. nonchromogenicum, and M. triviale ). 16S rRNA and 65-kDa heat shock protein (hsp 65) gene sequences of mycobacteria were determined and aligned with eleven other references for the comparison using similarity search against the GenBank and Ribosomal Database Project II (RDP) databases. 16S rRNA and hsp 65 genes of M. terrae complex showed genomic heterogeneity. Amongst the eighteen clinical isolates, nine were identified as M. nonchromogenicum, eight as M. terrae, one as M. mucogenicum with the molecular characteristic of rapid growth. M. nonchromogenicum could be subdivided into three subgroups, while M. terrae could be subdivided into two subgroups using a 5 bp criterion (>1% difference). Seven isolates in two subgroups of M. nonchromogenicum were Mycobacterium sp. strain MCRO 6, which was closely related to M. nonchromogenicum. The hsp 65 gene could not differentiate one M. nonchromogenicum from M. avium or one M. terrae from M. intracellulare. The nucleotide sequence analysis of 16S rRNA and hsp 65 genes was shown to be useful in identifying the M. terrae complex, but hsp 65 was less discriminating than 16S rRNA.

Bacterial Proteins↗

The ICOH and IUPAC international programme for establishing reference values of metals.

In cooperation with the ICOH Scientific Committee on the Toxicology of metals and IUPAC Commission on Toxicology, we have developed evaluation criteria for derivation of reference values for metal concentrations in human tissues and fluids. In a first attempt to illustrate how these criteria may be used, tentative reference values for mercury in human blood were derived. For persons who do not eat fish, a mean value of 10 mumol/1 (2 micrograms/1) was suggested. It was pointed out, however, that this value was based on information that did not meet the desired quality requirements, which, unfortunately were not met by any of the published reports.

Animals↗

Separation of human erythrocyte membrane associated proteins with one-dimensional and two-dimensional gel electrophoresis followed by identification with matrix-assisted laser desorption/ionization-time of flight mass spectrometry.

A classical proteomic analysis was used to establish a reference map of proteins associated with healthy human erythrocyte ghosts. Following osmotic lysis and differential centrifugation, ghost proteins were separated by either one-dimensional gel electrophoresis (1-DE) or two-dimensional gel electrophoresis (2-DE). Selected protein bands or spots were excised and trypsinized before mass spectrometric analyses and data mining was performed using the SWISS-PROT and NCBI nonredundant databases. A total of 102 protein spots from a 2-D gel were successfully identified. These corresponded to 59 distinct polypeptides with the remaining 43 being isoforms. As for the 1-D gel, 44 polypeptides were identified, of which 19 were also found on the 2-D gel. Most of the 19 common polypeptides were membrane cytoskeletal proteins that are often referred to as the "band" proteins. The remaining 25 polypeptides that were found exclusively on 1-D gels were proteins with high hydrophobicity (e.g., sorbitol dehydrogenase and glucose transporter) and high molecular mass (e.g., Kell blood group glycoprotein and Janus-kinase 2). A higher number of signaling proteins was also identified on 1-D gels compared to 2-D gels. These included Ras, cAMP dependent protein kinase and TGF-beta receptor type 1 precursor.

Centrifugation↗

Sequence database search using jumping alignments.

We describe a new algorithm for amino acid sequence classification and the detection of remote homologues. The rationale is to exploit both vertical and horizontal information of a multiple alignment in a well balanced manner. This is in contrast to established methods like profiles and hidden Markov models which focus on vertical information as they model the columns of the alignment independently. In our setting, we want to select from a given database of "candidate sequences" those proteins that belong to a given superfamily. In order to do so, each candidate sequence is separately tested against a multiple alignment of the known members of the superfamily by means of a new jumping alignment algorithm. This algorithm is an extension of the Smith-Waterman algorithm and computes a local alignment of a single sequence and a multiple alignment. In contrast to traditional methods, however, this alignment is not based on a summary of the individual columns of the multiple alignment. Rather, the candidate sequence at each position is aligned to one sequence of the multiple alignment, called the "reference sequence". In addition, the reference sequence may change within the alignment, while each such jump is penalized. To evaluate the discriminative quality of the jumping alignment algorithm, we compared it to hidden Markov models on a subset of the SCOP database of protein domains. The discriminative quality was assessed by counting the number of false positives that ranked higher than the first true positive (FP-count). For moderate FP-counts above five, the number of successful searches with our method was considerably higher than with hidden Markov models.

Animals↗

CCR9A and CCR9B: two receptors for the chemokine CCL25/TECK/Ck beta-15 that differ in their sensitivities to ligand.

We isolated cDNAs for a chemokine receptor-related protein having the database designation GPR-9-6. Two classes of cDNAs were identified from mRNAs that arose by alternative splicing and that encode receptors that we refer to as CCR9A and CCR9B. CCR9A is predicted to contain 12 additional amino acids at its N terminus as compared with CCR9B. Cells transfected with cDNAs for CCR9A and CCR9B responded to the chemokine CC chemokine ligand 25 (CCL25)/thymus-expressed chemokine (TECK)/chemokine beta-15 (CK beta-15) in assays for both calcium flux and chemotaxis. No other chemokines tested produced responses specific for the cDNA-transfected cells. mRNA for CCR9A/B is expressed predominantly in the thymus, coincident with the expression of CCL25, and highest expression for CCR9A/B among thymocyte subsets was found in CD4+CD8+ cells. mRNAs encoding the A and B forms of the receptor were expressed at a ratio of approximately 10:1 in immortalized T cell lines, in PBMC, and in diverse populations of thymocytes. The EC50 of CCL25 for CCR9A was lower than that for CCR9B, and CCR9A was desensitized by doses of CCL25 that failed to silence CCR9B. CCR9 is the first example of a chemokine receptor in which alternative mRNA splicing leads to proteins of differing activities, providing a mechanism for extending the range of concentrations over which a cell can respond to increments in the concentration of ligand. The study of CCR9A and CCR9B should enhance our understanding of the role of the chemokine system in T cell biology, particularly during the stages of thymocyte development.

Alternative Splicing↗

African swine fever virus-induced polypeptides in porcine alveolar macrophages and in Vero cells: two-dimensional gel analysis.

High-resolution two-dimensional electrophoresis followed by computer analysis has been used to study quantitatively the patterns of protein synthesis produced in porcine alveolar macrophages and in Vero cells infected with African swine fever virus (ASFV). Initially, a protein database for each cell type was constructed. The porcine alveolar macrophage database includes 995 polypeptides (818 acidic, isoelectric focusing (IEF) and 177 basic, nonequilibrium pH gradient electrophoresis (NEPHGE)) whereas the Vero database contains 1,398 polypeptides (1,127 acidic, IEF and 271 basic, NEPHGE). Taking these databases as reference, ASFV highly virulent strain E70 induces 57 acid and 43 basic polypeptides in porcine alveolar macrophages, which account for most of the information content of the virus DNA. The kinetics of synthesis of the virus-induced polypeptides showed the existence of three classes of proteins: one whose synthesis starts early after infection, continues for a period and then switches off; another whose synthesis also starts early but continues for prolonged periods; and a third which requires DNA replication. The attenuated, cell adapted, strain BA71V induces 92 acidic and 37 basic proteins in Vero cells. Significant differences were observed when comparing the patterns of polypeptides induced by the two viral strains. In both cell systems studied, ASFV infection produces a general shutoff of protein synthesis that affects up to 65% of the cellular proteins. Interestingly, 28 proteins of porcine alveolar macrophages and 48 proteins of Vero cells are stimulated at least two times by ASFV infection.

African Swine Fever Virus↗

SCOP: a structural classification of proteins database.

The Structural Classification of Proteins (SCOP) database provides a detailed and comprehensive description of the relationships of all known proteins structures. The classification is on hierarchical levels: the first two levels, family and superfamily, describe near and far evolutionary relationships; the third, fold, describes geometrical relationships. The distinction between evolutionary relationships and those that arise from the physics and chemistry of proteins is a feature that is unique to this database, so far. SCOP also provides for each structure links to atomic co-ordinates, images of the structures, interactive viewers, sequence data, data on any conformational changes related to function and literature references. The database is freely accessible on the World Wide Web (WWW) with an entry point at URL http://scop.mrc-lmb.cam.ac.uk/scop/

Amino Acid Sequence↗

Establishment of the human reflex tear two-dimensional polyacrylamide gel electrophoresis reference map: new proteins of potential diagnostic value.

To understand the changes in protein expression associated with various physiological states as well as the development of pathological eye disease, we have begun to map the protein components of normal human reflex tears. An analytical reference map of normal human reflex tears was created using two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) with pH 3.5-10 immobilized pH gradients (IPGs). Micropreparatively loaded gels were transferred to polyvinylidene difluoride (PVDF) and analysed by a combination of N-terminal sequence tagging and amino acid compositional analysis. Thirty spots were sequence tagged, resulting in identification of six different proteins (lipocalin, lysozyme, lactotransferrin, zinc-alpha-2 glycoprotein, cystatin S, cystatin SN) that matched to entries in the SWISS-PROT database. A group of N-terminally blocked proteins was clearly identified from SWISS-PROT by amino acid analysis, isoelectric point (pI) and molecular weight (Mr). A number of highly expressed protein components remain unidentified despite being subjected to amino acid analysis and Edman sequencing. A majority of the abundant proteins showed varying degrees of charge heterogeneity attributed to post-translational processing such as glycosylation and N-terminal truncation. We have identified a previously undescribed protein that we have named lacryglobin. This protein displays strong homology with mammaglobin, a protein overexpressed in breast cancer. The discovery of this homologue in tears offers the potential for disease diagnosis by screening tear fluid proteins.

Databases, Factual↗

Biological ontologies in rice databases. An introduction to the activities in Gramene and Oryzabase.

An enormous amount of information and materials in the field of biology has been accumulating, such as nucleotide and amino acid sequences, gene and protein functions, mutants and their phenotypes, and literature references, produced by the rapid development in this field. Effective use of the information may strongly promote biological studies, and may lead to many important findings. It is, however, time-consuming and laborious for individual researchers to collect information from individual original sites and to rearrange it for their own purpose. A concept, ontology, has been introduced in biology to support and encourage researchers to share and reuse information among biological databases. Ontology has a glossary, named dynamic controlled vocabulary, in which relationships between terms are defined. Since each term is strictly defined and identified with an ID number, a set of data represented in biological ontology is easily accessible to automated information processing, even if the data sets are across several databases and/or different organisms. In this mini-review, we introduce activities in Gramene and Oryzabase, which provide biological ontologies for Oryza sativa (rice).

Databases, Genetic↗

Gene Ontology annotation status of the fission yeast genome: preliminary coverage approaches 100%.

In this review, we present an overview of the Gene Ontology (GO) structure and describe how the GO is implemented for Sz. pombe and made available via Sz. pombe GeneDB (http://www.genedb.org/genedb/pombe/). We give a detailed progress report of Sz. pombe GO annotation, providing the current status of both manual and automatic annotations. Fission yeast has at least one GO annotation for 98.3% of its genes (excluding annotations to 'unknown' terms), greater than the current percentage coverage for any other organism. Approximately 65% (3225 gene products) have at least one annotation to each of the three ontologies (biological process, cellular component and molecular function). Approximately 30% (1443 gene products) have GO terms derived directly from small-scale experiments in fission yeast, supporting the validity of fission yeast as a model eukaryote and a reference organism.

Computational Biology↗

FlyBase: a Drosophila database.

FlyBase (http://flybase.bio.indiana.edu/) is a comprehensive database of genetic and molecular data concerning Drosophila . FlyBase is maintained as a relational database (in Sybase) and is made available as html documents and flat files. The scope of FlyBase includes: genes, alleles (with phenotypes), aberrations, transposons, pointers to sequence data, gene products, maps, clones, stock lists, Drosophila workers and bibliographic references.

Animals↗

Analysis of the HUPO Brain Proteome reference samples using 2-D DIGE and 2-D LC-MS/MS.

Within the Human Proteome Organization (HUPO) Brain Proteome Project, a pilot study was launched with reference samples shipped to nine international laboratories (see Hamacher et al., this Special Issue) to evaluate different proteome approaches in neuroscience and to build up a first version of a brain protein database. One part of the study addresses quantitative proteome alterations between three developmental stages (embryonic day 16; postnatal day 7; 8 weeks) of mouse brains. Five brains per stage were differentially analyzed by 2-D DIGE using internal standardization and overlapping pH gradients (pH 4-7 and 6-9). In total, 214 protein spots showing stage-dependent intensity alterations (> two-fold) were detected, 56 of which were identified. Several of them, e.g. members of the dihydropyrimidinase family, are known to be associated with brain development. To feed the HUPO BPP brain protein database, a robust 2-D LC-MS/MS method was applied to murine postnatal day 7 and human post-mortem brain samples. Using MASCOT and the IPI database, 350 human and 481 mouse proteins could be identified by at least two different peptides. The data are accessible through the PRIDE database (http://www.ebi.ac.uk/pride/).

Aging↗

The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1999.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: cross-references to additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT. The URLs for SWISS-PROT on the WWW are: http://www.expasy.ch/sprot and http://www. ebi.ac.uk/sprot

Amino Acid Sequence↗

Microsequences of 145 proteins recorded in the two-dimensional gel protein database of normal human epidermal keratinocytes.

Microsequencing of proteins recovered from two-dimensional (2-D) gels is being used systematically to identify proteins in the master human keratinocyte 2-D gel database. To date, about 250 protein spots recorded in human 2-D gel databases have been microsequenced and, of these, 145 are recorded in the keratinocyte database under the entry partial amino acid sequence. Coomassie Brilliant Blue-stained protein spots cut from several (up to 40) dry gels were concentrated by elution-concentration gel electrophoresis, electroblotted onto PVDF membranes and digested in situ with trypsin. Eluting peptides were separated by reversed-phase HPLC, collected individually and sequenced. Computer search using the FASTA and TFASTA programs from Genetics Computer Group indicated that 110 of the microsequenced polypeptides shared significant similarity with proteins contained in the PIR, Mipsx or GenEMBL databases. Only 35 polypeptides corresponded to hitherto unknown proteins. Peptide sequences of all 145 proteins are listed together with their coordinates (apparent molecular weight and pI) in the keratinocyte database.

Amino Acid Sequence↗

An evaluation, comparison, and accurate benchmarking of several publicly available MS/MS search algorithms: sensitivity and specificity analysis.

MS/MS and associated database search algorithms are essential proteomic tools for identifying peptides. Due to their widespread use, it is now time to perform a systematic analysis of the various algorithms currently in use. Using blood specimens used in the HUPO Plasma Proteome Project, we have evaluated five search algorithms with respect to their sensitivity and specificity, and have also accurately benchmarked them based on specified false-positive (FP) rates. Spectrum Mill and SEQUEST performed well in terms of sensitivity, but were inferior to MASCOT, X!Tandem, and Sonar in terms of specificity. Overall, MASCOT, a probabilistic search algorithm, correctly identified most peptides based on a specified FP rate. The rescoring algorithm, PeptideProphet, enhanced the overall performance of the SEQUEST algorithm, as well as provided predictable FP error rates. Ideally, score thresholds should be calculated for each peptide spectrum or minimally, derived from a reversed-sequence search as demonstrated in this study based on a validated data set. The availability of open-source search algorithms, such as X!Tandem, makes it feasible to further improve the validation process (manual or automatic) on the basis of "consensus scoring", i.e., the use of multiple (at least two) search algorithms to reduce the number of FPs. complement.

Algorithms↗

Fast structure alignment for protein databank searching.

A fast method is described for searching and analyzing the protein structure databank. It uses secondary structure followed by residue matching to compare protein structures and is developed from a previous structural alignment method based on dynamic programming. Linear representations of secondary structures are derived and their features compared to identify equivalent elements in two proteins. The secondary structure alignment then constrains the residue alignment, which compares only residues within aligned secondary structures and with similar buried areas and torsional angles. The initial secondary structure alignment improves accuracy and provides a means of filtering out unrelated proteins before the slower residue alignment stage. It is possible to search or sort the protein structure databank very quickly using just secondary structure comparisons. A search through 720 structures with a probe protein of 10 secondary structures required 1.7 CPU hours on a Sun 4/280. Alternatively, combined secondary structure and residue alignments, with a cutoff on the secondary structure score to remove pairs of unrelated proteins from further analysis, took 10.1 CPU hours. The method was applied in searches on different classes of proteins and to cluster a subset of the databank into structurally related groups. Relationships were consistent with known families of protein structure.

Amino Acid Sequence↗

Relationships between bacterial drug resistance pumps and other transport proteins.

We have used three reference sequences representative of bacterial drug resistance pumps and sugar transport proteins to collect the 91 most closely related sequences from a composite, nonredundant protein sequence database. Having eliminated certain very close relatives, the remainder were subjected to analysis and alignment by using two different similarity matrices: one of these was a matrix based on structural conservation of amino acid residues in proteins of known conformation and the other was based on the more familiar mutational matrix. Unrooted similarity trees for these proteins were constructed for each matrix and compared. A systematic analysis of the differences between these trees was undertaken and the sequences were analyzed for the presence or absence of certain sequence motifs. The results show that the clades created by the two methods are broadly comparable but that there are some clusters of sequences that are significantly different. Further analysis confirmed that (1) the sequences collected by this objective method are all known or putative 12-helix (in some cases reported as 14-helix) transmembrane proteins, (2) there is evidence for few cases of an origin based on gene duplication, (3) the bacterial drug resistance pumps are distributed in more than one clade and cannot be regarded as a definitive subset of these proteins, and that (4) the diversity is such that there is no evidence of a single ancestral protein. The possible extension of the methods to other cases of divergent protein sequences is discussed.

Amino Acid Sequence↗

Fission yeast tor1 functions in response to various stresses including nitrogen starvation, high osmolarity, and high temperature.

A target of rapamycin (TOR) protein is a protein kinase that exerts cellular signal transduction to regulate cell growth in response to extracellular nutrient conditions. In the Schizosaccharomyces pombe genome database, there are two genes encoding TOR-related proteins, but their functions have not been analyzed. Here we report that one of the genes, referred to as tor1+, is required for sexual development induced by nitrogen starvation. Ste11 is a key transcription factor for the initiation of sexual development. The expression of ste11+ is normally regulated in tor1- cells; and overexpression of ste11+ hardly rescues the defect in fertility in tor1-. Upon nitrogen starvation, tor1+ cells promote two rounds of the cell cycle to become arrested at the G1 phase before initiation of sexual development. The tor1- cells do not promote such a cell cycle, suggesting that Tor1 is necessary for the response to nitrogen starvation. The tor1- cells show no growth or very slow growth under various stress conditions, including external high pH, high concentrations of salts or sorbitol, and high temperature. These results suggest that Tor1 is necessary for any response to a wide range of stresses. The vegetative growth of tor1- cells is inhibited by rapamycin, although tor1+ cells are resistant to the drug. The tor1- cells are hypersensitive to fluphenazine and cyclosporin A, which specifically inhibit calmodulin and calcineurin, respectively.

Amino Acid Sequence↗