Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

NCI Thesaurus: a semantic model integrating cancer-related clinical and molecular information.

Over the last 8 years, the National Cancer Institute (NCI) has launched a major effort to integrate molecular and clinical cancer-related information within a unified biomedical informatics framework, with controlled terminology as its foundational layer. The NCI Thesaurus is the reference terminology underpinning these efforts. It is designed to meet the growing need for accurate, comprehensive, and shared terminology, covering topics including: cancers, findings, drugs, therapies, anatomy, genes, pathways, cellular and subcellular processes, proteins, and experimental organisms. The NCI Thesaurus provides a partial model of how these things relate to each other, responding to actual user needs and implemented in a deductive logic framework that can help maintain the integrity and extend the informational power of what is provided. This paper presents the semantic model for cancer diseases and its uses in integrating clinical and molecular knowledge, more briefly examines the models and uses for drug, biochemical pathway, and mouse terminology, and discusses limits of the current approach and directions for future work.

Biomedical Research↗

Quantification of the variation in percentage identity for protein sequence alignments.

BACKGROUND: Percentage Identity (PID) is frequently quoted in discussion of sequence alignments since it appears simple and easy to understand. However, although there are several different ways to calculate percentage identity and each may yield a different result for the same alignment, the method of calculation is rarely reported. Accordingly, quantification of the variation in PID caused by the different calculations would help in interpreting PID values in the literature. In this study, the variation in PID was quantified systematically on a reference set of 1028 alignments generated by comparison of the protein three-dimensional structures. Since the alignment algorithm may also affect the range of PID, this study also considered the effect of algorithm, and the combination of algorithm and PID method. RESULTS: The maximum variation in PID due to the calculation method was 11.5% while the effect of alignment algorithm on PID was up to 14.6% across three popular alignment methods. The combined effect of alignment algorithm and PID calculation gave a variation of up to 22% on the test data, with an average of 5.3% +/- 2.8% for sequence pairs with < 30% identity. In order to see which PID method was most highly correlated with structural similarity, four different PID calculations were compared to similarity scores (Sc) from the comparison of the corresponding protein three-dimensional structures. The highest correlation coefficient for a PID calculation was 0.80. In contrast, the more sophisticated Z-score calculated by reference to randomized sequences gave a correlation coefficient of 0.84. CONCLUSION: Although it is well known amongst expert sequence analysts that PID is a poor score for discriminating between protein sequences, the apparent simplicity of the percentage identity score encourages its widespread use in establishing cutoffs for structural similarity. This paper illustrates that not only is PID a poor measure of sequence similarity when compared to the Z-score, but that there is also a large uncertainty in reported PID values. Since better alternatives to PID exist to quantify sequence similarity, these should be quoted where possible in preference to PID. The findings presented here should prove helpful to those new to sequence analysis, and in warning those who seek to interpret the value of a PID reported in the literature.

Algorithms↗

Industrial-scale proteomics: from liters of plasma to chemically synthesized proteins.

Human blood plasma is a useful source of proteins associated with both health and disease. Analysis of human blood plasma is a challenge due to the large number of peptides and proteins present and the very wide range of concentrations. In order to identify as many proteins as possible for subsequent comparative studies, we developed an industrial-scale (2.5 liter) approach involving sample pooling for the analysis of smaller proteins (M(r) generally < ca. 40 000 and some fragments of very large proteins). Plasma from healthy males was depleted of abundant proteins (albumin and IgG), then smaller proteins and polypeptides were separated into 12 960 fractions by chromatographic techniques. Analysis of proteins and polypeptides was performed by mass spectrometry prior to and after enzymatic digestion. Thousands of peptide identifications were made, permitting the identification of 502 different proteins and polypeptides from a single pool, 405 of which are listed here. The numbers refer to chromatographically separable polypeptide entities present prior to digestion. Combining results from studies with other plasma pools we have identified over 700 different proteins and polypeptides in plasma. Relatively low abundance proteins such as leptin and ghrelin and peptides such as bradykinin, all invisible to two-dimensional gel technology, were clearly identified. Proteins of interest were synthesized by chemical methods for bioassays. We believe that this is the first time that the small proteins in human blood plasma have been separated and analyzed so extensively.

Amino Acid Sequence↗

Database resources of the National Center for Biotechnology.

In addition to maintaining the GenBank(R) nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval resources for the data in GenBank and other biological data made available through NCBI's Web site. NCBI resources include Entrez, PubMed, PubMed Central (PMC), LocusLink, the NCBITaxonomy Browser, BLAST, BLAST Link (BLink), Electronic PCR (e-PCR), Open Reading Frame (ORF) Finder, References Sequence (RefSeq), UniGene, HomoloGene, ProtEST, Database of Single Nucleotide Polymorphisms (dbSNP), Human/Mouse Homology Map, Cancer Chromosome Aberration Project (CCAP), Entrez Genomes and related tools, the Map Viewer, Model Maker (MM), Evidence Viewer (EV), Clusters of Orthologous Groups (COGs) database, Retroviral Genotyping Tools, SAGEmap, Gene Expression Omnibus (GEO), Online Mendelian Inheritance in Man (OMIM), the Molecular Modeling Database (MMDB), the Conserved Domain Database (CDD), and the Conserved Domain Architecture Retrieval Tool (CDART). Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized data sets. All of the resources can be accessed through the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Animals↗

Analysis of the compositional biases in Plasmodium falciparum genome and proteome using Arabidopsis thaliana as a reference.

Comparative genomic analysis of the malaria causative agent, Plasmodium falciparum, with other eukaryotes for which the complete genome is available, revealed that the genome from P. falciparum was more similar to the genome of a plant, Arabidopsis thaliana, than to other non-apicomplexan taxa. Plant-like sequences are thought to result from horizontal gene transfers after a secondary endosymbiosis involving an algal ancestor. The use of the A. thaliana genome and proteome as a reference gives an opportunity to refine our understanding of the extreme compositional bias in the P. falciparum genome that leads to a proteome-wide amino acid bias. A set of pairs of non-redundant protein homologues was selected owing to rigorous genome-wide sequence comparison methods. The introduction of A. thaliana as a reference was a mean to weight the magnitude of the protein evolutionary divergence in P. falciparum. The correlation of the amino acid proportions with evolutionary time supports the hypothesis that amino acids encoded by GC-rich codons are directionally substituted into amino acids encoded by AT-rich codons in the P. falciparum proteome. The long-term deviation of codons in malarial sequences appears as a possible consequence of a genome-wide tri-nucleotidic signature imprinting. Additionally, this study suggests possible working guidelines to improve the accuracy of P. falciparum sequence comparisons, for homology searches and phylogenetic studies.

AT Rich Sequence↗

Ontology annotation: mapping genomic regions to biological function.

With numerous whole genomes now in hand, and experimental data about genes and biological pathways on the increase, a systems approach to biological research is becoming essential. Ontologies provide a formal representation of knowledge that is amenable to computational as well as human analysis, an obvious underpinning of systems biology. Mapping function to gene products in the genome consists of two, somewhat intertwined enterprises: ontology building and ontology annotation. Ontology building is the formal representation of a domain of knowledge; ontology annotation is association of specific genomic regions (which we refer to simply as 'genes', including genes and their regulatory elements and products such as proteins and functional RNAs) to parts of the ontology. We consider two complementary representations of gene function: the Gene Ontology (GO) and pathway ontologies. GO represents function from the gene's eye view, in relation to a large and growing context of biological knowledge at all levels. Pathway ontologies represent function from the point of view of biochemical reactions and interactions, which are ordered into networks and causal cascades. The more mature GO provides an example of ontology annotation: how conclusions from the scientific literature and from evolutionary relationships are converted into formal statements about gene function. Annotations are made using a variety of different types of evidence, which can be used to estimate the relative reliability of different annotations.

Animals↗

Statistical limits to the identification of ion channel domains by sequence similarity.

The study of ion channel function is constrained by the availability of structures for only a small number of channels. A commonly used bioinformatics technique is to assert, based on sequence similarity, that a domain within a channel of interest has the same structure as a reference domain for which the structure is known. This technique, while useful, is often employed when there is only a slight similarity between the channel of interest and the domain of known structure. In this study, we exploit recent advances in structural genomics to calculate the sequence-based probability of the presence of putative domains in a number of ion channels. We find strong support for the presence of many domains that have been proposed in the literature. For example, eukaryotic and prokaryotic CLC proteins almost certainly share a common structure. A number of proposed domains, however, are not as well supported. In particular, for the COOH terminus of the BK channel we find a number of literature proposed domains for which the assertion of common structure based on common sequence has a nontrivial probability of error.

Animals↗

TRANSPATH: an information resource for storing and visualizing signaling pathways and their pathological aberrations.

TRANSPATH is a database about signal transduction events. It provides information about signaling molecules, their reactions and the pathways these reactions constitute. The representation of signaling molecules is organized in a number of orthogonal hierarchies reflecting the classification of the molecules, their species-specific or generic features, and their post-translational modifications. Reactions are similarly hierarchically organized in a three-layer architecture, differentiating between reactions that are evidenced by individual publications, generalizations of these reactions to construct species-independent 'reference pathways' and the 'semantic projections' of these pathways. A number of search and browse options allow easy access to the database contents, which can be visualized with the tool PathwayBuildertrade mark. The module PathoSign adds data about pathologically relevant mutations in signaling components, including their genotypes and phenotypes. TRANSPATH and PathoSign can be used as encyclopaedia, in the educational process, for vizualization and modeling of signal transduction networks and for the analysis of gene expression data. TRANSPATH Public 6.0 is freely accessible for users from non-profit organizations under http://www.gene-regulation.com/pub/databases.html.

Computer Graphics↗

Retinol and alpha-tocopherol in hemodialysis patients.

The biological effects of reactive oxygen species and other radicals controlled by antioxidant mechanisms are modified by various enzymes and other substrates. Antioxidant substrates are divided into those with lipophilic and hydrophilic groups. Retinol and tocopherol are the main representations of lipophilic antioxidants. The aim of the present study was to describe the changes of retinol and alpha-tocopherol which occurred in hemodialysis (HD) patients in respect to the influence of antioxidant systems. The experimental group consisted of 14 patients on regular HD treatment. The control group consisted of 14 healthy blood donors. HPLC was used to measure retinol and alpha-tocopherol in serum. We found that the retinol concentration was significantly higher in HD patients compared to controls (2.35 +/- 0.95 versus 0.90 +/- 0.23 mg/L, p < 0.0001). The concentration of alpha-tocopherol in serum was not different in both study groups (7.32 +/- 3.01 versus. 8.94 +/- 3.57 mg/L). A review of the MEDLINE database since 1985 found a few references concerning these important antioxidant vitamins in HD patients and these contained contrasting results. It has been suggested that some of the complications related to HD including cardiovascular complications, anemia and atherosclerosis may be due to ineffective antioxidant systems and/or to increased free oxygen radical production. The question about supplementation of antioxidants in HD patients is open although there are some positive data regarding the use of moderate and safe selenium supplementation in HD patients. HD patients treated by erythropoietin had increased plasma concentration of retinol and normal level of alpha-tocopherol compared to healthy group. However, this positive finding did not affect lipid peroxidation, which is increased in HD patients and leads to some complications during HD treatment.

Adult↗

Mapping and structural dissection of human 20 S proteasome using proteomic approaches.

The proteasome, a proteolytic complex present in all eukaryotic cells, is part of the ATP-dependent ubiquitin/proteasome pathway. It plays a critical role in the regulation of many physiological processes. The 20 S proteasome, the catalytic core of the 26 S proteasome, is made of four stacked rings of seven subunits each (alpha7beta7beta7alpha7). Here we studied the human 20 S proteasome using proteomics. This led to the establishment of a fine subunit reference map and to the identification of post-translational modifications. We found that the human 20 S proteasome, purified from erythrocytes, exhibited a high degree of structural heterogeneity, characterized by the presence of multiple isoforms for most of the alpha and beta subunits, including the catalytic ones, resulting in a total of at least 32 visible spots after Coomassie Blue staining. The different isoforms of a given subunit displayed shifted pI values, suggesting that they likely resulted from post-translational modifications. We then took advantage of the efficiency of complementary mass spectrometric approaches to investigate further these protein modifications at the structural level. In particular, we focused our efforts on the alpha7 subunit and characterized its N-acetylation and its phosphorylation site localized on Ser(250).

Amino Acid Sequence↗

An internet database of crotaline venom found in the United States.

Many snake venoms have been shown to be complex mixtures of pharmacologically important molecules, some of which have potential therapeutic value in the treatment of clot-induced ischemia, cancer and other human disorders. The literature contains many references on how venom and/or venom components are being used in medicine. Within the United States, there are 44 subspecies of poisonous snakes. Despite this rather vast diversity, 90% of the venom-related biomedical research conducted on native snakes found in the United States has been done on a limited number of the more common species. Since the venoms from most of the native species are not available or characterized, their composition and potential usefulness in medicine and applied biomedical research has not been explored. The Natural Toxins Research Center (NTRC) at Texas A&M University-Kingsville has developed a serpentarium that presently houses a population of over 250 snakes composed of 11 species and 20 subspecies. These snakes are cataloged on the Internet database along with their geographical location data, proteolytic activities, high performance liquid chromatography (HPLC) and electrophoretic titration (ET) profiles. Many of these snake venoms have never been characterized and few locale-specific differences within a species have been examined. These venoms can be queried through an on-line search routine. The database will be a useful starting point for anyone interested in isolating fibrinolytic enzymes, specific toxins, hemorrhagins, or other pharmacologically active proteins from snake venoms.

Animals↗

Phylogenetic tree information aids supervised learning for predicting protein-protein interaction based on distance matrices.

BACKGROUND: Protein-protein interactions are critical for cellular functions. Recently developed computational approaches for predicting protein-protein interactions utilize co-evolutionary information of the interacting partners, e.g., correlations between distance matrices, where each matrix stores the pairwise distances between a protein and its orthologs from a group of reference genomes. RESULTS: We proposed a novel, simple method to account for some of the intra-matrix correlations in improving the prediction accuracy. Specifically, the phylogenetic species tree of the reference genomes is used as a guide tree for hierarchical clustering of the orthologous proteins. The distances between these clusters, derived from the original pairwise distance matrix using the Neighbor Joining algorithm, form intermediate distance matrices, which are then transformed and concatenated into a super phylogenetic vector. A support vector machine is trained and tested on pairs of proteins, represented as super phylogenetic vectors, whose interactions are known. The performance, measured as ROC score in cross validation experiments, shows significant improvement of our method (ROC score 0.8446) over that of using Pearson correlations (0.6587). CONCLUSION: We have shown that the phylogenetic tree can be used as a guide to extract intra-matrix correlations in the distance matrices of orthologous proteins, where these correlations are represented as intermediate distance matrices of the ancestral orthologous proteins. Both the unsupervised and supervised learning paradigms benefit from the explicit inclusion of these intermediate distance matrices, and particularly so in the latter case, which offers a better balance between sensitivity and specificity in the prediction of protein-protein interactions.

Computational Biology↗

Foot-and-mouth disease type O viruses exhibit genetically and geographically distinct evolutionary lineages (topotypes).

Serotype O is the most prevalent of the seven serotypes of foot-and-mouth disease (FMD) virus and occurs in many parts of the world. The UPGMA method was used to construct a phylogenetic tree based on nucleotide sequences at the 3' end of the VP1 gene from 105 FMD type O viruses obtained from samples submitted to the OIE/FAO World Reference Laboratory for FMD. This analysis identified eight major genotypes when a value of 15% nucleotide difference was used as a cut-off. The validity of these groupings was tested on the complete VP1 gene sequences of 23 of these viruses by bootstrap resampling and construction of a neighbour-joining tree. These eight genetic lineages fell within geographical boundaries and we have used the term topotype to describe them. Using a large sequence database, the distribution of viruses belonging to each of the eight topotypes has been determined. These phylogenetically based epidemiological studies have also been used to identify viruses that have transgressed their normal ecological niches. Despite the high rate of mutation during replication of the FMD virus genome, the topotypes appear to represent evolutionary cul-de-sacs.

Africa, Eastern↗

Proteomic characterization of human normal articular chondrocytes: a novel tool for the study of osteoarthritis and other rheumatic diseases.

Articular cartilage is composed of cells and an extracellular matrix. The chondrocyte is the only cell type present in mature cartilage, and it is important in the control of cartilage integrity. There is currently a great lack of knowledge about the chondrocyte proteome. To solve this deficiency, we have obtained the first reference map of the human normal articular chondrocyte. Cells were isolated from cartilages obtained from autopsies without history of joint disease. Cultured cells were used to obtain protein extracts which were resolved by 2-DE and visualized by silver nitrate or CBB staining. Almost 200 spots were excised from the gels and analyzed using MALDI-TOF or MALDI-TOF/TOF MS. The analysis leads to the identification of 136 spots that represent 93 different proteins. A significant proportion of proteins are involved in cell organization (26%), energy (16%), protein fate (14%), metabolism (12%), and cell stress (12%). From all the identified proteins, annexins, vimentin, transgelin, destrin, cathepsin D, heat shock protein 47, and mitochondrial superoxide dismutase were more abundant in chondrocytes than in other types of mesenchymal cells such as Jurkat-T cells. As metabolic program of chondrocytes is altered in osteoarthritis and other rheumatic diseases, this proteomic map is an important tool for future studies on these pathologies.

Adolescent↗

Detection of MEF-1 laboratory reference strain of poliovirus type 2 in children with poliomyelitis in India in 2002 & 2003.

BACKGROUND & OBJECTIVES: Significant progress has been made towards eradication of poliomyelitis in India. Surveillance for acute flaccid paralysis (AFP) has reached high standards. Among the 3 types of polioviruses, type 2 had been eliminated in India and eradicated globally as of October 1999. However, we isolated wild poliovirus type 2 from a small number of polio cases in northern India in 2000 and again during December 2002 to February 2003. Using molecular tools the origin, of the wild type 2 poliovirus was investigated. METHODS: Polioviruses isolated from stool samples collected from patients with AFP were differentiated as wild virus or Sabin vaccine-like by ELISA and probe hybridization assays. Complete VP1 gene nucleotide sequences of the wild type 2 poliovirus isolates were determined by reverse transcriptase polymerase chain reaction (RT-PCR), followed by cycle sequencing. VP1 nucleotide sequences were compared with those of wild type 2 polioviruses that were indigenous in India in the past as well as prototype/laboratory strains and the GenBank database. RESULTS: Wild poliovirus type 2 was detected in stool samples from 6 patients with AFP in western Uttar Pradesh and 1 in Gujarat. In addition, the virus was isolated from one healthy contact child and from environmental sewage sample in Moradabad where three of these patients were reported. These isolates were identified as genetically closely related to laboratory reference strain MEF-1. Molecular characterization of the isolates confirmed that there was no evidence of extensive person-to-person transmission of the virus in the community. INTERPRETATION & CONCLUSION: Laboratory reference strain (MEF-1) of poliovirus type 2 caused paralytic poliomyelitis in 10 patients in September 2000 and November 2002 to February 2003. The origin of the virus was some laboratory as yet not identified. This episode highlights the urgent need for stringent containment of wild poliovirus containing materials in the laboratories across the country in order to prevent recurrence of such incidents.

Capsid Proteins↗

[Identification and analysis of a mouse gene homologous to human hepatitis B virus pre-S1 protein-binding protein using the bioinformatics method].

OBJECTIVE: To clone and identify the mouse gene homologous to human hepatitis B virus (HBV) pre-S1 protein-binding protein (PS1BP). METHODS: The human PS1BP cDNA sequence was used as the reference sequence to search homologous mouse cDNA sequence from GenBank established by National Center for Biotechnology (NCBI), National Institute of Health (NIH), for its homologous cDNA sequences of mouse by BLASTn tool. The characteristics of mouse PS1BP protein primary structure were predicted by online software. Finally the genomic DNA structure of mouse PS1BP was deduced and compared. RESULTS: The mouse PS1BP was identified and consisted of 1455 nt, coding a protein of 484 aa. The identity of human and mouse PS1BP protein is 84.92% (411/484). The genomic DNA of mouse PS1BP consisted of 3 exons and 2 introns. CONCLUSION: The identification and characterization of mouse PS1BP cDNA and genomic DNA pave a way for further study of their structures and functions.

Amino Acid Sequence↗

Detecting uber-operons in prokaryotic genomes.

We present a study on computational identification of uber-operons in a prokaryotic genome, each of which represents a group of operons that are evolutionarily or functionally associated through operons in other (reference) genomes. Uber-operons represent a rich set of footprints of operon evolution, whose full utilization could lead to new and more powerful tools for elucidation of biological pathways and networks than what operons have provided, and a better understanding of prokaryotic genome structures and evolution. Our prediction algorithm predicts uber-operons through identifying groups of functionally or transcriptionally related operons, whose gene sets are conserved across the target and multiple reference genomes. Using this algorithm, we have predicted uber-operons for each of a group of 91 genomes, using the other 90 genomes as references. In particular, we predicted 158 uber-operons in Escherichia coli K12 covering 1830 genes, and found that many of the uber-operons correspond to parts of known regulons or biological pathways or are involved in highly related biological processes based on their Gene Ontology (GO) assignments. For some of the predicted uber-operons that are not parts of known regulons or pathways, our analyses indicate that their genes are highly likely to work together in the same biological processes, suggesting the possibility of new regulons and pathways. We believe that our uber-operon prediction provides a highly useful capability and a rich information source for elucidation of complex biological processes, such as pathways in microbes. All the prediction results are available at our Uber-Operon Database: http://csbl.bmb.uga.edu/uber, the first of its kind.

Algorithms↗

A nonsense polymorphism (Y319X) of the solute carrier family 6 member 18 (SLC6A18) gene is not associated with hypertension and blood pressure in Japanese.

We investigated the possible association of solute carrier family 6 member 18 (SLC6A18) with hypertension and blood pressure in Japanese, since the homologous murine XT2 gene was recently reported to be associated with hypertension. The entire coding region of SLC6A18 was sequenced in 30 unrelated Japanese subjects. The deleterious effects of the observed nonsynonymous single nucleotide polymorphisms (SNPs) on the phenotype were predicted using bioinformatics software. We tested the associations of one deleterious SNP (Y319X) with blood pressure and hypertension in a general population of 1,004 subjects in one area of Japan. Both quantitative and qualitative analyses adjusting for age and body mass index (BMI) as covariates were undertaken. Four synonymous (P7P, T32T, G37G and V387V), three missense (S12C, I32T and L478P) and one nonsense (Y319X: g1230757 C > G) polymorphisms were found. One of the synonymous polymorphisms was novel (V387V) by reference to the dbSNP database. The Y319X genotype distribution of CC:CG:GG in this population showed frequencies of 0.382, 0.461 and 0.156, respectively, which followed Hardy-Weinberg equilibrium. The nonsense polymorphism had odds ratios of 0.83 (confidence interval [CI] = 0.59-1.15, p = 0.26) in males and 0.96 (CI = 0.72-1.29, p = 0.80) in females with hypertension or current medication for hypertension. For the quantitative analysis, we excluded the current medication subgroup. The nonsense allele was not a significant predictor for systolic or diastolic blood pressure. This is the first report showing that a single polymorphism in SLC6A18 is not associated with hypertension or blood pressure in Japanese.

Adult↗