Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Repositories”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

The Human Proteome Organization Plasma Proteome Project pilot phase: reference specimens, technology platform comparisons, and standardized data submissions and analyses.

A comprehensive, systematic characterization of cirolating proteins in health and disease will greatly facilitate development of biomarkers for prevention, diagnosis, and therapy of cancers and other diseases. The Human Proteome Organization Plasma Proteome Project pilot phase aims to (1) compare the advantages and limitations of many technology platforms; (2) contrast reference specimens of human plasma (ethylenediaminetetra acetic acid, heparin, citrate-anticoagulated) and serum, in terms of numbers of proteins identified and any interferences with various technology platforms; and (3) create a global knowledge base/data repository.

Biomarkers↗

Advances in the development of common interchange standards for proteomic data.

The generation of proteomics data is increasingly high-throughput and high volume. Both experimental design and the technologies used to produce and subsequently analyze the data are becoming ever more complex. An increasing need for methods by which such data can be accurately described, stored and exchanged between experimenters and data repositories has been recognised. Work by the Proteomics Standards Initiative of the Human Proteome Organisation has laid the foundation for the development of standards by which experimental design can be described and data exchange facilitated. At a recent workshop in Nice, participants gathered to review the progress made to date and assist in pushing the process still further forward.

Humans↗

Further steps towards data standardisation: the Proteomic Standards Initiative HUPO 3(rd) annual congress, Beijing 25-27(th) October, 2004.

The increasing volume of proteomics data currently being generated by increasingly high-throughput methodologies has led to an increasing need for methods by which such data can be accurately described, stored and exchanged between experimental researchers and data repositories. Work by the Proteomics Standards Initiative of the Human Proteome Organisation has laid the foundation for the development of standards by which experimental design can be described and data exchange facilitated. The progress of these efforts, and the direct benefits already accruing from them, were described at a plenary session of the 3(rd) Annual HUPO congress. Parallel sessions allowed the three work groups to present their progress to interested parties and to collect feedback from groups already implementing the available formats.

China↗

Centralized data analysis of a large interlaboratory proteomics project: a feasibility study.

The human Plasma Proteome Project (PPP) is a large-scale collaboration between many laboratories. One of the most demanding tasks in the PPP involved the analysis of very large amounts of raw MS/MS data produced by the participants. The main approach for managing this task was letting the participants analyze their own data and submit the results to the central PPP repository as lists of identified proteins and peptides. To complement this distributed approach, we also performed centralized analysis of the raw MS/MS data provided by the participants. Due to the data redundancy inherent in such a project, centralized analysis has the potential to reduce the computational effort by reducing redundancy before the analysis. Centralized analysis can also unify the process and take advantage of data sharing among laboratories to improve protein identification and validation. The process we employed included removing low-quality spectra, clustering spectra by mutual similarity, and applying uniform peptide and protein identification procedures. To demonstrate the process, we analyzed 5.28 million MS/MS spectra derived by eight laboratories from tryptic peptides of serum and plasma proteins.

Blood Proteins↗

SPLASH: systematic proteomics laboratory analysis and storage hub.

In the field of proteomics, the increasing difficulty to unify the data format, due to the different platforms/instrumentation and laboratory documentation systems, greatly hinders experimental data verification, exchange, and comparison. Therefore, it is essential to establish standard formats for every necessary aspect of proteomics data. One of the recently published data models is the proteomics experiment data repository [Taylor, C. F., Paton, N. W., Garwood, K. L., Kirby, P. D. et al., Nat. Biotechnol. 2003, 21, 247-254]. Compliant with this format, we developed the systematic proteomics laboratory analysis and storage hub (SPLASH) database system as an informatics infrastructure to support proteomics studies. It consists of three modules and provides proteomics researchers a common platform to store, manage, search, analyze, and exchange their data. (i) Data maintenance includes experimental data entry and update, uploading of experimental results in batch mode, and data exchange in the original PEDRo format. (ii) The data search module provides several means to search the database, to view either the protein information or the differential expression display by clicking on a gel image. (iii) The data mining module contains tools that perform biochemical pathway, statistics-associated gene ontology, and other comparative analyses for all the sample sets to interpret its biological meaning. These features make SPLASH a practical and powerful tool for the proteomics community.

Database Management Systems↗

Further steps in standardisation. Report of the second annual Proteomics Standards Initiative Spring Workshop (Siena, Italy 17-20th April 2005).

The spring workshop of the HUPO-PSI convened in Siena to further progress the data standards which are already making an impact on data exchange and deposition in the field of proteomics. Separate work groups pushed forward existing XML standards for the exchange of Molecular Interaction data (PSI-MI, MIF) and Mass Spectrometry data (PSI-MS, mzData) whilst significant progress was made on PSI-MS' mzIdent, which will allow the capture of data from analytical tools such as peak list search engines. A new focus for PSI (GPS, gel electrophoresis) was explored; as was the need for a common representation of protein modifications by all workers in the field of proteomics and beyond. All these efforts are contextualised by the work of the General Proteomics Standards workgroup; which in addition to the MIAPE reporting guidelines, is continually evolving an object model (PSI-OM) from which will be derived the general standard XML format for exchanging data between researchers, and for submission to repositories or journals.

Mass Spectrometry↗

Mass spectrometry-based proteomics for the detection of plant pathogens.

Plant diseases caused by fungi, oomycetes, viruses, and bacteria are devastating both to the economy and to the food supply of a nation. Therefore, the development of new, rapid methods to identify these pathogens is a highly important area of research that is of international concern. MS-based proteomics has become a powerful and increasingly popular approach to not only identify these pathogens, but also to better understand their biology. However, there is a distinction between identifying a pathogen protein and identifying a pathogen based upon the detection of one of its proteins and this must be considered before the general application of MS for plant pathogen detection is made. There has been a recent push in the proteomics community to make data from large-scale proteomics experiments publicly available in the form of a centralized repository. Such a resource could enable the use of MS as a universal plant pathogen detection technology.

Bacterial Proteins↗

Characterization and comparative analyses of transcriptomes from the normal and neoplastic human prostate.

BACKGROUND: The prostate gland is a highly specialized organ with functional attributes that serve to enhance the fertility of mammalian species. Pathological processes affecting the prostate include benign prostate hypertrophy and prostate carcinoma; diseases that account for major morbidity and mortality in middle-aged and elderly men. To facilitate studies of biological processes uniquely represented in the prostate and assess molecular alterations associated with prostate carcinoma, we sought to establish the diversity of gene expression in the normal and neoplastic prostate through the compilation and analysis of a prostate transcriptome. METHODS: We assembled and annotated ESTs derived from prostate cDNA libraries that were either produced in our laboratory or available from public sequence repositories such as CGAP, dbEST, and Unigene. Determinations of differential gene expression between the normal prostate, other normal tissues, and neoplastic prostate tissues was performed using statistical algorithms. Confirmation of differential expression was performed by quantitative PCR and Northern analysis. RESULTS: A total of 99,448 high-quality ESTs were assembled and annotated to produce a prostate transcriptome comprised of 24,580 distinct TUs. Comparative analyses of gene expression levels identified 61 TUs with exclusive expression in the prostate and 45 TUs with high levels of expression in the prostate relative to at least 25 other normal tissues (P > 0.99). Comparative analyses of ESTs derived from neoplastic prostate tissues identified 75 genes with dysregulated expression in cancer (P > 0.99). CONCLUSIONS: The human prostate expresses a diverse repertoire of genes that reflect a functionally complex organ. The identification of genes with prostate-restricted or enhanced expression may provide additional insights into the biochemical processes that interact to form the developmental, signaling, and metabolic pathways of the normal and neoplastic gland.

Algorithms↗

PhosphaBase: an ontology-driven database resource for protein phosphatases.

PhosphaBase is an ontology-driven database resource containing information on the protein phosphatase family. It is the first public resource dedicated to protein phosphatases, which are enzymes that perform dephosphorylation reactions. In conjunction with the phosphorylation action of protein kinases, phosphatases are involved in important control and communication mechanisms in the cell. They have also been implicated in many human diseases, including diabetes and obesity, cancers, and neurodegenerative conditions. PhosphaBase aims to centralize the growing base of knowledge in the phosphatase research domain. The resource is built around a formal, domain-specific DAML+OIL ontology, and the data are collected from heterogeneous biological sources using Gene Ontology terms as a means of data extraction. The overall ontology-driven architecture provides a robust structure with distinct advantages for sustainability and provides the potential for the development of diagnostic tools, as well as a data repository.

Animals↗

The Protein Coil Library: a structural database of nonhelix, nonstrand fragments derived from the PDB.

Approximately half the structure of folded proteins is either alpha-helix or beta-strand. We have developed a convenient repository of all remaining structure after these two regular secondary structure elements are removed. The Protein Coil Library (http://roselab.jhu.edu/coil/) allows rapid and comprehensive access to non-alpha-helix and non-beta-strand fragments contained in the Protein Data Bank (PDB). The library contains both sequence and structure information together with calculated torsion angles for both the backbone and side chains. Several search options are implemented, including a query function that uses output from popular PDB-culling servers directly. Additionally, several popular searches are stored and updated for immediate access. The library is a useful tool for exploring conformational propensities, turn motifs, and a recent model of the unfolded state.

Algorithms↗

Antifeedant and toxicity effects of thiophenes from four Echinops species against the Formosan subterranean termite, Coptotermes formosanus.

Over 220 crude extracts from repositories generated from plants native to Greece and Kazakhstan were evaluated for termiticidal activity against the Formosan subterranean termite, Coptotermes formosanus Shiraki (Isoptera: Rhinotermitidae). Emerging from this screening effort were bioactive extracts from two Greek species (Echinops ritro L. and Echinops spinosissimus Turra subsp. spinosissimus) and extracts from two Kazakhstan species (Echinops albicaulis Kar. & Kir. and Echinops transiliensis Golosh.). Fractionation and isolation of constituents from the most active extracts from each of the four species has been completed, resulting in the isolation of eight thiophenes possessing varying degrees of termiticidal activity. 2,2':5',2"-Terthiophene and 5'-(3-buten-1-ynyl)-2,2'-bithiophene demonstrated 100% mortality against C. formosanus within 9 days at 1 and 2 wt% concentrations respectively. In addition, all but two of the eight compounds tested were significantly different from the solvent controls in the filter paper consumption bioassay.

Animals↗

Moving from fast to ballistic gradient in liquid chromatography/tandem mass spectrometry pharmaceutical bioanalysis: Matrix effect and chromatographic evaluations.

The paper describes the steps taken by the authors to move from a fast to a ballistic gradient in routine liquid chromatography/tandem mass spectrometry (LC/MS/MS) analysis of plasma samples from pharmacokinetic (PK) profiling of new chemical entities. The reduction of column dimensions from 50 x 4.6 mm to 30 x 2.1 mm followed by optimization of chromatographic separation led to a decrease in the typical runtime from 5 (fast) to 2 min (ballistic) using an API4000 tandem mass spectrometer in Turbo Ionspray mode for detection. Three analytical standards representing typical molecular structures from our sample repository were used to spike plasma from four different species (rat, dog, human and mouse). Two different approaches were used to evaluate matrix effect: post-column infusion and comparison of the peak areas of neat standards and standards spiked after extraction into different pools of plasma; the influence of PEG400 as a typical dosing vehicle was also considered. Two different protein precipitation procedures were taken into account for sample extraction prior to injection. Peak shape, width and height, selectivity and sensitivity of the method were taken into account for chromatographic evaluation. The ballistic method was successfully cross-validated with the conventional fast gradient chromatographic assay.

Animals↗

Interfacing U.S. census map files with statistical graphics software: application and use in epidemiology.

In 1990, the United States Bureau of the Census released detailed geographic map files known as TIGER/Line (Topologically Integrated Geographic Encoding and Referencing). The TIGER files, accessible through purchase or federal repository libraries, contain 24 billion characters of data describing various geographic features including coastlines, hydrography, transportation networks, political boundaries, etc. for the entire United States. Many of these physical features are of potential interest in epidemiological case studies. Unfortunately, the TIGER data base only provides raw alphanumeric data; no utility software, graphical or otherwise, is included. Recently, the S statistical software package has been extended to include a map display function. The map function augments S's high-level approach towards statistical analysis and graphical display of data. Coupling this statistical software with the map data base developed for U.S. Census data collection will facilitate epidemiological research. We discuss the technical background necessary to utilize the TIGER data base for mapping with S. Two types of S maps, segment-based and polygon-based, are discussed along with methods to construct them from TIGER data. Polygon-based maps are useful for displaying regional statistical data, such as disease rates or incidence at the census tract level. Segment-based maps are easier to assemble and are appropriate when the data are not regionalized. Census tract data of AIDS incidence in San Francisco and lung cancer case locations relative to petrochemical refinery sites in Contra Costa County are used to illustrate the methods and potential uses of interfacing the TIGER data base with S.

Acquired Immunodeficiency Syndrome↗

Saccharomyces cerevisiae S288C genome annotation: a working hypothesis.

The S. cerevisiae genome is the most well-characterized eukaryotic genome and one of the simplest in terms of identifying open reading frames (ORFs), yet its primary annotation has been updated continually in the decade since its initial release in 1996 (Goffeau et al., 1996). The Saccharomyces Genome Database (SGD; www.yeastgenome.org) (Hirschman et al., 2006), the community-designated repository for this reference genome, strives to ensure that the S. cerevisiae annotation is as accurate and useful as possible. At SGD, the S. cerevisiae genome sequence and annotation are treated as a working hypothesis, which must be repeatedly tested and refined. In this paper, in celebration of the tenth anniversary of the completion of the S. cerevisiae genome sequence, we discuss the ways in which the S. cerevisiae sequence and annotation have changed, consider the multiple sources of experimental and comparative data on which these changes are based, and describe our methods for evaluating, incorporating and documenting these new data.

Base Sequence↗

Erythrocyte membrane ATP binding cassette (ABC) proteins: MRP1 and CFTR as well as CD39 (ecto-apyrase) involved in RBC ATP transport and elevated blood plasma ATP of cystic fibrosis.

In addition to the better-known roles of the erythrocyte in the transport of oxygen and carbon dioxide, the concept that the red blood cell is involved in the transport and release of ATP has been evolving (J. Luthje, Blut 59, 367, 1989; G. R. Bergfeld and T. Forrester, Cardiovasc. Res. 26, 40, 1992; M. L. Ellsworth et al., Am. J. Physiol. 269, H2155, 1995; R. S. Sprague et al., Am. J. Physiol. 275, H1726, 1998). Membrane proteins involved in the release of ATP from erythrocytes now appear to include members of the ATP binding cassette (ABC) family (C. F. Higgins, Annu. Rev. Cell Biol. 8, 67, 1992; C. F. Higgins, Cell 82, 693, 1995). In addition to defining physiologically the presence of ABC proteins in RBCs, accumulating gel electrophoretic evidence suggests that the cystic fibrosis transmembrane conductance regulator (CFTR) and the multidrug resistance-associated protein (MRP1), respectively, constitute significant proteins in the red blood cell membrane. As such, this finding makes the mature erythrocyte compartment a major mammalian repository of these important ABC proteins. Because of its relative structural simplicity and ready accessibility, the erythrocyte offers an ideal system to explore details of the physiological functions of ABC proteins. Moreover, the presence of different ABC proteins in a single membrane implies that interaction among these proteins and with other membrane proteins may be the norm and not the exception in terms of modulation of their functions.

ATP-Binding Cassette Transporters↗

On the role of the World Health Organization in the development of Sabin vaccines.

The World Health Organization has played a major part in the development, surveillance and distribution of attenuated poliovirus vaccines. At a time when most of the United States' efforts concerned the introduction of Salk-type vaccines, WHO initiated studies that set standards and permitted the large scale trials of Sabin and other attenuated vaccines. Independent expert review validated studies in countries such as the U.S.S.R. which helped lead to the adoption of Sabin vaccines for worldwide usage. Surveillance by WHO Collaborative Centres established the safety of Sabin vaccines and identified issues of reversion primarily concerning type 3 viruses, initiating studies which have elucidated the molecular mechanism of reversion. Efforts by the Biological Unit of the World Health Organization have ensured worldwide acceptable standards to control the safety and manufacture of vaccines. Revision of neurovirulence test methods has ensured adequate safety testing of vaccine lots, reduced the costs of such studies and the numbers of primates needed, important ethical and conservation issues. Finally, the World Health Organization has played a major part in the worldwide supply of vaccines at affordable prices and has been the repository of, and had the exclusive license, to Sabin vaccines since 1972.

Animals↗

CHRONOMERGE: an application for the merging and display of multiple time-stamped data streams.

CHRONOMERGE is a database application that facilitates merging and display of multiple time-stamped data streams. Each stream is a table containing time-stamped values of one or more parameters (such as a panel of laboratory tests) for multiple patients, and is typically created by querying a clinical data respiratory. The data within a single stream therefore represent a pool of multiple time series. The merge operation is complex because of the numerous options to be considered, such as the granularity of the time interval for merge, and the choice of statistical aggregates. CHRONOMERGE combines multiple streams into a single stream based on patient and time, or time alone (if aggregates are to be computed across patients). It allows specification of various options through a graphical user interface and generates appropriate SQL code (or invokes procedural routines) to perform the merge. The resultant stream, or subsets of it, can then be displayed graphically. CHRONOMERGE is intended to facilitate the analysis of time-stamped data that have been extracted from repositories when standard tools (such as the time-series modules of statistics packages) are inadequate.

Algorithms↗

The assessment of biomarkers to detect nephrotoxicity using an integrated database.

Groups of industrial workers exposed to heavy metals (cadmium, mercury, and lead) or solvents were studied together with corresponding control groups. The cohorts were collected from several European centers (countries). Eighty-one measurements were carried out on urine, blood, and serum samples and the results of these analyses together with questionnaire information on each individual were entered into a central database using the relational database package Rbase. After the completion of the database construction phase, the data were exported in a format suitable for analysis by the statistical package SAS. The potential value of each test as an indicator of nephrotoxicity was then assessed. Rigorous exclusion criteria were applied which resulted in the elimination of some tests and samples from the dataset. The measurable contributions of smoking, gender, metal exposure, and site were either singly or in combination assessed by biomarkers for nephrotoxicity. The parameters measured included three urinary enzymes, six specific proteins, total protein, two extracellular matrix markers, four prostaglandins and anti-GBM antibodies, and beta 2-microglobulin in serum. The most sensitive renal tests included the urinary enzymes N-acetyl-beta-D-glucosaminidase (NAG) and intestinal alkaline phosphatase (IAP), brush border antigens, and urinary low-molecular-weight proteins. Of the newer tests investigated the prostaglandins were the most promising. Different patterns of biomarker excretion were observed following exposure to lead, cadmium, or mercury. The dataset provides a unique repository of data which could provide the basis of an enlarging source of information on normal human reference ranges and on the effects of exposure to toxins and the use of biomarkers for monitoring nephrotoxicity.

Biomarkers↗