Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Curation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Osteoid osteoma of the spine: a novel technique using combined computer-assisted and gamma probe-guided high-speed intralesional drill excision.

STUDY DESIGN: A report of five cases of thoracolumbar osteoid osteoma treated with combined computer-assisted and gamma probe-guided high-speed drill excision. OBJECTIVES: To document the surgical technique consisting of a combination of both computer-assisted and gamma probe-guided high-speed drill excision for osteoid osteoma of the spine. SUMMARY OF BACKGROUND DATA: Curative treatment of spinal osteoid osteoma is performed by surgical intralesional excision of the nidus, but intraoperative localization of the nidus is often difficult. Although intraoperative gamma-probe guidance facilitates accurate localization of the nidus, wide surgical resection of the bony structure is still mandatory to ensure removal of the nidus. Computer-assisted surgery has been proven to facilitate surgical intervention in spinal surgery. However, there is no clinical report regarding the application and usefulness of computer-assisted intralesional excision of the osteoid nidus. Excision of the nidus with a computer-assisted high-speed drill and intraoperative gamma probe control may result in complete intralesional excision without sacrificing more bone than necessary. METHODS: One day before surgery, patients were injected with radioactive mTc-oxidronate. With a computed tomography-based electro-optical navigation system, real-time virtual images of the osteoid osteoma were generated by matching the intraoperative surface with preoperative computed tomography images. The osteoid osteoma was excised with the use of an image-guided high-speed drill, and complete excision was controlled with a gamma detection probe. RESULTS: Excision of the nidus was confirmed by relief of symptoms, postexcision computed tomography scans, and histologic evaluation on clinical and radiographic follow-up observation. All five patients reported immediate complete relief of characteristic pain and no evidence of recurrence after 6 to 33 months of follow-up observation. There were no complications. CONCLUSIONS: The combination of both computer-assisted surgery and gamma probe-guided high-speed drill excision for osteoid osteoma of the spine helps to localize and excise the nidus of the osteoid osteoma with minimal bone resection of the posterior spinal structures.

Adolescent↗

Linking biomedical language information and knowledge resources: GO and UMLS.

Integration of various informatics terminologies will be an essential activity towards supporting the advancement of both the biomedical and clinical sciences. The GO consortium has developed an impressive collection of biomedical terms specific to genes and proteins in a variety of organisms. The UMLS is a composite collection of various medical terminologies, pioneered by the National Library of Medicine. In the present study, we examine a variety of techniques for mapping terms from one terminology (GO) to another (UMLS), and describe their respective performances for a small, curated data set attained from the National Cancer Institute, which had precision values ranging from 30% (100% recall) to 95% (74% recall). Based on each technique's performance, we comment on how each can be used to enrich an existing terminology (UMLS) in future studies and how linking biological terminologies to UMLS differs from linking medical terminologies.

Algorithms↗

An integrated system for genetic analysis.

BACKGROUND: Large-scale genetic mapping projects require data management systems that can handle complex phenotypes and detect and correct high-throughput genotyping errors, yet are easy to use. DESCRIPTION: We have developed an Integrated Genotyping System (IGS) to meet this need. IGS securely stores, edits and analyses genotype and phenotype data. It stores information about DNA samples, plates, primers, markers and genotypes generated by a genotyping laboratory. Data are structured so that statistical genetic analysis of both case-control and pedigree data is straightforward. CONCLUSION: IGS can model complex phenotypes and contain genotypes from whole genome association studies. The database makes it possible to integrate genetic analysis with data curation. The IGS web site http://bioinformatics.well.ox.ac.uk/project-igs.shtml contains further information.

Chromosome Mapping↗

No simple dependence between protein evolution rate and the number of protein-protein interactions: only the most prolific interactors tend to evolve slowly.

BACKGROUND: It has been suggested that rates of protein evolution are influenced, to a great extent, by the proportion of amino acid residues that are directly involved in protein function. In agreement with this hypothesis, recent work has shown a negative correlation between evolutionary rates and the number of protein-protein interactions. However, the extent to which the number of protein-protein interactions influences evolutionary rates remains unclear. Here, we address this question at several different levels of evolutionary relatedness. RESULTS: Manually curated data on the number of protein-protein interactions among Saccharomyces cerevisiae proteins was examined for possible correlation with evolutionary rates between S. cerevisiae and Schizosaccharomyces pombe orthologs. Only a very weak negative correlation between the number of interactions and evolutionary rate of a protein was observed. Furthermore, no relationship was found between a more general measure of the evolutionary conservation of S. cerevisiae proteins, based on the taxonomic distribution of their homologs, and the number of protein-protein interactions. However, when the proteins from yeast were assorted into discrete bins according to the number of interactions, it turned out that 6.5% of the proteins with the greatest number of interactions evolved, on average, significantly slower than the rest of the proteins. Comparisons were also performed using protein-protein interaction data obtained with high-throughput analysis of Helicobacter pylori proteins. No convincing relationship between the number of protein-protein interactions and evolutionary rates was detected, either for comparisons of orthologs from two completely sequenced H. pylori strains or for comparisons of H. pylori and Campylobacter jejuni orthologs, even when the proteins were classified into bins by the number of interactions. CONCLUSION: The currently available comparative-genomic data do not support the hypothesis that the evolutionary rates of the majority of proteins substantially depend on the number of protein-protein interactions they are involved in. However, a small fraction of yeast proteins with the largest number of interactions (the hubs of the interaction network) tend to evolve slower than the bulk of the proteins.

Bacterial Proteins↗

Characteristics and regulatory elements defining constitutive splicing and different modes of alternative splicing in human and mouse.

Alternative splicing is a major contributor to genomic complexity, disease, and development. Previous studies have captured some of the characteristics that distinguish alternative splicing from constitutive splicing. However, most published work only focuses on skipped exons and/or a single species. Here we take advantage of the highly curated data in the MAASE database (see related paper in this issue) to analyze features that characterize different modes of splicing. Our analysis confirms previous observations about alternative splicing, including weaker splicing signals at alternative splice sites, higher sequence conservation surrounding orthologous alternative exons, shorter exon length, and more frequent reading frame maintenance in skipped exons. In addition, our study reveals potentially novel regulatory principles underlying distinct modes of alternative splicing and a role of a specific class of repeat elements (transposons) in the origin/evolution of alternative exons. These features suggest diverse regulatory mechanisms and evolutionary paths for different modes of alternative splicing.

Alternative Splicing↗

Meta-QTL Analysis Reveals Consensus Genomic Regions and Candidate Genes for Resistance to Sudden Death Syndrome in Soybean.

Sudden death syndrome (SDS), caused by Fusarium virguliforme, is one of the most economically important diseases limiting soybean production worldwide. Although numerous quantitative trait loci (QTL) associated with SDS resistance have been reported, inconsistencies among mapping populations, marker systems, and experimental conditions have hindered the identification of robust resistance loci for soybean improvement. In this study, a comprehensive meta-analysis was conducted to integrate published QTL and identify stable consensus genomic regions associated with SDS resistance. After a systematic literature survey and data curation, 153 QTL derived from 14 linkage-mapping studies were analyzed using a custom R-based workflow, resulting in the identification of 23 consensus meta-QTL (MQTL) distributed across 17 chromosomes. Several MQTL, particularly those located on chromosomes 6, 8, 18, and 20, were supported by multiple independent studies and represented major genomic hotspots for SDS resistance. Physical localization and functional annotation of these MQTL identified 217 candidate genes, including genes predicted to be involved in plant defense, signal transduction, transcriptional regulation, and secondary metabolism. Gene Ontology enrichment analysis identified response to salicylic acid as the only biological process that remained significant after FDR correction, whereas Kyoto Encyclopedia of Genes and Genomes pathway analysis did not identify significantly enriched pathways. Independent support using five published genome-wide association studies further supported several MQTL, especially those on chromosomes 6, 18, and 20, thereby increasing confidence in these genomic regions. The identified MQTL and prioritized candidate genes provide potential genomic resources for future marker development, improvement applications, and functional validation aimed at improving soybean resistance to SDS.

Fusarium virguliforme↗

International Life Sciences Institute North America Listeria monocytogenes strain collection: development of standard Listeria monocytogenes strain sets for research and validation studies.

Research and development efforts on bacterial foodborne pathogens, including the development of novel detection and subtyping methods, as well as validation studies for intervention strategies can greatly be enhanced through the availability and use of standardized strain collections. These types of strain collections are available for some foodborne pathogens, such as Salmonella and Escherichia coli. We have developed a standard Listeria monocytogenes strain collection that has not been previously available. The strain collection includes (i) a diversity set of 25 isolates chosen to represent a genetically diverse set of L. monocytogenes isolates as well as a single hemolytic Listeria innocua strain and (ii) an outbreak set, which includes 21 human and food isolates from nine major human listeriosis outbreaks that occurred between 1981 and 2002. The diversity set represents all three genetic L. monocytogenes lineages (I, n = 9; II, n = 9; and III, n = 6) as well as nine different serotypes. Molecular subtyping by EcoRI automated ribotyping and pulsed-field gel electrophoresis (PFGE) with AscI and ApaI separated the 25 isolates in the diversity set into 23 ribotypes and 25 PFGE types, confirming that this isolate set represents considerable genetic diversity. Molecular subtyping of isolates in the outbreak set confirmed that human and food isolates were identical by ribotype and PFGE, except for human and food isolates for two outbreaks, which displayed related but distinct PFGE patterns. Subtype and source data for all isolates in this strain collection are available on the Internet and are linked to the PathogenTracker database (www.pathogentracker.com), which allows the addition of new, relevant information on these isolates, including links to publications that have used isolates from this collection. We have thus developed a core L. monocytogenes strain collection, which will provide a resource for L. monocytogenes research and development efforts with centralized Internet-based data curation and integration.

Animals↗

Matching protein beta-sheet partners by feedforward and recurrent neural networks.

Predicting the secondary structure (alpha-helices, beta-sheets, coils) of proteins is an important step towards understanding their three dimensional conformations. Unlike alpha-helices that are built up from one contiguous region of the polypeptide chain, beta-sheets are more complex resulting from a combination of two or more disjoint regions. The exact nature of these long distance interactions remains unclear. Here we introduce two neural-network based methods for the prediction of amino acid partners in parallel as well as anti-parallel beta-sheets. The neural architectures predict whether two residues located at the center of two distant windows are paired or not in a beta-sheet structure. Variations on these architecture, including also profiles and ensembles, are trained and tested via five-fold cross validation using a large corpus of curated data. Prediction on both coupled and non-coupled residues currently approaches 84% accuracy, better than any previously reported method.

Animals↗

Sharing and community curation of mass spectrometry data with Global Natural Products Social Molecular Networking.

The potential of the diverse chemistries present in natural products (NP) for biotechnology and medicine remains untapped because NP databases are not searchable with raw data and the NP community has no way to share data other than in published papers. Although mass spectrometry (MS) techniques are well-suited to high-throughput characterization of NP, there is a pressing need for an infrastructure to enable sharing and curation of data. We present Global Natural Products Social Molecular Networking (GNPS; http://gnps.ucsd.edu), an open-access knowledge base for community-wide organization and sharing of raw, processed or identified tandem mass (MS/MS) spectrometry data. In GNPS, crowdsourced curation of freely available community-wide reference MS libraries will underpin improved annotations. Data-driven social-networking should facilitate identification of spectra and foster collaborations. We also introduce the concept of 'living data' through continuous reanalysis of deposited data.

Biological Products↗

Ontological visualization of protein-protein interactions.

BACKGROUND: Cellular processes require the interaction of many proteins across several cellular compartments. Determining the collective network of such interactions is an important aspect of understanding the role and regulation of individual proteins. The Gene Ontology (GO) is used by model organism databases and other bioinformatics resources to provide functional annotation of proteins. The annotation process provides a mechanism to document the binding of one protein with another. We have constructed protein interaction networks for mouse proteins utilizing the information encoded in the GO annotations. The work reported here presents a methodology for integrating and visualizing information on protein-protein interactions. RESULTS: GO annotation at Mouse Genome Informatics (MGI) captures 1318 curated, documented interactions. These include 129 binary interactions and 125 interaction involving three or more gene products. Three networks involve over 30 partners, the largest involving 109 proteins. Several tools are available at MGI to visualize and analyze these data. CONCLUSIONS: Curators at the MGI database annotate protein-protein interaction data from experimental reports from the literature. Integration of these data with the other types of data curated at MGI places protein binding data into the larger context of mouse biology and facilitates the generation of new biological hypotheses based on physical interactions among gene products.

Animals↗

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

MOTIVATION: An unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. RESULTS: Within 49 h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Machine Learning↗

The molecular biology database collection: an online compilation of relevant database resources.

The Molecular Biology Database Collection represents an effort geared at making molecular biology database resources more accessible to biologists. This online resource, available at http://www.oup.co.uk/nar/Volume_28/Issue_01/html /gkd115_gml.html, is intended to serve as a searchable, up-to-date, centralized jumping-off point to individual Web sites. An emphasis has also been placed on including databases where new value is added to the underlying data by virtue of curation, new data connections, or other innovative approaches.

Databases, Factual↗

The Molecular Biology Database Collection: an updated compilation of biological database resources.

The Molecular Biology Database Collection is an online resource listing key databases of value to the biological community. This Collection is intended to bring fellow scientists' attention to high-quality databases that are available throughout the world, rather than just be a lengthy listing of all available databases. As such, this up-to-date listing is intended to serve as the initial point from which to find specialized databases that may be of use in biological research. The databases included in this Collection provide new value to the underlying data by virtue of curation, new data connections or other innovative approaches. Short, searchable summaries of each of the databases included in the Collection are available through the Nucleic Acids Research Web site, at http://www. nar.oupjournals.org.

Animals↗

The Molecular Biology Database Collection: 2002 update.

The Molecular Biology Database Collection is an online resource listing key databases of value to the biological community. This Collection is intended to bring fellow scientists' attention to high-quality databases that are available throughout the world, rather than just be a lengthy listing of all available databases. As such, this up-to-date listing is intended to serve as the initial point from which to find specialized databases that may be of use in biological research. The databases included in this Collection provide new value to the underlying data by virtue of curation, new data connections or other innovative approaches. Short, searchable summaries and updates for each of the databases included in the Collection are available through the Nucleic Acids Research Web site at http://nar.oupjournals.org.

Computational Biology↗

The Molecular Biology Database Collection: 2003 update.

The Molecular Biology Database Collection is an online resource listing key databases of value to the biological community. This Collection is intended to bring fellow scientists' attention to high-quality databases that are available throughout the world, rather than just be a lengthy listing of all available databases. As such, this up-to-date listing is intended to serve as the jumping-off point from which to find specialized databases that may be of use in advancing biological research. The databases included in this Collection provide new value to the underlying data by virtue of curation, new data connections or other innovative approaches. Short, searchable summaries and updates for each of the databases included in this Collection are available through the Nucleic Acids Research Web site at http://nar.oupjournals.org.

Data Collection↗

Dynamic integration of gene annotation and its application to microarray analysis.

Comprehensive and structured annotations for all genes on a microarray chip are essential for the interpretation of its expression data. Currently, most chip gene annotations are one-line free text descriptions that are often partial, outdated and unsuitable for large-scale data analysis. Therefore the interpretation of microarray gene expression clusters is often limited. Although researchers can manually navigate a collection of databases for better annotations, it is only practical for limited number of genes. Existing meta-databases fail to provide comprehensive categorized annotations for hundreds of genes simultaneously. We have developed an automatic system to address this issue. GeneView system monitors various data sources, extracts gene information from a source whenever it is updated, comprehensively matches genes, and integrates them into a central database by categories, such as pathway, genetic mapping, phenotype, expression profile, domain structure, protein interaction, disease association, and references. The system consists of four major components: (1) relational database; (2) data processing; (3) user curation; (4) data query. We evaluated it by analyzing genes on cDNA and Affymetrix Oligo chips. In both cases, the system provided more accurate and comprehensive information than those provided by the vendors or the chip users, and helped identify new common functions among genes in the same expression clusters.

Cluster Analysis↗

A data integration methodology for systems biology.

Different experimental technologies measure different aspects of a system and to differing depth and breadth. High-throughput assays have inherently high false-positive and false-negative rates. Moreover, each technology includes systematic biases of a different nature. These differences make network reconstruction from multiple data sets difficult and error-prone. Additionally, because of the rapid rate of progress in biotechnology, there is usually no curated exemplar data set from which one might estimate data integration parameters. To address these concerns, we have developed data integration methods that can handle multiple data sets differing in statistical power, type, size, and network coverage without requiring a curated training data set. Our methodology is general in purpose and may be applied to integrate data from any existing and future technologies. Here we outline our methods and then demonstrate their performance by applying them to simulated data sets. The results show that these methods select true-positive data elements much more accurately than classical approaches. In an accompanying companion paper, we demonstrate the applicability of our approach to biological data. We have integrated our methodology into a free open source software package named POINTILLIST.

Informatics↗

Evaluation of text data mining for database curation: lessons learned from the KDD Challenge Cup.

MOTIVATION: The biological literature is a major repository of knowledge. Many biological databases draw much of their content from a careful curation of this literature. However, as the volume of literature increases, the burden of curation increases. Text mining may provide useful tools to assist in the curation process. To date, the lack of standards has made it impossible to determine whether text mining techniques are sufficiently mature to be useful. RESULTS: We report on a Challenge Evaluation task that we created for the Knowledge Discovery and Data Mining (KDD) Challenge Cup. We provided a training corpus of 862 articles consisting of journal articles curated in FlyBase, along with the associated lists of genes and gene products, as well as the relevant data fields from FlyBase. For the test, we provided a corpus of 213 new ('blind') articles; the 18 participating groups provided systems that flagged articles for curation, based on whether the article contained experimental evidence for gene expression products. We report on the evaluation results and describe the techniques used by the top performing groups.

Abstracting and Indexing↗