Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Curation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

VirGen: a comprehensive viral genome resource.

VirGen is a comprehensive viral genome resource that organizes the 'sequence space' of viral genomes in a structured fashion. It has been developed with the objective of serving as an annotated and curated database comprising complete genome sequences of viruses, value-added derived data and data mining tools. The current release (v1.1) contains 559 complete genomes in addition to 287 putative genomes of viruses belonging to eight viral families for which the host range includes animals and plants. Viral genomes in VirGen are annotated using sequence-based Bioinformatics approaches. The genomic data is also curated to identify 'alternate names' of viral proteins, where available. VirGen archives the results of comparisons of genomes, proteomes and individual proteins within and between viral species. It is the first resource to provide phylogenetic trees of viral species computed using whole-genome sequence data. The module of predicted B-cell antigenic determinants in VirGen is an attempt to link the genome to its vaccinome. Comparative genome analysis data facilitate the study of genome organization and evolution of viruses, which would have implications in applied research to identify candidates for the design of vaccines and antiviral drugs. VirGen is a relational database and is available at http://bioinfo. ernet.in/virgen/virgen.html.

Antigens, Viral↗

Prediction of survival in patients with esophageal carcinoma using artificial neural networks.

BACKGROUND: Accurate estimation of outcome in patients with malignant disease is an important component of the clinical decision-making process. To create a comprehensive prognostic model for esophageal carcinoma, artificial neural networks (ANNs) were applied to the analysis of a range of patient-related and tumor-related variables. METHODS: Clinical and pathologic data were collected from 418 patients with esophageal carcinoma who underwent resection with curative intent. A data base that included 199 variables was constructed. Using ANN-based sensitivity analysis, the optimal combination of variables was determined to allow creation of a survival prediction model. The accuracy (area under the receiver operator characteristic curve [AUR]) of this ANN model subsequently was compared with the accuracy of the conventional statistical technique: linear discriminant analysis (LDA). RESULTS: The optimal ANN models for predicting outcomes at 1 year and 5 years consisted of 65 variables (AUR = 0.883) and 60 variables (AUR = 0.884), respectively. These filtered, optimal data sets were significantly more accurate (P < 0.0001) than the original data set of 199 variables. The majority of ANN models demonstrated improved accuracy compared with corresponding LDA models for 1-year and 5-year survival predictions. Furthermore, ANN models based on the optimal data set were superior predictors of survival compared with a model based solely on TNM staging criteria (P < 0.0001). CONCLUSIONS: ANNs can be used to construct a highly accurate prognostic model for patients with esophageal carcinoma. Sensitivity analysis based on ANNs is a powerful tool for seeking optimal data sets.

Diagnosis, Computer-Assisted↗

Toward the development of a gene index to the human genome: an assessment of the nature of high-throughput EST sequence data.

A rigorous analysis of the Merck-sponsored EST data with respect to known gene sequences increases the utility of the data set and helps refine methods for building a gene index. A highly curated human transcript data base was used as a reference data set of known genes. A detailed analysis of EST sequences derived from known genes was performed to assess the accuracy of EST sequence annotation. The EST data was screened to remove low-quality and low-complexity sequences. A set of high-quality ESTs similar to the transcript data base was identified using BLAST; this subset of ESTs was compared with the set of known genes using the Smith-Waterman algorithm. Error rates of several types were assessed based on a flexible match criterion defining sequence identity. The rate of lane-tracking errors is very low, approximately 0.5%. Insert size data is accurate within approximately 20%. Reversed clone and internal priming error rates are approximately 5% and 2.5%, respectively, contributing to the incorrect identification of reads as 3' ends of genes. Follow-up investigation reveals that a significant number of clones, miscategorized as reversed, represent overlapping genes on the opposite strand of entries in the transcript data base. Relevance of these results to the creation of a high-quality index to the human genome capable of supporting diverse genomic investigations is discussed.

Algorithms↗

PATRIC: the VBI PathoSystems Resource Integration Center.

The PathoSystems Resource Integration Center (PATRIC) is one of eight Bioinformatics Resource Centers (BRCs) funded by the National Institute of Allergy and Infection Diseases (NIAID) to create a data and analysis resource for selected NIAID priority pathogens, specifically proteobacteria of the genera Brucella, Rickettsia and Coxiella, and corona-, calici- and lyssaviruses and viruses associated with hepatitis A and E. The goal of the project is to provide a comprehensive bioinformatics resource for these pathogens, including consistently annotated genome, proteome and metabolic pathway data to facilitate research into counter-measures, including drugs, vaccines and diagnostics. The project's curation strategy has three prongs: 'breadth first' beginning with whole-genome and proteome curation using standardized protocols, a 'targeted' approach addressing the specific needs of researchers and an integrative strategy to leverage high-throughput experimental data (e.g. microarrays, proteomics) and literature. The PATRIC infrastructure consists of a relational database, analytical pipelines and a website which supports browsing, querying, data visualization and the ability to download raw and curated data in standard formats. At present, the site warehouses complete sequences for 17 bacterial and 332 viral genomes. The PATRIC website (https://patric.vbi.vt.edu) will continually grow with the addition of data, analysis and functionality over the course of the project.

Bioterrorism↗

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics↗

[Major urological cancer excisions after the age of 75].

32 patients more than 75 years old had major surgical procedures for urologic neoplasms: radical nephrectomy for renal cell cancer (9 cases), nephroureterectomy for upper urinary tract tumor (3 cases), radical cystectomy for invasive bladder cancer (20 cases). Postoperative mortality was 12.5%. In the nephrectomy group, 3 palliative procedures gave a mean survival time of 7 months. On 9 curative procedures, 7 patients are alive and free of disease with a mean follow-up of 45.6 months. In the cystectomy group, 5 palliative procedures gave a mean survival time of 7 months. On 15 curative procedures, 6 patients are alive and free of disease with a mean follow-up of 18.6 months. Our data confirm that curative procedures can be performed in the elderly. Mean survival time and quality of life after palliative procedures suggest that only true comfort procedures have to be performed.

Age Factors↗

MetaCyc and AraCyc. Metabolic pathway databases for plant research.

MetaCyc (http://metacyc.org) contains experimentally determined biochemical pathways to be used as a reference database for metabolism. In conjunction with the Pathway Tools software, MetaCyc can be used to computationally predict the metabolic pathway complement of an annotated genome. To increase the breadth of pathways and enzymes, more than 60 plant-specific pathways have been added or updated in MetaCyc recently. In contrast to MetaCyc, which contains metabolic data for a wide range of organisms, AraCyc is a species-specific database containing only enzymes and pathways found in the model plant Arabidopsis (Arabidopsis thaliana). AraCyc (http://arabidopsis.org/tools/aracyc/) was the first computationally predicted plant metabolism database derived from MetaCyc. Since its initial computational build, AraCyc has been under continued curation to enhance data quality and to increase breadth of pathway coverage. Twenty-eight pathways have been manually curated from the literature recently. Pathway predictions in AraCyc have also been recently updated with the latest functional annotations of Arabidopsis genes that use controlled vocabulary and literature evidence. AraCyc currently features 1,418 unique genes mapped onto 204 pathways with 1,156 literature citations. The Omics Viewer, a user data visualization and analysis tool, allows a list of genes, enzymes, or metabolites with experimental values to be painted on a diagram of the full pathway map of AraCyc. Other recent enhancements to both MetaCyc and AraCyc include implementation of an evidence ontology, which has been used to provide information on data quality, expansion of the secondary metabolism node of the pathway ontology to accommodate curation of secondary metabolic pathways, and enhancement of the cellular component ontology for storing and displaying enzyme and pathway locations within subcellular compartments.

4-Hydroxyphenylpyruvate Dioxygenase↗

Antibiotic-impregnated bone graft to prevent infection after total hip arthroplasty (ABOGRAFT): protocol for a randomised, double-blind, placebo-controlled trial.

INTRODUCTION: Studies have shown promising results using bone graft as a carrier for local administration of antibiotics to reduce the risk of prosthetic joint infection (PJI). The objective of this clinical trial is to determine if tobramycin and vancomycin-impregnated bone graft is safe and effective in reducing the rate of PJI after total hip arthroplasty (THA). METHODS AND ANALYSIS: This study is an international, randomised, double-blinded, placebo-controlled clinical drug trial. Patients scheduled for THA (n=1100) requiring bone grafting (excluding revisions due to an ongoing infection) are randomised in a 1:1 ratio to prophylactic treatment with tobramycin and vancomycin or placebo-impregnated bone graft.The primary outcome is the time to reoperation due to infection or diagnosis of PJI, expressed as a relative risk difference between the two groups. A risk reduction of at least 50% is considered clinically relevant. Secondary outcomes are time to and reason for reoperation and implant revision, type of micro-organism and antibiotic susceptibility pattern within 2 and 5 years after surgery. Safety outcomes are the number of adverse events and revision rate due to aseptic loosening. The primary analysis will be performed using proportional hazard models. ETHICS AND DISSEMINATION: The study has been approved under the Clinical Trial Regulation No 536/2014 (EU CT; 2024-510921-25-00). Results will be published in open-access peer-reviewed journals and disseminated to patient organisations and the media, and de-identified individual participant data will be curated and shared on reasonable request in accordance with the Findability, Accessibility, Interoperability and Reuse principles, subject to the laws and regulations governing data protection in each participating country. TRIAL REGISTRATION NUMBER: NCT05169229.

Humans↗

Regression analysis of prognostic factors in colorectal cancer after curative resections.

The clinical, laboratory, and pathologic data of 310 patients who had curative resections were prospectively collected and analyzed in a multiple stepwise regression model. Although several factors (i.e., venous invasion) were of importance in univariate analysis, the following conclusions reflect the outcome and relative importance of the regression analysis only. Blood loss as an initial symptom and duration of symptoms were associated with a better prognosis. Location of the primary tumor, age, and sex did not appear to have prognostic value. Observations during operation such as palpable lymph nodes, fixity to adjacent organs, and tumor spill were related to a diminished tumor-free survival. Laboratory data (hemoglobin, leukocytes, ESR, GGTP, SGOT, SGPT, LDH, total protein, CEA) were tested for their potential prognostic values. Only a preoperative low protein level or an elevated CEA level were associated with an increased risk of death due to recurrent tumor. The histopathologic features (stage and grade), with the exception of venous invasion, were of relative importance in the determination of prognosis. The aforementioned variables can be included in a prognostic index on the base of which high-risk groups suitable for adjuvant studies can be identified.

Adenocarcinoma↗

The COSMIC (Catalogue of Somatic Mutations in Cancer) database and website.

The discovery of mutations in cancer genes has advanced our understanding of cancer. These results are dispersed across the scientific literature and with the availability of the human genome sequence will continue to accrue. The COSMIC (Catalogue of Somatic Mutations in Cancer) database and website have been developed to store somatic mutation data in a single location and display the data and other information related to human cancer. To populate this resource, data has currently been extracted from reports in the scientific literature for somatic mutations in four genes, BRAF, HRAS, KRAS2 and NRAS. At present, the database holds information on 66 634 samples and reports a total of 10 647 mutations. Through the web pages, these data can be queried, displayed as figures or tables and exported in a number of formats. COSMIC is an ongoing project that will continue to curate somatic mutation data and release it through the website.

Databases, Factual↗

FDG-PET improves the staging and selection of patients with recurrent colorectal cancer.

Whole-body fluorine-18 fluorodeoxyglucose positron emission tomography (FDG-PET) has proved effective in the diagnosis and staging of recurrent colorectal cancer. In this study, we analysed how PET affects the management of patients with recurrent colorectal cancer by permitting more accurate selection of candidates for curative resection. The data of 79 patients with known or suspected recurrent colorectal cancer were analysed. Conventional imaging modalities (CIM) and PET results were compared with regard to their accuracy in determining the extent and the resectability of tumour recurrence. Recurrence was demonstrated in 68 of the 79 patients. The data indicate that PET was superior to CIM for detection of recurrence at all sites except the liver. Based on the CIM+PET staging, surgery with curative intent was proposed in 39 patients and was indeed achieved in 31 of them (80%). PET was more accurate than CIM alone in predicting the resectability or non-resectability of the recurrence (82% vs 68%, P=0.02). It is concluded that whole-body FDG-PET is highly sensitive for both the diagnosis and the staging of patients with recurrent colorectal cancer. Its use in conjunction with conventional imaging procedures results in a more accurate selection of patients for surgical treatment with curative intent.

Colorectal Neoplasms↗

Lymph node retrieval in a randomized trial on western-type versus Japanese-type surgery in gastric cancer.

PURPOSE: In the tumor-node-metastasis (TNM) staging system, no recommendations are provided on what lymph node retrieval technique is to be used to determine lymph node status, which leads to variability in nodal status assessment and TNM staging. PATIENT AND METHODS: Lymph node retrieval was quantitated using data from 237 curatively resected gastric cancer patients, from a prospective, randomized trial that compared the Western resection with limited (D1) and the Japanese resection with extended lymphadenectomy (D2), and compared data from the literature. Moreover, the efficacy of different lymph node retrieval techniques was determined. RESULTS: The mean yield of lymph nodes was 15 in D1 and 30 in D2, which is similar to results from German investigators, but substantially lower than results from Japanese investigators (60 in D2). Use of a fat-clearance technique significantly increased (P = .01) nodal yields compared with conventional retrieval. Significantly higher yields (P < .001) were obtained by a Japanese surgeon using conventional retrieval directly postoperatively. Experience of surgicopathologic teams with processing resection specimens did not influence nodal yields. Further analysis showed that reference values for nodal yields per anatomically defined station as reported in the literature were contradicted by our results and indicated the ambiguity of such standards. CONCLUSION: Despite some anatomical variability in the distribution of lymph nodes, advice on the number of nodes to examine per N level, feasible in all patients, should be incorporated into the TNM classification to standardize nodal status assessment. Based on our findings, we advocate retrieval of nodes immediately postoperatively by the surgeon.

Europe↗

Public informatics resources for rice and other grasses.

As an emerging model system, rice will benefit from an informatics infrastructure which organizes genome data and makes it available worldwide. RiceGenes and other Internet-accessible resources are evolving to meet these goals. Grass crops such as rice, maize, millet, sorghum and wheat are closely related but are represented by independent database projects; interlinking these resources would create a broad view of grass genetics and make it easier to compare data across genomes. The future success of grass informatics depends on the development of new comparative mapping displays as well as the participation of the research community in assembling and curating comparative map data.

Databases, Factual↗

Clinical audit in radiation oncology: results from one centre.

This study aims to determine workload statistics and to document patterns of fractionation in a single centre in two time periods separated by 4 years. Patient, tumour and treatment-related data were collected for courses of radiation treatment that were commenced within two 6-month periods in both 1988 and 1993. In both time periods, 45-49% of patients were treated with curative intent. Of these, one-third were irradiated definitively and two-thirds in an adjuvant setting. Most of the remainder were treated with palliative intent. A few were treated for non-neoplastic conditions. The re-treatment rate in 1993 was 13%. In both time periods, breast and lung tumours represented approximately 20% each of the total treatment courses. Skin, head and neck, gynaecological, urological and haematological primary tumours accounted for 5-10% each. Treatment intents differed markedly for different primary sites. For example, in 1993 65% of patients with breast primaries were treated curatively compared with 6% of patients with lung primaries. Treatment schedules for curative intent were similar in both time periods and for the majority of treatment sites. Median fraction numbers were 25 (excluding skin primaries), reflecting conventional daily fractionation. Treatment schedules for palliation showed greater variation and there was a trend towards shorter treatment courses in 1993. For palliative treatment of bone, brain and lung, from either primary or metastatic disease, treatment schedules with 10-15 fractions were used most frequently in 1988. In 1993, however, the majority of patients received 1-5 fractions. In 1993, the breakdown of techniques according to treatment intent showed that for treatment with curative intent, single, parallel opposed and more complex field arrangements were used in 27% (includes skin primaries), 12% and 61% of treatment courses, respectively, compared with 29%, 59% and 12%, respectively, for palliative treatment courses. In 1993, one-third of patients receiving radiation treatment lived in the local health area. Patients living in areas with rural postcodes were more likely to receive palliative irradiation and had a higher incidence of melanoma than patients living in areas with Sydney metropolitan postcodes. As approximately 50% of patients were treated with palliative intent, changes in the fractionation patterns used can alter significantly the utilization and availability of megavoltage equipment. However, any reduction in attendances caused by hypofractionation for palliation may be offset by the trend to use hyperfractionation for curative treatments. The data support the hypothesis of reduced availability and use of radiation therapy in patients with cancer from rural areas.

Adolescent↗

Whole genome sequencing reveals a specific microbiota in subglottic stenosis C. acnes may contribute to inflammation.

PURPOSE: Subglottic stenosis (SGS) progressively reduces the airway below the vocal folds. The cause is not known and there is a recurrent need of surgical treatment. Including all phenotypes, SGS affects 1/400 000/yr, with a female dominance. Previous studies have revealed a possible role of the Mycobacterium complex in SGS development. Our hypothesis is that microbiota is associated with the inflammation in SGS, if true it might affect the prevailing treatment options. METHODS: This prospective cross-sectional study included biopsies from 34 patients with subglottic stenosis, collected between 2020 and 2023. Nucleic acids were extracted from the tissue samples and analysed using whole genome sequencing. Microbial composition was characterized using taxonomic profiling of sequencing data. Species with sufficient read counts were selected for further validation using sequence alignment methods to ensure accuracy of identification. RESULTS: Using the most comprehensive form of genomic testing currently in clinical use, we present curated and stable data on the presence of Cutibacterium acnes in 28 out of the 34 cases. CONCLUSION: Cutibacterium acnes may serve as a driver of the inflammation characterizing SGS and should be considered in therapeutically oriented future studies.

Cutibacterium acnes↗

A7DB: a relational database for mutational, physiological and pharmacological data related to the alpha7 nicotinic acetylcholine receptor.

BACKGROUND: Nicotinic acetylcholine receptors (nAChRs) are pentameric proteins that are important drug targets for a variety of diseases including Alzheimer's, schizophrenia and various forms of epilepsy. One of the most intensively studied nAChR subunits in recent years has been alpha7. This subunit can form functional homomeric pentamers (alpha7)5, which can make interpretation of physiological and structural data much simpler. The growing amount of structural, pharmacological and physiological data for these receptors indicates the need for a dedicated and accurate database to provide a means to access this information in a coherent manner. DESCRIPTION: A7DB http://www.lgics.org/a7db/ is a new relational database of manually curated experimental physiological data associated with the alpha7 nAChR. It aims to store as much of the pharmacology, physiology and structural data pertaining to the alpha7 nAChR. The data is accessed via web interface that allows a user to search the data in multiple ways: 1) a simple text query 2) an incremental query builder 3) an interactive query builder and 4) a file-based uploadable query. It currently holds more than 460 separately reported experiments on over 85 mutations. CONCLUSIONS: A7DB will be a useful tool to molecular biologists and bioinformaticians not only working on the alpha7 receptor family of proteins but also in the more general context of nicotinic receptor modelling. Furthermore it sets a precedent for expansion with the inclusion of all nicotinic receptor families and eventually all cys-loop receptor families.

Cholinergic Agents↗

The genexpress IMAGE knowledge base of the human muscle transcriptome: a resource of structural, functional, and positional candidate genes for muscle physiology and pathologies.

Sequence, gene mapping, and expression data corresponding to 910 genes transcribed in human skeletal muscle have been integrated to form the muscle module of the Genexpress IMAGE Knowledge Base. Based on cDNA array hybridization, a set of 14 transcripts preferentially or specifically expressed in muscle have been selected and characterized in more detail: Their pattern of expression was confirmed by Northern blot analysis; their structure was further characterized by full-insert cDNA sequencing and cDNA extension; the map location of the corresponding genes was refined by radiation hybrid mapping. Five of the 14 selected genes appear as interesting positional and functional candidate genes to study in relation with muscle physiology and/or specific orphan muscular pathologies. One example is discussed in more detail. The expression profiling data and the associated Genexpress Index2 entries for the 910 genes and the detailed characterization of the 14 selected transcripts are available from a dedicated Web server at. The database has been organized to provide the users with a working space where they can find curated, annotated, integrated data for their genes of interest. Different navigation routes to exploit the resource are discussed.

Base Sequence↗

Meta2DB: curated shotgun metagenomic feature sets and metadata for health state prediction.

SUMMARY: Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13&#x2009;897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health. AVAILABILITY: Data is available at https://gdo-meta2db.llnl.gov/ and https://zenodo.org/records/17315984.

Metadata↗