Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Curation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis↗

Non-curated distributed databases for experimental data and models in neuroscience.

Neuroscience is generating vast amounts of highly diverse data which is of potential interest to researchers beyond the laboratories in which it is collected. In particular, quantitative neuroanatomical data is relevant to a wide variety of areas, including studies of development, aging, pathology and in biophysically oriented computational modelling. Moreover, the relatively discrete and well-defined nature of the data make it an ideal application for developing systems designed to facilitate data archiving, sharing and reuse. At present, the only widely used forms of dissemination are figures and tables in published papers which suffer from inaccessibility and the loss of machine readability. They may also present only an averaged or otherwise selected subset of the available data. Numerous database projects are in progress to address these shortcomings. They employ a variety of architectures and philosophies, each with its own merits and disadvantages. One axis on which they may be distinguished is the degree of top-down control, or curation, involved in data entry. Here we consider one extreme of this scale in which there is no curation, minimal standardization and a wide degree of freedom in the form of records used to document data. Such a scheme has advantages in the ease of database creation and in the equitable assignment of perceived intellectual property by keeping the control of data in the hands of the experts who collected it. It does, however, require a more sophisticated infrastructure than conventional databases since the software must be capable of organizing diverse and differently documented data sets in an effective way. Several components of a software system to provide this infrastructure are now in place. Examples are presented, showing how these tools can be used to archive and publish neuronal morphology data, and how they can give an integrated view of data stored at many different sites.

Animals↗

Prostate cancer progression after therapy of primary curative intent: a review of data from prostate-specific antigen era.

BACKGROUND: Radical prostatectomy and radiotherapy (RT), both radical therapies, are the standard treatments of curative intent for early prostate cancer. However, these therapies are not curative in all patients and, consequently, a substantial proportion of treated patients remain at risk of disease progression and/or cancer-related death. METHODS: This article presents contemporary data on the incidence of prostate-specific antigen (PSA) and clinical disease progression after primary therapy of curative intent in relation to commonly assessed pretreatment or pathologic disease characteristics. RESULTS: The data highlight the substantial risk of progression for certain patient groups, such as those with Gleason score 8-10, cT3 disease, lymph node metastases, and/or pretreatment PSA levels > 20 ng/mL. CONCLUSIONS: Improved and/or additional treatment options are needed for these patient groups.

Brachytherapy↗

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational↗

FINDbase: a relational database recording frequencies of genetic defects leading to inherited disorders worldwide.

Frequency of INherited Disorders database (FINDbase) (http://www.findbase.org) is a relational database, derived from the ETHNOS software, recording frequencies of causative mutations leading to inherited disorders worldwide. Database records include the population and ethnic group, the disorder name and the related gene, accompanied by links to any corresponding locus-specific mutation database, to the respective Online Mendelian Inheritance in Man entries and the mutation together with its frequency in that population. The initial information is derived from the published literature, locus-specific databases and genetic disease consortia. FINDbase offers a user-friendly query interface, providing instant access to the list and frequencies of the different mutations. Query outputs can be either in a table or graphical format, accompanied by reference(s) on the data source. Registered users from three different groups, namely administrator, national coordinator and curator, are responsible for database curation and/or data entry/correction online via a password-protected interface. Databaseaccess is free of charge and there are no registration requirements for data querying. FINDbase provides a simple, web-based system for population-based mutation data collection and retrieval and can serve not only as a valuable online tool for molecular genetic testing of inherited disorders but also as a non-profit model for sustainable database funding, in the form of a 'database-journal'.

Databases, Genetic↗

GOBASE: the organelle genome database.

GOBASE (http://megasun.bch.umontreal.ca/gobase/) is a network-accessible biological database, which is unique in bringing together diverse biological data on organelles with taxonomically broad coverage, and in furnishing data that have been exhaustively verified and completed by experts. So far, we have focused on mitochondrial data: GOBASE contains all published nucleotide and protein sequences encoded by mitochondrial genomes, selected RNA secondary structures of mitochondria-encoded molecules, genetic maps of completely sequenced genomes, taxonomic information for all species whose sequences are present in the database and organismal descriptions of key protistan eukaryotes. All of these data have been integrated and organized in a formal database structure to allow sophisticated biological queries using terms that are inherent in biological concepts. Most importantly, data have been validated, completed, corrected and standardized, a prerequisite of meaningful analysis. In addition, where critical data are lacking, such as genetic maps and RNA secondary structures, they are generated by the GOBASE team and collaborators, and added to the database. The database is implemented in a relational database management system, but features an object-oriented view of the biological data through a Web/Genera-generated World Wide Web interface. Finally, we have developed software for database curation (i.e. data updates, validation and correction), which will be described in some detail in this paper.

Animals↗

ASAP: a resource for annotating, curating, comparing, and disseminating genomic data.

ASAP is a comprehensive web-based system for community genome annotation and analysis. ASAP is being used for a large-scale effort to augment and curate annotations for genomes of enterobacterial pathogens and for additional genome sequences. New tools, such as the genome alignment program Mauve, have been incorporated into ASAP in order to improve display and analysis of related genomes. Recent improvements to the database and challenges for future development of the system are discussed. ASAP is available on the web at https://asap.ahabs.wisc.edu/asap/logon.php.

Databases, Nucleic Acid↗

The Biomolecular Interaction Network Database and related tools 2005 update.

The Biomolecular Interaction Network Database (BIND) (http://bind.ca) archives biomolecular interaction, reaction, complex and pathway information. Our aim is to curate the details about molecular interactions that arise from published experimental research and to provide this information, as well as tools to enable data analysis, freely to researchers worldwide. BIND data are curated into a comprehensive machine-readable archive of computable information and provides users with methods to discover interactions and molecular mechanisms. BIND has worked to develop new methods for visualization that amplify the underlying annotation of genes and proteins to facilitate the study of molecular interaction networks. BIND has maintained an open database policy since its inception in 1999. Data growth has proceeded at a tremendous rate, approaching over 100 000 records. New services provided include a new BIND Query and Submission interface, a Standard Object Access Protocol service and the Small Molecule Interaction Database (http://smid.blueprint.org) that allows users to determine probable small molecule binding sites of new sequences and examine conserved binding residues.

Animals↗

Curative resection for rectal carcinoma: definition influences outcome in terms of local recurrence.

OBJECTIVE: There are wide variations in local recurrence rate following curative surgery for rectal cancer and there are substantial inconsistencies among surgeons regarding the method of defining curative resection. This paper seeks to explore whether defining criteria is one of the important factors driving the variations in outcome. METHOD: A literature review was undertaken to find all UK-based studies that had data on curative resection and local recurrence rates. The studies were divided into groups with distinct definitions of curative resection for rectal cancer. Meta-analyses were performed to pool the risks of local recurrence by group definition. Statistical tests were used to explore the variation in local recurrence by group. Confounding relationships of age, sex, Dukes stage, length of follow-up and year of study were explored as far as possible given the limitations of the available data. RESULTS: For rectal cancers significant differences were found between the pooled local recurrence risks by group definition (P < 0.01). Meta-regression tests including all the studies indicate that the definition of curative resection is an important predictor of local recurrence. CONCLUSION: It is suggested that a standardized approach towards defining curative resection and local recurrence may have a significant effect on outcomes in colorectal cancer surgery and would enable comparisons to be made between different series.

Journal Article↗

High-accuracy SNV calling for bacterial isolates using deep learning with AccuSNV.

Accurate detection of mutations within bacterial species is critical for fundamental studies of microbial evolution, reconstruction of transmission events, and identification of antimicrobial resistance mutations. Although many tools have been developed to identify single-nucleotide variants (SNVs) from whole-genome sequencing, they often suffer from high false-positive rates owing to the complexity of bacterial genomes and the need for different filtering cutoffs across sample types and sequencing depths. As data sets increase in size, the manual filtering required for high accuracy presents a significant obstacle. Here, we present AccuSNV, a novel deep learning-based tool for high-precision and automated bacterial SNV calling. Unlike traditional methods that process one sample at a time, AccuSNV leverages a convolutional neural network (CNN) that integrates alignment information across multiple samples, enhancing precision through learned across-sample patterns. We evaluate AccuSNV against seven popular SNV-calling tools using simulated data from six bacterial species with varied sequencing depths, numbers of isolates, mutations, and divergence levels. To further validate its real-world utility, we test AccuSNV on multiple curated bacterial data sets containing reported SNVs. In both simulated and real-world scenarios, AccuSNV consistently achieves the best performance. Moreover, AccuSNV provides comprehensive user-friendly downstream analysis modules and outputs, including mutation annotation information, phylogenetic inference, d N/d S calculations, and optional manual filtering. Together with the automated deep learning-based calling, these features make AccuSNV broadly accessible to users with different levels of computational expertise.

Deep Learning↗

The Mouse Genome Database (MGD): genetic and genomic information about the laboratory mouse. The Mouse Genome Database Group.

The Mouse Genome Database (MGD) focuses on the integration of mapping, homology, polymorphism and molecular data about the laboratory mouse. Detailed descriptions of genes including their chromosomal location, gene function, disease associations, mutant phenotypes, molecular polymorphisms and links to representative sequences including ESTs are integrated within MGD. The association of information from experiment to gene to genome requires careful coordination and implementation of standardized vocabularies, unique nomenclature constructions, and detailed information derived from multiple sources. This information is linked to other public databases that focus on additional information such as expression patterns, sequences, bibliographic details and large mapping panel data. Scientists participate in the curation of MGD data by generating the Chromosome Committee Reports, consulting on gene family nomenclature revisions, and providing descriptions of mouse strain characteristics and of new mutant phenotypes. MGD is accessible at http://www.informatics.jax.org

Animals↗

PMkbase (version 1.0): an interactive web-based tool for tracking bacterial metabolic traits using phenotype microarrays made interoperable with sequence information and visualizing/processing PM data.

Bacteria showcase remarkable metabolic diversity and traits, even among strains of the same species. In recent years, a large number of bacterial genomes have been sequenced, leading to the elucidation and documentation of genomic differences and commonalities across and within species. Genome-scale metabolic reconstructions, which are often defined and curated using data from phenotype microarrays, elucidate the differences in metabolic traits resulting from genomic diversity. These microarrays measure cellular respiration on a variety of carbon, nitrogen, phosphorus, and sulfur sources and various stressors and inhibitors over a period of time to determine the metabolic activity of a given strain. Despite their popularity in measuring bacterial metabolic activity and traits, no public databases that allow researchers to warehouse, access, and analyze this information currently exist. Additionally, there are no publicly available tools that allow researchers to view the variance of these metabolic traits across bacterial strains. To address this need, we present Phenotype Microarray Knowledgebase (PMkbase [version 1.0], https://pmkbase.com/), an interactive database that acts as a repository of phenotype microarray (PM) data with integrated sequence information. Binarized activity calls, along with associated kinetic parameters, are made for all metabolic substrates and inhibitors. Users can upload their own data for analysis and visualization and to perform quality checks on their experiments. PMkbase will address an unmet need to track and view bacterial metabolic traits and provide researchers with valuable information to develop metabolic models, enrich pangenomic analyses, and design new experiments.IMPORTANCEBacterial species can be differentiated by their metabolic profiles or the type of nutrients they consume. Interestingly, strains within the same species also display differences in nutrient consumption. Phenotype microarrays are a high-throughput, widely used technology to measure which substrates can be metabolized by various microbial strains and the extent to which inhibitors can affect it. Despite their widespread use, public databases to parse and access this data type at scale do not exist. PMkbase, which contains 9,024 data points for nitrogen substrate utilization, 41,664 data points for carbon substrate utilization, 8,448 data points for phosphorus/sulfur substrate utilization, and 27,264 data points on various antibiotics across three species (Escherichia coli, Pseudomonas putida, and Staphylococcus aureus), has been developed to allow researchers to freely access PM data, along with enriching the data with sequence information.

Bacteria↗

FLAGdb++: a database for the functional analysis of the Arabidopsis genome.

FLAGdb++ is dedicated to the integration and visualization of data for high-throughput functional analysis of a fully sequenced genome, as illustrated for Arabidopsis. FLAGdb++ displays the predicted or experimental data in a position-dependent way and displays correlations and relationships between different features. FLAGdb++ provides for a given genome region, summarized characteristics of experimental materials like probe lengths, locations and specificities having an impact upon the confidence we will put in the experimental results. A selected subset of the available information is linked to a locus represented on an easy-to-interpret and memorable graphical display. Data are curated, processed and formatted before their integration into FLAGdb++. FLAGdb++ contains different options for easy back and forth navigation through many loci selected at the start of a session. It includes an original two-component visualization of the data, a genome-wide and a local view, which are permanently linked and display complementary information. Density curves along the chromosomes may be displayed in parallel for suggesting correlations between different structural and functional data. FLAGdb++ is fully accessible at http://genoplante-info.infobiogen.fr/FLAGdb/.

Arabidopsis↗

Hepatic resection for metastatic renal tumors: is it worthwhile?

BACKGROUND: Liver metastases of malignant renal tumors are regarded as having an ominous prognosis because they are infrequently amenable to radical surgery and respond poorly to chemotherapy. Little is known of the outcome of isolated metastases to the liver for which resection is potentially curative. METHODS: Data on 14 patients with liver metastases from renal tumors who underwent a liver resection in a single center between 1982 and 2001 were analyzed retrospectively. RESULTS: There was no operative or postoperative mortality. The median survival was 26 months, with a survival rate of 69% at 1 year and 26% at 3 years. The curative pattern of hepatectomy (2-year survival, 69% vs. 0%; P =.001), an interval between the nephrectomy and the diagnosis of liver metastases in excess of 24 months (2-year survival, 71% vs. 25%; P =.05), tumor size <50 mm (2-year survival, 83% vs. 17%; P =.006), and the possibility of achieving a repeat hepatectomy in the case of recurrence (2-year survival, 100% vs. 21%; P =.02) were associated with a better outcome after the liver resection. Four patients were alive without evidence of disease at 6, 12, 26, and 96 months after the first hepatic resection, and one was alive with hepatic recurrence 18 months after resection. CONCLUSIONS: In patients with liver metastases of malignant renal tumors, an aggressive policy for achieving tumor eradication seems to offer a chance for long-term survival, especially after a long disease-free interval from the nephrectomy. However, despite an aggressive policy for achieving tumor eradication, recurrence frequently occurs after liver resection.

Adult↗

[Adjuvant therapy in gastric cancer].

In Western countries gastric cancer represents the third cause of death even if in the last twenty years the epidemiology of disease has changed. Surgery remains the treatment of choice and overall survival is still 7-15%. Survival data after curative resection are higher in Japan than in Western countries due to a substantial different surgical approach and "early" diagnosis. In both countries adjuvant treatment has been developed to increase the survival rate and different schedules and polipharmacological schemes have been tested. In Japanese trials a statistical significance in survival was observed with chemoimmunotherapy using chemotherapy as control arm. In Western countries data are not conclusive: most trial used surgery as control arm and sample size was not sufficient to show a significant difference between the two arms. The meta-analysis performed up to now have shown a trend of advantage in survival with adjuvant chemotherapy and many objections can be raised concerning the methodology of the same. In fact there are different types of meta-analyses according to whether they are based on the literature (MAL) or individual patient data (MAP or IPD meta-analysis). With an IPD meta-analysis a search is not only done in the literature for all relevant published trials, but also in the scientific community unpublished trials. For all trials, whether published or not, individual patient data on the endpoint of interest are obtained from the investigators. No meta-analysis performed up to now has adopted this methodology. Recently, combined therapy (CT/RT) has shown interesting results with an increase in DFS and OS. At the moment trials results are not sufficient to consider adjuvant chemotherapy the standard treatment: other large trials are required and a combined approach such as RT, IP chemotherapy, neoadjuvant plus adjuvant chemotherapy may be a future research possibility.

Antineoplastic Combined Chemotherapy Protocols↗

Parasite genome databases and web-based resources.

In the last decade, high-throughput genome sequencing and complementary techniques such as microarray and proteomics have generated, and will continue to generate, ever-increasing amounts of data. These technologies of gene discovery, expression, and functional analysis have been applied to a vast array of organisms, including parasites. In most instances, the data are freely available via the Internet, and researchers are becoming increasingly reliant on up-to-date, centralized data repositories to complement wet bench science. This chapter presents an overview of resources relevant to researchers with an interest in para-site genomics and biology. After briefly touching on some of the publicly available nucleotide and protein sequence as well as domain databases, the focus turns to parasite genome projects and associated Web-based resources. A list of parasite sequencing projects current at the time of writing, including relevant Web site addresses, is provided. The available resources range from network sites and project pages at sequencing institutes to databases that integrate and curate sequence data and associated annotation with diverse biological datasets. Particular attention is given to three databases, GeneDB (http://www.genedb.org/), PlasmoDB (http://plasmodb. org/), and tigr db, detailing the scope of each database and the tools available for data querying and retrieval.

Animals↗

Mouse Tumor Biology Database (MTB): status update and future directions.

The Mouse Tumor Biology (MTB) database provides access to data about endogenously arising tumors (both spontaneous and induced) in genetically defined mice (inbred, hybrid, mutant and genetically engineered mice). Data include information on the frequency and latency of mouse tumors, pathology reports and images, genomic changes occurring in the tumors, genetic (strain) background and literature or contributor citations. Data are curated from the primary literature or submitted directly from researchers. MTB is accessed via the Mouse Genome Informatics web site (http://www.informatics.jax.org). Integrated searches of MTB are enabled through use of multiple controlled vocabularies and by adherence to standardized nomenclature, when available. Recently MTB has been redesigned and its database infrastructure replaced with a robust relational database management system (RDMS). Web interface improvements include a new advanced query form and enhancements to already existing search capabilities. The Tumor Frequency Grid has been revised to enhance interactivity, providing an overview of reported tumor incidence across mouse strains and an entrée into the database. A new pathology data submission tool allows users to submit, edit and release data to the MTB system.

Animals↗

Gene expression in stem cell-supporting stromal cell lines.

Cells in the immediate microenvironment together with hematopoietic stem cells (HSCs) constitute the stem cell niche. The microenvironmental or stromal cells provide a complex molecular milieu that helps mediate and balance the self-renewal and commitment potentials of stem cells. The molecules in this milieu are not well defined. In this study, we have intersected previous cDNA subtraction studies with array expression methodologies to define and categorize known gene products expressed by HSC-supportive stromal cell lines. Data were curated from our previously released Stromal Cell Database (StroCDB) containing a set of gene products enriched for expression in the fetal liver stromal cell line AFT024. Global expression analyses were extended to other stem cell-supporting and -nonsupporting fetal liver stromal cell lines using commercially available microarrays. Known and previously described gene products from selected categories were studied: transcription factors, cell membrane proteins, cytoskeleton and related proteins, extracellular matrix proteins, cell adhesion molecules and their corresponding signaling molecules, and cytokines and their related mediators. More than 300 known gene products were selected for expression in HSC-supporting stromal cells compared to nonsupporting lines. Analyses of the data suggest that HSC-supportive cells are immature, sessile, and highly reactive after binding to integrin ligands and cytokines. Therefore, they provide a dynamic space poised to respond to molecular cues elaborated within the stem cell niche. The study provides a survey of known proteins that play key roles in the support of HSCs by fetal liver stromal cells. It also provides insights into the biology of the stem cell niche by highlighting the complex network of intercellular signaling and communication involved in the organization of the niche space.

Animals↗