Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

MedImg: An Integrated Database for Public Medical Images.

The advancements in deep learning algorithms for medical image analysis have garnered significant attention in recent years. While several studies have shown promising results, with models achieving or even surpassing human performance, translating these advancements into clinical practice is still accompanied by various challenges. A primary obstacle lies in the availability of large-scale, well-characterized datasets for validating the generalization of approaches. To address this challenge, we curated a diverse collection of medical image datasets from multiple public sources, containing 105 datasets and a total of 1,995,671 images. These images span 14 modalities, including X-ray, computed tomography, magnetic resonance imaging, optical coherence tomography, ultrasound, and endoscopy, and originate from 13 organs, such as the lung, brain, eye, and heart. Subsequently, we constructed an online database, MedImg, which incorporates and systematically organizes these medical images to facilitate data accessibility. MedImg serves as an intuitive and open-access platform for facilitating research in deep learning-based medical image analysis, accessible at https://www.cuilab.cn/medimg/.

Humans↗

The prevalence of dementia in Europe: a collaborative study of 1980-1990 findings. Eurodem Prevalence Research Group.

To obtain age- and gender-specific estimates of the prevalence of dementia in Europe and to study differences in prevalence across countries, we pooled and re-analysed original data of prevalence studies of dementia carried out in some European countries between 1980 and 1990. The study followed these steps: census of existing datasets, collection of data in a standardized format, selection of datasets suitable for comparison, comparison of age and gender patterns. From the 23 datasets of European surveys considered, 12 were selected for comparison. Only population-based studies in which dementia was defined by DSM-III or equivalent criteria and in which all subjects were examined personally were included. Studies in which institutionalized subjects were not investigated were excluded. Age- and gender-specific prevalences were compared within and across studies and overall prevalences were computed. Although prevalence estimates differed across studies, the general age- and gender-distribution was similar for all studies. The overall European prevalences for the five-year age groups from 60 to 94 years, were 1.0, 1.4, 4.1, 5.7, 13.0, 21.6 and 32.2%, respectively. In subjects under 75 years the prevalence of dementia was slightly higher in men than in women; in those aged 75 years or over the prevalence was higher in women. The prevalence figures nearly doubled with every five years of increase in age.

Adult↗

What lessons can be learned for cancer registration quality assurance from data users? Skin cancer as an example.

BACKGROUND: In cancer registration, data cleaning (i.e. amendments made by data users to datasets released by registries) is potentially informative for quality assurance, but generally underreported. AIM: To assess the scope for learning lessons about cancer registration quality assurance from a data user (using skin cancer as the example). METHODS: The main design features were: (i) A descriptive study identifying, qualitatively and quantitatively, the breadth, depth, and impact of quality assurance issues raised by a user cleaning Merseyside and Cheshire Cancer Registry skin cancer data. Errors were rectified and pitfalls for interpretation were identified. (ii) A nested validation of morphology and site coding on random samples of cutaneous malignant melanomas, basal cell carcinomas (BCC), and squamous cell carcinomas. The 33132-record dataset comprised: all registered skin lesions, except metastases; most recorded variables (about patient, lesion, treatment, outcome); for Merseyside and Cheshire residents diagnosed 1970-1991. RESULTS: (i) Ineligible cases represented 0.3% (97/33132), and were detected best by morphology checks. Most quality assurance issues identified related to local custom and practice, staff training, and computerization, being particularly illustrated by problematic BCC registration practice (e.g. records written over unchallenged by range checks; and idiosyncratic use of variables). (ii) Post-cleaning, morphology coding errors were minimal in the random samples. CONCLUSION: There is great scope for data users to contribute to cancer registration quality assurance. Ultimately, the study dataset appeared fit for epidemiological analysis and important quality assurance messages emerged. Shared explicit standard guidelines for data preparation and validation are needed by users, whose insights could and should be better recognized by cancer registries.

Adult↗

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals↗

Limits of predictive models using microarray data for breast cancer clinical treatment outcome.

Data from microarray studies have been used to develop predictive models for treatment outcome in breast cancer, such as a recently proposed predictive model for antiestrogen response after tamoxifen treatment that was based on the expression ratio of two genes. We attempted to validate this model on an independent cohort of 58 patients with resectable estrogen receptor-positive breast cancer. We measured expression of the genes HOXB13 and IL17BR with real time-quantitative polymerase chain reaction and assessed the association between their expression and outcome by use of univariate logistic regression, area under the receiver-operating-characteristic curve (AUC), a two-sample t test, and a Mann-Whitney test. We also applied standard supervised methods to the original microarray dataset and to another independent dataset from similar patients to estimate the classification accuracy obtainable by using more than two genes in a microarray-based predictive model. We could not validate the performance of the two-gene predictor on our cohort of samples (relation between outcome and the following genes estimated by logistic regression: for HOXB13, odds ratio [OR] = 1.04, 95% confidence interval [CI] = 0.92 to 1.16, P = .54; for IL17BR, OR = 0.69, 95% CI = 0.40 to 1.20, P = .18; and for HOXB13/IL17BR, OR = 1.30, 95% CI = 0.88 to 1.93, P = .18). Similar results were obtained with the AUC, a two-sample two-sided t test, and a Mann-Whitney test. In addition, estimates of classification accuracies applied to two independent microarray datasets highlighted the poor performance of treatment-response predictive models that can be achieved with the sample sizes of patients and informative genes to date.

Adult↗

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans↗

Identifying constraints on the higher-order structure of RNA: continued development and application of comparative sequence analysis methods.

Comparative sequence analysis addresses the problem of RNA folding and RNA structural diversity, and is responsible for determining the folding of many RNA molecules, including 5S, 16S, and 23S rRNAs, tRNA, RNAse P RNA, and Group I and II introns. Initially this method was utilized to fold these sequences into their secondary structures. More recently, this method has revealed numerous tertiary correlations, elucidating novel RNA structural motifs, several of which have been experimentally tested and verified, substantiating the general application of this approach. As successful as the comparative methods have been in elucidating higher-order structure, it is clear that additional structure constraints remain to be found. Deciphering such constraints requires more sensitive and rigorous protocols, in addition to RNA sequence datasets that contain additional phylogenetic diversity and an overall increase in the number of sequences. Various RNA databases, including the tRNA and rRNA sequence datasets, continue to grow in number as well as diversity. Described herein is the development of more rigorous comparative analysis protocols. Our initial development and applications on different RNA datasets have been very encouraging. Such analyses on tRNA, 16S and 23S rRNA are substantiating previously proposed associations and are now beginning to reveal additional constraints on these molecules. A subset of these involve several positions that correlate simultaneously with one another, implying units larger than a basepair can be under a phylogenetic constraint.

Base Sequence↗

A study into the effects of protein binding on nucleotide conformation.

In this study, we examine the effects of binding to protein upon nucleotide conformation, by the comparison of X-ray crystal structures of free and protein-bound nucleotides. A dataset of structurally non-homologous protein-nucleotide complexes was derived from the Brookhaven Protein Data Bank by a novel protocol of dual sequential and structural alignments, and a dataset of native nucleotide structures was obtained from the Cambridge Structural Database. The nucleotide torsion angles and sugar puckers, which describe nucleotide conformation, were analysed in both datasets and compared. Differences between them are described and discussed. Overall, the nucleotides were found to bind in low energy conformations, not significantly different from their 'free' conformations except that they adopted an extended conformation in preference to the 'closed' structure predominantly observed by free nucleotide. The archetypal conformation of a protein-bound nucleotide is derived from these observations.

Crystallography↗

MODBASE, a database of annotated comparative protein structure models.

MODBASE (http://guitar.rockefeller.edu/modbase) is a relational database of annotated comparative protein structure models for all available protein sequences matched to at least one known protein structure. The models are calculated by MODPIPE, an automated modeling pipeline that relies on PSI-BLAST, IMPALA and MODELLER. MODBASE uses the MySQL relational database management system for flexible and efficient querying, and the MODVIEW Netscape plugin for viewing and manipulating multiple sequences and structures. It is updated regularly to reflect the growth of the protein sequence and structure databases, as well as improvements in the software for calculating the models. For ease of access, MODBASE is organized into different datasets. The largest dataset contains models for domains in 304 517 out of 539 171 unique protein sequences in the complete TrEMBL database (23 March 2001); only models based on significant alignments (PSI-BLAST E-value < 10(-4)) and models assessed to have the correct fold are included. Other datasets include models for target selection and structure-based annotation by the New York Structural Genomics Research Consortium, models for prediction of genes in the Drosophila melanogaster genome, models for structure determination of several ribosomal particles and models calculated by the MODWEB comparative modeling web server.

Animals↗

Human non-synonymous SNPs: server and survey.

Human single nucleotide polymorphisms (SNPs) represent the most frequent type of human population DNA variation. One of the main goals of SNP research is to understand the genetics of the human phenotype variation and especially the genetic basis of human complex diseases. Non-synonymous coding SNPs (nsSNPs) comprise a group of SNPs that, together with SNPs in regulatory regions, are believed to have the highest impact on phenotype. Here we present a World Wide Web server to predict the effect of an nsSNP on protein structure and function. The prediction method enabled analysis of the publicly available SNP database HGVbase, which gave rise to a dataset of nsSNPs with predicted functionality. The dataset was further used to compare the effect of various structural and functional characteristics of amino acid substitutions responsible for phenotypic display of nsSNPs. We also studied the dependence of selective pressure on the structural and functional properties of proteins. We found that in our dataset the selection pressure against deleterious SNPs depends on the molecular function of the protein, although it is insensitive to several other protein features considered. The strongest selective pressure was detected for proteins involved in transcription regulation.

Databases, Genetic↗

WormBase: a comprehensive data resource for Caenorhabditis biology and genomics.

WormBase (http://www.wormbase.org), the model organism database for information about Caenorhabditis elegans and related nematodes, continues to expand in breadth and depth. Over the past year, WormBase has added multiple large-scale datasets including SAGE, interactome, 3D protein structure datasets and NCBI KOGs. To accommodate this growth, the International WormBase Consortium has improved the user interface by adding new features to aid in navigation, visualization of large-scale datasets, advanced searching and data mining. Internally, we have restructured the database models to rationalize the representation of genes and to prepare the system to accept the genome sequences of three additional Caenorhabditis species over the coming year.

Animals↗

PartiGeneDB--collating partial genomes.

Owing to the high costs involved, only 28 eukaryotic genomes have been fully sequenced to date. On the other hand, an increasing number of projects have been initiated to generate survey sequence data for a large number of other eukaryotic organisms. For the most part, these data are poorly organized and difficult to analyse. Here, we present PartiGeneDB (http://www.partigenedb.org), a publicly available database resource, which collates and processes these sequence datasets on a species-specific basis to form non-redundant sets of gene objects-which we term partial genomes. Users may query the database to identify particular genes of interest either on the basis of sequence similarity or via the use of simple text searches for specific patterns of BLAST annotation. Alternatively, users can examine entire partial genome datasets on the basis of relative expression of gene objects or by the use of an interactive Java-based tool (SimiTri), which displays sequence similarity relationships for a large number of sequence objects in a single graphic. PartiGeneDB facilitates regular incremental updates of new sequence datasets associated with both new and exisitng species. PartiGeneDB currently contains the assembled partial genomes derived from 1.83 million sequences associated with 247 different eukaryotes.

Databases, Nucleic Acid↗

MPact: the MIPS protein interaction resource on yeast.

In recent years, the Munich Information Center for Protein Sequences (MIPS) yeast protein-protein interaction (PPI) dataset has been used in numerous analyses of protein networks and has been called a gold standard because of its quality and comprehensiveness [H. Yu, N. M. Luscombe, H. X. Lu, X. Zhu, Y. Xia, J. D. Han, N. Bertin, S. Chung, M. Vidal and M. Gerstein (2004) Genome Res., 14, 1107-1118]. MPact and the yeast protein localization catalog provide information related to the proximity of proteins in yeast. Beside the integration of high-throughput data, information about experimental evidence for PPIs in the literature was compiled by experts adding up to 4300 distinct PPIs connecting 1500 proteins in yeast. As the interaction data is a complementary part of CYGD, interactive mapping of data on other integrated data types such as the functional classification catalog [A. Ruepp, A. Zollner, D. Maier, K. Albermann, J. Hani, M. Mokrejs, I. Tetko, U. Güldener, G. Mannhaupt, M. Münsterkötter and H. W. Mewes (2004) Nucleic Acids Res., 32, 5539-5545] is possible. A survey of signaling proteins and comparison with pathway data from KEGG demonstrates that based on these manually annotated data only an extensive overview of the complexity of this functional network can be obtained in yeast. The implementation of a web-based PPI-analysis tool allows analysis and visualization of protein interaction networks and facilitates integration of our curated data with high-throughput datasets. The complete dataset as well as user-defined sub-networks can be retrieved easily in the standardized PSI-MI format. The resource can be accessed through http://mips.gsf.de/genre/proj/mpact.

Databases, Protein↗

VOMBAT: prediction of transcription factor binding sites using variable order Bayesian trees.

Variable order Markov models and variable order Bayesian trees have been proposed for the recognition of transcription factor binding sites, and it could be demonstrated that they outperform traditional models, such as position weight matrices, Markov models and Bayesian trees. We develop a web server for the recognition of DNA binding sites based on variable order Markov models and variable order Bayesian trees offering the following functionality: (i) given datasets with annotated binding sites and genomic background sequences, variable order Markov models and variable order Bayesian trees can be trained; (ii) given a set of trained models, putative DNA binding sites can be predicted in a given set of genomic sequences and (iii) given a dataset with annotated binding sites and a dataset with genomic background sequences, cross-validation experiments for different model combinations with different parameter settings can be performed. Several of the offered services are computationally demanding, such as genome-wide predictions of DNA binding sites in mammalian genomes or sets of 10(4)-fold cross-validation experiments for different model combinations based on problem-specific data sets. In order to execute these jobs, and in order to serve multiple users at the same time, the web server is attached to a Linux cluster with 150 processors. VOMBAT is available at http://pdw-24.ipk-gatersleben.de:8080/VOMBAT/.

Algorithms↗

The Pathway Tools cellular overview diagram and Omics Viewer.

The Pathway Tools cellular overview diagram is a visual representation of the biochemical network of an organism. The overview is automatically created from a Pathway/Genome Database describing that organism. The cellular overview includes metabolic, transport and signaling pathways, and other membrane and periplasmic proteins. Pathway Tools supports interrogation and exploration of cellular biochemical networks through the overview diagram. Furthermore, a software component called the Omics Viewer provides visual analysis of whole-organism datasets using the overview diagram as an organizing framework. For example, gene expression and metabolomics measurements, alone or in combination, can be painted onto the overview, as can computed whole-organism datasets, such as predicted reaction-flux values. The cellular overview and Omics Viewer provide a mechanism whereby biologists can apply the pattern-recognition capabilities of the human visual system to analyze large-scale datasets in a biologically meaningful context. SRI's BioCyc.org website provides overview diagrams for more than 200 organisms. This article describes enhancements to the overview made since a 1999 publication, including the automatic layout capability, expansion of the cellular machinery that it includes, new semantic zooming and poster-generating capabilities, and extension of the Omics Viewer to support painting of metabolites, animations and zooming to individual pathway diagrams.

Computer Graphics↗

A Protein Classification Benchmark collection for machine learning.

Protein classification by machine learning algorithms is now widely used in structural and functional annotation of proteins. The Protein Classification Benchmark collection (http://hydra.icgeb.trieste.it/benchmark) was created in order to provide standard datasets on which the performance of machine learning methods can be compared. It is primarily meant for method developers and users interested in comparing methods under standardized conditions. The collection contains datasets of sequences and structures, and each set is subdivided into positive/negative, training/test sets in several ways. There is a total of 6405 classification tasks, 3297 on protein sequences, 3095 on protein structures and 10 on protein coding regions in DNA. Typical tasks include the classification of structural domains in the SCOP and CATH databases based on their sequences or structures, as well as various functional and taxonomic classification problems. In the case of hierarchical classification schemes, the classification tasks can be defined at various levels of the hierarchy (such as classes, folds, superfamilies, etc.). For each dataset there are distance matrices available that contain all vs. all comparison of the data, based on various sequence or structure comparison methods, as well as a set of classification performance measures computed with various classifier algorithms.

Algorithms↗

CSRDB: a small RNA integrated database and browser resource for cereals.

Plant small RNAs (smRNAs), which include microRNAs (miRNAs), short interfering RNAs (siRNAs) and trans-acting siRNAs (ta-siRNAs), are emerging as significant components of epigenetic processes and of gene networks involved in development and in homeostasis. Here we present a bioinformatics resource for cereal crops, the Cereal Small RNA Database (CSRDB), consisting of large-scale datasets of maize and rice smRNA sequences generated by high-throughput pyrosequencing. The smRNA sequences have been mapped to the rice genome and to the available maize genome sequence and these results are presented in two genome browser datasets using the Generic Genome Browser. Potential RNA targets for the smRNAs have been predicted and access to the resulting smRNA/RNA target pair dataset has been made available through a MySQL based relational database. Various ways to access the data are provided including links from the genome browser to the target database. Data linking and integration are the main focus for this interface, and internal as well as external links are present. The resource is available at http://sundarlab.ucdavis.edu/smrnas/ and will be updated as more sequences become available.

Databases, Nucleic Acid↗

Transformation of expression intensities across generations of Affymetrix microarrays using sequence matching and regression modeling.

The utility of previously generated microarray data is severely limited owing to small study size, leading to under-powered analysis, and failure of replication. Multiplicity of platforms and various sources of systematic noise limit the ability to compile existing data from similar studies. We present a model for transformation of data across different generations of Affymetrix arrays, developed using previously published datasets describing technical replicates performed with two generations of arrays. The transformation is based upon a probe set-specific regression model, generated from replicate measurements across platforms, performed using correlation coefficients. The model, when applied to the expression intensities of 5069 shared, sequence-matched probe sets in three different generations of Affymetrix Human oligonucleotide arrays, showed significant improvement in inter generation correlations between sample-wide means and individual probe set pairs. The approach was further validated by an observed reduction in Euclidean distance between signal intensities across generations for the predicted values. Finally, application of the model to independent, but related datasets resulted in improved clustering of samples based upon their biological, as opposed to technical, attributes. Our results suggest that this transformation method is a valuable tool for integrating microarray datasets from different generations of arrays.

Algorithms↗