Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

NMRb: a web-site repository for raw NMR datasets.

UNLABELLED: The development of NMR in structural proteomics requires the availability of automatic structure determination methods. Many researchers are commonly confronted with the lack of raw datasets during the validation step of such methods. In order to increase test possibilities, the NMRb web-site offers a database of NMR raw datasets, ordered by spectral characteristics. AVAILABILITY: NMRb is available from: http://nmrb.cbs.cnrs.fr. SUPPLEMENTARY INFORMATION: General organization of NMRb figure, relational model organization, and XML structure files are available from http://nmrb.cbs.cnrs.fr/nmrb-doc.html.

Database Management Systems↗

Discovery of stable and significant binding motif pairs from PDB complexes and protein interaction datasets.

MOTIVATION: Discovery of binding sites is important in the study of protein-protein interactions. In this paper, we introduce stable and significant motif pairs to model protein-binding sites. The stability is the pattern's resistance to some transformation. The significance is the unexpected frequency of occurrence of the pattern in a sequence dataset comprising known interacting protein pairs. Discovery of stable motif pairs is an iterative process, undergoing a chain of changing but converging patterns. Determining the starting point for such a chain is an interesting problem. We use a protein complex dataset extracted from the Protein Data Bank to help in identifying those starting points, so that the computational complexity of the problem is much released. RESULTS: We found 913 stable motif pairs, of which 765 are significant. We evaluated these motif pairs using comprehensive comparison results against random patterns. Wet-experimentally discovered motifs reported in the literature were also used to confirm the effectiveness of our method. SUPPLEMENTARY INFORMATION: http://sdmc.i2r.a-star.edu.sg/BindingMotifPairs.

Algorithms↗

Meta-analysis based on control of false discovery rate: combining yeast ChIP-chip datasets.

MOTIVATION: High-throughput microarray technology can be used to examine thousands of features, such as all the genes of an organism, and measure their expression. Two important issues of microarray bioinformatics are first, how to combine the significance values for each feature across experiments with high statistical power, and second, how to control the proportion of false positives. Existing methods address these issues separately, in spite of their linked usage. RESULTS: We present a novel method (ESP) to address the two requirements in an interdependent way. It generalizes the truncated product method of Zaykin et al. to combine only those significance values which clear their respective experiment-specific false discovery restrictive thresholds, thus allowing us to control the false discovery rate (FDR) for the final combined result. Further, we introduce several concepts that together offer FDR control, high power, quality control and speed-up in meta-analysis as done by our algorithm. Computational and statistical methods of research synthesis like the one described here will be increasingly important as additional genome-wide datasets accumulate in databases. We apply our method to combine three well-known ChIP-chip transcription factor binding datasets for budding yeast to identify significant intergenic regulatory sequences for nine cell cycle regulating transcription factors, both with high power and controlled FDR.

Algorithms↗

An anaesthetic minimum dataset and report format. Society for Computing and Technology in Anaesthesia (SCATA). European Society for Computing and Technology in Anaesthesia (ESCTAIC).

The dataset necessary to produce reports for anaesthetic training purposes is described, together with appropriate definitions. The format for a standard report that may be used in a logbook is also described. These have been accepted by the Royal College of Anaesthetists. The German Anaesthetic Society (Deutsche Gesellschaft für Anaesthesiologie und Intensivmedizin, DGAI) has accepted the dataset and definitions.

Anesthesiology↗

Measuring quality of care: fundamental information from administrative datasets.

Under proposals for national health insurance reform in the USA, employers and purchasing cooperatives will have to measure the quality of health care services. Their need for data systems upon which to base their decisions has stimulated dramatic innovation and rapid change in how health care information is collected, integrated from multiple sources, and reported. To make administrative data useful for quality measurement, careful attention must be given to information about: medical care utilization; patient characteristics; provider characteristics; and health plans. In this paper, we describe the extent to which this information is included in existing administrative datasets. We then suggest how planned datasets should be designed so they can be used to assess the quality of health care.

Databases, Factual↗

Medullary serotonergic network deficiency in the sudden infant death syndrome: review of a 15-year study of a single dataset.

The sudden infant death syndrome (SIDS) is the leading cause of postneonatal infant mortality in the United States today, despite a dramatic 38% decrease in incidence due to a national risk reduction campaign advocating the supine sleep position. Our research in SIDS brainstems, beginning in 1985 and involving a single, large dataset, has become increasingly focused upon a specific neurotransmitter (serotonin) and specific territories (ventral medulla and regions of the medullary reticular formation that contain secrotonergic neurons). Based on this research, we propose that SIDS, or a subset of SIDS, is due to a developmental abnormality in a medullary network composed of (at least in part) rhombic lip-derived, serotonergic neurons, including in the caudal raphé and arcuate nucleus (putative human homologue of the cat respiratory chemosensitive fields); and this abnormality results in a failure of protective responses to life-threatening stressors (e.g. asphyxia, hypoxia, hypercapnia) during sleep as the infant passes through a critical period in homeostatic control. We call this the medullary serotonergic network deficiency hypothesis. We review the triple-risk model for SIDS, the development of the dataset using tissue autoradiography for analyzing neurotransmitter receptor binding; age-dependent baseline neurochemical findings in the human brainstem during early life; the evidence for serotonergic, rhombic lip, and ventral medullary deficits in at least some SIDS victim; possible mechanisms of sudden infant death related to these deficits; and potential causes of the deficits in the medullary serotonergic network in SIDS victims. We conclude with a summary of future directions in SIDS brainstem research.

Animals↗

BioGRID: a general repository for interaction datasets.

Access to unified datasets of protein and genetic interactions is critical for interrogation of gene/protein function and analysis of global network properties. BioGRID is a freely accessible database of physical and genetic interactions available at http://www.thebiogrid.org. BioGRID release version 2.0 includes >116 000 interactions from Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster and Homo sapiens. Over 30 000 interactions have recently been added from 5778 sources through exhaustive curation of the Saccharomyces cerevisiae primary literature. An internally hyper-linked web interface allows for rapid search and retrieval of interaction data. Full or user-defined datasets are freely downloadable as tab-delimited text files and PSI-MI XML. Pre-computed graphical layouts of interactions are available in a variety of file formats. User-customized graphs with embedded protein, gene and interaction attributes can be constructed with a visualization system called Osprey that is dynamically linked to the BioGRID.

Animals↗

Combining experimental and predicted datasets for determination of the subcellular location of proteins in Arabidopsis.

Substantial experimental datasets defining the subcellular location of Arabidopsis (Arabidopsis thaliana) proteins have been reported in the literature in the form of organelle proteomes built from mass spectrometry data (approximately 2,500 proteins). Subcellular location for specific proteins has also been published based on imaging of chimeric fluorescent fusion proteins in intact cells (approximately 900 proteins). Further, the more diverse history of biochemical determination of subcellular location is stored in the entries of the Swiss-Prot database for the products of many Arabidopsis genes (approximately 1,800 proteins). Combined with the range of bioinformatic targeting prediction tools and comparative genomic analysis, these experimental datasets provide a powerful basis for defining the final location of proteins within the wide variety of subcellular structures present inside Arabidopsis cells. We have analyzed these published experimental and prediction data to answer a range of substantial questions facing researchers about the veracity of these approaches to determining protein location and their interrelatedness. We have merged these data to form the subcellular location database for Arabidopsis proteins (SUBA), providing an integrated understanding of protein location, encompassing the plastid, mitochondrion, peroxisome, nucleus, plasma membrane, endoplasmic reticulum, vacuole, Golgi, cytoskeleton structures, and cytosol (www.suba.bcs.uwa.edu.au). This includes data on more than 4,400 nonredundant Arabidopsis protein sequences. We also provide researchers with an online resource that may be used to query protein sets or protein families and determine whether predicted or experimental location data exist; to analyze the nature of contamination between published proteome sets; and/or for building theoretical subcellular proteomes in Arabidopsis using the latest experimental data.

Arabidopsis↗

Three-dimensional texture analysis of MRI brain datasets.

A method is proposed for three-dimensional (3-D) texture analysis of magnetic resonance imaging brain datasets. It is based on extended, multisort co-occurrence matrices that employ intensity, gradient and anisotropy image features in a uniform way. Basic properties of matrices as well as their sensitivity and dependence on spatial image scaling are evaluated. The ability of the suggested 3-D texture descriptors is demonstrated on nontrivial classification tasks for pathologic findings in brain datasets.

Brain↗

Bayesian class discovery in microarray datasets.

A novel approach to class discovery in gene expression datasets is presented. In the context of clinical diagnosis, the central goal of class discovery algorithms is to simultaneously find putative (sub-)types of diseases and to identify informative subsets of genes with disease-type specific expression profile. Contrary to many other approaches in the literature, the method presented implements a wrapper strategy for feature selection, in the sense that the features are directly selected by optimizing the discriminative power of the used partitioning algorithm. The usual combinatorial problems associated with wrapper approaches are overcome by a Bayesian inference mechanism. On the technical side, we present an efficient optimization algorithm with guaranteed local convergence property. The only free parameter of the optimization method is selected by a resampling-based stability analysis. Experiments with Leukemia and Lymphoma datasets demonstrate that our method is able to correctly infer partitions and corresponding subsets of genes which both are relevant in a biological sense. Moreover, the frequently observed problem of ambiguities caused by different but equally high-scoring partitions is successfully overcome by the model selection method proposed.

Algorithms↗

Features affecting Cas9-induced editing efficiency and patterns in tomato: evidence from a large CRISPR dataset.

CRISPR/Cas9 is a cornerstone of plant genome editing, yet the determinants of editing efficiency for a given single-guide RNAs (sgRNAs) and DNA double-strand break (DSB) repair outcomes remain poorly understood, particularly in plants. Here, we generated a large experimental dataset comprising 420 sgRNAs targeting promoters, exons, and introns of 137 genes in tomato protoplasts, and quantified editing efficiency and repair footprints together with chromatin accessibility and transcriptional state in the same cellular context. Editing efficiency was consistently higher at targets in accessible chromatin and modestly higher in promoters and introns than in exons, whereas transcriptional activity had no detectable effect. Editing efficiencies were more similar among sgRNAs targeting the same gene than among different genes, revealing a local genomic influence on Cas9 activity. A distinct subset of sgRNAs achieved near-complete editing and produced characteristic repair footprints dominated by long deletions with extended microhomology tracts, indicative of microhomology-mediated end joining (MMEJ), resembling patterns associated with high-efficiency guides in human cells, and suggesting conserved sequence-driven repair biases across species. In contrast, widely used human-trained prediction models failed to accurately rank sgRNA performance in plants, highlighting the limits of cross-species predictability. Together, this dataset provides a resource for improving guide design and mechanistic understanding of plant DNA repair.

Solanum lycopersicum↗

Automated seed localization from CT datasets of the prostate.

With the increasing utilization of permanent brachytherapy implants for treating carcinoma of the prostate, the importance of accurate post-treatment dose calculation also increases for assessing patient outcome and planning future treatments. An automatic method for seed localization of permanent brachytherapy implants, using CT datasets of the prostate, has been developed and tested on a phantom using an actual patient planned seed distribution. This method was also compared to results with the three-film technique for three patient datasets. The automatic method is as accurate or more accurate than the three film technique for 1 mm, 3 mm, and 5 mm contiguous CT slices, and eliminates the inter- and intra-observer variability of the manual methods. The automated method improves the localization of brachytherapy seeds while reducing the time required for the user to input information, and is demonstrated to be less operator dependent, less time consuming, and potentially more accurate than the three-film technique.

Biophysical Phenomena↗

Upper gastrointestinal cancer pathology reporting: a regional audit to compare standards with minimum datasets.

AIMS: Accurate pathological (pTNM) staging of oesophageal and gastric cancer provides important prognostic information. The aim of this study was to compare the standard of pathology reporting of oesophageal and gastric cancer resections from a cancer network with standards set by the Royal College of Pathologists. METHODS: All reports for oesophageal and gastric cancer resections from the five hospitals in the cancer network in 2001 were collected. Individual items of information were compared with minimum datasets provided by the Royal College of Pathologists. Items were classified as "complete", "partially complete", or "absent". RESULTS: One hundred and ten reports were audited (54 oesophageal and 56 gastric). Fourteen gastric and 17 oesophagectomy reports were over 75% complete. Clinically important missing data occurred most frequently for the pM component of TNM staging (pMx omitted in 87 reports) and completeness of resection expressed as a bold statement (absent in 50 reports). Twelve reports could not be classified because the specimen contained no residual tumour after neoadjuvant treatment. CONCLUSION: The use of a standard proforma for reporting upper gastrointestinal cancers based on a minimum dataset provided by the Royal College of Pathologists is recommended, with modifications to allow for specimens with no tumour after neoadjuvant treatment.

England↗

Essential dataset for ambulatory ear, nose, and throat care in general practice: an aid for quality assessment.

OBJECTIVE: To describe the documentation of care for the usual range of ear, nose, and throat (ENT) problems seen in primary care as a basis for developing a computerised information system to aid quality assessment. DESIGN: Descriptive study of the pattern of ENT problems and diagnoses and treatment as recorded in individual case notes. SETTING: The primary health care centre in Mjölby, Sweden. PATIENTS: Consultations for ENT problems from a 10% sample randomly selected from all consultations (n = 22,600) in one year. From this sample 375 consultations for ENT problems (16% of all consultations) by 272 patients were identified. MAIN MEASURES: The detailed documentation of each consultation was retrieved from the individual records and compared with the data required for a computer based information system designed to help in quality management. RESULTS: Although the overall picture gained from the data retrieved from the notes suggested that ENT care was probably adequate, the recorded details were limited. The written case notes were insufficient when compared with the details required for a computerised system based on an essential dataset designed to allow assessment of diagnostic accuracy and appropriateness of treatment of ENT problems in primary care. CONCLUSION: There is a gap between the amount and the type of information needed for accurate and useful quality assessment and that which is normally included in case notes. More detailed information is needed if general practitioners' notes are to be used for regular quality assessment of ENT problems but that would mean more time spent on keeping notes. This would be difficult to justify. IMPLICATIONS: The routine information systems used at this primary healthcare centre did not produce sufficient documentation for quality assessment of ENT care. This dilemma might be resolved by specially designed desktop computer software accessed through an essential dataset.

Adolescent↗

A comparison of the general linear mixed model and repeated measures ANOVA using a dataset with multiple missing data points.

Longitudinal methods are the methods of choice for researchers who view their phenomena of interest as dynamic. Although statistical methods have remained largely fixed in a linear view of biology and behavior, more recent methods, such as the general linear mixed model (mixed model), can be used to analyze dynamic phenomena that are often of interest to nurses. Two strengths of the mixed model are (1) the ability to accommodate missing data points often encountered in longitudinal datasets and (2) the ability to model nonlinear, individual characteristics. The purpose of this article is to demonstrate the advantages of using the mixed model for analyzing nonlinear, longitudinal datasets with multiple missing data points by comparing the mixed model to the widely used repeated measures ANOVA using an experimental set of data. The decision-making steps in analyzing the data using both the mixed model and the repeated measures ANOVA are described.

Analysis of Variance↗

Large datasets: common uses and caveats.

BACKGROUND: Increasingly, large collections of pre-existing data are being used to analyze the occurrence, burden, and health care resources directed to the management of various skin diseases. OBJECTIVE: This article discusses a number of different types of large datasets along with their common uses. Various concerns about the use of this information are also discussed. CONCLUSION: Although large datasets provide significant statistical power with readily available data, there are significant concerns, particularly regarding data quality and statistical analysis. Readers need to be aware of how an investigator has addressed these issues. Furthermore, the profession needs to be cognizant of very legitimate public concerns regarding confidentiality of personal information.

Databases, Factual↗

A prediction-based resampling method for estimating the number of clusters in a dataset.

BACKGROUND: Microarray technology is increasingly being applied in biological and medical research to address a wide range of problems, such as the classification of tumors. An important statistical problem associated with tumor classification is the identification of new tumor classes using gene-expression profiles. Two essential aspects of this clustering problem are: to estimate the number of clusters, if any, in a dataset; and to allocate tumor samples to these clusters, and assess the confidence of cluster assignments for individual samples. Here we address the first of these problems. RESULTS: We have developed a new prediction-based resampling method, Clest, to estimate the number of clusters in a dataset. The performance of the new and existing methods were compared using simulated data and gene-expression data from four recently published cancer microarray studies. Clest was generally found to be more accurate and robust than the six existing methods considered in the study. CONCLUSIONS: Focusing on prediction accuracy in conjunction with resampling produces accurate and robust estimates of the number of clusters.

Algorithms↗

The GRID: the General Repository for Interaction Datasets.

We have developed a relational database, called the General Repository for Interaction Datasets (The GRID) to archive and display physical, genetic and functional interactions. The GRID displays data-rich interaction tables for any protein of interest, combines literature-derived and high-throughput interaction datasets, and is readily accessible via the web. Interactions parsed in The GRID can be viewed in graphical form with a versatile visualization tool called Osprey.

DNA, Fungal↗