Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Prediction of 3-yr cadaveric graft survival based on pre-transplant variables in a large national dataset.

Pre- and post-transplant predictive factors of graft survival for optimal and expanded criteria grafts have been studied in the past. The goal of our study was to evaluate the recent large set of United Network of Organ Sharing records (1990-1998) to generate a prediction algorithm of 3-yr graft survival based on pre-transplant variables alone. The dataset of patients with end-stage renal disease and cadaveric kidney or kidney-pancreas transplantation (1990-1998) used in the study consisted of 37,407 records. Logistic regression (LM) and a tree-based model (TBM) were used to identify predictors of 3-yr allograft survival and to generate prediction algorithm. Donor and recipient demographic characteristics (age, race, and gender) and body mass index showed non-linear, while human leukocyte antigen match showed strong linear relationships with 3-yr graft survival. Prediction of the probability of graft survival from the model, achieved a good match with the observed survival of the separate dataset, with a correlation of r = 0.998 for LM and r = 0.984 for TBM. The positive predictive value (PV) of allograft survival with LM and TBM was 76.0% and the negative PV was 63 and 53.8% for LM and TBM, respectively. Both LM and the TBM can potentially be used in clinical practice for long-term prediction of kidney allograft survival based on pre-transplant variables.

Algorithms↗

RR-interval-based atrial fibrillation detection and burden estimation: cross-dataset validation and calibration-aware probability analysis.

Objective.Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift.Approach.This study developed an interpretable-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error (MAE) and agreement analyses.Main results.The model showed high discrimination in both development and external evaluation, with external ROC-AUC ofand PR-AUC of. However, external calibration deteriorated despite preserved ranking performance, with Brier score of, ECE(15) of, calibration intercept of, and calibration slope of. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with MAE ofversus, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased fromwithout recalibration toafter Platt recalibration andafter isotonic recalibration, while median ECE(15) decreased fromtoand, respectively.Significance.These findings indicate that-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.

Atrial Fibrillation↗

Identification of characteristic oligonucleotides in the bacterial 16S ribosomal RNA sequence dataset.

MOTIVATION: The phylogenetic structure of the bacterial world has been intensively studied by comparing sequences of 16S ribosomal RNA (16S rRNA). This database of sequences is now widely used to design probes for the detection of specific bacteria or groups of bacteria one at a time. The success of such methods reflects the fact that there are local sequence segments that are highly characteristic of particular organisms or groups of organisms. It is not clear, however, the extent to which such signature sequences exist in the 16S rRNA dataset. A better understanding of the numbers and distribution of highly informative oligonucleotide sequences may facilitate the design of hybridization arrays that can characterize the phylogenetic position of an unknown organism or serve as the basis for the development of novel approaches for use in bacterial identification. RESULTS: A computer-based algorithm that characterizes the extent to which any individual oligonucleotide sequence in 16S rRNA is characteristic of any particular bacterial grouping was developed. A measure of signature quality, Q(s), was formulated and subsequently calculated for every individual oligonucleotide sequence in the size range of 5-11 nucleotides and for 15mers with reference to each cluster and subcluster in a 929 organism representative phylogenetic tree. Subsequently, the perfect signature sequences were compared to the full set of 7322 sequences to see how common false positives were. The work completed here establishes beyond any doubt that highly characteristic oligonucleotides exist in the bacterial 16S rRNA sequence dataset in large numbers. Over 16,000 15mers were identified that might be useful as signatures. Signature oligonucleotides are available for over 80% of the nodes in the representative tree.

Algorithms↗

NMRb: a web-site repository for raw NMR datasets.

UNLABELLED: The development of NMR in structural proteomics requires the availability of automatic structure determination methods. Many researchers are commonly confronted with the lack of raw datasets during the validation step of such methods. In order to increase test possibilities, the NMRb web-site offers a database of NMR raw datasets, ordered by spectral characteristics. AVAILABILITY: NMRb is available from: http://nmrb.cbs.cnrs.fr. SUPPLEMENTARY INFORMATION: General organization of NMRb figure, relational model organization, and XML structure files are available from http://nmrb.cbs.cnrs.fr/nmrb-doc.html.

Database Management Systems↗

Discovery of stable and significant binding motif pairs from PDB complexes and protein interaction datasets.

MOTIVATION: Discovery of binding sites is important in the study of protein-protein interactions. In this paper, we introduce stable and significant motif pairs to model protein-binding sites. The stability is the pattern's resistance to some transformation. The significance is the unexpected frequency of occurrence of the pattern in a sequence dataset comprising known interacting protein pairs. Discovery of stable motif pairs is an iterative process, undergoing a chain of changing but converging patterns. Determining the starting point for such a chain is an interesting problem. We use a protein complex dataset extracted from the Protein Data Bank to help in identifying those starting points, so that the computational complexity of the problem is much released. RESULTS: We found 913 stable motif pairs, of which 765 are significant. We evaluated these motif pairs using comprehensive comparison results against random patterns. Wet-experimentally discovered motifs reported in the literature were also used to confirm the effectiveness of our method. SUPPLEMENTARY INFORMATION: http://sdmc.i2r.a-star.edu.sg/BindingMotifPairs.

Algorithms↗

An anaesthetic minimum dataset and report format. Society for Computing and Technology in Anaesthesia (SCATA). European Society for Computing and Technology in Anaesthesia (ESCTAIC).

The dataset necessary to produce reports for anaesthetic training purposes is described, together with appropriate definitions. The format for a standard report that may be used in a logbook is also described. These have been accepted by the Royal College of Anaesthetists. The German Anaesthetic Society (Deutsche Gesellschaft für Anaesthesiologie und Intensivmedizin, DGAI) has accepted the dataset and definitions.

Anesthesiology↗

Measuring quality of care: fundamental information from administrative datasets.

Under proposals for national health insurance reform in the USA, employers and purchasing cooperatives will have to measure the quality of health care services. Their need for data systems upon which to base their decisions has stimulated dramatic innovation and rapid change in how health care information is collected, integrated from multiple sources, and reported. To make administrative data useful for quality measurement, careful attention must be given to information about: medical care utilization; patient characteristics; provider characteristics; and health plans. In this paper, we describe the extent to which this information is included in existing administrative datasets. We then suggest how planned datasets should be designed so they can be used to assess the quality of health care.

Databases, Factual↗

Medullary serotonergic network deficiency in the sudden infant death syndrome: review of a 15-year study of a single dataset.

The sudden infant death syndrome (SIDS) is the leading cause of postneonatal infant mortality in the United States today, despite a dramatic 38% decrease in incidence due to a national risk reduction campaign advocating the supine sleep position. Our research in SIDS brainstems, beginning in 1985 and involving a single, large dataset, has become increasingly focused upon a specific neurotransmitter (serotonin) and specific territories (ventral medulla and regions of the medullary reticular formation that contain secrotonergic neurons). Based on this research, we propose that SIDS, or a subset of SIDS, is due to a developmental abnormality in a medullary network composed of (at least in part) rhombic lip-derived, serotonergic neurons, including in the caudal raphé and arcuate nucleus (putative human homologue of the cat respiratory chemosensitive fields); and this abnormality results in a failure of protective responses to life-threatening stressors (e.g. asphyxia, hypoxia, hypercapnia) during sleep as the infant passes through a critical period in homeostatic control. We call this the medullary serotonergic network deficiency hypothesis. We review the triple-risk model for SIDS, the development of the dataset using tissue autoradiography for analyzing neurotransmitter receptor binding; age-dependent baseline neurochemical findings in the human brainstem during early life; the evidence for serotonergic, rhombic lip, and ventral medullary deficits in at least some SIDS victim; possible mechanisms of sudden infant death related to these deficits; and potential causes of the deficits in the medullary serotonergic network in SIDS victims. We conclude with a summary of future directions in SIDS brainstem research.

Animals↗

Three-dimensional texture analysis of MRI brain datasets.

A method is proposed for three-dimensional (3-D) texture analysis of magnetic resonance imaging brain datasets. It is based on extended, multisort co-occurrence matrices that employ intensity, gradient and anisotropy image features in a uniform way. Basic properties of matrices as well as their sensitivity and dependence on spatial image scaling are evaluated. The ability of the suggested 3-D texture descriptors is demonstrated on nontrivial classification tasks for pathologic findings in brain datasets.

Brain↗

Bayesian class discovery in microarray datasets.

A novel approach to class discovery in gene expression datasets is presented. In the context of clinical diagnosis, the central goal of class discovery algorithms is to simultaneously find putative (sub-)types of diseases and to identify informative subsets of genes with disease-type specific expression profile. Contrary to many other approaches in the literature, the method presented implements a wrapper strategy for feature selection, in the sense that the features are directly selected by optimizing the discriminative power of the used partitioning algorithm. The usual combinatorial problems associated with wrapper approaches are overcome by a Bayesian inference mechanism. On the technical side, we present an efficient optimization algorithm with guaranteed local convergence property. The only free parameter of the optimization method is selected by a resampling-based stability analysis. Experiments with Leukemia and Lymphoma datasets demonstrate that our method is able to correctly infer partitions and corresponding subsets of genes which both are relevant in a biological sense. Moreover, the frequently observed problem of ambiguities caused by different but equally high-scoring partitions is successfully overcome by the model selection method proposed.

Algorithms↗

Features affecting Cas9-induced editing efficiency and patterns in tomato: evidence from a large CRISPR dataset.

CRISPR/Cas9 is a cornerstone of plant genome editing, yet the determinants of editing efficiency for a given single-guide RNAs (sgRNAs) and DNA double-strand break (DSB) repair outcomes remain poorly understood, particularly in plants. Here, we generated a large experimental dataset comprising 420 sgRNAs targeting promoters, exons, and introns of 137 genes in tomato protoplasts, and quantified editing efficiency and repair footprints together with chromatin accessibility and transcriptional state in the same cellular context. Editing efficiency was consistently higher at targets in accessible chromatin and modestly higher in promoters and introns than in exons, whereas transcriptional activity had no detectable effect. Editing efficiencies were more similar among sgRNAs targeting the same gene than among different genes, revealing a local genomic influence on Cas9 activity. A distinct subset of sgRNAs achieved near-complete editing and produced characteristic repair footprints dominated by long deletions with extended microhomology tracts, indicative of microhomology-mediated end joining (MMEJ), resembling patterns associated with high-efficiency guides in human cells, and suggesting conserved sequence-driven repair biases across species. In contrast, widely used human-trained prediction models failed to accurately rank sgRNA performance in plants, highlighting the limits of cross-species predictability. Together, this dataset provides a resource for improving guide design and mechanistic understanding of plant DNA repair.

Solanum lycopersicum↗

Automated seed localization from CT datasets of the prostate.

With the increasing utilization of permanent brachytherapy implants for treating carcinoma of the prostate, the importance of accurate post-treatment dose calculation also increases for assessing patient outcome and planning future treatments. An automatic method for seed localization of permanent brachytherapy implants, using CT datasets of the prostate, has been developed and tested on a phantom using an actual patient planned seed distribution. This method was also compared to results with the three-film technique for three patient datasets. The automatic method is as accurate or more accurate than the three film technique for 1 mm, 3 mm, and 5 mm contiguous CT slices, and eliminates the inter- and intra-observer variability of the manual methods. The automated method improves the localization of brachytherapy seeds while reducing the time required for the user to input information, and is demonstrated to be less operator dependent, less time consuming, and potentially more accurate than the three-film technique.

Biophysical Phenomena↗

Upper gastrointestinal cancer pathology reporting: a regional audit to compare standards with minimum datasets.

AIMS: Accurate pathological (pTNM) staging of oesophageal and gastric cancer provides important prognostic information. The aim of this study was to compare the standard of pathology reporting of oesophageal and gastric cancer resections from a cancer network with standards set by the Royal College of Pathologists. METHODS: All reports for oesophageal and gastric cancer resections from the five hospitals in the cancer network in 2001 were collected. Individual items of information were compared with minimum datasets provided by the Royal College of Pathologists. Items were classified as "complete", "partially complete", or "absent". RESULTS: One hundred and ten reports were audited (54 oesophageal and 56 gastric). Fourteen gastric and 17 oesophagectomy reports were over 75% complete. Clinically important missing data occurred most frequently for the pM component of TNM staging (pMx omitted in 87 reports) and completeness of resection expressed as a bold statement (absent in 50 reports). Twelve reports could not be classified because the specimen contained no residual tumour after neoadjuvant treatment. CONCLUSION: The use of a standard proforma for reporting upper gastrointestinal cancers based on a minimum dataset provided by the Royal College of Pathologists is recommended, with modifications to allow for specimens with no tumour after neoadjuvant treatment.

England↗

Essential dataset for ambulatory ear, nose, and throat care in general practice: an aid for quality assessment.

OBJECTIVE: To describe the documentation of care for the usual range of ear, nose, and throat (ENT) problems seen in primary care as a basis for developing a computerised information system to aid quality assessment. DESIGN: Descriptive study of the pattern of ENT problems and diagnoses and treatment as recorded in individual case notes. SETTING: The primary health care centre in Mjölby, Sweden. PATIENTS: Consultations for ENT problems from a 10% sample randomly selected from all consultations (n = 22,600) in one year. From this sample 375 consultations for ENT problems (16% of all consultations) by 272 patients were identified. MAIN MEASURES: The detailed documentation of each consultation was retrieved from the individual records and compared with the data required for a computer based information system designed to help in quality management. RESULTS: Although the overall picture gained from the data retrieved from the notes suggested that ENT care was probably adequate, the recorded details were limited. The written case notes were insufficient when compared with the details required for a computerised system based on an essential dataset designed to allow assessment of diagnostic accuracy and appropriateness of treatment of ENT problems in primary care. CONCLUSION: There is a gap between the amount and the type of information needed for accurate and useful quality assessment and that which is normally included in case notes. More detailed information is needed if general practitioners' notes are to be used for regular quality assessment of ENT problems but that would mean more time spent on keeping notes. This would be difficult to justify. IMPLICATIONS: The routine information systems used at this primary healthcare centre did not produce sufficient documentation for quality assessment of ENT care. This dilemma might be resolved by specially designed desktop computer software accessed through an essential dataset.

Adolescent↗

A comparison of the general linear mixed model and repeated measures ANOVA using a dataset with multiple missing data points.

Longitudinal methods are the methods of choice for researchers who view their phenomena of interest as dynamic. Although statistical methods have remained largely fixed in a linear view of biology and behavior, more recent methods, such as the general linear mixed model (mixed model), can be used to analyze dynamic phenomena that are often of interest to nurses. Two strengths of the mixed model are (1) the ability to accommodate missing data points often encountered in longitudinal datasets and (2) the ability to model nonlinear, individual characteristics. The purpose of this article is to demonstrate the advantages of using the mixed model for analyzing nonlinear, longitudinal datasets with multiple missing data points by comparing the mixed model to the widely used repeated measures ANOVA using an experimental set of data. The decision-making steps in analyzing the data using both the mixed model and the repeated measures ANOVA are described.

Analysis of Variance↗

Large datasets: common uses and caveats.

BACKGROUND: Increasingly, large collections of pre-existing data are being used to analyze the occurrence, burden, and health care resources directed to the management of various skin diseases. OBJECTIVE: This article discusses a number of different types of large datasets along with their common uses. Various concerns about the use of this information are also discussed. CONCLUSION: Although large datasets provide significant statistical power with readily available data, there are significant concerns, particularly regarding data quality and statistical analysis. Readers need to be aware of how an investigator has addressed these issues. Furthermore, the profession needs to be cognizant of very legitimate public concerns regarding confidentiality of personal information.

Databases, Factual↗

A prediction-based resampling method for estimating the number of clusters in a dataset.

BACKGROUND: Microarray technology is increasingly being applied in biological and medical research to address a wide range of problems, such as the classification of tumors. An important statistical problem associated with tumor classification is the identification of new tumor classes using gene-expression profiles. Two essential aspects of this clustering problem are: to estimate the number of clusters, if any, in a dataset; and to allocate tumor samples to these clusters, and assess the confidence of cluster assignments for individual samples. Here we address the first of these problems. RESULTS: We have developed a new prediction-based resampling method, Clest, to estimate the number of clusters in a dataset. The performance of the new and existing methods were compared using simulated data and gene-expression data from four recently published cancer microarray studies. Clest was generally found to be more accurate and robust than the six existing methods considered in the study. CONCLUSIONS: Focusing on prediction accuracy in conjunction with resampling produces accurate and robust estimates of the number of clusters.

Algorithms↗

The GRID: the General Repository for Interaction Datasets.

We have developed a relational database, called the General Repository for Interaction Datasets (The GRID) to archive and display physical, genetic and functional interactions. The GRID displays data-rich interaction tables for any protein of interest, combines literature-derived and high-throughput interaction datasets, and is readily accessible via the web. Interactions parsed in The GRID can be viewed in graphical form with a versatile visualization tool called Osprey.

DNA, Fungal↗