Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

A symbolic environment for visualizing activated foci in functional neuroimaging datasets.

This paper presents a symbolic visualization environment known as the Corner Cube environment, which was developed to facilitate rapid examination and comparison of activated foci defined by analyses of functional neuroimaging datasets. We have performed a comparative evaluation of this environment against maximum-intensity projection and 'gallery of slices' displays, and the results suggest that the Corner Cube environment has definite advantages over both conventional display techniques. We conclude that the Corner Cube is an effective tool for summarizing the spatial characteristics of activated foci within an easily understood visual context and is especially useful for displaying the similarities and differences in functional neuroimaging datasets.

Brain↗

Comparing functional genomic datasets: lessons from DNA microarray analyses of host-pathogen interactions.

Functional genomic technologies such as high density DNA microarrays allow biologists to study the structure and behavior of thousands of genes in a single experiment. One of the fields in which microarrays have had an increasingly important impact is host-pathogen interactions. Early investigations in this area over the past two years not only emphasize the utility of this approach, but also highlight the stereotyped gene expression responses of different host cells to diverse infectious stimuli, and the potential value of broad dataset comparisons in revealing fundamental features of innate immunity. The comparative analysis of recently published datasets involving human gene expression responses to two bacterial respiratory pathogens illustrates many of these points. Comparisons between these large, highly parallel sets of experimental observations also emphasize important technical and experimental design issues as future challenges.

Bordetella pertussis↗

Rotavirus vaccine effectiveness against rotavirus and acute gastroenteritis mortality: an analysis of pooled case-control studies from the MNSSTER-V dataset.

BACKGROUND: Rotavirus accounts for an estimated 25% of diarrhoea deaths in children under 5 years globally, and more than 140 countries have included rotavirus vaccines in their routine national infant vaccination programmes. We aimed to calculate rotavirus vaccine effectiveness against rotavirus-positive and all-cause acute gastroenteritis deaths. METHODS: The Multi-National Subpopulations Study to Evaluate Rotavirus Vaccines (MNSSTER-V) dataset combines child-level data from test-negative case-control studies of rotavirus vaccine effectiveness that enrolled children under 5 years of age seeking care for acute gastroenteritis at hospitals or emergency departments in 24 countries between July 1, 2007, and Aug 24, 2023. Children were included in this study if they were: younger than 5 years, met the acute gastroenteritis case definition (had at least three episodes of diarrhoea in a 24-h period, had non-bloody and non-chronic diarrhoea, and were enrolled within 7 days of diarrhoea onset), met vaccine card quality metrics, had vaccine delivery dates if the child was reported to have received a rotavirus vaccine, and had a reported outcome of death or discharge. In-hospital acute gastroenteritis deaths were characterised, and rotavirus vaccine effectiveness against all-cause and rotavirus-positive acute gastroenteritis mortality was calculated using an unconditional logistic regression model with adjustment for national under-5 mortality strata and child's age. Vaccine effectiveness analyses against all-cause and rotavirus-positive acute gastroenteritis mortality were restricted to children aged at least 3 months who received any routine vaccines from countries reporting at least one acute gastroenteritis death. FINDINGS: From the MNSSTER-V dataset, we included 27 252 children younger than 5 years enrolled from 22 countries; outcomes of patients were not available for two countries. At least one in-hospital acute gastroenteritis death was reported from 16 countries including 21 522 children; in total, 183 all-cause acute gastroenteritis deaths and 25 rotavirus-positive deaths were reported. Among children aged at least 3 months who had received any routine vaccines, receiving at least one dose of a rotavirus vaccine had an adjusted vaccine effectiveness of 75·8% (95% CI 28·4 to 91·8; n=13 630) against rotavirus-positive acute gastroenteritis mortality and 20·8% (-47·0 to 57·3; n=20 005) against all-cause acute gastroenteritis mortality. INTERPRETATION: Rotavirus vaccines are effective in preventing rotavirus-positive acute gastroenteritis mortality. Continued efforts to improve vaccine delivery could help to reduce acute gastroenteritis mortality due to rotavirus worldwide. FUNDING: None.

Humans↗

Stenosis quantification from post-stenotic signal loss in phase-contrast MRA datasets of flow phantoms and renal arteries.

In this study a semi-automated and observer-independent algorithm for quantifying post-stenotic signal loss (PSL) in 3D phase-contrast (PC) magnetic resonance angiography (MRA) of patients with renal artery stenosis is presented. This algorithm was developed on MRA datasets of stenotic phantoms, which were included in a flow circuit with stationary flows. The length and the severity of the PSL (incorporating both length and degree of PSL) in the maximum intensity projections (MIPs) of MRA datasets were proposed for quantifying stenoses. The algorithm was tested in renal arteries of ten patients with renal artery stenosis and seven healthy volunteers. Digital subtraction angiography (DSA) was performed in the patients and served as the gold standard. Stenosis severity showed better correlation with the severity of the PSL than with the length, both for in vitro as in vivo. Spearman correlation coefficients (rS) showed statistically significant correlations between the severity of the PSL and parameters determined by DSA, i.e. percent diameter stenosis (rS = 0.90). The length of the PSL showed no correlation with the diameter stenosis (rS = 0.37).

Adult↗

Evaluation of a novel molecular vibration-based descriptor (EVA) for QSAR studies: 2. Model validation using a benchmark steroid dataset.

The EVA molecular descriptor derived from calculated molecular vibrational frequencies is validated for use in QSAR studies. EVA provides a conformationally sensitive but, unlike 3D-QSAR methods such as CoMFA, superposition-free descriptor that has been shown to perform well with a wide range of datasets and biological endpoints. A detailed study is made using a benchmark steroid dataset with a training/test set division of structures. Intensive statistical validation tests are undertaken including various forms of crossvalidation and repeated random permutation testing. Latent variable score plots show that the distribution of structures in reduced dimensional space can be rationalized in terms of activity classes and that EVA is sensitive to structural inconsistencies. Together, the findings indicate that EVA is a statistically robust means of detecting structure-activity correlations with performance entirely comparable to that of analogous CoMFAs. The EVA descriptor is shown to be conformationally sensitive and as such can be considered to be a 3D descriptor but with the advantage over CoMFA that structural superposition is not required. EVA has the property that in certain situations the conformational sensitivity can be altered through the appropriate choice of the EVA sigma parameter.

Models, Molecular↗

RR-interval-based atrial fibrillation detection and burden estimation: cross-dataset validation and calibration-aware probability analysis.

Objective.Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift.Approach.This study developed an interpretable-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error (MAE) and agreement analyses.Main results.The model showed high discrimination in both development and external evaluation, with external ROC-AUC ofand PR-AUC of. However, external calibration deteriorated despite preserved ranking performance, with Brier score of, ECE(15) of, calibration intercept of, and calibration slope of. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with MAE ofversus, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased fromwithout recalibration toafter Platt recalibration andafter isotonic recalibration, while median ECE(15) decreased fromtoand, respectively.Significance.These findings indicate that-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.

Atrial Fibrillation↗

Identification of characteristic oligonucleotides in the bacterial 16S ribosomal RNA sequence dataset.

MOTIVATION: The phylogenetic structure of the bacterial world has been intensively studied by comparing sequences of 16S ribosomal RNA (16S rRNA). This database of sequences is now widely used to design probes for the detection of specific bacteria or groups of bacteria one at a time. The success of such methods reflects the fact that there are local sequence segments that are highly characteristic of particular organisms or groups of organisms. It is not clear, however, the extent to which such signature sequences exist in the 16S rRNA dataset. A better understanding of the numbers and distribution of highly informative oligonucleotide sequences may facilitate the design of hybridization arrays that can characterize the phylogenetic position of an unknown organism or serve as the basis for the development of novel approaches for use in bacterial identification. RESULTS: A computer-based algorithm that characterizes the extent to which any individual oligonucleotide sequence in 16S rRNA is characteristic of any particular bacterial grouping was developed. A measure of signature quality, Q(s), was formulated and subsequently calculated for every individual oligonucleotide sequence in the size range of 5-11 nucleotides and for 15mers with reference to each cluster and subcluster in a 929 organism representative phylogenetic tree. Subsequently, the perfect signature sequences were compared to the full set of 7322 sequences to see how common false positives were. The work completed here establishes beyond any doubt that highly characteristic oligonucleotides exist in the bacterial 16S rRNA sequence dataset in large numbers. Over 16,000 15mers were identified that might be useful as signatures. Signature oligonucleotides are available for over 80% of the nodes in the representative tree.

Algorithms↗

An anaesthetic minimum dataset and report format. Society for Computing and Technology in Anaesthesia (SCATA). European Society for Computing and Technology in Anaesthesia (ESCTAIC).

The dataset necessary to produce reports for anaesthetic training purposes is described, together with appropriate definitions. The format for a standard report that may be used in a logbook is also described. These have been accepted by the Royal College of Anaesthetists. The German Anaesthetic Society (Deutsche Gesellschaft für Anaesthesiologie und Intensivmedizin, DGAI) has accepted the dataset and definitions.

Anesthesiology↗

Measuring quality of care: fundamental information from administrative datasets.

Under proposals for national health insurance reform in the USA, employers and purchasing cooperatives will have to measure the quality of health care services. Their need for data systems upon which to base their decisions has stimulated dramatic innovation and rapid change in how health care information is collected, integrated from multiple sources, and reported. To make administrative data useful for quality measurement, careful attention must be given to information about: medical care utilization; patient characteristics; provider characteristics; and health plans. In this paper, we describe the extent to which this information is included in existing administrative datasets. We then suggest how planned datasets should be designed so they can be used to assess the quality of health care.

Databases, Factual↗

Medullary serotonergic network deficiency in the sudden infant death syndrome: review of a 15-year study of a single dataset.

The sudden infant death syndrome (SIDS) is the leading cause of postneonatal infant mortality in the United States today, despite a dramatic 38% decrease in incidence due to a national risk reduction campaign advocating the supine sleep position. Our research in SIDS brainstems, beginning in 1985 and involving a single, large dataset, has become increasingly focused upon a specific neurotransmitter (serotonin) and specific territories (ventral medulla and regions of the medullary reticular formation that contain secrotonergic neurons). Based on this research, we propose that SIDS, or a subset of SIDS, is due to a developmental abnormality in a medullary network composed of (at least in part) rhombic lip-derived, serotonergic neurons, including in the caudal raphé and arcuate nucleus (putative human homologue of the cat respiratory chemosensitive fields); and this abnormality results in a failure of protective responses to life-threatening stressors (e.g. asphyxia, hypoxia, hypercapnia) during sleep as the infant passes through a critical period in homeostatic control. We call this the medullary serotonergic network deficiency hypothesis. We review the triple-risk model for SIDS, the development of the dataset using tissue autoradiography for analyzing neurotransmitter receptor binding; age-dependent baseline neurochemical findings in the human brainstem during early life; the evidence for serotonergic, rhombic lip, and ventral medullary deficits in at least some SIDS victim; possible mechanisms of sudden infant death related to these deficits; and potential causes of the deficits in the medullary serotonergic network in SIDS victims. We conclude with a summary of future directions in SIDS brainstem research.

Animals↗

Three-dimensional texture analysis of MRI brain datasets.

A method is proposed for three-dimensional (3-D) texture analysis of magnetic resonance imaging brain datasets. It is based on extended, multisort co-occurrence matrices that employ intensity, gradient and anisotropy image features in a uniform way. Basic properties of matrices as well as their sensitivity and dependence on spatial image scaling are evaluated. The ability of the suggested 3-D texture descriptors is demonstrated on nontrivial classification tasks for pathologic findings in brain datasets.

Brain↗

Features affecting Cas9-induced editing efficiency and patterns in tomato: evidence from a large CRISPR dataset.

CRISPR/Cas9 is a cornerstone of plant genome editing, yet the determinants of editing efficiency for a given single-guide RNAs (sgRNAs) and DNA double-strand break (DSB) repair outcomes remain poorly understood, particularly in plants. Here, we generated a large experimental dataset comprising 420 sgRNAs targeting promoters, exons, and introns of 137 genes in tomato protoplasts, and quantified editing efficiency and repair footprints together with chromatin accessibility and transcriptional state in the same cellular context. Editing efficiency was consistently higher at targets in accessible chromatin and modestly higher in promoters and introns than in exons, whereas transcriptional activity had no detectable effect. Editing efficiencies were more similar among sgRNAs targeting the same gene than among different genes, revealing a local genomic influence on Cas9 activity. A distinct subset of sgRNAs achieved near-complete editing and produced characteristic repair footprints dominated by long deletions with extended microhomology tracts, indicative of microhomology-mediated end joining (MMEJ), resembling patterns associated with high-efficiency guides in human cells, and suggesting conserved sequence-driven repair biases across species. In contrast, widely used human-trained prediction models failed to accurately rank sgRNA performance in plants, highlighting the limits of cross-species predictability. Together, this dataset provides a resource for improving guide design and mechanistic understanding of plant DNA repair.

Solanum lycopersicum↗

Automated seed localization from CT datasets of the prostate.

With the increasing utilization of permanent brachytherapy implants for treating carcinoma of the prostate, the importance of accurate post-treatment dose calculation also increases for assessing patient outcome and planning future treatments. An automatic method for seed localization of permanent brachytherapy implants, using CT datasets of the prostate, has been developed and tested on a phantom using an actual patient planned seed distribution. This method was also compared to results with the three-film technique for three patient datasets. The automatic method is as accurate or more accurate than the three film technique for 1 mm, 3 mm, and 5 mm contiguous CT slices, and eliminates the inter- and intra-observer variability of the manual methods. The automated method improves the localization of brachytherapy seeds while reducing the time required for the user to input information, and is demonstrated to be less operator dependent, less time consuming, and potentially more accurate than the three-film technique.

Biophysical Phenomena↗

Essential dataset for ambulatory ear, nose, and throat care in general practice: an aid for quality assessment.

OBJECTIVE: To describe the documentation of care for the usual range of ear, nose, and throat (ENT) problems seen in primary care as a basis for developing a computerised information system to aid quality assessment. DESIGN: Descriptive study of the pattern of ENT problems and diagnoses and treatment as recorded in individual case notes. SETTING: The primary health care centre in Mjölby, Sweden. PATIENTS: Consultations for ENT problems from a 10% sample randomly selected from all consultations (n = 22,600) in one year. From this sample 375 consultations for ENT problems (16% of all consultations) by 272 patients were identified. MAIN MEASURES: The detailed documentation of each consultation was retrieved from the individual records and compared with the data required for a computer based information system designed to help in quality management. RESULTS: Although the overall picture gained from the data retrieved from the notes suggested that ENT care was probably adequate, the recorded details were limited. The written case notes were insufficient when compared with the details required for a computerised system based on an essential dataset designed to allow assessment of diagnostic accuracy and appropriateness of treatment of ENT problems in primary care. CONCLUSION: There is a gap between the amount and the type of information needed for accurate and useful quality assessment and that which is normally included in case notes. More detailed information is needed if general practitioners' notes are to be used for regular quality assessment of ENT problems but that would mean more time spent on keeping notes. This would be difficult to justify. IMPLICATIONS: The routine information systems used at this primary healthcare centre did not produce sufficient documentation for quality assessment of ENT care. This dilemma might be resolved by specially designed desktop computer software accessed through an essential dataset.

Adolescent↗

Large datasets: common uses and caveats.

BACKGROUND: Increasingly, large collections of pre-existing data are being used to analyze the occurrence, burden, and health care resources directed to the management of various skin diseases. OBJECTIVE: This article discusses a number of different types of large datasets along with their common uses. Various concerns about the use of this information are also discussed. CONCLUSION: Although large datasets provide significant statistical power with readily available data, there are significant concerns, particularly regarding data quality and statistical analysis. Readers need to be aware of how an investigator has addressed these issues. Furthermore, the profession needs to be cognizant of very legitimate public concerns regarding confidentiality of personal information.

Databases, Factual↗

A prediction-based resampling method for estimating the number of clusters in a dataset.

BACKGROUND: Microarray technology is increasingly being applied in biological and medical research to address a wide range of problems, such as the classification of tumors. An important statistical problem associated with tumor classification is the identification of new tumor classes using gene-expression profiles. Two essential aspects of this clustering problem are: to estimate the number of clusters, if any, in a dataset; and to allocate tumor samples to these clusters, and assess the confidence of cluster assignments for individual samples. Here we address the first of these problems. RESULTS: We have developed a new prediction-based resampling method, Clest, to estimate the number of clusters in a dataset. The performance of the new and existing methods were compared using simulated data and gene-expression data from four recently published cancer microarray studies. Clest was generally found to be more accurate and robust than the six existing methods considered in the study. CONCLUSIONS: Focusing on prediction accuracy in conjunction with resampling produces accurate and robust estimates of the number of clusters.

Algorithms↗

The GRID: the General Repository for Interaction Datasets.

We have developed a relational database, called the General Repository for Interaction Datasets (The GRID) to archive and display physical, genetic and functional interactions. The GRID displays data-rich interaction tables for any protein of interest, combines literature-derived and high-throughput interaction datasets, and is readily accessible via the web. Interactions parsed in The GRID can be viewed in graphical form with a versatile visualization tool called Osprey.

DNA, Fungal↗

UnionLoops: a workflow for calling chromatin loops across related Hi-C datasets with improved specificity, precision, and sensitivity.

Chromatin loop calling from chromatin interaction data often exhibits substantial variability across related samples. We present UnionLoops, a computational workflow for chromatin loop calling across multiple related samples. UnionLoops integrates information across datasets to determine positions and dataset-specificity of looping interactions. It constructs a unified candidate loop set, applies consistent filtering and aggregation, and evaluates loop support across samples. We demonstrate that UnionLoops increases sensitivity for detecting shared chromatin loops, reduces spurious sample-specific calls, and improves concordance with independent genomic features, including CTCF and cohesin occupancy. UnionLoops enables improved biological interpretation of chromatin loop organization and dynamics across related conditions.

Chromatin↗