Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Three-dimensional computed tomography bronchoscopy using clinical datasets: a comparison with fibreoptic bronchoscopy.

OBJECTIVE: To assess three-dimensional computed tomography 'bronchoscopic' (3-DCTB) reconstruction of routine CT data as a non-invasive method of airway visualization, and compare it with fibreoptic bronchoscopy (FOB). METHODS: Fourteen datasets were acquired from 13 patients undergoing both FOB and CT examination of the chest. Standard continuous volume CT using 6 mm collimation and clinical FOB techniques were employed. Images were obtained from 3-DCTB reconstructions by segmentation and surface recognition algorithms generating surface rendered 'bronchoscopic views'. These were scored for technical quality and anatomical detail. The most distal bronchi seen in left upper and right lower lobes were recorded for FOB and 3-DCTB. RESULTS: On FOB, the subsegmental bronchi were seen in right lower and in left upper lobe in 10/14 cases and 4/14 cases, respectively. Visualization of the subsegmental airways was not achieved with 3-DCTB, as they could not be identified with confidence. 3-DCTB never achieved a more distal view than obtained by FOB. Using 3-DCT, the right, lower lobe segmental bronchi were seen in 10/14 cases, and lobar bronchus in 14/14 cases (two occluded). In the left upper lobe, 3-DCT showed segmental bronchi in 6/14 cases, lobar bronchus in 11/14 cases (one occluded) and the left main bronchus appeared occluded in 3/14 cases. Overall, technical quality and anatomical detail scores of the carina and proximal bronchi ranked significantly higher than views of segmental bronchi. CONCLUSIONS: 3-DCTB cannot routinely replace FOB for inspection of major and segmental bronchi. Subsegmental bronchi cannot be adequately demonstrated by 3-DCTB using 6 mm collimation datasets.

Adult↗

Knowledge discovery with classification rules in a cardiovascular dataset.

In this paper we study an evolutionary machine learning approach to data mining and knowledge discovery based on the induction of classification rules. A method for automatic rules induction called AREX using evolutionary induction of decision trees and automatic programming is introduced. The proposed algorithm is applied to a cardiovascular dataset consisting of different groups of attributes which should possibly reveal the presence of some specific cardiovascular problems in young patients. A case study is presented that shows the use of AREX for the classification of patients and for discovering possible new medical knowledge from the dataset. The defined knowledge discovery loop comprises a medical expert's assessment of induced rules to drive the evolution of rule sets towards more appropriate solutions. The final result is the discovery of a possible new medical knowledge in the field of pediatric cardiology.

Algorithms↗

Time factors in larynx tumor radiotherapy: lag times and intertumor heterogeneity in clinical datasets from four centers.

PURPOSE: To use the time-dependent linear-quadratic model, both in the standard form and in a form modified to incorporate intertumor heterogeneity, in a reanalysis of 4 datasets for larynx tumor control, to provide more representative and direct estimates of the lag period, the time factor (lambda/alpha), and the clonogen population inactivation dose ([lnk]/alpha). METHODS AND MATERIALS: The data comprised 2,225 patients treated in Edinburgh (UK), Glasgow (UK), Manchester (UK), or Toronto (Canada), with tumor control assessed after at least 2 years. Heterogeneity in each series was taken into account using the coefficient of variation (CV) of the clonogen radiosensitivity (alpha). Maximum likelihood techniques were used to provide best estimates of the parameters, and also direct estimation of the more stable parameter ratios of interest. RESULTS: The use of different heterogeneity factors for the different series allowed common dose/time parameters to be fitted across all four series in a way not possible using the standard model, enabling the inherent effect of heterogeneity in flattening dose-response curves and in reducing time factors to be separated from the underlying more-representative values. Radiosensitivity CVs were calculated to be 30% (Edinburgh), 36% (Glasgow), 40% (Manchester), and 71% (Toronto). The lag phase was 32 days (95% CL 20-38 days) which was longer than the value of 23 days (11-36 days) deduced using the standard model without the heterogeneity parameter. The time factor was 1.2 (0.8-2.2) Gy/day, again greater than the value of 0.80 (0.54-1.41) Gy/day derived using the standard model. Similar larger time factors and longer lag periods could be reproduced using the standard model either by using a parameterization based on parameter ratios, or by omitting the discordant Toronto data and refitting just the data from the three UK centers. CONCLUSION: It was concluded that the heterogeneity model provides a better representation of the time factor for tumor control when data are analyzed comprising different stages of disease treated at different centers. The model allows different amounts of heterogeneity in different series, which tend to flatten dose-responses curves and reduce time factors, to be taken in to account. Also, direct maximum likelihood estimates can be made of the lag period, the time factor (lambda/alpha), and the fractionation sensitivity (beta/alpha), as well as the clonogen population inactivation dose (lnk)/alpha. Values of these parameter ratios are more robust and stable than the individual parameter values. The results of the present analysis using a total of 2,225 patients from four centers indicate that the average lag period may be somewhat longer and the average time factor somewhat greater (and the 95% confidence limits of the time factor exclude previous estimates), than the values deduced previously using simpler models and more diverse multi-center datasets.

Dose-Response Relationship, Radiation↗

Objective stenosis quantification from post-stenotic signal loss in phase-contrast magnetic resonance angiographic datasets of flow phantoms and renal arteries.

In this study a semi-automated and observer-independent algorithm for quantifying post-stenotic signal loss (PSL) in three-dimensional phase-contrast (PC) magnetic resonance angiography (MRA) of patients with renal artery stenosis is presented. This algorithm was developed on MRA datasets of stenotic phantoms, included in a flow circuit with stationary flows. The length and the severity of the PSL (incorporating both the length and the degree of PSL) in the MRA datasets were proposed for quantifying the stenoses. The algorithm was tested in renal arteries; ten patients with renal artery stenosis and seven healthy volunteers were investigated. Digital subtraction angiography was performed in the patients and served as the gold standard. Stenosis severity showed better correlation with the severity of the PSL than with the length, both for in vitro and in vivo measurements. Spearman correlation coefficients (rs) showed statistically significant correlations between the severity of the PSL and parameters determined by digital subtraction angiography, i.e., percent diameter stenosis (rs = 0.90). The length of the PSL showed no correlation with the diameter stenosis (rs = 0.37). In conclusion, this study presents a semi-automated and observer-independent way of quantifying signal loss, and the severity of the PSL is proposed for quantifying stenoses, rather than the length of PSL.

Aged↗

Virtual endoscopy: development and evaluation using the Visible Human datasets.

Virtual endoscopy (VE) is a new method of diagnosis using computer processing of 3D image datasets (such as CT or MRI scans) to provide simulated visualizations of patient specific organs similar or equivalent to those produced by standard endoscopic procedures. Conventional endoscopy is invasive and often uncomfortable for patients. It sometimes has serious side effects such as perforation, infection and hemorrhage. VE visualization avoids these risks and can minimize difficulties and decrease morbidity when used before actual endoscopic procedures. In addition, there are many body regions not compatible with real endoscopy that can be explored with VE. Eventually, VE may replace many forms of real endoscopy. There remains a critical need to refine and validate VE visualizations for routine clinical use. We have used the Visible Human Dataset from the National Library of Medicine to develop and test these procedures and to evaluate their use in a variety of clinical applications. We have developed specific clinical protocols to compare virtual endoscopy with real endoscopy. We have developed informative and dynamic on-screen navigation guides to help the surgeon or physician interactively determine body orientation and precise anatomical localization while performing the VE procedures. Additionally, the adjunctive value of full 3D imaging (e.g. looking "outside" of the normal field of view) during the VE exam is being evaluated. Quantitative analyses of local geometric and densitometric properties obtained from the virtual procedures ("virtual biopsy") are being developed and compared with other direct measures. Preliminary results suggest that these virtual procedures can provide accurate, reproducible and clinically useful visualizations and measurements. These studies will help drive improvements in and lend credibility to VE procedures and simulations as routine clinical tools. VE holds significant promise for optimizing endoscopic diagnostic procedures, minimizing patient risk and morbidity, and reducing health care costs.

Anatomy, Cross-Sectional↗

Using dependency/association rules to find indications for computed tomography in a head trauma dataset.

Analysis of a clinical head trauma dataset was aided by the use of a new, binary-based data mining technique, termed Boolean analyzer (BA), which finds dependency/association rules. With initial guidance from a domain user or domain expert, the BA algorithm is given one or more metrics to partition the entire dataset. The weighted rules are in the form of Boolean expressions. To augment the analysis of the rules produced, we applied a probabilistic interestingness measure (PIM) to order the generated rules based on event dependency, where events are combinations of primed and unprimed variables. Interpretation of the dependency rules generated on the clinical head trauma data resulted in a set of criteria that identified minor head trauma patients needing computed tomography (CT) scans. The BA criteria contained fewer variables than were found using recursive partitioning of Chi-square values (five variables versus seven variables, respectively). The BA five-variable criteria set was more sensitive but less specific than the seven-variable Chi-square criteria set. We believe that the BA method has broad applicability in the medical domain, and hope that this paper will stimulate other creative applications of the technique.

Algorithms↗

How do insect nuclear and mitochondrial gene substitution patterns differ? Insights from Bayesian analyses of combined datasets.

We analyzed 12 combined mitochondrial and nuclear gene datasets in seven orders of insects using both equal weights parsimony (to evaluate phylogenetic utility) and Bayesian methods (to investigate substitution patterns). For the Bayesian analyses we used relatively complex models (e.g., general time reversible models with rate variation) that allowed us to quantitatively compare relative rates among genes and codon positions, patterns of rate variation among genes, and substitution patterns within genes. Our analyses indicate that nuclear and mitochondrial genes differ in a number of important ways, some of which are correlated with phylogenetic utility. First and most obviously, nuclear genes generally evolve more slowly than mitochondrial genes (except in one case), making them better markers for deep divergences. Second, nuclear genes showed universally high values of CI and (generally) contribute more to overall tree resolution than mitochondrial genes (as measured by partitioned Bremer support). Third, nuclear genes show more homogeneous patterns of among-site rate variation (higher values of alpha than mitochondrial genes). Finally, nuclear genes show more symmetrical transformation rate matrices than mitochondrial genes. The combination of low values of alpha and highly asymmetrical transformation rate matrices may explain the overall poor performance of mitochondrial genes when compared to nuclear genes in the same analysis. Our analyses indicate that some parameters are highly correlated. For example, A/T bias was positively and significantly associated with relative rate and CI was positively and significantly associated with alpha (the shape of the gamma distribution). These results provide important insights into the substitution patterns that might characterized high quality genes for phylogenetic analysis: high values of alpha, unbiased base composition, and symmetrical transformation rate matrices. We argue that insect molecular systematists should increasingly focus on nuclear rather than mitochondrial gene datasets because nuclear genes do not suffer from the same substitutional biases that characterize mitochondrial genes.

Animals↗

A symbolic environment for visualizing activated foci in functional neuroimaging datasets.

This paper presents a symbolic visualization environment known as the Corner Cube environment, which was developed to facilitate rapid examination and comparison of activated foci defined by analyses of functional neuroimaging datasets. We have performed a comparative evaluation of this environment against maximum-intensity projection and 'gallery of slices' displays, and the results suggest that the Corner Cube environment has definite advantages over both conventional display techniques. We conclude that the Corner Cube is an effective tool for summarizing the spatial characteristics of activated foci within an easily understood visual context and is especially useful for displaying the similarities and differences in functional neuroimaging datasets.

Brain↗

Comparing functional genomic datasets: lessons from DNA microarray analyses of host-pathogen interactions.

Functional genomic technologies such as high density DNA microarrays allow biologists to study the structure and behavior of thousands of genes in a single experiment. One of the fields in which microarrays have had an increasingly important impact is host-pathogen interactions. Early investigations in this area over the past two years not only emphasize the utility of this approach, but also highlight the stereotyped gene expression responses of different host cells to diverse infectious stimuli, and the potential value of broad dataset comparisons in revealing fundamental features of innate immunity. The comparative analysis of recently published datasets involving human gene expression responses to two bacterial respiratory pathogens illustrates many of these points. Comparisons between these large, highly parallel sets of experimental observations also emphasize important technical and experimental design issues as future challenges.

Bordetella pertussis↗

Rotavirus vaccine effectiveness against rotavirus and acute gastroenteritis mortality: an analysis of pooled case-control studies from the MNSSTER-V dataset.

BACKGROUND: Rotavirus accounts for an estimated 25% of diarrhoea deaths in children under 5 years globally, and more than 140 countries have included rotavirus vaccines in their routine national infant vaccination programmes. We aimed to calculate rotavirus vaccine effectiveness against rotavirus-positive and all-cause acute gastroenteritis deaths. METHODS: The Multi-National Subpopulations Study to Evaluate Rotavirus Vaccines (MNSSTER-V) dataset combines child-level data from test-negative case-control studies of rotavirus vaccine effectiveness that enrolled children under 5 years of age seeking care for acute gastroenteritis at hospitals or emergency departments in 24 countries between July 1, 2007, and Aug 24, 2023. Children were included in this study if they were: younger than 5 years, met the acute gastroenteritis case definition (had at least three episodes of diarrhoea in a 24-h period, had non-bloody and non-chronic diarrhoea, and were enrolled within 7 days of diarrhoea onset), met vaccine card quality metrics, had vaccine delivery dates if the child was reported to have received a rotavirus vaccine, and had a reported outcome of death or discharge. In-hospital acute gastroenteritis deaths were characterised, and rotavirus vaccine effectiveness against all-cause and rotavirus-positive acute gastroenteritis mortality was calculated using an unconditional logistic regression model with adjustment for national under-5 mortality strata and child's age. Vaccine effectiveness analyses against all-cause and rotavirus-positive acute gastroenteritis mortality were restricted to children aged at least 3 months who received any routine vaccines from countries reporting at least one acute gastroenteritis death. FINDINGS: From the MNSSTER-V dataset, we included 27 252 children younger than 5 years enrolled from 22 countries; outcomes of patients were not available for two countries. At least one in-hospital acute gastroenteritis death was reported from 16 countries including 21 522 children; in total, 183 all-cause acute gastroenteritis deaths and 25 rotavirus-positive deaths were reported. Among children aged at least 3 months who had received any routine vaccines, receiving at least one dose of a rotavirus vaccine had an adjusted vaccine effectiveness of 75·8% (95% CI 28·4 to 91·8; n=13 630) against rotavirus-positive acute gastroenteritis mortality and 20·8% (-47·0 to 57·3; n=20 005) against all-cause acute gastroenteritis mortality. INTERPRETATION: Rotavirus vaccines are effective in preventing rotavirus-positive acute gastroenteritis mortality. Continued efforts to improve vaccine delivery could help to reduce acute gastroenteritis mortality due to rotavirus worldwide. FUNDING: None.

Humans↗

Stenosis quantification from post-stenotic signal loss in phase-contrast MRA datasets of flow phantoms and renal arteries.

In this study a semi-automated and observer-independent algorithm for quantifying post-stenotic signal loss (PSL) in 3D phase-contrast (PC) magnetic resonance angiography (MRA) of patients with renal artery stenosis is presented. This algorithm was developed on MRA datasets of stenotic phantoms, which were included in a flow circuit with stationary flows. The length and the severity of the PSL (incorporating both length and degree of PSL) in the maximum intensity projections (MIPs) of MRA datasets were proposed for quantifying stenoses. The algorithm was tested in renal arteries of ten patients with renal artery stenosis and seven healthy volunteers. Digital subtraction angiography (DSA) was performed in the patients and served as the gold standard. Stenosis severity showed better correlation with the severity of the PSL than with the length, both for in vitro as in vivo. Spearman correlation coefficients (rS) showed statistically significant correlations between the severity of the PSL and parameters determined by DSA, i.e. percent diameter stenosis (rS = 0.90). The length of the PSL showed no correlation with the diameter stenosis (rS = 0.37).

Adult↗

Evaluation of a novel molecular vibration-based descriptor (EVA) for QSAR studies: 2. Model validation using a benchmark steroid dataset.

The EVA molecular descriptor derived from calculated molecular vibrational frequencies is validated for use in QSAR studies. EVA provides a conformationally sensitive but, unlike 3D-QSAR methods such as CoMFA, superposition-free descriptor that has been shown to perform well with a wide range of datasets and biological endpoints. A detailed study is made using a benchmark steroid dataset with a training/test set division of structures. Intensive statistical validation tests are undertaken including various forms of crossvalidation and repeated random permutation testing. Latent variable score plots show that the distribution of structures in reduced dimensional space can be rationalized in terms of activity classes and that EVA is sensitive to structural inconsistencies. Together, the findings indicate that EVA is a statistically robust means of detecting structure-activity correlations with performance entirely comparable to that of analogous CoMFAs. The EVA descriptor is shown to be conformationally sensitive and as such can be considered to be a 3D descriptor but with the advantage over CoMFA that structural superposition is not required. EVA has the property that in certain situations the conformational sensitivity can be altered through the appropriate choice of the EVA sigma parameter.

Models, Molecular↗

Integrated image analysis solutions for PET datasets in damaged brain.

OBJECTIVE: To identify and discuss the problems inherent in the processing of multiparametric functional imaging datasets from patients with acute brain injury, using "triple oxygen" positron emission tomography (PET) as an example. To present an integrated imaging solution for analysis of such datasets and report on its use in practice. METHODS: Structured analysis of the steps in image analysis for triple oxygen studies in patients with acute brain injury, subarachnoid haemorrhage and carotid artery disease was undertaken. Systematic identification of the drawbacks of conventional manual processing techniques, which make use of modules from different commercially available software packages was carried out. This analysis informed the construction of an image analysis gateway that addressed these issues. RESULTS: We describe an integrated image analysis gateway (PETAn) that provides for the automated integration of all processing steps, along with outputs that detail analysis stages, error logs and results. Implementation of PETAn allowed batch processing of clinical PET data sets, reduced the level of technical expertise required for analysis, improved the quality of results and achieved significant reductions in operator analysis time. CONCLUSIONS: While modules from several imaging software suites can be usefully combined to allow analysis of multiparametric PET data sets from patients with acute brain injury, sequential use of these modules for various steps in image analysis is fraught with difficulty. Careful automation and integration of these steps greatly facilitates the interrogation of information-rich data sets and increases research productivity.

Algorithms↗

Analysis of the human protein interactome and comparison with yeast, worm and fly interaction datasets.

We present the first analysis of the human proteome with regard to interactions between proteins. We also compare the human interactome with the available interaction datasets from yeast (Saccharomyces cerevisiae), worm (Caenorhabditis elegans) and fly (Drosophila melanogaster). Of >70,000 binary interactions, only 42 were common to human, worm and fly, and only 16 were common to all four datasets. An additional 36 interactions were common to fly and worm but were not observed in humans, although a coimmunoprecipitation assay showed that 9 of the interactions do occur in humans. A re-examination of the connectivity of essential genes in yeast and humans indicated that the available data do not support the presumption that the number of interaction partners can accurately predict whether a gene is essential. Finally, we found that proteins encoded by genes mutated in inherited genetic disorders are likely to interact with proteins known to cause similar disorders, suggesting the existence of disease subnetworks. The human interaction map constructed from our analysis should facilitate an integrative systems biology approach to elucidating the cellular networks that contribute to health and disease states.

Animals↗

Prediction of 3-yr cadaveric graft survival based on pre-transplant variables in a large national dataset.

Pre- and post-transplant predictive factors of graft survival for optimal and expanded criteria grafts have been studied in the past. The goal of our study was to evaluate the recent large set of United Network of Organ Sharing records (1990-1998) to generate a prediction algorithm of 3-yr graft survival based on pre-transplant variables alone. The dataset of patients with end-stage renal disease and cadaveric kidney or kidney-pancreas transplantation (1990-1998) used in the study consisted of 37,407 records. Logistic regression (LM) and a tree-based model (TBM) were used to identify predictors of 3-yr allograft survival and to generate prediction algorithm. Donor and recipient demographic characteristics (age, race, and gender) and body mass index showed non-linear, while human leukocyte antigen match showed strong linear relationships with 3-yr graft survival. Prediction of the probability of graft survival from the model, achieved a good match with the observed survival of the separate dataset, with a correlation of r = 0.998 for LM and r = 0.984 for TBM. The positive predictive value (PV) of allograft survival with LM and TBM was 76.0% and the negative PV was 63 and 53.8% for LM and TBM, respectively. Both LM and the TBM can potentially be used in clinical practice for long-term prediction of kidney allograft survival based on pre-transplant variables.

Algorithms↗

Comprehensive analysis of a multidimensional liquid chromatography mass spectrometry dataset acquired on a quadrupole selecting, quadrupole collision cell, time-of-flight mass spectrometer: I. How much of the data is theoretically interpretable by search engines?

An in-depth analysis of a multidimensional chromatography-mass spectrometry dataset acquired on a quadrupole selecting, quadrupole collision cell, time-of-flight (QqTOF) geometry instrument was carried out. A total of 3269 CID spectra were acquired. Through manual verification of database search results and de novo interpretation of spectra 2368 spectra could be confidently determined as predicted tryptic peptides. A detailed analysis of the non-matching spectra was also carried out, highlighting what the non-matching spectra in a database search typically are composed of. The results of this comprehensive dataset study demonstrate that QqTOF instruments produce information-rich data of which a high percentage of the data is readily interpretable.

Chromatography, Liquid↗

RR-interval-based atrial fibrillation detection and burden estimation: cross-dataset validation and calibration-aware probability analysis.

Objective.Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift.Approach.This study developed an interpretable-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error (MAE) and agreement analyses.Main results.The model showed high discrimination in both development and external evaluation, with external ROC-AUC ofand PR-AUC of. However, external calibration deteriorated despite preserved ranking performance, with Brier score of, ECE(15) of, calibration intercept of, and calibration slope of. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with MAE ofversus, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased fromwithout recalibration toafter Platt recalibration andafter isotonic recalibration, while median ECE(15) decreased fromtoand, respectively.Significance.These findings indicate that-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.

Atrial Fibrillation↗

Identification of characteristic oligonucleotides in the bacterial 16S ribosomal RNA sequence dataset.

MOTIVATION: The phylogenetic structure of the bacterial world has been intensively studied by comparing sequences of 16S ribosomal RNA (16S rRNA). This database of sequences is now widely used to design probes for the detection of specific bacteria or groups of bacteria one at a time. The success of such methods reflects the fact that there are local sequence segments that are highly characteristic of particular organisms or groups of organisms. It is not clear, however, the extent to which such signature sequences exist in the 16S rRNA dataset. A better understanding of the numbers and distribution of highly informative oligonucleotide sequences may facilitate the design of hybridization arrays that can characterize the phylogenetic position of an unknown organism or serve as the basis for the development of novel approaches for use in bacterial identification. RESULTS: A computer-based algorithm that characterizes the extent to which any individual oligonucleotide sequence in 16S rRNA is characteristic of any particular bacterial grouping was developed. A measure of signature quality, Q(s), was formulated and subsequently calculated for every individual oligonucleotide sequence in the size range of 5-11 nucleotides and for 15mers with reference to each cluster and subcluster in a 929 organism representative phylogenetic tree. Subsequently, the perfect signature sequences were compared to the full set of 7322 sequences to see how common false positives were. The work completed here establishes beyond any doubt that highly characteristic oligonucleotides exist in the bacterial 16S rRNA sequence dataset in large numbers. Over 16,000 15mers were identified that might be useful as signatures. Signature oligonucleotides are available for over 80% of the nodes in the representative tree.

Algorithms↗