Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Texture mapping based visualisation methods for the manipulation of CT data: interaction and ergonomics.

The manipulation of large CT datasets needs fast visualisation methods for a comfortable user interaction. Modern visualisation techniques make use texture hardware in graphics workstations extensively. In the following we will present an interactive tool for the positioning of anatomical landmarks in CT datasets of non-pathological children. The tool includes a fast visualisation of CT cross sections based on a texture mapping technique and an interactive three-dimensional view of the segmented CT dataset.

Cephalometry↗

Manipulation of volumetric patient data in a distributed virtual reality environment.

Due to increases in network speed and bandwidth, distributed exploration of medical data in immersive Virtual Reality (VR) environments is becoming increasingly feasible. The volumetric display of radiological data in such environments presents a unique set of challenges. The shear size and complexity of the datasets involved not only make them difficult to transmit to remote sites, but these datasets also require extensive user interaction in order to make them understandable to the investigator and manageable to the rendering hardware. A sophisticated VR user interface is required in order for the clinician to focus on the aspects of the data that will provide educational and/or diagnostic insight. We will describe a software system of data acquisition, data display, Tele-Immersion, and data manipulation that supports interactive, collaborative investigation of large radiological datasets. The hardware required in this strategy is still at the high-end of the graphics workstation market. Future software ports to Linux and NT, along with the rapid development of PC graphics cards, open the possibility for later work with Linux or NT PCs and PC clusters.

Computer Communication Networks↗

Calculation of the left ventricular ejection fraction without edge detection: application to small hearts.

UNLABELLED: Quantitative gated SPECT (QGS) software has been reported to overestimate the left ventricular ejection fraction (LVEF) in patients with small hearts. This finding is caused by the inaccurate detection of the endocardial surface of the left ventricle (LV) due to low resolution and partial-volume effects. In this article we develop a method to calculate the LVEF from gated SPECT data without edge detection and compare it with the QGS method of calculating the LVEF. METHODS: The short-axis images were transformed to the prolate spheroid coordinate system, and detection of the layer of maximum counts (a surface area of maximum counts) was made. First, the volume enclosed by the layer of maximum counts (V(max)) was calculated; then the corresponding ejection fraction [(LVEF)(max)] was calculated. The LVEF was calculated by multiplying the (LVEF)(max) by a constant factor, which was determined from a series of calculations made using QGS on larger hearts. In computer simulations the end-diastolic left ventricular volume (EDV) and the targeted LVEF (tLVEF) were varied to produce LVs of different sizes. The LVs were modeled by 2 confocal hemiellipsoids with 7 different EDVs. The tLVEF was increased from 25% to 75%, in 5% step-size increments, for a total of 11 different ejection fractions. These datasets were then smoothed, creating a total of 77 smoothed sets. The smoothed images were processed by the QGS method and by our method. In patient studies, 58 patient datasets were processed by the QGS method and by our method. No attenuation correction was performed on these datasets. The patients were divided into 2 groups: 44 patients with large hearts (EDV > or = 80 mL) and 14 patients with small hearts (EDV < 80 mL). RESULTS: In computer simulations, the QGS method and our method performed well when imaging large EDVs (EDV > or = 80 mL). Our method derived better results than did the QGS method for small EDVs. In patient studies the LVEF calculated by our method matched well with the QGS LVEF in the 44 patients with large hearts. The correlation coefficient between them was found to be 0.957. Of the 14 patients with small hearts, the LVEFs of 5 patients were severely overestimated by the QGS method compared with the results obtained with our method. CONCLUSION: It is possible to calculate the LVEF without edge detection. Compared with QGS LVEF, our method gave better results for small LVs in computer simulations.

Algorithms↗

Meta-analysis of microarrays: interstudy validation of gene expression profiles reveals pathway dysregulation in prostate cancer.

The increasing availability and maturity of DNA microarray technology has led to an explosion of cancer profiling studies. To extract maximum value from the accumulating mass of publicly available cancer gene expression data, methods are needed to evaluate, integrate, and intervalidate multiple datasets. Here we demonstrate a statistical model for performing meta-analysis of independent microarray datasets. Implementation of this model revealed that four prostate cancer gene expression datasets shared significantly similar results, independent of the method and technology used (i.e., spotted cDNA versus oligonucleotide). This interstudy cross-validation approach generated a cohort of genes that were consistently and significantly dysregulated in prostate cancer. Bioinformatic investigation of these genes revealed a synchronous network of transcriptional regulation in the polyamine and purine biosynthesis pathways. Beyond the specific implications for prostate cancer, this work establishes a much-needed model for the evaluation, cross-validation, and comparison of multiple cancer profiling studies.

Adenosine Monophosphate↗

Establishing connections between microarray expression data and chemotherapeutic cancer pharmacology.

We have investigated three different microarray datasets of approximately 6 K gene expressions across the National Cancer Institute's panel of 60 tumor cell lines. Initial assessments of reproducibility for gene expressions within each dataset, as derived from sequence analysis of full-length sequences as well as expressed sequence tags (EST), found statistically significant results for no more than 36% of those cases where at least one replicate of a gene appears on the array. Filtering the data based only on pairwise comparisons among these three datasets creates a list of approximately 400 significant concordant expression patterns. The expression profiles of these smaller sets of genes were used to locate similar expression profiles of synthetic agents screened against these same 60 tumor cell lines. A correspondence was found between mRNA expression patterns and 50% growth inhibition response patterns of screened agents for 11 cases that were subsequently verifiable from ligand-target crystallographic data. Notable amongst these cases are genes encoding a variety of kinases, which were also found to be targets of small drug-like molecules within the database of protein structures. These 11 cases lend support to the premise that similarities between expression patterns and chemical responses for the National Cancer Institute's tumor panel can be related to known cases of molecular structure and putative cellular function. The details of the 11 verifiable cases and the concordant gene subsets are provided. Discussions about the prospects of using this approach as a data mining tool are included.

Algorithms↗

[Describing language of spectra and rough set].

It is the traditional way to analyze spectra by experiences in astronomical field. And until now there has never been a suitable theoretical frame to describe spectra, which is may be owing to small spectra datasets that astronomers can get by low-level instruments. With the high-speed development of telescopes, especially on behalf of LAMOST, a large telescope which can collect more than 20,000 spectra in an observing night, spectra datasets are becoming larger and larger very fast. Facing these voluminous datasets, the traditional spectra-processing way simply depending on experiences becomes unfit. In this paper, we develop a brand-new language--describing language of spectra (DLS) to describe spectra of celestial bodies by defining BE (Basic element). And based on DLS, we introduce the method of RSDA (Rough set and data analysis), which is a technique of data mining. By RSDA we extract some rules of stellar spectra, and this experiment can be regarded as an application of DLS.

Algorithms↗

Predicting ADME properties and side effects: the BioPrint approach.

Computational methods are increasingly used to streamline and enhance the lead discovery and optimization process. However, accurate prediction of absorption, distribution, metabolism and excretion (ADME) and adverse drug reactions (ADR) is often difficult, due to the complexity of underlying physiological mechanisms. Modeling approaches have been hampered by the lack of large, robust and standardized training datasets. In an extensive effort to build such a dataset, the BioPrint database was constructed by systematic profiling of nearly all drugs available on the market, as well as numerous reference compounds. The database is composed of several large datasets: compound structures and molecular descriptors, in vitro ADME and pharmacology profiles, and complementary clinical data including therapeutic use information, pharmacokinetics profiles and ADR profiles. These data have allowed the development of computational tools designed to integrate a program of computational chemistry into library design and lead development. Models based on chemical structure are strengthened by in vitro results that can be used as additional compound descriptors to predict complex in vivo endpoints. The BioPrint pharmacoinformatics platform represents a systematic effort to accelerate the process of drug discovery, improve quantitative structure-activity relationships and develop in vitro/in vivo associations. In this review, we will discuss the importance of training set size and diversity in model development, the implementation of linear and neighborhood modeling approaches, and the use of in silico methods to predict potential clinical liabilities.

Animals↗

Expression of matrix metalloproteinase 9 (MMP-9/gelatinase B) in adenocarcinomas strongly correlated with expression of immune response genes.

Matrix metalloproteinases (MMPs) are endopeptidases considered to be important regulators of the microenvironment of cancer. While MMPs are traditionally associated with the extracellular matrix (ECM), here we provide new evidence from an analysis of gene expression profiles from human tumor tissue that MMP-9 (gelatinase B) is associated with elements of the immune system in a way analogous to the association of other MMPs, such as MMP-2 (gelatinase A), with components of the ECM. An analysis of three independent microarray datasets of lung adenocarcinomas from previous studies [Nat. Med. 8, 816-824 (2002); Proc. Natl. Acad. Sci. USA 98, 13790-13795 (2001); Proc. Natl. Acad. Sci. USA, 98, 13784-13789 (2001)] showed that, in each dataset, out of the set of genes with significant correlations in mRNA expression to the expression of MMP9 (P < 0.005), a highly disproportionate number were found to be annotated in the Locuslink database as having a role in the anti-pathogen response. By comparison, out of the set of genes significantly correlated with the expression of MMP2, a highly disproportionate number were known components of the ECM. The same patterns observed in the lung data for both MMP2 and MMP9 were found as well in an additional published dataset of colon and ovarian adenocarcinomas [Am. J. Pathol. 159, 1231-1238 (2001)]. The results of this study suggest a greater functional role for MMP-9 in the immune response to cancer than what may previously have been thought.

Adenocarcinoma↗

Temporal correlation between opiate seizures in East/Southeast Asia and B.C. heroin deaths: a transoceanic model of heroin death risk.

BACKGROUND: Because heroin supply changes cannot be measured directly, their impact on populations is poorly understood. British Columbia has experienced an injection drug use epidemic since the 1980s that resulted in 2,590 illicit drug deaths from 1990-1999. Since previous work indicates heroin seizures can correlate with supply and B.C. receives heroin only from Southeast Asia, this study examined B.C. heroin deaths against opiate seizures in East/Southeast Asia. METHODS: Opiate seizures in East/Southeast Asia and data from two B.C. mortality datasets containing heroin deaths were examined. The Pearson correlation coefficient for seizures against each mortality dataset was determined. RESULTS: Opiate seizures, all illicit drug deaths and all opiate deaths concurrently increased twice and decreased twice from 1989-1999, and all reached new peak values in 1993. Three B.C. sub-regions exhibited illicit drug deaths rate trends concurrent with the three principal datasets studied. The Pearson correlation coefficient for opiate-induced deaths against opiate seizures from 1980-1999 was R=0.915 (p<0.0001), and for illicit drug deaths against opiate seizures from 1987-1999 was R=0.896 (p<0.0001). CONCLUSIONS: From 1980-1999, opiate seizures in East/Southeast Asia were very strongly correlated with B.C. opiate and illicit drug deaths. The number of B.C. heroin-related deaths may be strongly linked to heroin supply. Enforcement services are not effective in preventing harm caused by heroin in B.C.; therefore, Canada should examine other methods to prevent harm. The case for harm reduction is strengthened by the ineffectiveness of enforcement and the unlikelihood of imminent eradication of heroin production in Southeast Asia.

Asia, Southeastern↗

Integrating syndromic surveillance data across multiple locations: effects on outbreak detection performance.

Syndromic surveillance systems are being deployed widely to monitor for signals of covert bioterrorist attacks. Regional systems are being established through the integration of local surveillance data across multiple facilities. We studied how different methods of data integration affect outbreak detection performance. We used a simulation relying on a semi-synthetic dataset, introducing simulated outbreaks of different sizes into historical visit data from two hospitals. In one simulation, we introduced the synthetic outbreak evenly into both hospital datasets (aggregate model). In the second, the outbreak was introduced into only one or the other of the hospital datasets (local model). We found that the aggregate model had a higher sensitivity for detecting outbreaks that were evenly distributed between the hospitals. However, for outbreaks that were localized to one facility, maintaining individual models for each location proved to be better. Given the complementary benefits offered by both approaches, the results suggest building a hybrid system that includes both individual models for each location, and an aggregate model that combines all the data. We also discuss options for multi-level signal integration hierarchies.

Bioterrorism↗

Text mining of DNA sequence homology searches.

Primary tasks in analysis and annotation of expressed sequence tag (EST) datasets are to identify similarity among sequences by unsupervised clustering and assign putative function based on BLAST homology searches. We investigated the usefulness of text mining as a simple approach for further higher-level clustering of EST datasets using IBM Intelligent Miner for Text v2.3 tools. Agglomerative and k-means clustering tools were used to cluster BLASTx homology search documents from two onion EST datasets and optimised by pre-processing and pruning. Subjective evaluation confirmed that these tools provided biologically useful and complementary views of the two libraries, provided new insights into their composition and revealed clusters previously identified by human experts. We compared BLASTx textual clusters for two gene families with their DNA sequence-based clusters and confirmed that these shared similar morphology.

Abstracting and Indexing↗

Abductive network committees for improved classification of medical data.

OBJECTIVES: To introduce abductive network classifier committees as an ensemble method for improving classification accuracy in medical diagnosis. While neural networks allow many ways to introduce enough diversity among member models to improve performance when forming a committee, the self-organizing, automatic-stopping nature, and learning approach used by abductive networks are not very conducive for this purpose. We explore ways of over-coming this limitation and demonstrate improved classification on three standard medical datasets. METHODS: Two standard 2-class medical datasets (Pima Indians Diabetes and Heart Disease) and a 6-class dataset (Dermatology) were used to investigate ways of training abductive networks with adequate independence, as well as methods of combining their outputs to form a network that improves performance beyond that of single models. RESULTS: Two- or three-member committees of models trained on completely or partially different subsets of training data and using simple output combination methods achieve improvements between 2 and 5 percentage points in the classification accuracy over the best single model developed using the full training set. CONCLUSIONS: Varying model complexity alone gives abductive network models that are too correlated to ensure enough diversity for forming a useful committee. Diversity achieved through training member networks on independent subsets of the training data outweighs limitations of the smaller training set for each, resulting in net gain in committee performance. As such models train faster and can be trained in parallel, this can also speed up classifier development.

Data Collection↗

A technique for standardized central analysis of 6-(18)F-fluoro-L-DOPA PET data from a multicenter study.

UNLABELLED: We have recently completed a large 6-(18)F-fluoro-L-DOPA ((18)F-DOPA) PET study comparing rates of loss of dopamine terminal function in Parkinson's disease (PD) patients taking either the dopamine agonist ropinirole or L-DOPA. This trial involved a "distributed acquisition/centralized analysis" method, in which (18)F-DOPA images were acquired at 6 different PET centers around the world and then analyzed at a single site. To our knowledge, this is the first time such a centralized approach has been employed with (18)F-DOPA PET and this descriptive basic science article outlines the methods used. METHODS: One hundred eighty-six PD patients were randomized (1:1) to ropinirole or L-DOPA therapy, and (18)F-DOPA PET was performed at baseline and again at 2 y. The primary outcome measure was the percentage change in putamen (18)F-DOPA influx rate constant (K(i)) from Patlak graphical analysis. Dynamic images were acquired and reconstructed using each center's individual protocols before being transferred to the site performing the central analysis. Once there, individual parametric K(i) images were created using a single analysis program without file formats being transformed from the original. Parametric images were then normalized to standard space and K(i) values extracted with a region of interest analysis. Significant K(i) changes were also localized at a voxel level with statistical parametric mapping. These processes required numerous checks to ensure the integrity of each dataset. RESULTS: Three hundred twenty-five (170 baseline, 155 follow-up) dynamic PET datasets were acquired, of which 12 were considered uninterpretable due to missing time frames, radiopharmaceutical problems, lack of measured attenuation correction, or excessive head movement. In those datasets suitable for central analysis, after quality control and spatial normalization of the images had been applied, putamen (18)F-DOPA signal decline was found to be significantly (one third) slower in the ropinirole group compared with that of the L-DOPA group. CONCLUSION: Paired (18)F-DOPA-PET images acquired from multiple sites can be successfully analyzed centrally to assess the efficacy of potential disease-modifying therapies in PD. However, numerous options must be considered and data checks put in place before adopting such an approach. Centralized analysis offers the potential for improved detection of outcomes due to the standardization of the analytic approach and allows the analysis of large numbers of PET studies.

Algorithms↗

A novel algorithm for scalable and accurate Bayesian network learning.

Bayesian Networks (BN) is a knowledge representation formalism that has been proven to be valuable in biomedicine for constructing decision support systems and for generating causal hypotheses from data. Given the emergence of datasets in medicine and biology with thousands of variables and that current algorithms do not scale more than a few hundred variables in practical domains, new efficient and accurate algorithms are needed to learn high quality BNs from data. We present a new algorithm called Max-Min Hill-Climbing (MMHC) that builds upon and improves the Sparse Candidate (SC) algorithm; a state-of-the-art algorithm that scales up to datasets involving hundreds of variables provided the generating networks are sparse. Compared to the SC, on a number of datasets from medicine and biology, (a) MMHC discovers BNs that are structurally closer to the data-generating BN, (b) the discovered networks are more probable given the data, (c) MMHC is computationally more efficient and scalable than SC, and (d) the generating networks are not required to be uniformly sparse nor is the user of MMHC required to guess correctly the network connectivity

Algorithms↗

Comparison of machine learning techniques with classical statistical models in predicting health outcomes.

Several machine learning techniques (multilayer and single layer perceptron, logistic regression, least square linear separation and support vector machines) are applied to calculate the risk of death from two biomedical data sets, one from patient care records, and another from a population survey. Each dataset contained multiple sources of information: history of related symptoms and other illnesses, physical examination findings, laboratory tests, medications (patient records dataset), health attitudes, and disabilities in activities of daily living (survey dataset). Each technique showed very good mortality prediction in the acute patients data sample (AUC up to 0.89) and fair prediction accuracy for six year mortality (AUC from 0.70 to 0.76) in individuals from epidemiological database surveys. The results suggest that the nature of data is of primary importance rather than the learning technique. However, the consistently superior performance of the artificial neural network (multi-layer perceptron) indicates that nonlinear relationships (which cannot be discerned by linear separation techniques) can provide additional improvement in correctly predicting health outcomes.

Aged↗

Adaptive hybrid interpolation techniques for direct Haptic rendering of isosurfaces.

Direct Haptic rendering of voxels from an anatomical dataset provides patient specific haptic feedback vital for diagnosis and surgical planning. Our algorithm uses zero sets of scalar trivariate function for polynomial interpolation with sixty-four neighborhood points to generate isosurfaces on the fly for haptic rendering. This approach gives continuity in surfaces as well as better capture of isosurface features of the medical dataset. The detailed algorithm is presented along with the description of results from haptically rendering medical datasets.

Algorithms↗

[Theory, method and application of method R on estimation of (co)variance components].

Theory, method and application of Method R on estimation of (co)variance components were reviewed in order to make the method be reasonably used. Estimation requires R values,which are regressions of predicted random effects that are calculated using complete dataset on predicted random effects that are calculated using random subsets of the same data. By using multivariate iteration algorithm based on a transformation matrix,and combining with the preconditioned conjugate gradient to solve the mixed model equations, the computation efficiency of Method R is much improved. Method R is computationally inexpensive,and the sampling errors and approximate credible intervals of estimates can be obtained. Disadvantages of Method R include a larger sampling variance than other methods for the same data,and biased estimates in small datasets. As an alternative method, Method R can be used in larger datasets. It is necessary to study its theoretical properties and broaden its application range further.

Algorithms↗

Automatic method to assess local CT-MR imaging registration accuracy on images of the head.

BACKGROUND AND PURPOSE: Precise registration of CT and MR images is crucial in many clinical cases for proper diagnosis, decision making or navigation in surgical interventions. Various algorithms can be used to register CT and MR datasets, but prior to clinical use the result must be validated. To evaluate the registration result by visual inspection is tiring and time-consuming. We propose a new automatic registration assessment method, which provides the user a color-coded fused representation of the CT and MR images, and indicates the location and extent of poor registration accuracy. METHODS: The method for local assessment of CT-MR registration is based on segmentation of bone structures in the CT and MR images, followed by a voxel correspondence analysis. The result is represented as a color-coded overlay. The algorithm was tested on simulated and real datasets with different levels of noise and intensity non-uniformity. RESULTS: Based on tests on simulated MR imaging data, it was found that the algorithm was robust for noise levels up to 7% and intensity non-uniformities up to 20% of the full intensity scale. Due to the inability to distinguish clearly between bone and cerebro-spinal fluids in the MR image (T1-weighted), the algorithm was found to be optimistic in the sense that a number of voxels are classified as well-registered although they should not. However, nearly all voxels classified as misregistered are correctly classified. CONCLUSION: The proposed algorithm offers a new way to automatically assess the CT-MR image registration accuracy locally in all the areas of the volume that contain bone and to represent the result with a user-friendly, intuitive color-coded overlay on the fused dataset.

Algorithms↗