Unlocking datasets.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
A nonlinear 0/1 mixed integer programming model is presented for a constrained discriminant analysis problem. The model enables controlling misclassification probabilities by placing restrictions on the numbers of misclassifications allowed among the training entities and incorporating a "reserved-judgment" region to which entities whose classifications are difficult to determine may be allocated. A linearization of the model is given, and preliminary numerical results for two medical and one non medical domain are presented.
Explore the source record for details and available documents.
The objective of this study was to compare and contrast two techniques of modeling mortality in a 30 bed multi-disciplinary ICU; neural networks and logistic regression. Fifteen physiological variables were recorded on day 3 for 422 consecutive patients whose duration of stay was over 72 hours. Two separate models were built using each technique. First, logistic and neural network models were constructed on the complete 422 patient dataset and discrimination was compared. Second, the database was randomly divided into a 284 patient developmental dataset and a 138 patient validation dataset. The developmental dataset was used to construct logistic and neural net models and the predictive power of these models was verified on the validation dataset. On the complete dataset, the neural network clearly outperformed the logistic model (sensitivity and specificity of 1 and .997 vs. .525 and .966, area under ROC curve .9993 vs. .9259), while both performed equally well on the validation dataset (area under ROC of .82). The excellent performance of the neural net on the complete dataset reveals that the problem is classifiable. Since our dataset only contained 40 mortality events, it is highly likely that the validation dataset was not representative of the developmental dataset, which led to a decreased predictive performance by both the neural net and the logistic regression models. Theoretically, given an extensive dataset, the neural network should be able to perform mortality prediction with a sensitivity and a specificity approaching 95%. Clinically, this would be an extremely important achievement.(ABSTRACT TRUNCATED AT 250 WORDS)
BACKGROUND: Microarray technology has made it possible to simultaneously measure the expression levels of large numbers of genes in a short time. Gene expression data is information rich; however, extensive data mining is required to identify the patterns that characterize the underlying mechanisms of action. Clustering is an important tool for finding groups of genes with similar expression patterns in microarray data analysis. However, hard clustering methods, which assign each gene exactly to one cluster, are poorly suited to the analysis of microarray datasets because in such datasets the clusters of genes frequently overlap. RESULTS: In this study we applied the fuzzy partitional clustering method known as Fuzzy C-Means (FCM) to overcome the limitations of hard clustering. To identify the effect of data normalization, we used three normalization methods, the two common scale and location transformations and Lowess normalization methods, to normalize three microarray datasets and three simulated datasets. First we determined the optimal parameters for FCM clustering. We found that the optimal fuzzification parameter in the FCM analysis of a microarray dataset depended on the normalization method applied to the dataset during preprocessing. We additionally evaluated the effect of normalization of noisy datasets on the results obtained when hard clustering or FCM clustering was applied to those datasets. The effects of normalization were evaluated using both simulated datasets and microarray datasets. A comparative analysis showed that the clustering results depended on the normalization method used and the noisiness of the data. In particular, the selection of the fuzzification parameter value for the FCM method was sensitive to the normalization method used for datasets with large variations across samples. CONCLUSION: Lowess normalization is more robust for clustering of genes from general microarray data than the two common scale and location adjustment methods when samples have varying expression patterns or are noisy. In particular, the FCM method slightly outperformed the hard clustering methods when the expression patterns of genes overlapped and was advantageous in finding co-regulated genes. Thus, the FCM approach offers a convenient method for finding subsets of genes that are strongly associated to a given cluster.
Sixteen reference strains and thirteen fresh isolates of three putatively novel Streptomyces species were examined six times over twenty months using pyrolysis mass spectrometry to examine the long-term reproducibility of the procedure. The reference strains and new isolates were correctly identified using information in each of the datasets and operational fingerprinting, but direct statistical comparison of the datasets for strain identification was unsuccessful between datasets. Artificial neural networks were also used to identify the strains held in the datasets. Neural networks trained with pyrolysis mass spectra from a single dataset were found to successfully identify the reference strains and fresh isolates in that dataset but were unable to identify many of the strains in the other datasets. However, a neural network trained on representative pyrolysis mass spectra from each of the first three datasets were found to identify the reference strains and fresh isolates in those three datasets and in the three subsequent datasets. Therefore, artificial neural network analysis of pyrolysis mass spectrometric data can provide a rapid, cost-effective, accurate and long-term reproducible way of identifying and typing microorganisms.
BACKGROUND: Previously, we reported significant linkage of body mass index (BMI) to chromosomes 6 and 11 across six examinations, covering 28 years, of the Framingham Heart Study. These results were on all individuals available at each exam, thus the sample size varied from exam to exam. To remove any effect of sample size variation we have now constructed six subsets; for each exam individuals were only included if they were measured at every exam, i.e. for each exam, included individuals comprise the intersection of the original six exams. This strategy preferentially removed older individuals who died before reaching the sixth exam, thus the intersection datasets are smaller (n = 1114) and significantly younger than the full datasets. We performed variance components linkage analysis on these intersection datasets and on their sex-specific subsets. RESULTS: Results from the sex-specific genome scans revealed 11 regions in which a sex-specific maximum lodscore was at least 2.0 for at least one dataset. Randomization tests indicated that all 11 regions had significant (p < 0.05) differences in sex-specific maximum lodscores for at least three datasets. The strongest sex-specific linkage was for men on chromosome 16 with maximum lodscores 2.70, 3.00, 3.42, 3.61, 2.56 and 1.93 for datasets 1-6 respectively. Results from the full genome scans revealed that linked regions on chromosomes 6 and 11 remained significantly and consistently linked in the intersection datasets. Surprisingly, the maximum lodscore on chromosome 10 for dataset 1 increased from 0.97 in the older original dataset to 4.23 in the younger smaller intersection dataset. This difference in maximum lodscores was highly significant (p < 0.0001), implying that the effect of this chromosome may vary with age. Age effects may also exist for the linked regions on chromosomes 6 and 11. CONCLUSION: Sex specific effects of chromosomal regions on BMI are common in the Framingham study. Some evidence also exists for age-specific effects of chromosomal regions.
A bivariate threshold-linear (TL) and a bivariate linear-linear (LL) model were assessed for the genetic analysis of 56-d nonreturn (NR56) and interval from calving to first insemination (CFI) in first-lactation Norwegian Red (former Norwegian Dairy Cattle) (NRF). Three different datasets were used to infer genetic parameters and to predict transmitting abilities for NRF sires. Mean progeny group sizes were 147.8, 102.7, and 56.5 daughters, and the corresponding number of sires were 746, 743, and 742 in the 3 datasets. Otherwise, the structures of the 3 datasets were similar. When the TL model was used, heritability of liability to NR56 was 2.8% in the 2 larger datasets and 3.8% in the smallest dataset. In the LL model, the heritability of NR56 in the largest dataset and in the 2 smaller datasets was 1.2 and 0.9%, respectively. For CFI, the heritability was similar in TL and LL models, ranging from 2.4 to 2.7%. The small heritability of the 2 reproductive traits implies that most of the variation is environmental and that large progeny groups are required to get accurate sire PTA. The point estimates of the genetic correlation between NR56 and CFI were near zero in both models. The 2 bivariate models were compared in terms of predictive ability using logistic regression and a chi2 statistic based on differences between observed and predicted outcomes for NR56 in a separate dataset. Comparison was also with respect to ranking of sires and correlations between sire posterior means (TL model) and PTA (LL model). We found very small differences in ability to predict NR56 between the 2 bivariate models, regardless of the dataset used. Correlations between sire posterior means (TL) and sire PTA (LL) and rank correlations between sire evaluations were all >0.98 in the 3 datasets. At present, the LL model is preferred for sire evaluations of NR56 and CFI in NRF. This is because the LL model is less computationally demanding and more robust with respect to the structure of the data than TL.
With an increasing number of publicly available microarray datasets, it becomes attractive to borrow information from other relevant studies to have more reliable and powerful analysis of a given dataset. We do not assume that subjects in the current study and other relevant studies are drawn from the same population as assumed by meta-analysis. In particular, the set of parameters in the current study may be different from that of the other studies. We consider sample classification based on gene expression profiles in this context. We propose two new methods, a weighted partial least squares (WPLS) method and a weighted penalized partial least squares (WPPLS) method, to build a classifier by a combined use of multiple datasets. The methods can weight the individual datasets depending on their relevance to the current study. A more standard approach is first to build a classifier using each of the individual datasets, then to combine the outputs of the multiple classifiers using a weighted voting. Using two quite different datasets on human heart failure, we show first that WPLS/WPPLS, by borrowing information from the other dataset, can improve the performance of PLS/PPLS built on only a single dataset. Second, WPLS/WPPLS performs better than the standard approach of combining multiple classifiers. Third, WPPLS can improve over WPLS, just as PPLS does over PLS for a single dataset.
PURPOSE: Motion artifacts can significantly deteriorate the precision of a computer-assisted surgical intervention because they destroy the isometric representation of tomographic pictures. In the context of a study, the influence of typical motion artifacts on the precision of markerless laser registration in image-guided oral and maxillofacial surgery was analyzed, and quality factors for evaluation of the isometry of a computed tomography (CT) dataset were determined. PATIENTS AND METHODS: Twenty patients underwent markerless registration, the precision being determined by means of intraoral evaluation markers. Then the 20 CT datasets were used for simulation of a typical motion artifact. The precision of the overlay of the dataset was checked again on the navigation workstation, in absence of the patient. The navigation system used was the Surgical Segment Navigator SSN++ (University of Heidelberg, Heidelberg, Germany). RESULTS: The motion artifacts reduced the average patient registration from 1 to 4 mm. Quality factors for the isometry of a CT dataset were: the volume enclosed between the soft tissue mantles of the preoperative CT dataset and the intraoperative laser scan dataset, as well as the orientation of the normal vectors on the 3-dimensional reconstruction of the CT dataset. CONCLUSION: The isometry of a CT dataset should always be checked before performance of a computer assisted surgical intervention because anisometric datasets result in inaccurate patient registration and navigation.
BACKGROUND AND PURPOSE: Comparable, standardised data on the quality and efficiency of stroke care in Germany are lacking. The Arbeitsgemeinschaft Deutscher Schlaganfall-Register (ADSR--German Stroke Registries Study Group) has defined a "Minimum DataSet" for the evaluation of quality indicators of stroke treatment in Germany. METHODS: The ADSR is a voluntary network of current regional stroke registries aiming at a standardisation in the use of stroke terminology and methods of data collection for German stroke databases. Currently six regional stroke registries are cooperating in the ADSR, combining data from about 18,000 stroke patients in 110 hospitals annually. RESULTS: In the design of the ADSR DataSet a modular approach was chosen. The ADSR "Minimum DataSet" was adapted for a wide use in different health care facilities. In addition to the "Minimum DataSet" an "Advanced DataSet" will be developed to document additional items of stroke care in specialised stroke centres. The ADSR DataSet collection will be completed by special "Extended DataSets", designed for answering centre-specific and research questions. CONCLUSION: The ADSR "Minimum DataSet" allows a standardised assessment of stroke care in Germany. It is the first questionnaire that provides valid and reliable comparisons between different clinical settings as well as regional stroke databases. The ADSR "Minimum DataSet" defines core items for a future National German Health Report on stroke care.
OBJECTIVE: To evaluate automatic vessel tracking techniques in the course of preoperative planning prior to transluminal aortic endograft implantation by comparing accuracy, reproducibility, and postprocessing time with source image and volume-rendered analysis methods. METHODS: Multislice computed tomography datasets of 5 patients with abdominal aortic aneurysms were preoperatively examined, performing volumetric analysis of diameter and position of renal artery orifices, aneurysmal neck, maximal aneurysmal extension, aortic bifurcation, and iliac arteries and bifurcation. Analysis was realized by utilizing transverse datasets, volume rendering, and automated vessel tracking strategies (MxView, Philips, Best, The Netherlands). Measurement techniques were evaluated by 2 independent readers 3 times for each patient and measurement modality. Statistical analysis evaluated accuracy of the measurements and intra- and interobserver reliability. Postprocessing time was documented. RESULTS: Using transverse source datasets, intraobserver reliability ranged from 0.49 to 0.58. Intraobserver reliability improved to 0.7 to 0.98 when volume-rendered datasets were evaluated. Interobserver variability for transverse and volume-rendered datasets ranged from 0.49 to 0.76 and 0.70 to 0.96, respectively. Automated vessel tracking datasets did not demonstrate any intra- or interobserver variability. Based on transverse datasets, the length and diameter of iliac arteries and location and diameter of the aneurysmal neck were measured as statistically different in all cases in contrast to volume rendering and automated segmentation techniques. Postprocessing time consumption for measurements based on transverse, volume-rendered, and automated tracking segmentation datasets averaged 3.32 minutes, 25.43 minutes, and 2.24 minutes, respectively. CONCLUSIONS: Preoperative measurements improve significantly if datasets are evaluated based on volume-rendering techniques. This time-consuming procedure can be shortened, while further reducing observer variability, with automatic segmentation techniques.
BACKGROUND: Precise classification of cancer types is critically important for early cancer diagnosis and treatment. Numerous efforts have been made to use gene expression profiles to improve precision of tumor classification. However, reliable cancer-related signals are generally lacking. METHOD: Using recent datasets on colon and prostate cancer, a data transformation procedure from single gene expression to pair-wise gene expression ratio is proposed. Making use of the internal consistency of each expression profiling dataset this transformation improves the signal to noise ratio of the dataset and uncovers new relevant cancer-related signals (features). The efficiency in using the transformed dataset to perform normal/tumor classification was investigated using feature partitioning with informative features (gene annotation) as discriminating axes (single gene expression or pair-wise gene expression ratio). Classification results were compared to the original datasets for up to 10-feature model classifiers. RESULTS: 82 and 262 genes that have high correlation to tissue phenotype were selected from the colon and prostate datasets respectively. Remarkably, data transformation of the highly noisy expression data successfully led to lower the coefficient of variation (CV) for the within-class samples as well as improved the correlation with tissue phenotypes. The transformed dataset exhibited lower CV when compared to that of single gene expression. In the colon cancer set, the minimum CV decreased from 45.3% to 16.5%. In prostate cancer, comparable CV was achieved with and without transformation. This improvement in CV, coupled with the improved correlation between the pair-wise gene expression ratio and tissue phenotypes, yielded higher classification efficiency, especially with the colon dataset - from 87.1% to 93.5%. Over 90% of the top ten discriminating axes in both datasets showed significant improvement after data transformation. The high classification efficiency achieved suggested that there exist some cancer-related signals in the form of pair-wise gene expression ratio. CONCLUSION: The results from this study indicated that: 1) in the case when the pair-wise expression ratio transformation achieves lower CV and higher correlation to tissue phenotypes, a better classification of tissue type will follow. 2) the comparable classification accuracy achieved after data transformation suggested that pair-wise gene expression ratio between some pairs of genes can identify reliable markers for cancer.
Low-cost and high-throughput RNA sequencing data for barley RILs achieved GP performance comparable to or better than traditional SNP array datasets when combined with parental whole-genome sequencing SNP data. The field of genomic selection (GS) is advancing rapidly on many fronts including the utilization of multi-omics datasets with the goal of increasing prediction ability and becoming an integral part of an increasing number of breeding programs ensuring future food security. In this study, we used RNA sequencing (RNA-Seq) data to perform genomic prediction (GP) on three related barley RIL populations. We investigated the potential of increasing prediction ability by combining genomic and transcriptomic datasets, adding whole-genome sequencing (WGS) SNP data, functional annotation-based filtering, and empirical quality filtering. Our RNA-Seq data were generated cost-efficiently using small-footprint plant cultivation, high-throughput RNA extraction, and Library preparation miniaturization. We also examined sequencing depth reduction as an additional cost-saving measure. We used fivefold cross-validation to evaluate the prediction ability of the gene expression dataset, the RNA-Seq SNP dataset, and the consensus SNP dataset between the RNA-Seq and parental WGS data, resulting in prediction abilities between 0.73 and 0.78. The consensus SNP dataset performed best, with five out of eight traits performing significantly better compared to a 50K SNP array, which served as a benchmark. The advantage of the consensus SNP dataset was most prominent in the inter-population predictions, in which the training and validation sets originated from different RIL sub-populations. We were therefore able to not only show that RNA-Seq data alone are able to predict various complex traits in barley using RILs, but also that the performance can be further increased with WGS data for which the public availability will steadily increase.
It has been difficult to isolate the factors that limit contrast discrimination, one of the most fundamental aspects of the visual system. claim to have found a method that can answer the question of why discrimination thresholds increase with reference contrast. Is it because of a saturating contrast response function or because of increasing (multiplicative) noise? Based on four datasets they conclude that multiplicative noise is the controlling factor. disagree and claim the jury is still out because only one of the four datasets has sufficiently good statistics to support the KCT claim. I reanalyze the KCT data and come to a different conclusion. I agree with GM that two of the four datasets have thresholds that are too low to be useful in discriminating models and that one dataset supports the KCT claim. The fourth dataset is the most interesting one in that it provides the strongest support for the KCT claim, but GM throw it out because the chi(2) of the best fit is high. The present paper makes a number of points: (1) two novel methods are used to fit the fourth dataset. One pair of models is based on the strong "finger error" asymmetry between the high and low contrast asymptotes of the psychometric function in the fourth dataset. I find that some version of multiplicative noise is needed. However, it may be multiplicative noise that depends on prior trials rather than just on the present trial. (2) Another model that allows the contrast response function to have maximal freedom fits the fourth dataset with a reasonable chi square, and with a need for multiplicative noise. (3) I examine alternative parameterizations of the model functions used by KCT and GM that provide a more intuitive interpretation of the parameters. In summary, although I find the data do support a generalized form of multiplicative noise, I agree with GM that the jury remains out about what are the factors that limit contrast discrimination.
Neonatal-lamb mortality represents an economic loss and welfare concern. Two factors often associated with the risk of mortality are birth-weight and serum immunoglobulin concentration. We used data from two studies to investigate risk factors for mortality between 2 and 14 days of age and factors affecting birth-weight and serum immunoglobulin concentration at 48h of age. Dataset 1 included 1339 lambs born on eight farms during the 1995 spring lambing season; dataset 2 included 3172 lambs on seven farms during the 1991 spring lambing season. To account for some of the potential clustering within the data, multilevel models were used. Most (>75%) of the variation in the risk of mortality was at the lamb level. In dataset 1, factors significantly associated with increased odds of lamb mortality included low birth-weight and low serum immunoglobulin concentration. In dataset 2, significant risk factors for mortality included low birth-weight, ewe body-condition score, being born late in the season (relative to other lambs on the farm) and being born in multiple litters. There was a significant interaction between the effects of litter size and birth-weight. (Serum immunoglobulin concentration was not available for dataset 2.) More than half of the variation in birth-weight was at the ewe level, 27% at the lamb level, and 18% at the farm level (dataset 1). Single birth and being male were associated with increased birth-weight in both datasets. In dataset 2 only, increasing ewe condition score and birth early in the study period were also associated with increased birth-weight. Fifty-six percent of the variation in immunoglobulin concentration was at the lamb level, 36% at the ewe level and only 7% at the farm level. Factors associated with reduced serum immunoglobulin concentration included early or late birth in the lambing season, being born later than 14 days after the first lamb born on the farm, multiple-birth litters and maternal mastitis.
The aim of this report is to describe the use of WinBUGS for two datasets that arise from typical population pharmacokinetic studies. The first dataset relates to gentamicin concentration-time data that arose as part of routine clinical care of 55 neonates. The second dataset incorporated data from 96 patients receiving enoxaparin. Both datasets were originally analyzed by using NONMEM. In the first instance, although NONMEM provided reasonable estimates of the fixed effects parameters it was unable to provide satisfactory estimates of the between-subject variance. In the second instance, the use of NONMEM resulted in the development of a successful model, albeit with limited available information on the between-subject variability of the pharmacokinetic parameters. WinBUGS was used to develop a model for both of these datasets. Model comparison for the enoxaparin dataset was performed by using the posterior distribution of the log-likelihood and a posterior predictive check. The use of WinBUGS supported the same structural models tried in NONMEM. For the gentamicin dataset a one-compartment model with intravenous infusion was developed, and the population parameters including the full between-subject variance-covariance matrix were available. Analysis of the enoxaparin dataset supported a two compartment model as superior to the one-compartment model, based on the posterior predictive check. Again, the full between-subject variance-covariance matrix parameters were available. Fully Bayesian approaches using MCMC methods, via WinBUGS, can offer added value for analysis of population pharmacokinetic data.
AIMS: To compare the frequencies of adverse drug reactions (ADRs) to amiodarone from three separate datasets: (i) a meta-analysis of clinical trials, (ii) spontaneous reports published in medical journals, and (iii) spontaneous reports sent to the World Health Organization (WHO). METHODS: We classified the ADRs into eight categories, based on the site involved, and built a rank order of the ADRs by category (from most to least commonly reported) for each dataset. We also calculated the relative proportions for all eight ADR categories within each dataset, in order to be able to compare the distributions of ADR frequencies: we assigned an index value of 1.0 to the frequency of respiratory toxicity in each set and calculated values for the other ADRs relative to respiratory toxicity. RESULTS: Thyroid disorders were the most commonly reported ADRs in the WHO dataset. In contrast, published case reports showed a preponderance of respiratory disorders, while in the meta-analysis cardiac conduction problems were the most frequent. The rank orders of ADRs differed among the three datasets, as did the index values of specific ADR categories with respect to the respiratory category. CONCLUSIONS: The distributions of ADR rank order and relative frequencies are dissimilar among the three datasets, as each dataset compiles information in a different way. Nevertheless, each dataset has its own specific strengths, and all three should be used together in obtaining a complete picture of a drug's safety profile. Important therapeutic and regulatory decisions should not simply be based on one source of data.