Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Are statewide trauma registries comparable? Reaching for a national trauma dataset.

BACKGROUND: Statewide trauma registries have proliferated in the last decade, suggesting that information could be aggregated to provide an accurate depiction of serious injury in the United States. OBJECTIVES: To determine whether variability exists in the composition and content of statewide trauma registries, specifically addressing case-acquisition, case-definition (inclusion criteria), and registry-coding conventions. METHODS: A cross-sectional, two-part survey was administered to managers of all statewide trauma registries. State trauma registrars also provided inclusion and exclusion criteria from their state registry and abstracted a clinical vignette designed to identify coding inconsistencies. RESULTS: Thirty-two states maintain a centralized registry, but requirements for data submission vary significantly. Inclusion and exclusion criteria also vary, particularly for nontraumatic injuries. Coding conventions adopted by states for vague or missing information are dissimilar. When abstractions of the clinical vignette are compared, only 19% and 47% of states provided similar quantity or content for injury e-coding and diagnostic coding, respectively. Injury severity scores (based on diagnostic coding) demonstrated a range from 2 to 18. CONCLUSIONS: Statewide trauma registries are prevalent but vary significantly in composition and content. Standardizing inclusion criteria, variable definitions, and coding conventions would greatly enhance the usability of an aggregated, national trauma registry.

Cross-Sectional Studies↗

Recognition and management of depression in skilled-nursing and long-term care settings: evolving targets for quality improvement.

OBJECTIVE: Depression is a common disorder associated with suffering, morbidity, and mortality in nursing home residents. It is treatable, and improving the quality of treatment can have a major impact. METHODS: MPRO, Michigan's Quality Improvement Organization, initiated a quality-improvement project in 14 nursing facilities to improve the accuracy of assessments, targeting, and monitoring of care. Electronic Minimum Data Set (MDS) data and medical-record abstraction results were combined to form the analytic dataset. RESULTS: Findings from the baseline phase demonstrated that, according to medical and administrative records, 26% of newly admitted nursing home residents had symptoms of depression that were apparent at admission, and an additional 12% were recognized early in their stay. Eighty-one percent of residents with depression were receiving treatment on admission to the facility, and 79% of those with depression recognized by Day 14 were treated by then. CONCLUSIONS: These data demonstrate progress toward improving the initiation of treatment for depression in nursing homes; however, there are still opportunities for improving the quality of care and, especially, the quality of assessments. The authors recommend the addition of the Geriatric Depression Scale to the federally mandated MDS for cognitively intact patients. There could also be mechanisms to ensure that providers and facilities follow recommended practice guidelines. Initiating treatment with antidepressant medications should be followed with monitoring of residents to identify those who still have depressive symptoms and to modify or intensify their treatment.

Aged↗

A Bayesian system integrating expression data with sequence patterns for localizing proteins: comprehensive application to the yeast genome.

We develop a probabilistic system for predicting the subcellular localization of proteins and estimating the relative population of the various compartments in yeast. Our system employs a Bayesian approach, updating a protein's probability of being in a compartment, based on a diverse range of 30 features. These range from specific motifs (e.g. signal sequences or the HDEL motif) to overall properties of a sequence (e.g. surface composition or isoelectric point) to whole-genome data (e.g. absolute mRNA expression levels or their fluctuations). The strength of our approach is the easy integration of many features, particularly the whole-genome expression data. We construct a training and testing set of approximately 1300 yeast proteins with an experimentally known localization from merging, filtering, and standardizing the annotation in the MIPS, Swiss-Prot and YPD databases, and we achieve 75 % accuracy on individual protein predictions using this dataset. Moreover, we are able to estimate the relative protein population of the various compartments without requiring a definite localization for every protein. This approach, which is based on an analogy to formalism in quantum mechanics, gives better accuracy in determining relative compartment populations than that obtained by simply tallying the localization predictions for individual proteins (on the yeast proteins with known localization, 92% versus 74%). Our training and testing also highlights which of the 30 features are informative and which are redundant (19 being particularly useful). After developing our system, we apply it to the 4700 yeast proteins with currently unknown localization and estimate the relative population of the various compartments in the entire yeast genome. An unbiased prior is essential to this extrapolated estimate; for this, we use the MIPS localization catalogue, and adapt recent results on the localization of yeast proteins obtained by Snyder and colleagues using a minitransposon system. Our final localizations for all approximately 6000 proteins in the yeast genome are available over the web at: http://bioinfo.mbb.yale. edu/genome/localize.

Amino Acid Motifs↗

Oblique decision trees for spatial pattern detection: optimal algorithm and application to malaria risk.

BACKGROUND: In order to detect potential disease clusters where a putative source cannot be specified, classical procedures scan the geographical area with circular windows through a specified grid imposed to the map. However, the choice of the windows' shapes, sizes and centers is critical and different choices may not provide exactly the same results. The aim of our work was to use an Oblique Decision Tree model (ODT) which provides potential clusters without pre-specifying shapes, sizes or centers. For this purpose, we have developed an ODT-algorithm to find an oblique partition of the space defined by the geographic coordinates. METHODS: ODT is based on the classification and regression tree (CART). As CART finds out rectangular partitions of the covariate space, ODT provides oblique partitions maximizing the interclass variance of the independent variable. Since it is a NP-Hard problem in RN, classical ODT-algorithms use evolutionary procedures or heuristics. We have developed an optimal ODT-algorithm in R2, based on the directions defined by each couple of point locations. This partition provided potential clusters which can be tested with Monte-Carlo inference. We applied the ODT-model to a dataset in order to identify potential high risk clusters of malaria in a village in Western Africa during the dry season. The ODT results were compared with those of the Kulldorff' s SaTScan. RESULTS: The ODT procedure provided four classes of risk of infection. In the first high risk class 60%, 95% confidence interval (CI95%) [52.22-67.55], of the children was infected. Monte-Carlo inference showed that the spatial pattern issued from the ODT-model was significant (p < 0.0001). Satscan results yielded one significant cluster where the risk of disease was high with an infectious rate of 54.21%, CI95% [47.51-60.75]. Obviously, his center was located within the first high risk ODT class. Both procedures provided similar results identifying a high risk cluster in the western part of the village where a mosquito breeding point was located. CONCLUSION: ODT-models improve the classical scanning procedures by detecting potential disease clusters independently of any specification of the shapes, sizes or centers of the clusters.

Africa, Western↗

EST mining of the UniGene dataset to identify retina-specific genes.

Age-related macular degeneration (AMD) is a multifactorial disorder affecting the visual system with a high prevalence among the elderly population but with no effective therapy available at present. To better understand the pathogenesis of this disorder, the identification of the genetic factors and the determination of their contribution to AMD is needed. Towards this goal, we are pursuing a strategy that makes use of the EST data processed in the UniGene database and aims at the generation of a comprehensive catalogue of genes preferentially active in the human retina. Subsequently, these genes will be systematically assessed in AMD. We performed a retina EST sampling and obtained a total of 673 clusters containing only retina ESTs as well as 568 clusters with at least 30% of the ESTs in each cluster originating from retina cDNA libraries. Of these, 180 representative EST clusters with varying retina and non-retina EST contents were analyzed for their in vitro expression. This approach identified 39 transcripts with retina-specific expression. One of these genes (C18orf2) mapping to chromosome 18 was further characterized. Multiple C18orf2 transcripts display a complex pattern of differential splicing in the human retina. The various isoforms encode hypothetical polypeptides with no homologies to known proteins or protein motifs.

Alternative Splicing↗

Statin treatment and adherence to national cholesterol guidelines after ischemic stroke.

BACKGROUND: National cholesterol guidelines have defined high vascular risk individuals as those who could potentially benefit most from statin therapy. The authors aimed to determine the rate of statin use, its predictors, and the achievement of national guideline target lipid goals among ischemic stroke survivors. METHODS: The authors abstracted data from the Vitamin Intervention for Stroke Prevention (VISP) study database from the United States and Canada to incorporate into algorithms for initiating statin therapy according to the National Cholesterol Education Program (NCEP) guidelines for high-risk individuals. The authors applied these algorithms to all study subjects. Univariate as well as multivariate associations for target lipid levels and statin implementation were then evaluated utilizing pertinent demographic, clinical, and laboratory data. RESULTS: Of 2,894 subjects in the analysis dataset, 38% were women; 71% were recruited in the United States and 29% in Canada. Of 769 high-risk subjects, 262 (34%) had a low-density lipoprotein (LDL) level > or =130 mg/dL and 124 of these (47%) were not on statin. Among those high-risk persons on statin treatment, only 42% had an LDL < or =100 mg/dL. Subjects in the overall cohort were more likely to be on a statin if they were treated in the United States or had a history of hypertension or coronary artery disease. CONCLUSIONS: Approximately one out of three guideline-eligible high vascular risk ischemic stroke patients in this study had low-density lipoprotein cholesterol concentrations above qualifying levels for pharmacologic therapy, but half of these patients were not taking a statin, and of those receiving statin treatment, less than half were within recommended lipid goals.

Adult↗

Proof-of-principle phase II MRI studies in stroke: sample size estimates from dichotomous and continuous data.

BACKGROUND AND PURPOSE: Since the failure of a number of phase III trials of neuroprotection in ischemic stroke, the need for smaller phase II studies with MRI surrogates has emerged. There is, however, little information available about sample size requirements for such phase II trials and rarely enough patients in single studies to make robust estimates. We have formed an international collaborative group to assemble larger datasets and from these have generated sample size tables for MRI-based infarct expansion as the outcome measure. METHODS: Twelve centers from Australia, Europe, and North America contributed data from patients with hemispheric ischemic stroke. Infarct expansion was defined from initial diffusion-weighted images and later fluid-attenuated inversion recover or T2 images. Sample size estimates were calculated from data on infarct expansion ratios treated as dichotomous or continuous variables. A nonparametric approach was used because the distribution of infarct expansion was resistant to all forms of transformation. RESULTS: As an example, a 20% absolute reduction in infarct expansion ratio (< or = 1), 80% power, and alpha = 0.05 requires 99 patients in each arm. To achieve an equivalent effect size with a continuous approach requires 61 patients. CONCLUSIONS: These tables will be useful in planning phase II trials of therapy with the use of MRI outcome measures. For positive studies, biologically plausible surrogates such as these may provide a rationale for proceeding to phase III trials.

Australia↗

Modeling shape and topology of low-resolution density maps of biological macromolecules.

In the present work we develop an efficient way of representing the geometry and topology of volumetric datasets of biological structures from medium to low resolution, aiming at storing and querying them in a database framework. We make use of a new vector quantization algorithm to select the points within the macromolecule that best approximate the probability density function of the original volume data. Connectivity among points is obtained with the use of the alpha shapes theory. This novel data representation has a number of interesting characteristics, such as 1) it allows us to automatically segment and quantify a number of important structural features from low-resolution maps, such as cavities and channels, opening the possibility of querying large collections of maps on the basis of these quantitative structural features; 2) it provides a compact representation in terms of size; 3) it contains a subset of three-dimensional points that optimally quantify the densities of medium resolution data; and 4) a general model of the geometry and topology of the macromolecule (as opposite to a spatially unrelated bunch of voxels) is easily obtained by the use of the alpha shapes theory.

Algorithms↗

In vitro validation of right ventricular volume measurement by three dimensional echocardiography.

OBJECTIVE: Evaluation of ability of three dimensional echocardiography to accurately assess right ventricular volumes in vitro. METHODS: Silicone casts of normal human right ventricles were examined. Each was filled with three different volumes of water to yield 15 different measurements. The casts were examined in a waterbath with three dimensional echocardiography using a 7.5 MHz ultrasound probe mounted in a scan frame. It was steered by a stepper motor, which moved the probe in steps of 0.25 mm over a distance of 5.9 cm inside the frame, acquiring an image at each step. 236 parallel slices of the cast were thus obtained, forming the three dimensional dataset. The longest axis of the right ventricular volume was defined and the area of perpendicular 1 mm thick slices was outlined manually to calculate the area of each slice. This was multiplied by the slice thickness to obtain the volume of each slice; the respective volumes were added to obtain the volume of the whole cast. RESULTS: The casts had a median volume of 31.1 (23) ml (range 15-100); three dimensional echocardiography gave a median volume of 29.0 (21.7) ml (15.7-91.7). Interobserver variability was 4.5% (0.4%-13.6%) and intraobserver variability 4.3% (0.2%-9.3%). Correlation between real cast volumes and volumes measured by three dimensional echocardiography was 0.99 (y = 1.08 x -0.16) with an SEE of 2.7 ml. Limits for agreement between methods ranged from -3.1 ml to 8.3 ml. In 14 of the 15 measurements, volume by three dimensional echocardiography was smaller than real volume, with the mean difference being 7.4% (2.8%-19.5%). This may be due to the thickening of surfaces of structures when imaged by ultrasonography. CONCLUSION: Right ventricular volumes can accurately be determined by three dimensional echocardiography.

Echocardiography↗

404 not found: the stability and persistence of URLs published in MEDLINE.

MOTIVATION: The advent of the World Wide Web has enabled unprecedented supplementation of traditional journal publications, allowing access to resources, such as video, sound, software, databases, datasets too large to publish, and even supplementary information and discussion. However, unlike traditional publications, continued availability of these online resources is not guaranteed. An automated survey was conducted to quantify the growth in Uniform Resource Locators (URLs) published to date in MEDLINE abstracts, their current availability and distribution by journal. RESULTS: Of 1630 unique URLs identified, formatting and/or spelling errors were detected within 201 (12%) of them as published. After corrections were made, a survey revealed that approximately 63% of these URLs were consistently available, and another 19% were available intermittently. The rate of failure was far worse for anonymous login to FTP sites, with only 12 of 33 sites (36%) responding. This survey also shows that journals vary disproportionately in the number of web citations published, suggesting policy implementation among a few could have a profound impact overall. Out of the 306 journals with a URL published in an abstract, Bioinformatics published the most (12% of total). AVAILABILITY: URL database and program available by request.

Abstracting and Indexing↗

Multiple imputation for model checking: completed-data plots with missing and latent data.

In problems with missing or latent data, a standard approach is to first impute the unobserved data, then perform all statistical analyses on the completed dataset--corresponding to the observed data and imputed unobserved data--using standard procedures for complete-data inference. Here, we extend this approach to model checking by demonstrating the advantages of the use of completed-data model diagnostics on imputed completed datasets. The approach is set in the theoretical framework of Bayesian posterior predictive checks (but, as with missing-data imputation, our methods of missing-data model checking can also be interpreted as "predictive inference" in a non-Bayesian context). We consider the graphical diagnostics within this framework. Advantages of the completed-data approach include: (1) One can often check model fit in terms of quantities that are of key substantive interest in a natural way, which is not always possible using observed data alone. (2) In problems with missing data, checks may be devised that do not require to model the missingness or inclusion mechanism; the latter is useful for the analysis of ignorable but unknown data collection mechanisms, such as are often assumed in the analysis of sample surveys and observational studies. (3) In many problems with latent data, it is possible to check qualitative features of the model (for example, independence of two variables) that can be naturally formalized with the help of the latent data. We illustrate with several applied examples.

Animals↗

Establishing connections between microarray expression data and chemotherapeutic cancer pharmacology.

We have investigated three different microarray datasets of approximately 6 K gene expressions across the National Cancer Institute's panel of 60 tumor cell lines. Initial assessments of reproducibility for gene expressions within each dataset, as derived from sequence analysis of full-length sequences as well as expressed sequence tags (EST), found statistically significant results for no more than 36% of those cases where at least one replicate of a gene appears on the array. Filtering the data based only on pairwise comparisons among these three datasets creates a list of approximately 400 significant concordant expression patterns. The expression profiles of these smaller sets of genes were used to locate similar expression profiles of synthetic agents screened against these same 60 tumor cell lines. A correspondence was found between mRNA expression patterns and 50% growth inhibition response patterns of screened agents for 11 cases that were subsequently verifiable from ligand-target crystallographic data. Notable amongst these cases are genes encoding a variety of kinases, which were also found to be targets of small drug-like molecules within the database of protein structures. These 11 cases lend support to the premise that similarities between expression patterns and chemical responses for the National Cancer Institute's tumor panel can be related to known cases of molecular structure and putative cellular function. The details of the 11 verifiable cases and the concordant gene subsets are provided. Discussions about the prospects of using this approach as a data mining tool are included.

Algorithms↗

Protein classification based on text document classification techniques.

The need for accurate, automated protein classification methods continues to increase as advances in biotechnology uncover new proteins. G-protein coupled receptors (GPCRs) are a particularly difficult superfamily of proteins to classify due to extreme diversity among its members. Previous comparisons of BLAST, k-nearest neighbor (k-NN), hidden markov model (HMM) and support vector machine (SVM) using alignment-based features have suggested that classifiers at the complexity of SVM are needed to attain high accuracy. Here, analogous to document classification, we applied Decision Tree and Naive Bayes classifiers with chi-square feature selection on counts of n-grams (i.e. short peptide sequences of length n) to this classification task. Using the GPCR dataset and evaluation protocol from the previous study, the Naive Bayes classifier attained an accuracy of 93.0 and 92.4% in level I and level II subfamily classification respectively, while SVM has a reported accuracy of 88.4 and 86.3%. This is a 39.7 and 44.5% reduction in residual error for level I and level II subfamily classification, respectively. The Decision Tree, while inferior to SVM, outperforms HMM in both level I and level II subfamily classification. For those GPCR families whose profiles are stored in the Protein FAMilies database of alignments and HMMs (PFAM), our method performs comparably to a search against those profiles. Finally, our method can be generalized to other protein families by applying it to the superfamily of nuclear receptors with 94.5, 97.8 and 93.6% accuracy in family, level I and level II subfamily classification respectively.

Algorithms↗

Implications of chance baseline differences in repeated measurement designs.

Datasets representing randomized, parallel-groups designs were analyzed by repeated measurements ANOVA and linear trend analysis with and without baseline values being covaried. ANOVA tests for the between-groups main effect, groups X times interaction, and differences in linear trends across time periods are shown to be seriously conservative or seriously nonconservative, depending on the direction and significance of chance baseline mean difference. Inclusion of baseline scores as a covariate in the repeated measurements ANOVA provides appropriate correction for the between-groups (average) effect across time, but the covariate provides no correction for the within-subject effects that are concerned with differences in the rates or patterns of change across time. If one desires to evaluate differences between patterns of treatment-induced change, tests of significance for differences in group means on composite trend scores with covariance correction for baseline are recommended. If covariance correction is not or cannot be employed, the potentially "favorable" or "unfavorable" influence of chance baseline differences on tests of significance needs to be explicitly recognized.

Analysis of Variance↗

A new approach for filtering noise from high-density oligonucleotide microarray datasets.

Although DNA microarrays are powerful tools for profiling gene expression, the dynamic range and the sheer number of signals produced require efficient procedures for distinguishing false positive results (noise) from changes in expression that are 'real' (independently reproducible). We have developed an approach to filter noise from datasets generated when high density oligonucleotide-based microarrays are used to compare two distinct RNA populations. First, we performed comparisons between chips hybridized with cRNAs prepared from an identical starting RNA population; an 'Increase' or 'Decrease' call in such a comparison was defined as a false positive. Plotting the average distribution of these false positive signal intensities across 18 such comparisons of nine independent RNA preparations allowed us to develop a series of noise-filtering look-up tables (LUTs). Using a database of 70 separate chip-to-chip comparisons between distinct RNA preparations prepared by different workers at different sites and at different times, we show that the LUTs can be used to predict the likelihood that a given transcript called Increased or Decreased in one comparison will again be called Increased or Decreased in a replicate comparison. Evidence is presented that this LUT-based scoring system provides greater predictive value for reproducible microarray results than imposition of arbitrary fold-change thresholds and accurately predicts which microarray-identified changes will be validated by independent assays such as quantitative real-time PCR.

Databases as Topic↗

Evaluation of a new measure of blood glucose variability in diabetes.

OBJECTIVE: Recent studies show the importance of controlling blood glucose variability in relationship to both reducing hypoglycemia and attenuating the risk for cardiovascular and behavioral complications due to hyperglycemia. It is therefore important to design variability measures that are equally predictive of low and high blood glucose excursions. RESEARCH DESIGN AND METHODS: We introduce the average daily risk range (ADRR), a variability measure computed from routine self-monitored blood glucose (SMBG) data. The ADRR was constructed using a development dataset for 39 and 31 adults with type 1 and type 2 diabetes, respectively. The formula was then fixed, and the ADRR was compared against other variability measures using an independent validation dataset containing approximately 4 months of SMBG for 254 and 81 adults with type 1 and type 2 diabetes. RESULTS: From the 1st month of validation SMBG data, we computed the ADRR, blood glucose SD and coefficient of variation, daily blood glucose range and interquartile range, mean amplitude of glycemic excursion, M-value, and lability index. Then all measures were tested as predictors of low blood glucose (<2.2 mmol/l; <3.9 mmol/l) and high (>10 mmol/l; >22.2 mmol/l) events in the subsequent 3 months. The ADRR was the best predictor of both hypoglycemia and hyperglycemia, with a 6-fold increase in the likelihood of hypoglycemia and 3.5-fold increase in the likelihood of hyperglycemia across its risk categories. CONCLUSIONS: In a large SMBG database, the ADRR showed strong association with subsequent out-of-control glucose readings. Compared with other variability measures, the ADRR demonstrated a superior balance of sensitivity to predicting both hypoglycemia and hyperglycemia. This prediction was independent from type of diabetes.

Adult↗

Does elimination of placebo responders in a placebo run-in increase the treatment effect in randomized clinical trials? A meta-analytic evaluation.

The use of a placebo run-in phase, in which placebo responders are withdrawn from a study before random assignment to treatment condition, has been criticized as favoring the active treatment in clinical trials. We compared the effect size of randomized, placebo-controlled clinical trials (in the treatment of depression with selective serotonin reuptake inhibitors [SSRIs]) that include a placebo run-in phase with those that do not, using a meta-analytic approach. This study differed from earlier meta-analytic studies in that it considered only SSRIs and included only studies using continuous measures of depression, allowing for a more refined assessment of effect size. An extensive literature search identified 43 datasets published between 1980 and 2000 comparing placebo with SSRI and using a continuous measure of depression (usually the Hamilton Depression Rating Scale). We included only studies of at least 6 weeks' duration focusing on treatment for primary acute major depression in adults 18-65 years of age. Studies focusing on depression in specific medical illnesses were not included. Analysis of efficacy was based on 3047 subjects treated with an SSRI antidepressant and 3740 subjects treated with a placebo. There was no statistically significant difference in effect size between the clinical trials that had a placebo run-in phase followed by withdrawal of placebo responders and those trials that did not. Despite the lack of a statistically significant difference between studies of withdrawing early placebo responders and those not using this procedure, this approach is likely to continue to be used widely because it produces large absolute effect sizes. It is recommended that future studies clearly describe these procedures and report the number of subjects dropped from the study for early placebo response and other reasons.

Antidepressive Agents, Second-Generation↗

Comparative phylogeographic summary statistics for testing simultaneous vicariance.

Testing for simultaneous vicariance across comparative phylogeographic data sets is a notoriously difficult problem hindered by mutational variance, the coalescent variance, and variability across pairs of sister taxa in parameters that affect genetic divergence. We simulate vicariance to characterize the behaviour of several commonly used summary statistics across a range of divergence times, and to characterize this behaviour in comparative phylogeographic datasets having multiple taxon-pairs. We found Tajima's D to be relatively uncorrelated with other summary statistics across divergence times, and using simple hypothesis testing of simultaneous vicariance given variable population sizes, we counter-intuitively found that the variance across taxon pairs in Nei and Li's net nucleotide divergence (pi(net)), a common measure of population divergence, is often inferior to using the variance in Tajima's D across taxon pairs as a test statistic to distinguish ancient simultaneous vicariance from variable vicariance histories. The opposite and more intuitive pattern is found for testing more recent simultaneous vicariance, and overall we found that depending on the timing of vicariance, one of these two test statistics can achieve high statistical power for rejecting simultaneous vicariance, given a reasonable number of intron loci (> 5 loci, 400 bp) and a range of conditions. These results suggest that components of these two composite summary statistics should be used in future simulation-based methods which can simultaneously use a pool of summary statistics to test comparative the phylogeographic hypotheses we consider here.

Classification↗