Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Treating missing data in a clinical neuropsychological dataset--data imputation.

Missing data frequently reduce the applicability of clinically collected data in research requiring multivariate statistics. In data imputation, missing values are replaced by predicted values obtained from models based on auxiliary information. Our aim was to complete a clinical child neuropsychological data set containing 5.2% of missing observations. This was to be used in research requiring multivariate statistics. We compared four data imputation methods by artificially deleting some data. A real-donor imputation method which preserved the parameter estimates and which predicted the observed values with acceptable accuracy was used to complete the data set. In addressing the lack of studies with regard to treatment of missing data in neuropsychological data sets, this study presents information on the outcomes of applying data imputation methods to such data. The imputation modeling described can be applied to a variety of clinical neuropsychological data sets.

Child↗

Musculoskeletal model of the upper limb based on the visible human male dataset.

A mathematical model of the human upper limb was developed based on high-resolution medical images of the muscles and bones obtained from the Visible Human Male (VHM) project. Three-dimensional surfaces of the muscles and bones were reconstructed from Computed Tomography (CT) images and Color Cryosection images obtained from the VHM cadaver. Thirteen degrees of freedom were used to describe the orientations of seven bones in the model: clavicle, scapula, humerus, radius, ulna, carpal bones, and hand. All of the major articulations from the shoulder girdle down to the wrist were included in the model. The model was actuated by 42 muscle bundles, which represented the actions of 26 muscle groups in the upper limb. The paths of the muscles were modeled using a new approach called the Obstacle-set Method [33]. The calculated paths of the muscles were verified by comparing the muscle moment arms computed in the model with the results of anatomical studies reported in the literature. In-vivo measurements of maximum isometric muscle torques developed at the shoulder, elbow, and wrist were also used to estimate the architectural properties of each musculotendon actuator in the model. The entire musculoskeletal model can be reconstructed using the data given in this paper, along with information presented in a companion paper which defines the kinematic structure of the model [26].

Adult↗

Care after the onset of serious illness: a novel claims-based dataset exploiting substantial cross-set linkages to study end-of-life care.

To date, there has not been a study using a large, nationally representative group of patients with serious illness who are at risk for hospice use and who are followed forward in time to understand the determinants of hospice use. In this paper, we outline the development of a large new cohort of 1,221,153 Medicare beneficiaries newly diagnosed with 1 of 13 serious conditions in 1993, a cohort that can be used to study end-of-life care in the United States. In describing our methods, we illustrate the possible utility of Medicare claims for end-of-life research. The members of our cohort are followed forward for hospice and other health care use through December 1997, and for mortality through June 1999. Medicare claims data on their inpatient and outpatient hospital use is also collected. Based on the ZIP Codes and counties in which cohort members lived, we were also able to characterize the health care markets of cohort members, as well as obtain other socioeconomic information about them. Information about cohort member's health care providers is also available. Detailed health information about cohort members' spouses was also collected. We conclude by highlighting the types of analyses that can be conducted in this data set.

Cohort Studies↗

tidk: a toolkit to rapidly identify telomeric repeats from genomic datasets.

SUMMARY: "tidk" (short for telomere identification toolkit) uses a simple, fast algorithm to scan long DNA reads for the presence of short tandemly repeated DNA in runs, and to aggregate them based on canonical DNA string representation. These are telomeric repeat candidates. Our algorithm is shown to be accurate in genomes for which the telomeric repeat unit is known and is tested across a wide variety of newly assembled genomes to uncover new telomeric repeat units. Tools are provided to identify telomeric repeats de novo, scan genomes for known telomeric repeats, and to visualize telomeric repeats on the assembly. "tidk" is implemented in Rust and is available as a command line tool which can be compiled using the Rust toolchain or downloaded as a binary from bioconda. AVAILABILITY AND IMPLEMENTATION: The "tidk" Rust crate is freely available under the MIT license (https://crates.io/crates/tidk), and the source code is available at https://github.com/tolkit/telomeric-identifier.

Telomere↗

TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets.

TGICL is a pipeline for analysis of large Expressed Sequence Tags (EST) and mRNA databases in which the sequences are first clustered based on pairwise sequence similarity, and then assembled by individual clusters (optionally with quality values) to produce longer, more complete consensus sequences. The system can run on multi-CPU architectures including SMP and PVM.

Cluster Analysis↗

Reproducibility of SELDI-TOF protein patterns in serum: comparing datasets from different experiments.

MOTIVATION: There has been much interest in using patterns derived from surface-enhanced laser desorption and ionization (SELDI) protein mass spectra from serum to differentiate samples from patients both with and without disease. Such patterns have been used without identification of the underlying proteins responsible. However, there are questions as to the stability of this procedure over multiple experiments. RESULTS: We compared SELDI proteomic spectra from serum from three experiments by the same group on separating ovarian cancer from normal tissue. These spectra are available on the web at http://clinicalproteomics.steem.com. In general, the results were not reproducible across experiments. Baseline correction prevents reproduction of the results for two of the experiments. In one experiment, there is evidence of a major shift in protocol mid-experiment which could bias the results. In another, structure in the noise regions of the spectra allows us to distinguish normal from cancer, suggesting that the normals and cancers were processed differently. Sets of features found to discriminate well in one experiment do not generalize to other experiments. Finally, the mass calibration in all three experiments appears suspect. Taken together, these and other concerns suggest that much of the structure uncovered in these experiments could be due to artifacts of sample processing, not to the underlying biology of cancer. We provide some guidelines for design and analysis in experiments like these to ensure better reproducible, biologically meaningfully results. AVAILABILITY: The MATLAB and Perl code used in our analyses is available at http://bioinformatics.mdanderson.org

Algorithms↗

The choice of optimal distance measure in genome-wide datasets.

MOTIVATION: Many types of genomic data are naturally represented as binary vectors. Numerous tasks in computational biology can be cast as analysis of relationships between these vectors, and the first step is, frequently, to compute their pairwise distance matrix. Many distance measures have been proposed in the literature, but there is no theory justifying the choice of distance measure. RESULTS: We examine the approaches to measuring distances between binary vectors and study the characteristic properties of various distance measures and their performance in several tasks of genome analysis. Most distance measures between binary vectors turn out to belong to a single parametric family, namely generalized average-based distance with different exponents. We show that descriptive statistics of distance distribution, such as skewness and kurtosis, can guide the appropriate choice of the exponent. On the contrary, the more familiar distance properties, such as metric and additivity, appear to have much less effect on the performance of distances. AVAILABILITY: R code GADIST and Supplementary material are available at http://research.stowers-institute.org/bioinfo/

Algorithms↗

Context-specific infinite mixtures for clustering gene expression profiles across diverse microarray dataset.

MOTIVATION: Identifying groups of co-regulated genes by monitoring their expression over various experimental conditions is complicated by the fact that such co-regulation is condition-specific. Ignoring the context-specific nature of co-regulation significantly reduces the ability of clustering procedures to detect co-expressed genes due to additional 'noise' introduced by non-informative measurements. RESULTS: We have developed a novel Bayesian hierarchical model and corresponding computational algorithms for clustering gene expression profiles across diverse experimental conditions and studies that accounts for context-specificity of gene expression patterns. The model is based on the Bayesian infinite mixtures framework and does not require a priori specification of the number of clusters. We demonstrate that explicit modeling of context-specificity results in increased accuracy of the cluster analysis by examining the specificity and sensitivity of clusters in microarray data. We also demonstrate that probabilities of co-expression derived from the posterior distribution of clusterings are valid estimates of statistical significance of created clusters. AVAILABILITY: The open-source package gimm is available at http://eh3.uc.edu/gimm.

Algorithms↗

Meta-analysis models with group structure for pleiotropy detection at gene and variant level using summary statistics from multiple datasets.

Genome-wide association studies (GWASs) have highlighted the importance of pleiotropy in human diseases, where one gene can impact 2 or more unrelated traits. Examining shared genetic risk factors across multiple diseases can enhance our understanding of these conditions by pinpointing new genes and biological pathways involved. Furthermore, with an increasing wealth of GWAS summary statistics available to the scientific community, leveraging these findings across multiple phenotypes could unveil novel pleiotropic associations. Existing selection methods examine pleiotropic associations one by one at a scale of either the genetic variant or the gene, and thus cannot consider all the genetic information at the same time. To address this limitation, we propose a new approach called MPSG (Meta-analysis model adapted for Pleiotropy Selection with Group structure). This method performs a penalized multivariate meta-analysis method adapted for pleiotropy and takes into account the group structure information nested in the data to select relevant variants and genes (or pathways) from all the genetic information. To do so, we implemented an alternating direction method of multipliers algorithm. We compared the performance of the method with other benchmark meta-analysis approaches such as GCPBayes, PLACO, and ASSET by considering as inputs different kinds of summary statistics. We provide an application of our method to the identification of potential pleiotropic genes between breast and thyroid cancers.

Humans↗

A novel method for estimating substitution rate variation among sites in a large dataset of homologous DNA sequences.

We present here a novel method to estimate the site-specific relative variability in large sets of homologous sequences. It is based on the simple idea that the more closely related are the compared sequences, the higher the probability of observing nucleotide changes at rapidly evolving sites. A simulation study has been carried out to support the reliability of the method, which has been applied also to analyzing the site variability of all available human sequences corresponding to the two hypervariable regions of the mitochondrial D-loop.

Computer Simulation↗

Adverse outcomes in Belgian acute hospitals: retrospective analysis of the national hospital discharge dataset.

OBJECTIVE: The prevalence and variability of adverse outcome rates in Belgian acute hospitals is examined by using the national hospital discharge database. DESIGN: setting, and participants. Retrospective analysis based on administrative data of all Belgian acute hospitals, covering the full medical (n = 1 024 743) and surgical (n = 633 027) in-patients population for the year 2000. MAIN OUTCOME MEASURES: For 11 adverse outcomes and failure-to-rescue, the rates and variability among hospitals were studied. The all patient refined diagnostic-related groups (APR-DRG) method was used for risk adjustment. RESULTS: The prevalence of adverse outcomes was 7.12% in the medical and 6.32% in the surgical group. Rates ranged from 6.25 (deep venous thrombosis) to 32.3 (urinary tract infection) outcomes per 1000 discharges in the medical group and from 3.39 (deep venous thrombosis) to 17.6 (urinary tract infection) outcomes per 1000 discharges in the surgical group. The failure-to-rescue rate was 240 and 211 per 1000 discharges, respectively. Except for pressure ulcers and hospital-acquired sepsis, the prevalence of adverse outcomes was significantly higher (P = 0.001) in the medical group. All adverse outcome rates varied substantially among the hospitals surveyed. CONCLUSIONS: This study identifies the occurrence of adverse outcomes in a national population. It adds information to the growing body of knowledge in predominantly Anglo-Saxon countries about adverse outcomes. Striking variation exists in the risk-adjusted adverse outcome rates across Belgian acute hospitals, revealing a large potential for quality gains that encourage further action.

Adult↗

Genome-related datasets within the E. coli Genetic Stock Center database.

The contents of the E. coli Genetic Stock Center database and the availability in electronic form of the subset of information most relevant to sequence databases are described. The database uses the long-standing Stock Center records (developed and curated by Dr B.J.Bachmann) in describing genotypes of mutant derivatives of E.coli K-12 in terms of alleles, structural mutations, mating type, and plasmids as well as the derivation, names and originators of the strain, and references. The database includes descriptions of mutations, mutation properties, genes, gene properties, and gene products, with EC number identifiers for enzymes. Sequence information is not included, but entries refer to sequence database accession numbers for sequenced regions. A gene is described as a subtype of a more general category of chromosome interval called Site. Since sites are used to describe any chromosomal interval, mapping information is associated with sites. Alleles are described as mutations of those sites and they are not primary map objects, but inherit map position information from the corresponding site description. The database design is intended to preserve richness of detail where it is known and uncertainty of measurements or information as it occurs in order to represent the stock center records as accurately as possible.

Bacterial Proteins↗

CONFAC: automated application of comparative genomic promoter analysis to DNA microarray datasets.

The advent of DNA microarray technology and the sequencing of multiple vertebrate genomes has provided a unique opportunity for the integration of comparative genomics with high-throughput gene expression analysis. Here we describe the conserved transcription factor binding site (CONFAC) software that enables the high-throughput identification of conserved transcription factor binding sites (TFBSs) in the regulatory regions of hundreds of genes at a time (http://morenolab.whitehead.emory.edu/cgi-bin/confac/login.pl). The CONFAC software compares non-coding regulatory sequences between human and mouse genomes to enable identification of conserved TFBSs that are significantly enriched in promoters of gene clusters from microarray analyses compared to sets of unchanging control genes using a Mann-Whitney U-test. Analysis of random gene sets demonstrated that using our approach, over 98% of TFBSs had false positive rates below 5%. As a proof-of-principle, we have validated the CONFAC software using gene sets from four separate microarray studies and identified TFBSs known to be functionally important for regulation of each of the four gene sets.

Animals↗

FLIGHT: database and tools for the integration and cross-correlation of large-scale RNAi phenotypic datasets.

FLIGHT (www.flight.licr.org) is a new database designed to help researchers browse and cross-correlate data from large-scale RNAi studies. To date, the majority of these functional genomic screens have been carried out using Drosophila cell lines. These RNAi screens follow 100 years of classical Drosophila genetics, but have already revealed their potential by ascribing an impressive number of functions to known and novel genes. This has in turn given rise to a pressing need for tools to simplify the analysis of the large amount of phenotypic information generated. FLIGHT aims to do this by providing users with a gene-centric view of screen results and by making it possible to cluster phenotypic data to identify genes with related functions. Additionally, FLIGHT provides microarray expression data for many of the Drosophila cell lines commonly used in RNAi screens. This, together with information about cell lines, protocols and dsRNA primer sequences, is intended to help researchers design their own cell-based screens. Finally, although the current focus of FLIGHT is Drosophila, the database has been designed to facilitate the comparison of functional data across species and to help researchers working with other systems navigate their way through the fly genome.

Animals↗