Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “datasets”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Parametric and nonparametric population methods: their comparative performance in analysing a clinical dataset and two Monte Carlo simulation studies.

BACKGROUND AND OBJECTIVES: This study examined parametric and nonparametric population modelling methods in three different analyses. The first analysis was of a real, although small, clinical dataset from 17 patients receiving intramuscular amikacin. The second analysis was of a Monte Carlo simulation study in which the populations ranged from 25 to 800 subjects, the model parameter distributions were Gaussian and all the simulated parameter values of the subjects were exactly known prior to the analysis. The third analysis was again of a Monte Carlo study in which the exactly known population sample consisted of a unimodal Gaussian distribution for the apparent volume of distribution (V(d)), but a bimodal distribution for the elimination rate constant (k(e)), simulating rapid and slow eliminators of a drug. METHODS: For the clinical dataset, the parametric iterative two-stage Bayesian (IT2B) approach, with the first-order conditional estimation (FOCE) approximation calculation of the conditional likelihoods, was used together with the nonparametric expectation-maximisation (NPEM) and nonparametric adaptive grid (NPAG) approaches, both of which use exact computations of the likelihood. For the first Monte Carlo simulation study, these programs were also used. A one-compartment model with unimodal Gaussian parameters V(d) and k(e) was employed, with a simulated intravenous bolus dose and two simulated serum concentrations per subject. In addition, a newer parametric expectation-maximisation (PEM) program with a Faure low discrepancy computation of the conditional likelihoods, as well as nonlinear mixed-effects modelling software (NONMEM), both the first-order (FO) and the FOCE versions, were used. For the second Monte Carlo study, a one-compartment model with an intravenous bolus dose was again used, with five simulated serum samples obtained from early to late after dosing. A unimodal distribution for V(d) and a bimodal distribution for k(e) were chosen to simulate two subpopulations of 'fast' and 'slow' metabolisers of a drug. NPEM results were compared with that of a unimodal parametric joint density having the true population parameter means and covariance. RESULTS: For the clinical dataset, the interindividual parameter percent coefficients of variation (CV%) were smallest with IT2B, suggesting less diversity in the population parameter distributions. However, the exact likelihood of the results was also smaller with IT2B, and was 14 logs greater with NPEM and NPAG, both of which found a greater and more likely diversity in the population studied. For the first Monte Carlo dataset, NPAG and PEM, both using accurate likelihood computations, showed statistical consistency. Consistency means that the more subjects studied, the closer the estimated parameter values approach the true values. NONMEM FOCE and NONMEM FO, as well as the IT2B FOCE methods, do not have this guarantee. Results obtained by IT2B FOCE, for example, often strayed visibly away from the true values as more subjects were studied. Furthermore, with respect to statistical efficiency (precision of parameter estimates), NPAG and PEM had good efficiency and precise parameter estimates, while precision suffered with NONMEM FOCE and IT2B FOCE, and severely so with NONMEM FO. For the second Monte Carlo dataset, NPEM closely approximated the true bimodal population joint density, while an exact parametric representation of an assumed joint unimodal density having the true population means, standard deviations and correlation gave a totally different picture. CONCLUSIONS: The smaller population interindividual CV% estimates with IT2B on the clinical dataset are probably the result of assuming Gaussian parameter distributions and/or of using the FOCE approximation. NPEM and NPAG, having no constraints on the shape of the population parameter distributions, and which compute the likelihood exactly and estimate parameter values with greater precision, detected the more likely greater diversity in the parameter values in the population studied. In the first Monte Carlo study, NPAG and PEM had more precise parameter estimates than either IT2B FOCE or NONMEM FOCE, as well as much more precise estimates than NONMEM FO. In the second Monte Carlo study, NPEM easily detected the bimodal parameter distribution at this initial step without requiring any further information. Population modelling methods using exact or accurate computations have more precise parameter estimation, better stochastic convergence properties and are, very importantly, statistically consistent. Nonparametric methods are better than parametric methods at analysing populations having unanticipated non-Gaussian or multimodal parameter distributions.

Aged↗

Overview of the HUPO Plasma Proteome Project: results from the pilot phase with 35 collaborating laboratories and multiple analytical groups, generating a core dataset of 3020 proteins and a publicly-available database.

HUPO initiated the Plasma Proteome Project (PPP) in 2002. Its pilot phase has (1) evaluated advantages and limitations of many depletion, fractionation, and MS technology platforms; (2) compared PPP reference specimens of human serum and EDTA, heparin, and citrate-anti-coagulated plasma; and (3) created a publicly-available knowledge base (www.bioinformatics.med.umich.edu/hupo/ppp; www.ebi.ac.uk/pride). Thirty-five participating laboratories in 13 countries submitted datasets. Working groups addressed (a) specimen stability and protein concentrations; (b) protein identifications from 18 MS/MS datasets; (c) independent analyses from raw MS-MS spectra; (d) search engine performance, subproteome analyses, and biological insights; (e) antibody arrays; and (f) direct MS/SELDI analyses. MS-MS datasets had 15 710 different International Protein Index (IPI) protein IDs; our integration algorithm applied to multiple matches of peptide sequences yielded 9504 IPI proteins identified with one or more peptides and 3020 proteins identified with two or more peptides (the Core Dataset). These proteins have been characterized with Gene Ontology, InterPro, Novartis Atlas, OMIM, and immunoassay-based concentration determinations. The database permits examination of many other subsets, such as 1274 proteins identified with three or more peptides. Reverse protein to DNA matching identified proteins for 118 previously unidentified ORFs. We recommend use of plasma instead of serum, with EDTA (or citrate) for anticoagulation. To improve resolution, sensitivity and reproducibility of peptide identifications and protein matches, we recommend combinations of depletion, fractionation, and MS/MS technologies, with explicit criteria for evaluation of spectra, use of search algorithms, and integration of homologous protein matches. This Special Issue of PROTEOMICS presents papers integral to the collaborative analysis plus many reports of supplementary work on various aspects of the PPP workplan. These PPP results on complexity, dynamic range, incomplete sampling, false-positive matches, and integration of diverse datasets for plasma and serum proteins lay a foundation for development and validation of circulating protein biomarkers in health and disease.

Algorithms↗

Comparison of sequence and structure-based datasets for nonredundant structural data mining.

Structural data mining studies attempt to deduce general principles of protein structure from solved structures deposited in the protein data bank (PDB). The entire database is unsuitable for such studies because it is not representative of the ensemble of protein folds. Given that novel folds continue to be unearthed, some folds are currently unrepresented in the PDB while other folds are overrepresented. Overrepresentation can easily be avoided by filtering the dataset. PDB_SELECT is a well-used representative subset of the PDB that has been deduced by sequence comparison. Specifically, structures with sequences that exhibit a pairwise sequence identity above a threshold value are weeded from the dataset. Although length criteria for pairwise alignments have a structural basis, this automated method of pruning is essentially sequence-based and runs into problems in the twilight zone, possibly resulting in some folds being overrepresented. The value-added structure databases SCOP and CATH are also a potential source of a nonredundant dataset. Here we compare the sequence-derived dataset PDB_SELECT with the structural databases SCOP (Structural Classification Of Proteins) and CATH (Class-Architecture-Topology-Homology). We show that some folds remain overrepresented in the PDB_SELECT dataset while other folds are not represented at all. However, SCOP and CATH also have their own problems such as the labor-intensiveness of the update process and the problem of determining whether all folds are equally or sufficiently distant. We discuss areas where further work is required.

Amino Acid Sequence↗

The use of mutual information in registration of CT and MRI datasets post permanent implant.

PURPOSE: To determine the feasibility of registration of MRI and CT datasets post permanent prostate implant by the use of mutual information. METHODS AND MATERIALS: Five patients who underwent permanent (125)I implant for prostate carcinoma were studied. Two weeks postimplant an axial CT, T2-weighted-axial, sagittal and coronal MRI, and T1-fat-saturation MRI scans were obtained. Registrations of MRI to CT and MRI to MRI datasets were performed by mutual information, an automated process of data registration matching all information in specified dataset regions of interest. Registration quality was evaluated by visual inspection, agreement with seed- to-seed registration, and histogram analysis. RESULTS: Rapid registration (<30 minutes) of CT and MRI datasets can be accomplished through the use of mutual information. All methods of registration evaluation confirmed excellent registration quality. Although D90 and V100 for the prostate were comparable between MRI- and CT-based dosimetry, dose to critical structures/microenvironments (anterior base, posterior base, bladder outlet, lower sphincter, bulbar urethra) defined on MRI varied widely. CONCLUSIONS: Efficient and accurate registration of MRI and CT datasets following prostate implant is possible, and improves the accuracy of postimplant dosimetry by superior definition of the prostate. Definition of critical microenvironments and adjacent structures will improve dose and toxicity correlation and ultimately improve planning strategies.

Brachytherapy↗

The effective use of a summary table and decision tree methodology to analyze very large healthcare datasets.

Very large datasets typically consists of millions of records, with many variables. Such datasets are stored and maintained by organizations because of the perceived potential information they contain. However, the problem with very large datasets is that traditional methods of data mining are not capable of retrieving this information because the software may be overwhelmed by the memory or computing requirements. In this article we outline a method that can analyze very large datasets. The method initially performs a data reduction step through the use of a summary table, which is then used as a reference dataset that is recursively partitioned to grow a decision tree.

Data Interpretation, Statistical↗

The development of a national agreed minimum diabetes dataset for New Zealand.

The development of a minimum diabetes dataset (MDD) for monitoring diabetes in New Zealand was intended to facilitate diabetes quality initiatives. Existing published datasets were reviewed and a draft MDD for New Zealand was distributed to all 147 specialist, general practice and relevant community groups. Data definitions were either identical or compatible with other datasets and dataset items included if there were at least six supportive replies. All groups were followed up by letter and telephone. A total of 26 (18%) replies were received. Comments were reviewed and the recommended MDD finalised. There was agreement that this dataset would be embedded into the software of at least three commercially available patient management systems. In conclusion, developling an agreed national MDD was difficult, in spite of its potential utility for local, regional and national collation of diabetes data allowing those involved to generate a picture of diabetes and its outcomes.

Database Management Systems↗

Asymmetric integration of various cancer datasets for identifying risk-associated variants and genes.

MOTIVATION: Cancer genomic research provides an opportunity to identify cancer risk-associated genes, but often suffers from undesirable low statistical power due to a limited sample size. Integrated analysis with different cancers has the potential to enhance statistical power for identifying pan-cancer risk genes. However, substantial heterogeneity across various cancers makes this challenging. RESULTS: Recently, a novel asymmetric integration method was developed that can deal with data heterogeneity and exclude unhelpful datasets from the analysis. We adapted and applied this method to integrate genotype datasets with matched case and control individuals from the Michigan Genomics Initiative, using each cancer as the primary dataset of interest and the other cancers as auxiliary datasets, respectively. Conditional logistic regression models were coupled with the asymmetric integrated framework to handle the matched case-control study design and permutation tests were performed to control for false discovery rates (FDRs). At the same FDR level, the integrated analysis found more potential genetic variants and genes that are associated with the risks of various cancers, showcasing the promise of the proposed approach for integrated analysis of cancer datasets. AVAILABILITY AND IMPLEMENTATION: Our method is available as source code at https://github.com/rxxwang/integrate_cancer.

Journal Article↗

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

MOTIVATION: An unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. RESULTS: Within 49&#x2009;h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Machine Learning↗

A statistical framework for combining and interpreting proteomic datasets.

MOTIVATION: To identify accurately protein function on a proteome-wide scale requires integrating data within and between high-throughput experiments. High-throughput proteomic datasets often have high rates of errors and thus yield incomplete and contradictory information. In this study, we develop a simple statistical framework using Bayes' law to interpret such data and combine information from different high-throughput experiments. In order to illustrate our approach we apply it to two protein complex purification datasets. RESULTS: Our approach shows how to use high-throughput data to calculate accurately the probability that two proteins are part of the same complex. Importantly, our approach does not need a reference set of verified protein interactions to determine false positive and false negative error rates of protein association. We also demonstrate how to combine information from two separate protein purification datasets into a combined dataset that has greater coverage and accuracy than either dataset alone. In addition, we also provide a technique for estimating the total number of proteins which can be detected using a particular experimental technique. AVAILABILITY: A suite of simple programs to accomplish some of the above tasks is available at www.unm.edu/~compbio/software/DatasetAssess

Algorithms↗

A scalable method for integration and functional analysis of multiple microarray datasets.

MOTIVATION: The diverse microarray datasets that have become available over the past several years represent a rich opportunity and challenge for biological data mining. Many supervised and unsupervised methods have been developed for the analysis of individual microarray datasets. However, integrated analysis of multiple datasets can provide a broader insight into genetic regulation of specific biological pathways under a variety of conditions. RESULTS: To aid in the analysis of such large compendia of microarray experiments, we present Microarray Experiment Functional Integration Technology (MEFIT), a scalable Bayesian framework for predicting functional relationships from integrated microarray datasets. Furthermore, MEFIT predicts these functional relationships within the context of specific biological processes. All results are provided in the context of one or more specific biological functions, which can be provided by a biologist or drawn automatically from catalogs such as the Gene Ontology (GO). Using MEFIT, we integrated 40 Saccharomyces cerevisiae microarray datasets spanning 712 unique conditions. In tests based on 110 biological functions drawn from the GO biological process ontology, MEFIT provided a 5% or greater performance increase for 54 functions, with a 5% or more decrease in performance in only two functions.

Algorithms↗

Sungear: interactive visualization and functional analysis of genomic datasets.

UNLABELLED: Sungear is a software system that supports a rapid, visually interactive and biologist-driven comparison of large datasets. The datasets can come from microarray experiments (e.g. genes induced in each experiment), from comparative genomics (e.g. genes present in each genome) or even from non-biological applications (e.g. demographics or baseball statistics). Sungear represents multiple datasets as vertices in a polygon. Each possible intersection among the sets is represented as a circle inside the polygon. The position of the circle is determined by the position of the vertices represented in the intersection and the area of the circle is determined by the number of elements in the intersection. Sungear shows which Gene Ontology terms are over-represented in a subset of circles or anchors. The intuitive Sungear interface has enabled biologists to determine quickly which dataset or groups of datasets play a role in a biological function of interest. AVAILABILITY: A live online version of Sungear can be found at http://virtualplant-prod.bio.nyu.edu/cgi-bin/sungear/index.cgi

Algorithms↗

Extending gerontological research through linking investigators' studies to public-use datasets.

PURPOSE: Public-use datasets can extend data collected by individual investigators in various ways: making external comparisons, providing additional data on individual respondents, and creating internal comparison groups. The authors describe the advantages and limitations of these methods and practical and conceptual issues in combining investigator-initiated and public-use datasets. DESIGN AND METHODS: These issues are illustrated with a study of functional decline among 674 patients following hospitalization for hip fracture that was augmented with data from a public-use dataset, the Established Populations for Epidemiologic Studies of the Elderly (EPESE). RESULTS: By creating an internal comparison group of EPESE respondents, frequency matched to hip fracture patients on age, sex, and baseline functional limitations, the authors formed a single dataset and performed multivariable analyses of factors associated with functional decline. IMPLICATIONS: Gerontological research may benefit by applying these methods to program evaluations and longitudinal analyses of health outcomes with numerous public-use datasets.

Activities of Daily Living↗

Bayesian model averaging of Bayesian network classifiers over multiple node-orders: application to sparse datasets.

Bayesian model averaging (BMA) can resolve the overfitting problem by explicitly incorporating the model uncertainty into the analysis procedure. Hence, it can be used to improve the generalization performance of Bayesian network classifiers. Until now, BMA of Bayesian network classifiers has only been performed in some restricted forms, e.g., the model is averaged given a single node-order, because of its heavy computational burden. However, it can be hard to obtain a good node-order when the available training dataset is sparse. To alleviate this problem, we propose BMA of Bayesian network classifiers over several distinct node-orders obtained using the Markov chain Monte Carlo sampling technique. The proposed method was examined using two synthetic problems and four real-life datasets. First, we show that the proposed method is especially effective when the given dataset is very sparse. The classification accuracy of averaging over multiple node-orders was higher in most cases than that achieved using a single node-order in our experiments. We also present experimental results for test datasets with unobserved variables, where the quality of the averaged node-order is more important. Through these experiments, we show that the difference in classification performance between the cases of multiple node-orders and single node-order is related to the level of noise, confirming the relative benefit of averaging over multiple node-orders for incomplete data. We conclude that BMA of Bayesian network classifiers over multiple node-orders has an apparent advantage when the given dataset is sparse and noisy, despite the method's heavy computational cost.

Algorithms↗

Application of machine learning and visualization of heterogeneous datasets to uncover relationships between translation and developmental stage expression of C. elegans mRNAs.

The relationships between genes in neighboring clusters in a self-organizing map (SOM) and properties attributed to them are sometimes difficult to discern, especially when heterogeneous datasets are used. We report a novel approach to identify correlations between heterogeneous datasets. One dataset, derived from microarray analysis of polysomal distribution, contained changes in the translational efficiency of Caenorhabditis elegans mRNAs resulting from loss of specific eIF4E isoform. The other dataset contained expression patterns of mRNAs across all developmental stages. Two algorithms were applied to these datasets: a classical scatter plot and an SOM. The outputs were linked using a two-dimensional color scale. This revealed that an mRNA's eIF4E-dependent translational efficiency is strongly dependent on its expression during development. This correlation was not detectable with a traditional one-dimensional color scale.

Algorithms↗

Nh3D: a reference dataset of non-homologous protein structures.

BACKGROUND: The statistical analysis of protein structures requires datasets in which structural features can be considered independently distributed, i.e. not related through common ancestry, and that fulfil minimal requirements regarding the experimental quality of the structures it contains. However, non-redundant datasets based on sequence similarity invariably contain distantly related homologues. Here we provide a reference dataset of non-homologous protein domains, assuming that structural dissimilarity at the topology level is incompatible with recognizable common ancestry. The dataset is based on domains at the Topology level of the CATH database which hierarchically classifies all protein structures. It contains the best refined representatives of each Topology level, validates structural dissimilarity and removes internally duplicated fragments. The compilation of Nh3D is fully scripted. RESULTS: The current Nh3D list contains 570 domains with a total of 90780 residues. It covers more than 70% of folds at the Topology level of the CATH database and represents more than 90% of the structures in the PDB that have been classified by CATH. We observe that even though all protein pairs are structurally dissimilar, some pairwise sequence identities after global alignment are greater than 30%. CONCLUSION: Nh3D is freely available as a reference dataset for the statistical analysis of sequence and structure features of proteins in the PDB. Regularly updated versions of Nh3D and the corresponding PDB-formatted coordinate sets are accessible from our Web site http://www.schematikon.org.

Algorithms↗

Coming to terms with datasets for diabetes care.

The development of datasets for the recording, transfer, monitoring and improvement of diabetes care has to take into account the perspectives of different stakeholders. Success of a dataset is dependent on shared ownership, professional support, clear objectives and development by consensus. Different types of concepts are discernible that share similar features, including the classes of disease, qualitative observations, quantitative observations, procedures and record specific concepts. The identification of these types assists the development of a dataset by clarifying the characteristics that need to be defined. The specification of the component data items ideally requires adherence to established terminology principles where each concept is uniquely labelled by an unambiguous term. These identified concepts should also be exclusive, stable over time, defined, reusable and relevant. Attention should also be given to the scope of each concept and consistency in handling additional detail. The dataset overall should adhere to principles of time representation and be manageable, implementable, regulated, extendible, piloted and ethical. Consideration of these principles should improve the likelihood that a dataset will be widely adopted, integrated and be further developed in existing patterns care.

Computer Communication Networks↗

On combining multiple microarray studies for improved functional classification by whole-dataset feature selection.

As microarray technologies become routinely applied in genome laboratories for studying gene expression, it is not uncommon that experiments on identical or similar sets of genes are conducted by multiple laboratories for various functional studies of these genes. Much of such data are often available to researchers for their data analysis, either through collaborators or from online gene expression databases. It will be useful to combine data from different microarray studies to improve the microarray data mining results. We show that the functional classification of genes from microarray data can be improved further by combining gene expression data from multiple microarray studies, even if the experimental focus or conditions for each experimental study may differ. However, blindly combining all available datasets may not always improve the analysis results---it is important to be selective of the datasets for inclusion. In our approach, we consider each dataset to be one feature, and then apply feature selection strategies to select appropriate datasets for training. With a simple hill-climbing method, we show that gene classification performances can be improved by whole-dataset feature selection.

Algorithms↗

How does spatial extent of fMRI datasets affect independent component analysis decomposition?

Spatial independent component analysis (sICA) of functional magnetic resonance imaging (fMRI) time series can generate meaningful activation maps and associated descriptive signals, which are useful to evaluate datasets of the entire brain or selected portions of it. Besides computational implications, variations in the input dataset combined with the multivariate nature of ICA may lead to different spatial or temporal readouts of brain activation phenomena. By reducing and increasing a volume of interest (VOI), we applied sICA to different datasets from real activation experiments with multislice acquisition and single or multiple sensory-motor task-induced blood oxygenation level-dependent (BOLD) signal sources with different spatial and temporal structure. Using receiver operating characteristics (ROC) methodology for accuracy evaluation and multiple regression analysis as benchmark, we compared sICA decompositions of reduced and increased VOI fMRI time-series containing auditory, motor and hemifield visual activation occurring separately or simultaneously in time. Both approaches yielded valid results; however, the results of the increased VOI approach were spatially more accurate compared to the results of the decreased VOI approach. This is consistent with the capability of sICA to take advantage of extended samples of statistical observations and suggests that sICA is more powerful with extended rather than reduced VOI datasets to delineate brain activity.

Acoustic Stimulation↗