Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Molecular classification of cancer: unsupervised self-organizing map analysis of gene expression microarray data.

An unsupervised self-organizing map-based clustering strategy has been developed to classify tissue samples from an oligonucleotide microarray patient database. Our method is based on the likelihood that a test data vector may have a gene expression fingerprint that is shared by more than one tumor class and as such can identify datasets that cannot be unequivocally assigned to a single tumor class. Our self-organizing map analysis completely separated the tumor from the normal expression datasets. Within the 14 different tumor types, classification accuracies on the order of approximately 80% correct were achieved. Nearly perfect classifications were found for leukemia, central nervous system, melanoma, uterine, and lymphoma tumor types, with very poor classifications found for colorectal, ovarian, breast, and lung tumors. Classification results were further analyzed to identify sets of differentially expressed genes between tumor and normal gene expressions and among each tumor class. Within the total pool of 1139 genes most differentially expressed in this dataset, subsets were found that could be vetted according to previously published literature sources to be specific tumor markers. Attempts to classify gene expression datasets from other sources found a wide range of classification accuracies. Discussions about the utility of this method and the quality of data needed for accurate tumor classifications are provided.

Algorithms↗

Analysis of field and laboratory data to derive selenium toxicity thresholds for birds.

In this paper, we critically evaluate the statistical approaches and datasets previously used to derive chronic egg selenium thresholds for mallard ducks (laboratory data) and black-necked stilts (field data). These effect concentration thresholds of 3%, 10% (EC10), or 20% have been used by regulatory agencies to set avian protection criteria and site remediation goals, thus the need for careful assessment of the data. The present review indicates that the stilt field dataset used to establish a frequently cited chronic avian egg selenium threshold of 6 mg/kg dry weight lacks statistical robustness (r2 = 0.19-0.28 based on generalized linear models), suggesting that stilt embryo sensitivity to selenium is highly variable or that factors other than selenium are principally responsible for the increase in effects observed at the lower range of this dataset. Hockey stick regressions used with the stilt field dataset improve the statistical relationship (r2 = 0.90-0.97) but result in considerably higher egg selenium thresholds (EC10 = 21-31 mg/kg dry wt). Laboratory-derived (for mallards) and field-derived (for stilts) teratogenicity EC10 values are quite similar (16-24 mg/kg dry wt). Laboratory data regarding mallard egg inviability and duckling mortality data provide the most sensitive and statistically robust chronic threshold (EC10) with logit, probit, and hockey stick regressions fitted to laboratory data, resulting in mean egg selenium EC10 values of 12 to 15 mg/kg dry weight (r2 = 0.75-0.90).

Animals↗

Causal discovery using a Bayesian local causal discovery algorithm.

This study focused on the development and application of an efficient algorithm to induce causal relationships from observational data. The algorithm, called BLCD, is based on a causal Bayesian network framework. BLCD initially uses heuristic greedy search to derive the Markov Blanket (MB) of a node that serves as the "locality" for the identification of pair-wise causal relationships. BLCD takes as input a dataset and outputs potential causes of the form variable X causally influences variable Y. Identification of the causal factors of diseases and outcomes, can help formulate better management, prevention and control strategies for the improvement of health care. In this study we focused on investigating factors that may contribute causally to infant mortality in the United States. We used the U.S. Linked Birth/Infant Death dataset for 1991 with more than four million records and about 200 variables for each record. Our sample consisted of 41,155 re-cords randomly selected from the whole dataset. Each record had maternal, paternal and child factors and the outcome at the end of the first year--whether the infant survived or not. Using the infant birth and death dataset as input, BLCD out-put six purported causal relationships. Three out of the six relationships seem plausible. Even though we have not yet discovered a clinically novel causal link, we plan to look for novel causal pathways using the full sample.

Algorithms↗

Estimating the number of clusters in DNA microarray data.

OBJECTIVES: The main objective of the research is an application of the clustering and cluster validity methods to estimate the number of clusters in cancer tumor datasets. A weighed voting technique is going to be used to improve the prediction of the number of clusters based on different data mining techniques. These tools may be used for the identification of new tumour classes using DNA microarray datasets. This estimation approach may perform a useful tool to support biological and biomedical knowledge discovery. METHODS: Three clustering and two validation algorithms were applied to two cancer tumor datasets. Recent studies confirm that there is no universal pattern recognition and clustering model to predict molecular profiles across different datasets. Thus, it is useful not to rely on one single clustering or validation method, but to apply a variety of approaches. Therefore, combination of these methods may be successfully used for the estimation of the number of clusters. RESULTS: The methods implemented in this research may contribute to the validation of clustering results and the estimation of the number of clusters. The results show that this estimation approach may represent an effective tool to support biomedical knowledge discovery and healthcare applications. CONCLUSION: The methods implemented in this research may be successfully used for the estimation of the number of clusters. The methods implemented in this research may contribute to the validation of clustering results and the estimation of the number of clusters. These tools may be used for the identification of new tumour classes using gene expression profiles.

Central Nervous System Neoplasms↗

Variability and reproducibility of rubidium-82 kinetic parameters in the myocardium of the anesthetized canine.

UNLABELLED: Kinetic analysis of 82Rb dynamic PET data produces quantitative measures which could be used to evaluate ischemic heart disease. These measures have the potential to generate objective comparisons of different patients or the same patient at different times. To achieve this potential, it is essential to determine the variability and reproducibility of the kinetic parameters. METHODS: A total of 48 82Rb dynamic PET datasets were acquired from two pure bred beagles. Each animal underwent eight 82Rb PET studies with essentially the same protocol for three successive weeks. Data were acquired with the Donner 600-Crystal Positron Tomograph (PET600). In each week, single-slice dynamic 82Rb PET datasets were collected with the animal at rest at three different gantry positions separated by 5 mm. Additional dataset were collected after dipyridamole infusion and after administration of aminophylline to induce a return to rest. A two-compartment kinetic model with correction for myocardial vasculature and spillover from the left ventricular blood pool was used to analyze the dynamic datasets. Model parameters for uptake (k1), washout (k2) and vascular fraction (fv) were estimated in 11-14 myocardial regions of interest (ROIs) using a weighted least-squares criterion. Statistical fluctuation due to the PET acquisition process was minimized by using a relatively high 82Rb dose (about 30 mCi) to take advantage of the high count rate capacity of the PET600. RESULTS: The variation in mean k1, where the mean is taken over the myocardial ROIs was 10%-20% (Dog 1) and 15%-50% (Dog 2) among the rest studies conducted on the same date. Similar variation was evident in comparing studies in the same animal for different weeks. CONCLUSION: Spatial and temporal variation in estimates of the uptake rate (k1) of 82Rb in the resting myocardium of the anesthetized canine are small in relation to the functional increase in k1 following dipyridamole infusion.

Aminophylline↗

Consistently processed RNA sequencing data from 50 sources enriched for pediatric data.

Larger cohorts improve the power of tumor gene expression analysis, but the signal is muddied if datasets are processed using different methods or have inaccurate metadata. Here we present five compendia containing consistently processed gene expression data derived from 16,446 diverse RNA sequencing datasets. To create the compendia, we obtained access to RNA sequence data from repositories containing public data as well as clinical partners with access to non-published data. We then assessed the quality, quantified gene expression, harmonized clinical metadata, and released the expression values and metadata without access restrictions. These datasets have been used for diverse projects ranging from identifying similarities between tumor types to assessing how well cell lines recapitulate tumors. They have also been used for n-of-1 analysis to identify genes with unusual expression patterns in a single sample and to infer molecular diagnosis. The comparison to new data is enabled by our dockerized, freely available pipeline. The compendia have been cited in at least 20 publications.

Humans↗

Temporal multiomics gene expression data of human embryonic stem cell-derived cardiomyocyte differentiation.

Human embryonic stem cells (hESCs) serve as a valuable in vitro model for studying early human developmental processes due to their ability to differentiate into all three germ layers. Here, we present a comprehensive multi-omics dataset generated by differentiating hESCs into cardiomyocytes via the mesodermal lineage, collecting samples at 10 distinct time points. We measured mRNA levels by mRNA sequencing (mRNA-seq), translation levels by ribosome profiling (Ribo-seq), and protein levels by quantitative mass spectrometry-based proteomics. Technical validation confirmed high quality and reproducibility across all datasets, with strong correlations between replicates. This extensive dataset provides critical insights into the complex regulatory mechanisms of cardiomyocyte differentiation and serves as a valuable resource for the research community, aiding in the exploration of mammalian development and gene regulation.

Humans↗

An integrated global resource of wetland microbiomes linking environmental metadata, community profiles, and genome-resolved metabolic traits.

Wetlands are biogeochemical hotspots pivotal to global carbon and nutrient cycling, yet genome-resolved studies across diverse wetland types remain limited. To address this, we constructed a global wetland metagenomic dataset, integrating environmental metadata, community profiles, and genome-resolved metabolic traits. This dataset comprises 1,962 samples-including 129 newly sequenced field-collected samples-from lakes, rivers, paddies, marshes, and coastal wetlands, spanning water, soil, and sediment habitats. We generated comprehensive taxonomic profiles for all 1,962 samples, and used 251 samples to reconstruct 5,704 sample-specific metagenome-assembled genomes (MAGs). These MAGs were subsequently dereplicated to establish a normalized, non-redundant catalog of 4,164 representative genomes. We further mapped gene repertoires to 549 KEGG modules to decode the metabolic potential of all 5,704 MAGs. This dataset depicts an overview of microbial genomic diversity across global wetlands and provides a comprehensive resource for understanding the metabolic capabilities, ecology, and evolution of wetland microbiomes.

Wetlands↗

Integration of single cell multiomics data by deep transfer hypergraph neural network.

Multi-omics characterization of individual cells offers remarkable potential for analyzing the dynamics and relationships of gene regulatory states across millions of cells. How to integrate multimodal data is an open problem, existing integration methods struggle with accuracy and modality-specific biological variation retention. In this paper, we present scHyper (scalable, interpretable machine learning for single cell integration), a low-code and data-efficient deep transfer model designed for integrating paired and unpaired single-cell multimodal data. We benchmark scHyper against datasets from different multimodal data. ScHyper learns a low-dimensional representation and aligns the covariance matrices of the measured modalities, achieving high accuracy even with large scale atlas-level datasets with low memory and computational time across different cell lines, shedding light on regulatory relationships between different types of omics. Altogether, we show that scHyper is a versatile and robust tool for cell-type label transfer and integration from multimodal single-cell datasets.

Single-Cell Analysis↗

Integrative multi-omics quantitative trait loci prioritize CASP7 as a candidate protective gene for cataract.

Cataracts are the leading cause of vision loss worldwide. Despite surgery being the only effective treatment, its economic burden highlights the necessity of exploring the pathogenesis of cataracts. In this study, we analyzed 4 large-scale GWAS (genome-wide association study) datasets for cataracts and performed SMR analysis along with heterogeneity in dependent instruments (HEIDI) testing to explore the effects of methylation, expression, and protein QTLs on cataracts. We further validated shared genetic variants through COLOC analysis. Additionally, we searched datasets related to cataracts from the Gene Expression Omnibus (GEO) database for differentially expressed genes (DEGs) and Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes pathway (KEGG) enrichment analyses. By integrating summary-based Mendelian randomization (SMR) results with bioinformatics findings, CASP7 showed a consistent protective-direction association with cataract risk (mQTL: OR [95% CI] = 0.959 [0.941-0.977], FDR-adjusted P = .039; eQTL: OR [95% CI] = 0.897 [0.860-0.937], FDR-adjusted P = .0046; pQTL: OR [95% CI] = 0.597 [0.483-0.738], FDR-adjusted P = .00083). GEO-based analyses provided transcriptomic support for CASP7 involvement in cataract-related lens biology. These findings prioritize CASP7 as a genetically supported candidate protective gene associated with cataract risk. Because this study is based on public summary-level and transcriptomic datasets, the results should be interpreted cautiously and require functional validation in human lens-relevant systems.

Quantitative Trait Loci↗

ChromaFactor: Deconvolution of single-molecule chromatin organization with non-negative matrix factorization.

The investigation of chromatin organization in single cells holds great promise for identifying causal relationships between genome structure and function. However, analysis of single-molecule data is hampered by extreme yet inherent heterogeneity, making it challenging to determine the contributions of individual chromatin fibers to bulk trends. To address this challenge, we propose ChromaFactor, a novel computational approach based on non-negative matrix factorization that deconvolves single-molecule chromatin organization datasets into their most salient primary components. ChromaFactor provides the ability to identify trends accounting for the maximum variance in the dataset while simultaneously describing the contribution of individual molecules to each component. Applying our approach to two single-molecule imaging datasets across different genomic scales, we find that these primary components demonstrate significant correlation with key functional phenotypes, including active transcription, enhancer-promoter distance, and genomic compartment. Also, we find that some bulk trends exist at the single-cell level, but only in a small fraction of cells, suggesting that critical changes in genome organization may be driven by specific rare subpopulations rather than occurring uniformly across all cells. ChromaFactor offers a robust tool for understanding the complex interplay between chromatin structure and function on individual DNA molecules, pinpointing which subpopulations drive functional changes and fostering new insights into cellular heterogeneity and its implications for bulk genomic phenomena.

Animals↗

Robust estimation of the probabilities of 3-D clusters in functional brain images: application to PET data.

Recently, we presented a method (the CS method) for estimating the probability distributions of the sizes of supra threshold clusters in functional brain images [Ledberg A, Akerman S, Roland PE. 1998. Estimating the significance of 3D clusters in functional brain images. NeuroImage 8:113-128]. In that method, the significance of the observed test statistic (cluster size) is assessed by comparing it with a sample of the test statistic obtained from simulated statistical images (SSIs). These images are generated to have the same spatial autocorrelation as the observed statistical image (t-image) would have under the null hypothesis. The CS method relies on the assumptions that the t-images are stationary and that they can be transformed to have a normal distribution. These assumptions are not always valid, and thus limit the applicability of the method. The purpose of this paper is to present a modification of the previous method, that does not depend on these assumptions. This modified CS method (MCS) uses the residuals in the linear model as a model of a dataset obtained under the null hypothesis. Subsequently, datasets with the same distribution as the residuals are generated, and from these datasets the SSIs are derived. These SSIs are t-distributed. Thus, a conversion to normal distribution is no longer needed. Furthermore, no assumptions concerning the stationarity of the statistical images are needed. The MCS method is validated on both synthetical images and PET images and is shown to give accurate estimates of the probability distribution of the cluster size statistic.

Brain↗

IARC p53 mutation database: a relational database to compile and analyze p53 mutations in human tumors and cell lines. International Agency for Research on Cancer.

The tumor suppressor p53 gene is the most frequently mutated gene in human cancer. To date, more than 10,000 mutations have been described in the literature, and these data are available in various electronic formats on the World Wide Web. Here we describe the structure and format of the different p53 datasets maintained and curated at the International Agency for Research on Cancer (IARC) in Lyon, France. These include p53 somatic mutations (more than 10,000 entries), p53 germline mutations (144 entries), and p53 polymorphisms (13 entries), with the somatic mutations organized into a relational database using AccessTM. The main features of these datasets are (1) controlled entry with standardized format and restricted vocabulary, (2) inclusion of annotations on individual characteristics and exposures, and (3) a classification of pathologies based on the International Classification of Diseases for Oncology (ICD-O). In addition, several interfaces have been developed to analyze the data in order to produce mutation spectra, codon analyses, or visualization of the mutation with the tertiary structure of the protein. All datasets and tools for analysis are available at http://www.iarc.fr/p53/homepage.

Databases, Factual↗

Comparison of variance components and sibpair-based approaches to quantitative trait linkage analysis in unselected samples.

We compared the statistical performance of sibpair-based and variance components approaches to multipoint linkage analysis of a quantitative trait in unselected samples. As a benchmark dataset, we used the simulated family data from Genetic Analysis Workshop 10 [Goldin et al., 1997], and each method was used to screen all 200 replications of the GAW10 genome for evidence of linkage to quantitative trait Q1. The sibpair and variance components methods were each applied to datasets comprising single-sibpairs and complete sibships, and for further comparison we also applied the variance components method to the nuclear family and extended pedigree datasets. For each analysis, the unbiasedness and efficiency of parameter estimation, the power to detect linkage, and the Type I error rate were estimated empirically. Sibpair and variance components methods exhibited comparable performance in terms of the unbiasedness of the estimate of QTL location and the Type I error rate. Within the single-sibpair and sibship sampling units, the variance components approach gave consistently superior power and efficiency of parameter estimation. Within each method, the statistical performance was improved by the use of the larger and more informative sampling units.

Chromosome Mapping↗

Linkage analysis of 150 high-risk prostate cancer families at 1q24-25.

Confirmation of linkage and estimation of the proportion of families who are linked in large independent datasets is essential to understanding the significance of cancer susceptibility genes. We report here on an analysis of 150 high-risk prostate cancer families (2,176 individuals) for potential linkage to the HPC1 prostate cancer susceptibility locus at 1q24-25. This dataset includes 640 affected men with an average age at prostate cancer diagnosis of 66. 8 years (range, 39-94), representing the largest collection of high-risk families analyzed for linkage in this region to date. Linkage to multiple 1q24-25 markers was strongly rejected for the sample as a whole (lod scores at theta = 0 ranged from -30.83 to -18. 42). Assuming heterogeneity, the estimated proportion of families linked (alpha) at HPC1 in the entire dataset was 2.6%, using multipoint analysis. Because locus heterogeneity may lead to false rejection of linkage, data were stratified based on homogeneous subsets. When restricted to 21 Caucasian families with five or more affected family members and mean age at diagnosis < = 65 years, the lod scores at theta = 0 remained less than -4.0. These results indicate that the overall portion of hereditary prostate cancer families whose disease is due to inherited variation in HPC1 may be less than originally estimated.

Adult↗

Towards a Standard Threshold for Genome Wide Significance in Dogs.

Genome-wide association studies (GWAS) are a foundational step in tying phenotype to genotype, relying on statistical significance thresholds to distinguish true- from false-positive signals of association. Dog genomics has long relied on per-study Bonferroni thresholds of significance, basing these on SNP chip levels of markers (~100&#x2009;k to >&#x2009;14&#x2009;M variable sites). However, as the field progresses into whole genome imputation analyses and more powerful meta-analyses, there is a clear need to develop a standard significance threshold for common-variant GWAS. Using 1591 dogs from the broad-ancestry Dog10K dataset, we performed permutation analysis and developed GWAS thresholds for datasets using either 1% or 5% minor allele frequencies. The resultant p-values, 4.2&#x2009;&#xd7;&#x2009;10-7 and 5.0&#x2009;&#xd7;&#x2009;10-7 respectively, are similar to previous Bonferroni levels (p-value ~6&#x2009;&#xd7;&#x2009;10-7), but less restrictive than the standard human p-value, 5&#x2009;&#xd7;&#x2009;10-8, which is sometimes used in dog studies. Given the diverse haplotypes from the >&#x2009;320 breeds in the Dog10K input dataset, we suggest a p-value of 4&#x2009;&#xd7;&#x2009;10-7 as a standard significance threshold that could be applied to any dog GWAS.

Animals↗

Congenital central hypoventilation syndrome: inheritance and relation to sudden infant death syndrome.

We evaluated the families of 50 children with idiopathic congenital central hypoventilation syndrome (CCHS) to 1) test genetic hypotheses, 2) explore the relationship to Hirschsprung disease (HD), and 3) examine other clinical findings including sudden infant death syndrome (SIDS) in relatives of CCHS patients. A questionnaire was administered to parents of each proband to determine a detailed pedigree and medical history for 3 generations including 1,482 relatives. The data were analyzed under the unified mixed model method (assumes individual genotype composed of multifactorial [MF] and major locus [ML] components). Analysis was made of the Total dataset and on subdivided data sets: HIR1 = families of probands with HD (n = 8) vs. HIR2 = families of probands without HD; then under a premise that severe, chronic constipation may be a milder form of HD (i.e., ganglion cells present but dysfunctional), CON1 = families of probands with HD or constipation (n = 13) vs. CON2 = families of probands without HD or constipation. By statistical genetic analysis of the Total, HIR1, and CON1 datasets, the MF and ML hypotheses were about equally likely, with the MF model slightly more parsimonious. Although HIR2 and CON2 datasets indicated no familiality, statistical evidence of heterogeneity between the results of HIR1 and HIR2, or between CON1 and CON2 was lacking. A SIDS incidence of 11.2/1,000 was documented among the relatives of CON1 vs. 1.8/1,000 among relatives of CON2. Our results are consistent with familiality by either MF or ML models. Recurrence risk is likely < 5%. The relationship of CCHS to the high familial incidence of SIDS is intriguing and demands further investigation.

Female↗

Joint modelling of repeated transitions in follow-up data--a case study on breast cancer data.

In longitudinal studies where time to a final event is the ultimate outcome often information is available about intermediate events the individuals may experience during the observation period. Even though many extensions of the Cox proportional hazards model have been proposed to model such multivariate time-to-event data these approaches are still very rarely applied to real datasets. The aim of this paper is to illustrate the application of extended Cox models for multiple time-to-event data and to show their implementation in popular statistical software packages. We demonstrate a systematic way of jointly modelling similar or repeated transitions in follow-up data by analysing an event-history dataset consisting of 270 breast cancer patients, that were followed-up for different clinical events during treatment in metastatic disease. First, we show how this methodology can also be applied to non Markovian stochastic processes by representing these processes as "conditional" Markov processes. Secondly, we compare the application of different Cox-related approaches to the breast cancer data by varying their key model components (i.e. analysis time scale, risk set and baseline hazard function). Our study showed that extended Cox models are a powerful tool for analysing complex event history datasets since the approach can address many dynamic data features such as multiple time scales, dynamic risk sets, time-varying covariates, transition by covariate interactions, autoregressive dependence or intra-subject correlation.

Algorithms↗