Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Independent methods for evolutionary genetic dating provide insights into Y-chromosomal STR mutation rates confirming data from direct father-son transmissions.

Five datasets consisting of samples jointly typed for Y-chromosomal Unique Event Polymorphism (UEP) and simple tandem repeat (STR) markers were re-examined with independent methods for dating the different UEP-defined lineages. We report on the results obtained with an original program which performs comparative dating (BARCODE) in comparison with coalescent analyses performed with BATWING under various prior conditions. For the first time these are equalized across datasets. We also report on the results concerning STR mutability as obtained with both methods. The dating results for the entire series of sub-haplogroups are highly correlated. Within coalescent analyses, dating-estimates under a wide range of priors tend to converge. As to STR mutation rates the main findings are: (1) large variations among loci within the same dataset with both methods, also when the same prior was used for all loci; (2) figures in most cases above 1x10(-3) and often above 2x10(-3); (3) a few loci that mutate differently across studies. These results closely match those obtained from direct observation of father-son transmissions. Overall, this work supports the use of genetic dating procedures that take into account the complexity of the phenomenon, with a repertoire of priors tailored on the particular dataset.

Chromosomes, Human, Y↗

MED12-STAT1-TAP2 axis regulates CD8 + T cell cytotoxicity and mediates immunotherapy outcome in non-small cell lung cancer.

Although immunotherapy for late-stage non-small cell lung carcinoma (NSCLC) has been clinically utilized, its prognosis remains highly heterogeneous, prompting us to investigate novel predictive immunotherapy biomarkers for NSCLC. We analyzed the correlations between MED12 nonsynonymous mutations and survival, clinical, genomic, transcriptomic information, and immune infiltration information through data mining across multiple datasets. We also investigated the mechanism of MED12 using luciferase assay, Western blot, ChIP-PCR, and siRNA. MED12 is significantly associated with survival in completely independent immunotherapy datasets, including MSKCC (N = 350), Naiyer2015 (N = 34), our own (N = 295) and the pan-cancer dataset, but not in the TCGA dataset, where patients received non-immunotherapy regimens. Mutations in MED12 showed no significant correlation with known metrics (TMB, IPS/CTLA4/PD1 status, PD-1/PD-L1 expression, and TCR/BCR status) or DNA Damage Repair (DDR) pathway mutations, yet they carried independent prognostic information according to the Cox multivariate regression. On the other hand, MED12 mutation is significantly associated with multiple immune-related pathways and immune infiltration of CD8 + T cells and activated NK cells. Lactate dehydrogenase assay revealed that knockdown of TAP2 restored the upregulation of CD8 + T cell cytotoxicity triggered by MED12 knockdown. ChIP-PCR, luciferase assay and siRNA knock down assay indicate that MED12 binds to the promoter region of STAT1 to suppress its transcription, while the transcription factor STAT1 promotes the transcription of TAP2, thus inhibiting the antigen processing and presentation. Collectively, MED12 mutation is an independent and valuable biomarker for predicting the response to immune checkpoint inhibitor (ICI)therapy in NSCLC by modulating CD8 + T cell cytotoxicity via the STAT1/TAP2 axis.

Humans↗

Accelerated dynamic Fourier velocity encoding by exploiting velocity-spatio-temporal correlations.

OBJECTIVE: To describe how the information content in a Fourier velocity encoding (FVE) scan can be transformed into a very sparse representation and to develop a method that exploits the compactness of the data to significantly accelerate the acquisition. MATERIALS AND METHODS: For validation, fully sampled FVE datasets were acquired in phantom and in vivo experiments. Fivefold and eightfold acceleration was simulated by using only one fifth or one eighth of the data for reconstruction in the proposed method based on the k-t BLAST framework. Reconstructed images were compared quantitatively to those from the fully sampled data. RESULTS: Velocity spectra in the accelerated datasets were comparable to the spectra from fully sampled datasets. The detected peak velocities remained accurate even at eightfold acceleration, and the overall shape of the spectra was well preserved. Slight temporal smoothing was seen in the accelerated datasets. CONCLUSION: A novel technique for accelerating time-resolved FVE scan is presented. It is possible to accelerate FVE to acquisition speeds comparable to a standard time-resolved phase-contrast scan.

Algorithms↗

Phylogenetic utility of protein (RPB2, beta-tubulin) and ribosomal (LSU, SSU) gene sequences in the systematics of Sordariomycetes (Ascomycota, Fungi).

The Sordariomycetes is an important group of fungi whose taxonomic relationships and classification is obscure. There is presently no multi-gene molecular phylogeny that addresses evolutionary relationships among different classes and orders. In this study, phylogenetic analyses with a broad taxon sampling of the Sordariomycetes were conducted to evaluate the utility of four gene regions (LSU rDNA, SSU rDNA, beta-tubulin and RPB2) for inferring evolutionary relationships at different taxonomic ranks. Single and multi-gene genealogies inferred from Bayesian and Maximum Parsimony analyses were compared in individual and combined datasets. At the subclass level, SSU rDNA phylogenies demonstrate their utility as a marker to infer phylogenetic relationships at higher levels. All analyses with SSU rDNA alone, combined LSU rDNA and SSU rDNA, and the combined 28 S rDNA, SSU rDNA and RPB2 datasets resulted in three subclasses: Hypocreomycetidae, Sordariomycetidae and Xylariomycetidae, which correspond well to established morphological classification schemes. At the ordinal level, the best resolved phylogeny was obtained from the combined LSU rDNA and SSU rDNA datasets. Individually, the RPB2 gene dataset resulted in significantly higher number of parsimony informative characters. Our results supported the recent separation of Boliniaceae, Chaetosphaeriaceae and Coniochaetaceae from Sordariales and placement of Coronophorales in Hypocreomycetidae. Microascales was found to be paraphyletic and Ceratocystis is phylogenetically associated to Faurelina, while Microascus and Petriella formed another clade and basal to other members of Halosphaeriales. In addition, the order Lulworthiales does not appear to fit in any of the three subclasses. Congruence between morphological and molecular classification schemes is discussed.

Ascomycota↗

Prediction of plasma protein binding of drugs using Kier-Hall valence connectivity indices and 4D-fingerprint molecular similarity analyses.

A 115 compound dataset for HSA binding is divided into the training set and the test set based on molecular similarity and cluster analyses. Both Kier-Hall valence connectivity indices and 4D-fingerprint similarity measures were applied to this dataset. Four different predictive schemes (SM, SA, SR, SC) were applied to the test set based on the similarity measures of each compound to the compounds in the training set. The first algorithmic scheme (SM) predicts the binding affinity of a test compound using only the most similar training set compound's binding affinity. This scheme has relatively poor predictivity based both on Kier-Hall valence connectivity indices similarity measures and 4D-fingerprints similarity analyses. The other three algorithmic schemes (SM SR, SC), which assign a weighting coefficient to each of the top-ten most similar training set compounds, have reasonable predictivity of a test set. The algorithmic scheme which categorizes the most similar compounds into different weighted clusters predicts the test set best. The 4D-fingerprints provide 36 different individual IPE/IPE type molecular similarity measures. This study supports that some types of similarity measures are highly similar to one another for this dataset. Both the Kier-Hall valence connectivity indices similarity measures and the 4D-fingerprints have nearly same predictivity for this particular dataset.

Algorithms↗

Novel approach to evolutionary neural network based descriptor selection and QSAR model development.

Capability of evolutionary neural network (ENN) based QSAR approach to direct the descriptor selection process towards stable descriptor subset (DS) composition characterized by acceptable generalization, as well as the influence of description stability on QSAR model interpretation have been examined. In order to analyze the DS stability and QSAR model generalization properties multiple random dataset partitions into training and test set were made. Acceptability criteria proposed by Golbraikh et al. [J. Comput.-Aided Mol. Des., 17 (2003) 241] have been chosen for selection of highly predictive QSAR models from a set of all models produced by ENN for each dataset splitting. All QSAR models that pass Golbraikh's filter generated by ENN for each dataset partition were collected. Two final DS forming principles were compared. Standard principle is based on selection of descriptors characterized by highest frequencies among all descriptors that appear in the pool [J. Chem. Inf. Comput. Sci., 43 (2003) 949]. Search across the model pool for DS that are stable against multiple dataset subsampling i.e. universal DS solutions is the basis of novel approach. Based on described principles benzodiazepine QSAR has been proposed and evaluated against results reported by others in terms of final DS composition and model predictive performance.

Benzodiazepines↗

Metrics for external model evaluation with an application to the population pharmacokinetics of gliclazide.

PURPOSE: The aim of this study is to define and illustrate metrics for the external evaluation of a population model. MATERIALS AND METHODS: In this paper, several types of metrics are defined: based on observations (standardized prediction error with or without simulation and normalized prediction distribution error); based on hyperparameters (with or without simulation); based on the likelihood of the model. All the metrics described above are applied to evaluate a model built from two phase II studies of gliclazide. A real phase I dataset and two datasets simulated with the real dataset design are used as external validation datasets to show and compare how metrics are able to detect and explain potential adequacies or inadequacies of the model. RESULTS: Normalized prediction errors calculated without any approximation, and metrics based on hyperparameters or on objective function have good theoretical properties to be used for external model evaluation and showed satisfactory behaviour in the simulation study. CONCLUSIONS: For external model evaluation, prediction distribution errors are recommended when the aim is to use the model to simulate data. Metrics through hyperparameters should be preferred when the aim is to compare two populations and metrics based on the objective function are useful during the model building process.

Algorithms↗

Network-based integration of metabolomics data from large-scale repositories.

INTRODUCTION: Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. OBJECTIVES: This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. METHODS: We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github.com/EloisaRL/Metabolomic-data-analysis-app/tree/main . RESULTS: As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. CONCLUSION: Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Metabolomics↗

Problems and suggested solutions in creating an archive of clinical trials data to permit later meta-analysis: an example of methotrexate trials in rheumatoid arthritis.

Because data archives contain patient-based rather than study-based data, they can address meta-analytic questions on uncommon outcomes and on predefined patient subsets, questions that are difficult to address using the traditional meta-analytic approach based on grouped data. We report the tasks involved in establishing the first data archive of rheumatoid arthritis trials. In general, problems stem from the heterogeneity of trials in the archive and we suggest some solutions. In the initial phases, difficulties include recruitment and incomplete participation of trial investigators, whereas later on, other issues arise, such as quality control, coping with different dataset designs, and incomplete documentation. Other issues include heterogeneous measures, missing variables, and comparing data across different visit intervals and trial lengths. Suggested solutions include requesting trial data in predefined archive-wide structures and asking for all possible documentation for each dataset. Data cleaning is necessary, as is rescaling of variables or developing unit-free outcomes, and estimating data for missing variables. Archive design should allow for referencing a patient's data among various datasets. Although one goal is to reduce the quantity of data in the archive while retaining information content, data from early stages of archive building must be accessible for developing new analysis datasets. Documentation of archive building and software choices are discussed. Our experience suggests data archiving for meta-analysis is time consuming and expensive, yet it provides a useful method for analyzing data from multiple trials.

Archives↗

Hybrid segmentation and virtual bronchoscopy based on CT images.

RATIONALE AND OBJECTIVES: Introduction of combination of the segmentation tool SegoMeTex and the virtual endoscopy system VIVENDI to perform virtual endoscopic inspections of the human lung. This virtual bronchoscopy system enables visualization of the tracheobronchial tree down to seventh generation. Furthermore, the modified virtual system visualizes hidden structures such as segmented vascular system or tumors. MATERIALS AND METHODS: The segmentation is based on image data acquired by a multislice computed tomography scanner. SegoMeTex is used to segment the tracheobronchial tree by a hybrid system with minimal user action. Similarly, the complementary pulmonary arterial can be segmented, whereas additional structures such as tumors are marked manually. On this dataset, subsequently, data structures of the inner surface for virtual endoscopy are generated. Finally, the dataset can be explored by a virtual bronchoscopy procedure using the VIVENDI system. RESULTS: The segmentation method was successfully tested on 22 patients. The hybrid segmentation system identified bronchi up to the sixth generation with a sensitivity of more than 58%, and a positive predictive value of more than 90%. After the segmentation, the datasets are explored interactively (>30 fps on a standard personal computer platform in real-time rendering) using the virtual endoscopy software. The exploration exposed a high-quality reconstruction, even of small structures throughout the dataset. CONCLUSION: Virtual bronchoscopy in combining with a highly sensitive segmentation is a valuable tool for the localization and measurement of stenosis for resection planning.

Bronchial Diseases↗

Estimation of demography and mutation rates from one million haploid genomes.

As genetic sequencing costs have plummeted, datasets with sizes previously unthinkable have begun to appear. Such datasets present opportunities to learn about evolutionary history, particularly via rare alleles that record the very recent past. However, beyond the computational challenges inherent in the analysis of many large-scale datasets, large population-genetic datasets present theoretical problems. In particular, the majority of population-genetic tools require the assumption that each mutant allele in the sample is the result of a single mutation (the "infinite-sites" assumption), which is violated in large samples. Here, we present DR EVIL, a method for estimating mutation rates and recent demographic history from very large samples. DR EVIL avoids the infinite-sites assumption by using a diffusion approximation to a branching-process model with recurrent mutation. This approach results in tractable likelihoods that are accurate for rare alleles. We show that DR EVIL performs well in simulations and apply it to rare-variant data from one million haploid samples. We identify mutation-rate heterogeneity even after accounting for trinucleotide context and methylation status. We also predict that at modern sample sizes, the alleles at most polymorphic sites with high mutation rates represent the descendants of multiple mutation events.

Haploidy↗

Increasing prevalence of HIV-1 protease inhibitor-associated mutations correlates with long-term non-suppressive protease inhibitor treatment.

Treatment of human immunodeficiency virus type 1 with protease inhibitors (PIs) is associated with the emergence of resistance-associated mutations. Treatment-characterized datasets have been used to identify novel treatment-associated protease mutations. In this study, we utilized two large reference laboratory databases (>115,000 viral sequences) to identify non-established resistance-associated protease mutations. We found 20 non-established protease mutations occurring in 82% of viruses with a PI resistance score of 4-7, 62% of viruses with a resistance score of 1-3, and 35% of viruses with no predicted PI resistance. We correlated mutational prevalence to treatment duration in a treatment-characterized dataset of 2161 patients undergoing non-suppressive PI therapy. In the non-suppressed dataset, 24 mutations became more prevalent and three mutations became less prevalent after more than 48 months of non-suppressive PI-therapy. Longer durations of non-suppressive treatment correlated with higher PI resistance scores. Mutations at eight non-established positions that were more common in viruses with the longest duration of non-suppressive therapy were also more common in viruses with the highest PI resistance score. Covariation analysis of 3036 protease amino acid substitutions identified 75 positive and nine negative correlations between resistance associated positions. Our findings support the utility of reference laboratory datasets for surveillance of mutation prevalence and covariation.

Amino Acid Sequence↗

Beyond multidimensionality: a systematic review of recurrent frailty archetypes in community-dwelling older adults.

BACKGROUND: Frailty is a clinically heterogeneous geriatric syndrome commonly summarised using physical or multidomain severity scores. Whether person-centred analyses identify recurring within-frailty configurations has not been systematically examined in community-dwelling older adults. METHODS: We searched PubMed, Embase, MEDLINE, and CINAHL (January 2000-November 2025) for cross-sectional studies using latent class, latent profile, or analogous clustering methods to derive frailty subgroups. Quality was assessed using the AHRQ checklist and a purpose-built appraisal of person-centred model reporting. Study-derived classes were mapped in duplicate to a structured archetype framework developed through comparison of class-defining features across studies. RESULTS: Fourteen reports representing 12 independent datasets from eight countries were included. Six configurations were identified: minimally impaired reference, mobility-physical, nutritional-metabolic, cognitive-predominant, combined cognitive-physical, and psychosocial/mood-predominant. Convergence was measurement-dependent. The reference and mobility-physical configurations recurred across physical-only and multidomain indicator sets, while the combined cognitive-physical configuration appeared across several multidomain frameworks but required cognition to be measured. The remaining configurations emerged only when their defining domains were included. Evidence of prognostic value beyond aggregate frailty severity came from one deficit-index study. Collapsing shared-provenance reports and excluding the boundary-eligible study did not alter recurrence; excluding the Croatian dataset left five configurations recurrent, with the cognitive-predominant configuration supported by one independent dataset. CONCLUSIONS: Person-centred analyses identify recurring within-frailty configurations, but their apparent stability is partly measurement-dependent. A five-configuration core persisted after exclusion of the Croatian dataset, whereas the cognitive-predominant configuration remained weakly replicated. Harmonised indicators and rigorous external validation are needed before clinical application.

Humans↗

Real time computation and temporal coherence of opacity transfer functions for direct volume rendering of ultrasound data.

Opacity transfer function (OTF) generation for direct volume rendering of medical image data is an intensely discussed subject. Several automatic methods exist for CT and MRI data, which are not apt for ultrasound data, mainly due to its low signal-to-noise ratio. Furthermore, ultrasound (US) imaging is able to produce time-varying 3D datasets in real time thus opening the door to 4D visualization. However, OTF design for 4D datasets has not been exhaustively discussed until now. We present an efficient solution to generate an optimized OTF for a given 3DUS dataset in real time. Our method results in excellent visualization which we demonstrate using 3D fetus datasets. Finally, we discuss the applicability of our method to 4DUS visualization.

Austria↗

Progressive lossless compression of volumetric data using small memory load.

Nowadays, applications dealing with volumetric datasets, Medical applications being a typical representative, have become possible even on low cost computers due to a rapid increase of computer memory and processing power. However, even today, dealing with volumetric datasets creates two considerable problems: slow visualization and large file sizes. While recently, due to significant progress in graphics hardware, real-time or near real-time volume visualization has become possible, volume compression still remains a problematic issue. This paper introduces a new method for lossless compression of volumetric datasets. It is based on quadtree encoding. The method consists of three steps: during initialization, so-called division quadtree is built. The smallest unit of the division quadtree is called basic macro-block. During the processing phase, Boolean intersection is built on pairs of quadtrees, and the differences are stored. In the last phase, the variable length encoding is applied to reduce the entropy among the differences. Proposed method supports progressive visualization, what is especially important when a transfer trough the internet is needed. To test the efficiency of this method it was compared to popular octree encoding scheme. The results proved that data coherence is exploited more sufficiently using proposed quadtree approach. Additional advantage of this approach is that the algorithm does not need a lot of memory space. Only two quadtrees of two consecutive slices need be loaded in the memory at the same time. This feature makes this algorithm extremely attractive for possible hardware implementation. This paper introduces a new method for the compression of volumetric datasets. It is based on quadtree encoding. This method consists of three steps: during initialization, a so-called division quadtree is built. The smallest, unit of the division quadtree is called a basic macro-block. A Boolean intersection is built on pairs of quadtrees during the processing phase and the differences are stored. In the last phase, variable length encoding is applied to reduce entropy among the differences. This method has been compared with the popular octree-based method and gives, in general, better compression results. In addition, this method can be realized using small on-board memory.

Data Compression↗

Descriptive study comparing routine hospital administrative data with the Vascular Society of Great Britain and Ireland's National Vascular Database.

OBJECTIVE: To compare patient volume and outcomes in vascular surgery between an administrative data set (Hospital Episode Statistics) and a clinical database (National Vascular Database). DESIGN: Descriptive study. METHODS: Volume of cases determined by age, sex, year and procedure and in-hospital mortality by procedure for both datasets for patients undergoing either repair of abdominal aortic aneurysm, carotid endarterectomy or infrainguinal bypass over a three year period between 1st April 2001 and 31st March 2004. RESULTS: There were 32,242 admissions with a mention of the three selected vascular procedures within the administrative data set compared to 8462 within the clinical database. For NHS trusts common to both datasets, there were twice as many procedures (16,923) recorded within the administrative dataset compared to the clinical database. Patient characteristics were similar across both databases. Further analysis limiting the administrative data to records attributed to consultants known to contribute to the clinical database showed much closer agreement with only 11% more repairs of abdominal aortic aneurysm recorded within the administrative dataset compared to the National Vascular Database. CONCLUSIONS: There are significant differences in total numbers between HES and the NVD. If the National Vascular Database is to become a credible source of information on activity and outcomes for vascular surgery, there is a clear need to increase the number of contributing surgeons and to increase the completeness of data submitted. Further analysis at individual record level is needed to identify other reasons for discrepancies which could help to enhance data quality, both within Hospital Episode Statistics and within the National Vascular Database.

Adult↗

Gene selection and classification from microarray data using kernel machine.

The discrimination of cancer patients (including subtypes) based on gene expression data is a critical problem with clinical ramifications. Central to solving this problem is the issue of how to extract the most relevant genes from the several thousand genes on a typical microarray. Here, we propose a methodology that can effectively select an informative subset of genes and classify the subtypes (or patients) of disease using the selected genes. We employ a kernel machine, kernel Fisher discriminant analysis (KFDA), for discrimination and use the derivatives of the kernel function to perform gene selection. Using a modified form of KFDA in the minimum squared error (MSE) sense and the gradients of the kernel functions, we construct an effective gene selection criterion. We assess the performance of the proposed methodology by applying it to three gene expression datasets: leukemia dataset, breast cancer dataset and colon cancer dataset. Using a few informative genes, the proposed method accurately and reliably classified cancer subtypes (or patients). Also, through a comparison study, we verify the reliability of the gene selection and discrimination results.

Algorithms↗

The impact of respiration on left atrial and pulmonary venous anatomy: implications for image-guided intervention.

BACKGROUND: Image-guided intervention using pre-acquired CT/MR 3-dimensional images is an emerging strategy for atrial fibrillation (AF) ablation but may be limited by its use of static images to depict dynamic physiology. The effect of biologic factors such as respiration on the left atrial-pulmonary venous (LA-PV) anatomy is not well understood but is likely to have important implications. Conventional CT/MR imaging is performed during an inspiratory breath-hold, while electroanatomical mapping (EAM) during "quiet" breathing approximates an expiratory breath-hold. This study examined the effects of respiration on LA-PV anatomy and the error introduced by respiration on the integration of EAM with 3D MR imaging. METHODS: Pre-procedural MRI angiography was performed at both end-expiration (EXP) and end-inspiration (INSP) in 20 patients undergoing AF catheter ablation. 3D INSP and EXP surface reconstructions of the LA-PVs were compared. In selected pts, EAM data acquired during the ablation procedure (n=7) were integrated with the 3D MRI datasets. RESULTS: Qualitative assessment of the INSP and EXP 3D images revealed splaying of the PVs and reduction in PV caliber of the right-sided PVs during held inspiration. After aligning these two datasets, the average surface-to-surface distance calculated by region ranged from 1.99mm (right middle PV) to 3.79mm (left superior PV). Registration of the EAM to the MRI models was better for the EXP dataset (2.30+/-0.73mm) than the INSP dataset (3.03+/-0.57mm; p=0.004). CONCLUSION: There are significant changes in LA-PV anatomy with respiration. MR images acquired during standard held inspiration may introduce unnecessary errors in registration during image-guided intervention.

Aged↗