Search PubMedSearch

SEARCH · Search PubMed

Results for “Data analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

[Improvement of resolution of positron annihilation radiation energy spectrum measurement by data analysis (author's transl)].

Recently much attention is being paid to positron annihilation radiation energy spectrum measurement as a simple method that gives information about the momentum distribution of an annihilating pair. However, this method has a disadvantage that it is sensitive to drifts of the measuring system and the resolution is still insufficient. We have succeeded to reduce the influence of the drifts to a negligibly small degree by using, in addition to the necessary procedure for a temperature control and stabilizing of AC power lines, a compensation technique in data analysis. We have further attempted to improve the resolution by deconvolution processing. By these procedures, we have been able to separate a narrow component which has not otherwise been resolved directly from the spectrum. The resolution attained by this way is 0.43 keV (FWHM), and is two- or three-fold superior to those ever reported.

Models, Theoretical

Physiological consequences of experimental cerebral missile injury and use of data analysis to predict survival.

The authors describe cerebrovascular and cerebral metabolic changes in monkeys, subjected to cerebral missile injury. After injury with BB pellet at 90 m/sec, there is a rapid rise in intracranial pressure (ICP), which reaches a peak 2 to 5 minutes posttrauma, and then falls to about 20 to 30 mm Hg. This, with a fall in mean blood pressure (MBP), results in a 50% reduction in cerebral perfusion pressure (CPP), Cerebral blood flow (CBF) is also reduced, although acutely there is no close relationship with (CPP). Cerebrovascular resistance falls initially and then at 30 minutes rises to very high values. Cerebral metabolic rates (CMR's) for oxygen fall after injury and remain low for the rest of the animal's life; CMR's for lactate rise immediately after injury and persists for 5 hours, then fall. After injury with a faster missile (180 m/sec), the ICP rises higher and faster, and the peak is shorter. The CCP is reduced in this injury to approximately 30 mm Hg, and only one animal survived more than 1 hour. With the conventional forms of data analysis, the length of survival after injury correlates well with MBP, ICP, and CBF, but separately they were completely unsatisfactory for prediction of an individuals prognosis. With the technique of multiple linear regression analysis, the survival of individual animals could be predicted with great accuracy. This is possible also when two postinjury parameters,CBF and MBP, are used.

Animals

Vancomycin Effectiveness in Reducing Surgical Site Infection in Posterior Spinal Fusion Surgery: A Retrospective Data Analysis of the STRIVE Trial.

STUDY DESIGN: Retrospective analysis of prospectively collected data. OBJECTIVE: To re-evaluate vancomycin as a preventive measure for surgical site infection (SSI). SUMMARY OF BACKGROUND DATA: Intrawound vancomycin powder is used to prevent SSIs in spinal surgery. Prior studies, often limited to single institutions or small samples, have shown mixed efficacy and potential increases in non- S. aureus and Gram-negative infections. We hypothesized that SSIs rates would be similar with and without intrawound vancomycin in posterior spinal fusion (PSF) surgery. METHODS: Prospectively collected data from the 3595 patients in the STaphylococcus aureus suRgical Inpatient Vaccine Efficacy (STRIVE) trial were stratified by intrawound antibiotic usage. Multivariate logistic regression assessed the effect of vancomycin use on SSI, adjusting for patient demographics and SSI-associated risk factors. Secondary outcomes included critical care stay, reoperation, sepsis, and hospital readmission. RESULTS: Of 3311 patients who underwent surgery, 847 (26%) received only intrawound vancomycin and 1534 (46%) received no intrawound antibiotics. Sixty (8%) patients developed postoperative SSI, of whom 20 (33%) had received intrawound vancomycin. Receiving intrawound vancomycin was not associated with SSI incidence versus no intrawound antibiotics [odds ratio (OR): 0.77; 95% CI: 0.42-1.42], critical care stay (OR: 0.94; 95% CI: 0.78-1.12), or sepsis (OR: 2.04; 95% CI: 0.62-6.73). However, intrawound vancomycin was associated with increased odds of hospital readmission (OR: 1.82; 95% CI: 1.28-2.6; P < 0.001) and reoperation (OR: 1.75; 95% CI: 1.18-2.6; P = 0.005). Factors significantly associated with intrawound vancomycin use included intraoperative antibiotic readministration (OR: 2.97; 95% CI: 1.36-6.5; P =0.006) and hospital location, lower odds in Europe (OR: 0.13; 95% CI: 0.06-0.29; P < 0.001) or Asia (OR: 0.02; 95% CI: 0-0.08; P < 0.001) versus North America. CONCLUSIONS: Intraoperative vancomycin use was not associated with reduced SSI incidence compared with no intrawound antibiotics after PSF surgery. LEVEL OF EVIDENCE: Level II.

Humans

Some applications of categorical data analysis to epidemiological studies.

Several examples of categorized data from epidemiological studies are analyzed to illustrate that more informative analysis than tests of independence can be performed by fitting models. All of the analyses fit into a unified conceptual framework that can be performed by weighted least squares. The methods presented show how to calculate point estimate of parameters, asymptotic variances, and asymptotically valid chi 2 tests. The examples presented are analysis of relative risks estimated from several 2 x 2 tables, analysis of selected features of life tables, construction of synthetic life tables from cross-sectional studies, and analysis of dose-response curves.

Actuarial Analysis

OmicsQ: a user-friendly platform for interactive quantitative omics data analysis.

MOTIVATION: High-throughput omics technologies generate complex datasets with thousands of features that are quantified across multiple experimental conditions, but often suffer from incomplete measurements, missing values, and individually fluctuating variances. This requires analytical tools for accurate, deep and insightful biological interpretation, capable of dealing with a large variety of data properties and different amounts of completeness. Software capable of handling such data complexity and integrating with external applications for downstream analysis remains rare and mostly relies on programming-based environments, limiting accessibility for researchers without computational expertise. RESULTS: We present OmicsQ, an interactive, web-based platform designed to streamline quantitative omics data analysis. OmicsQ provides an intuitive, browser-based visualization interface that integrates established statistical processing tools. Those include robust batch correction, automated experimental design annotation, and handling of missing data without imputation, which maintains data integrity and avoids artifacts from a priori assumptions. OmicsQ seamlessly interacts with external applications (e.g. PolySTest, VSClust, ComplexBrowser) for statistical testing, clustering, analysis of protein complex behavior, and pathway enrichment, offering a comprehensive and flexible workflow from data import to biological interpretation that is broadly applicable across domains. AVAILABILITY AND IMPLEMENTATION: OmicsQ is implemented in R and Shiny and is available at https://computproteomics.bmb.sdu.dk/app_direct/OmicsQ. Source code and installation instructions: https://github.com/computproteomics/OmicsQ, DOI: 10.5281/zenodo.17778420.

Software

Multimodal CustOmics: A unified and interpretable multi-task deep learning framework for multimodal integrative data analysis in oncology.

Characterizing cancer presents a delicate challenge as it involves deciphering complex biological interactions within the tumor's microenvironment. Clinical trials often provide histology images and molecular profiling of tumors, which can help understand these interactions. Despite recent advances in representing multimodal data for weakly supervised tasks in the medical domain, achieving a coherent and interpretable fusion of whole slide images and multi-omics data is still a challenge. Each modality operates at distinct biological levels, introducing substantial correlations between and within data sources. In response to these challenges, we propose a novel deep-learning-based approach designed to represent multi-omics & histopathology data for precision medicine in a readily interpretable manner. While our approach demonstrates superior performance compared to state-of-the-art methods across multiple test cases, it also deals with incomplete and missing data in a robust manner. It extracts various scores characterizing the activity of each modality and their interactions at the pathway and gene levels. The strength of our method lies in its capacity to unravel pathway activation through multimodal relationships and to extend enrichment analysis to spatial data for supervised tasks. We showcase its predictive capacity and interpretation scores by extensively exploring multiple TCGA datasets and validation cohorts. The method opens new perspectives in understanding the complex relationships between multimodal pathological genomic data in different cancer types and is publicly available on Github.

Deep Learning

A versatile physiological data analysis system using an Intel 8080 microprocessor.

A microprocessor based physiological data processor has been realised. The system controls the data flow from physiological experiments and performs on-line mean and variance calculations with an output in graphical form. The analyser accepts one data point every 0.5 ms and has a capacity of 128 records each containing 800 data points. Post stimulus histogram and interval histogram analysis programs have also been written and implemented.

Computers

AmpSeqR: an R package for amplicon deep sequencing&#xa0;data&#xa0;analysis.

Amplicon sequencing (AmpSeq) is a methodology that targets specific genomic regions of interest for polymerase chain reaction (PCR) amplification so that they can be sequenced to a high depth of coverage. Amplicons are typically chosen to be highly polymorphic, usually with several highly informative, high frequency single nucleotide polymorphisms (SNPs) segregating in an amplicon of 100-200 base pair (bp). This allows high sensitivity detection and quantification of the frequency of each sequence within each sample making it suitable for applications such as low frequency somatic mosaicism detection or minor clone detection in mixed samples. AmpSeq is being increasingly applied to both biological and medical studies, in applications such as cancer, infectious diseases and brain mosaicism studies. Current bioinformatics pipelines for AmpSeq data processing lack downstream analysis, have difficulty distinguishing between true sequences and PCR sequencing errors and artifacts, and often require bioinformatic expertise. We present a new R package: AmpSeqR, designed for the processing of deep short-read amplicon sequencing data, with a focus on infectious diseases. The pipeline integrates several existing R packages combining them with newly developed functions to perform optimal filtering of reads to remove noise and improve the accuracy of the detected sequences data, permitting detection of very low frequency clones in mixed samples. The package provides useful functions including data pre-processing, amplicon sequence variants (ASVs) estimation, data post-processing, data visualization, and automatically generates a comprehensive Rmarkdown report that contains all essential results facilitating easy inclusion into reports and publications. AmpSeqR is publicly available at https://github.com/bahlolab/AmpSeqR.

High-Throughput Nucleotide Sequencing

SimpleMicrobiome: An integrated web-based platform for streamlined microbiome data analysis and visualization.

Microbiome studies require multiple analytical steps after initial sequence processing. These steps commonly include data harmonization, preprocessing, taxonomic profiling, diversity analysis, differential abundance testing, predictive modeling, network inference, and preparation of publication-ready outputs. Although robust packages are available for many of these tasks, routine use often depends on command-line workflows, repeated data reformatting, and method-specific scripting. These requirements can limit accessibility for experimental researchers and complicate consistent analysis across interdisciplinary teams. We developed SimpleMicrobiome, a web-based R Shiny platform that integrates established microbiome analysis methods into a single interactive downstream workflow. The application accepts standard abundance, taxonomy, and metadata tables, supports interactive preprocessing and sample filtering, and provides modules for taxa profile visualization, alpha and beta diversity analysis, ANCOM-BC2 and MaAsLin2 differential abundance testing, Random Forest modeling with SHAP-based interpretation, microbial association network inference using SparCC and SPIEC-EASI through NetCoMi, correlation heatmaps, and dbRDA/CAP-style association biplots. The platform is implemented as a modular Shiny application so that preprocessing choices are propagated across downstream analyses, results can be exported as figures and tables, and the same application can be run through the public server, source-code installation, or a Docker image. SimpleMicrobiome consolidates major downstream microbiome analysis tasks in an accessible browser-based environment while retaining links to established analytical frameworks. The platform may reduce technical barriers for non-programming users, improve consistency across exploratory and reporting-oriented analyses, and support collaborative microbiome research. The public application is available at https://simplemicrobiome.mglab.org, the source code is available at https://github.com/yjcho2252/SimpleMicrobiome, and a Docker image for local deployment is available at https://hub.docker.com/r/mglab2252/simplemicrobiome.

differential abundance

Identification of Critical Genes for Recurrent Aphthous Ulcer by Transcriptome Data Analysis and Mendelian Randomization.

PURPOSE: Recurrent aphthous ulcer (RAU) is a common oral mucosal disorder with a poorly understood etiology, significantly affecting patients' quality of life. This study aims to investigate critical genes linked to RAU and explore their biological mechanisms using transcriptomic data and Mendelian randomization (MR) analysis. MATERIALS AND METHODS: RAU-related gene expression data from the GEO database (GSE37265) were analyzed to identify differentially expressed genes (DEGs). A two-sample MR approach was used to assess the causal impact of expression quantitative trait loci (eQTL) on RAU. Critical genes were identified by intersecting DEGs with significant MR findings. GO and KEGG pathway enrichment analyses were performed, along with GSEA and immune cell infiltration analysis, to investigate the functions and mechanisms of these genes in RAU. RESULTS: A total of 184 differentially expressed genes (DEGs) were identified, while 339 RAU-associated genes were screened through MR analysis. Cross-validation further identified 7 critical genes. Among these, CCR1, ERP27, HCK, MICB, and SLC2A3 showed protective associations with RAU risk, whereas CD177 and IFITM1 were positively associated with increased risk. Enrichment analysis revealed that these genes are involved in specific biological processes, including cell migration, immune response, and metabolic regulation, which are closely linked to RAU pathogenesis. CONCLUSION: This systematic study comprehensively investigates the critical causative genes underlying RAU, emphasizing the intricate relationships between immune regulation and metabolic disturbances in its pathology. These findings lay a solid foundation for the development of novel biomarkers and may inform future research on targeted therapeutic strategies for RAU.

Stomatitis, Aphthous

NanoASV: a snakemake workflow for reproducible field-based Nanopore full-length 16S metabarcoding amplicon data analysis.

SUMMARY: NanoASV is a conda environment and snakemake-based workflow using state-of-the-art bioinformatics software to process full-length SSU rRNA (16S/18S) amplicons acquired with Oxford Nanopore Sequencing technology. Its strength lies in reproducibility, portability, and the possibility to run offline, allowing in-field analysis. It can be installed on the Nanopore MK1C sequencing device and process data locally. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available at https://github.com/ImagoXV/NanoASV and Zenodo archive at https://doi.org/10.5281/zenodo.14730742.

Software

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics

TaxTriage: an open-source metagenomic sequencing data analysis pipeline enabling putative pathogen detection.

MOTIVATION: TaxTriage is a comprehensive pathogen identification workflow designed for both short- and long-read untargeted DNA and RNA sequencing data. Combining read classification, mapping, and de novo assembly approaches, putative pathogens are identified through comparisons to curated pathogens and abundance expectations from healthy cohort data. Flexible installation options are enabled using Nextflow&#x2122; (NF), including cloud deployment via NF Tower (Seqera Platform) and local installation on a variety of systems, including standalone installations without external internet access. Final analysis summaries are compiled into an Organism Discovery Report, which lists likely pathogens and supporting data, including a custom confidence score. RESULTS: Evaluation of published in silico, clinical, and outbreak datasets identified performance comparable to alternative cloud-based processing pipelines for expected pathogen and co-infection detection with similar sensitivity and increased specificity. To support both public health and veterinary diagnostics communities, customization options have been incorporated to enable improved performance for host species of interest. AVAILABILITY AND IMPLEMENTATION: Source code for TaxTriage is freely available at https://github.com/jhuapl-bio/taxtriage. TaxTriage v2.1.1 has been archived on Zenodo at https://zenodo.org/records/17081354 to permit reproducible analysis as described in this manuscript.

Software

Relationship Between Number of Acute Pancreatitis Episodes and Risk of New-onset Diabetes in the U.S.: A Real-world Data Analysis.

INTRODUCTION: Acute pancreatitis (AP) is a common inflammatory disorder that is associated with increased risk for diabetes mellitus (DM). It remains unclear whether recurrent acute pancreatitis (RAP) is associated with further increased risk of incident DM. This study aims to investigate the association between RAP and incident DM using real-world data. METHODS: We conducted a retrospective cohort study using the MerativeTM MarketScan&#xae; claims database (2016-2023), identifying patients with AP and no prior history of DM at baseline. The primary exposure of interest, RAP, was defined as one or more episodes of AP occurring &#x2265;90 days after the index AP diagnosis, whereas one episode of AP referred to a single episode of AP (SAP) with no subsequent recurrence within 90 days following the index event. A multivariable stratified Cox proportional hazards regression models were used to determine the association between RAP and incident DM, identified using ICD-10 codes. RESULTS: In total, 16,184 individuals with AP (mean [SD] age: 45.8 [12.3]) contributed 40,712 person-years of follow-up, during which 1,477 incident cases of DM were documented. Individuals with RAP had an increased risk of incident DM compared with those with a SAP(adjusted HR, 1.92; 95% CI, 1.61-2.29). The risk increased significantly with the frequency of RAP. In comparing the modifying effect of patient demographics and comorbidities, a stronger association between RAP and incident DM was observed in females (adjusted HR, 2.44; 95% CI, 1.87-3.19) than in males (adjusted HR, 1.64; 95% CI, 1.30-2.07; Pinteraction=0.03). Also, stronger associations were observed among younger patients (18-46&#xa0;y) (adjusted HR=2.56; 95% CI, 1.97-3.31) and among non-tobacco abuse (adjusted HR=2.19; 95% CI, 1.81-2.65), with significant interactions for all comparisons (Pinteraction<0.05. CONCLUSIONS: In this real-world study, RAP was associated with an increased risk of incident DM. Our findings highlight an opportunity for glycemic monitoring and proactive management of patients with RAP to mitigate their risk of developing DM.

AP

Scale reliant mixed effects models enhance microbiome data analysis.

Linear models, including those used for differential abundance analyses, are frequently used in microbiome research to assess how experimental conditions (e.g., disease state or age) affect microbial abundance. Linear mixed-effects models (MEMs) extend linear models to accommodate complex designs, such as longitudinal sampling or hierarchical study structures. However, when applied to microbiome data, existing MEM approaches suffer from high false positive and false negative rates because sequence counts are compositional - they reflect relative rather than absolute abundances. Current methods attempt to overcome this limitation through normalization, but these approaches rely on strong, often unrealistic assumptions about the unmeasured biological scale (e.g., total microbial load). Here we introduce scale-reliant mixed-effects models (SR-MEM), which extend our earlier scale-reliant inference framework by explicitly modeling uncertainty in the unmeasured scale via user-defined probability distributions. By treating scale as a latent variable rather than fixing it through normalization, SR-MEM enables robust inference for complex experimental designs. SR-MEM can incorporate external scale measurements (e.g., flow cytometry, qPCR) or leverage scale information from independent studies to further improve inference. Across simulations and multiple real-world case studies, SR-MEM consistently controls the false discovery rate while maintaining comparable or higher power than standard approaches relying on normalization or bias correction. In reanalyses of published datasets, SR-MEM yields results that are more reproducible across studies and more consistent with known biological and pharmacological effects. SR-MEM provides a principled and practical framework for mixed-effects modeling of microbiome sequence count data in the presence of unmeasured biological scale. By avoiding normalization-based assumptions and instead propagating scale uncertainty through inference, SR-MEM improves error control and reproducibility in longitudinal and hierarchical studies. An accessible implementation is provided in the ALDEx3 R package.

Microbiota

Uncertainty Modeling Outperforms Machine Learning for Microbiome Data Analysis.

Microbiome sequencing measures relative rather than absolute abundances, providing no direct information about total microbial load. Normalization methods attempt to compensate, but rely on strong, often untestable assumptions that can bias inference. Experimental measurements of load (e.g., qPCR, flow cytometry) offer a solution, but remain costly and uncommon. A recent high-profile study proposed that machine learning could bypass this limitation by predicting microbial load from sequencing data alone. To evaluate this claim, we assembled mutt, the largest public database of paired sequencing and load measurements, spanning 35 studies and over 15,000 samples. Using mutt, we show that published machine learning models fail to generalize: on average they perform worse than a naive baseline that always predicted the training set mean. These failures stem from covariate shift-limited shared taxa between studies, differences in community composition, and differences in preprocessing pipelines-that silently derail model inputs. In contrast, Bayesian partially identified models do not attempt to impute microbial load, but instead propagate scale uncertainty through downstream analyses. Across 30 benchmark datasets, Bayesian partially identified models consistently outperformed normalization and machine learning approaches, providing a principled and reproducible foundation for microbiome inference.

16S rRNA-seq