Search PubMedSearch

SEARCH · Search PubMed

Results for “data analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Discussion of PET workshop reports, including recommendations of PET Data Analysis Working Group.

On May 1-2, 1989, a PET Data Analysis Working Group convened to consider positron emission tomography (PET) methodology and data analysis. The papers presented and the recommendations of the Group are reviewed. The Group recommended that a standard phantom of the human brain be used by different institutions to examine machine and data reconstruction PET variables. Interinstitutional comparisons could be aided by using a standard three-dimensional coordinate system. Deformations within individual diseased or atypical brains would require nonlinear as well as linear transformations to the standard space, using magnetic resonance images in register with the PET images. Methods for intersubject averaging of pixel-by-pixel or region-of-interest data, as well as appropriate statistical methods, need to be developed. PET data may first be exploratory and hypothesis-generating (with less stringent statistical theory), then later used to test hypotheses (with more stringent statistical criteria). Common databases, obtained by computer simulation models with known inherent structure, or directly by PET measurements on different groups, could be used to compare analytical and statistical methods among institutions.

Brain

SpectroPipeR-a streamlining post Spectronaut® DIA-MS data analysis R package.

SUMMARY: Proteome studies frequently encounter challenges in down-stream data analysis due to limited bioinformatics resources, rapid data generation, and variations in analytical methods. To address these issues, we developed SpectroPipeR, an R package designed to streamline data analysis tasks and provide a comprehensive, standardized pipeline for Spectronaut® DIA-MS data. This novel package automates various analytical processes, including XIC plots, ID rate summary, normalization, batch and covariate adjustment, relative protein quantification, multivariate analysis, and statistical analysis, while generating interactive HTML reports for e.g. ELN systems. AVAILABILITY AND IMPLEMENTATION: The SpectroPipeR package (manual: https://stemicha.github.io/SpectroPipeR/) was written in R and is freely available on GitHub (https://github.com/stemicha/SpectroPipeR).

Software

Evaluation of Bayesian estimation in comparison to NONMEM for population pharmacokinetic data analysis: application to pefloxacin in intensive care unit patients.

The pharmacokinetics of pefloxacin (PF) were investigated in a population of 74 intensive care unit patients receiving 400 mg bid as 1-hr infusion using (i) Bayesian estimation (BE) of individual patient parameters followed by multiple linear regression (MLR) analysis and (ii) NONMEM analysis. The data consisted of 3 to 9 PF plasma levels per patient measured over 1 to 3 dosage intervals (total 113) according to four different limited (suboptimal) sampling 3-point protocols. Twenty-nine covariates (including 15 comedications) were considered to explain the interpatient variability. Predicted PF CL for a patient with median covariates values was similar in both BE/MLR and NONMEM analysis (4.02 and 3.92 L/hr, respectively). Bilirubin level and age were identified as the major determinants of PF CL by both approaches with similar predicted magnitude of effects (about 40 and 30% decrease of median CL, respectively). Confounding effects were observed between creatinine clearance (26% decrease of PF CL in the BE/MLR model), simplified acute physiology score (a global score based on 14 biological and clinical variables) (18% decrease of median CL in the NONMEM model) and age (entered in both models) which were highly correlated in our data base. However, both models predicted similar PF CL for actual subpopulations by using actual covariate values. Finally, the NONMEM analysis allowed identification of an effect of weight on CL (decrease of CL for weight < 65 kg) whereas the BE/MLR analysis predicted an increase of CL in patients treated with phenobarbital. In conclusion, both approaches allowed identification of the major risk factors of PF pharmacokinetics in ICU patients. Their potential use at different stages of drug development is discussed.

Adolescent

Distributed data analysis in a multicenter study: the CARDIA Study.

Unlike distributed data entry, which is used in many large epidemiologic studies and multicenter clinical trials, distributed data analysis is a relatively new concept. This paper reports on the usefulness of such a system in the Coronary Artery Risk Development in Young Adults (CARDIA) Study. CARDIA distributes the entire examination dataset to participating centers soon after completion of each round of data collection. The process was designed to encourage more numerous, diverse, and rapid publications, and to allow for more efficient use of the manpower and expertise in centers. Responsibilities of the coordinating center have changed from a conventional coordinating center but remain substantial due to the need for collating, monitoring, verifying, and documenting the distributed data analysis (DDA) system. DDA is successful from the standpoint of implementation and operation--21 manuscripts representing work analyzed at six participating centers had been submitted for publication within 3.5 years of the completion of the baseline examination.

Adolescent

Methods of data analysis in the emergency medicine literature.

The authors hypothesized that data analysis in the current emergency medicine literature uses relatively few methods and sought to determine the frequency distributions of each method of analysis. The authors defined their population as original contributions in three refereed emergency medicine journals from September, 1985 through July, 1989. Letters to the editor, brief reports, reviews, and case reports were excluded. The authors reviewed 250 randomly selected articles and identified the method(s) of data analysis in each. The absolute frequency distribution of statistics were as follows: descriptive statistics only, 31%; contingency tables, 35% (chi 2, 28.4%; Fisher's exact test, 13.2%; McNemar's test, 0.4%); Student's t-test, 34%; ANOVA/ANCOVA, 12%; regression techniques, 8% (simple linear regression, 4.0%; multiple regression, 3.6%; logistic regression, 1.6%); nonparametric tests, 7% (Mann-Whitney, 2.8%; Wilcoxon, 2.4%; Dunnett, 0.8%; Kolmogorv-Smirnov, 0.4%; Kruskal-Wallis, 0.4%); multiple comparisons, 6% (Scheffé, 4.4%; Newman-Keuls, 2.0%); correlation techniques, 4% (Pearson product-moment correlation coefficient, 2.8%; Kendall's tau, 0.8%; Spearman's rho, 0.4%); confidence intervals, 2%. Correction techniques were used in 9% (Dunn-Bonnferoni, 4.8%; Yates correction, 4.4%). No statistics were found in 2% of the articles reviewed. Five statistical methods account for the vast majority (97% cumulative) of statistical uses in emergency medicine literature. This information should prove useful in deciding which tests should be emphasized in educating emergency physicians.

Curriculum

UALCAN Mobile, an app for cancer proteogenomic data analysis.

Cancer is a complex disease affecting various organs and is a major cause of death worldwide. During cancer initiation, disease progression, and tumor metastasis, various genomic and proteomic alterations are observed. Recent technological advances have led to the generation of large amounts of molecular data, including genomics and transcriptomics. These large-scale datasets can be utilized to analyze and identify sub-class-specific cancer biomarkers and targets. However, there is a need for the development of user-friendly tools for large-scale data analysis, disseminating the analyzed data in a visualizable format to cancer researchers with no programming skills. We developed UALCAN, a comprehensive platform that allows users to integrate disparate data to better understand the genes, proteins, and pathways perturbed in cancer and make discoveries of potential biomarkers and targets. In the current study, we describe the development of the UALCAN Mobile application (app) that will provide cancer transcriptomic data obtained from The Cancer Genome Atlas (TCGA) project to evaluate protein-coding gene expression based on various stratifications, including stage, grade, race, gender, and molecular-subtypes across over 30 types of cancers. In addition, the UALCAN mobile provides data analysis options for epigenetic changes due to DNA promoter methylation and Clinical Proteomic Tumor Analysis Consortium (CPTAC) cancer proteomic data. The app provides access to large cancer molecular datasets on the go. To find changes in the expression of causative genes and proteins and to identify biomarkers and therapeutic targets, UALCAN mobile app will be extremely valuable. The "UALCAN Mobile" app is free to use and can be downloaded from both the iOS/Apple and the Android Play Store and has been downloaded over 100 times in each of iOS and android app stores.

app

Research in physical medicine and rehabilitation. V. Data entry and early exploratory data analysis.

The process of data entry and initial analysis to locate data errors is described. Basic terms are defined and a simple method of entering data by using word processing software is illustrated. Data checking is done by using visual check of the raw data. Statistical programs are then used to locate possible data errors by finding data points (outliers) that are very different from the average. Special graphic output of statistical programs, scatterplots and box and whisker plots can be used to further locate questionable data. Examples of data entry forms and annotated step by step data cleaning with the use of inexpensive programs for personal computers are presented.

Computers

Uncovering psychiatric test information with graphical techniques of Exploratory Data Analysis.

This article illustrates how Exploratory Data Analysis (EDA) can complement conventional statistical methods in evaluating psychiatric tests. Using one recent EDA computer program, we evaluated the ability of repeated psychiatric screening tests (the General Health Questionnaire [GHQ]) to predict medical and psychiatric service use in a Health Maintenance Organization (HMO), the Harvard Community Health Plan (HCHP). Using a stratified random sample of 244 new HCHP enrollees and viewing three-dimensional graphs of their data from multiple perspectives, we found two subpopulations: low GHQ scorers, for whom the tests did not predict service use; and high scorers, for whom they did. Surprisingly, improving scores forecast increased use and chronically high scores predicted diminished use. Using another stratified random sample of 213 new HCHP enrollees, and with scatterplot matrices from another interactive computer program, we found that high and unchanging GHQ scores forecast HMO dropout. We examine possible interpretations--for example, that chronically distressed patients may become immobilized, diminish service use, and ultimately leave the HMO. We also explain how EDA methods may help uncover elusive results in other data (e.g., mental health outcomes).

Adult

Research in physical medicine and rehabilitation. VIII. Preliminary data analysis.

This paper describes important aspects of preliminary data analysis to be taken after data are checked for clerical entry errors and before the primary statistical analysis is performed. These include description and graphic display of each variable, recoding categorical data, transforming continuous data into another continuous variable and recoding continuous to categorical data. Missing values and outlying data points are identified and several techniques are recommended to minimize mistakes in variable recoding. Related variables measured with different units may be combined by using the z transformation and converted back to one of the original units for ease of interpretation. Finally, both categorical and continuous variables are checked for reliability by using kappa or the intraclass R.

Data Collection

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score&#xa0;=&#xa0;0.67-0.90) in genus diversity and showed a high correlation (rSpearman&#xa0;=&#xa0;0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics

Computer-assisted diagnosis by a model-free system of direct data analysis.

The basis of the method of data analysis presented is, in the case of any diagnostic test, the automatic compilation of separate frequency distributions for each diagnostic classification. The distinction of different test results for different diseases (the correlation for which the tests are used) can thus be quantitatively monitored. This offers opportunities for more specific control of the accuracy of the data base. Measurements of relative frequencies obtained from the frequency distributions of individuals with and without a given disease can serve as a quantitative handle for the selection of the combination of tests, and for adjustments of individual parameters, which will maximize the discrimination. The usual cutoffs are not used. A data-processing system can serve for the direct incorporation of patient chart data (including test results), and for the automation of the analysis described, with pattern recognition or cluster-seeking techniques. The ability of this system of analysis to minimize some of the problems associated with methods utilizing mathematical models is discussed.

Diagnosis, Computer-Assisted

Microcomputer-assisted univariate survival data analysis using Kaplan-Meier life table estimators.

We describe a microcomputer program (KMSURV) for exploratory univariate statistical analysis of survival data which is directly applicable to the evaluation of clinical trials and to retrospective epidemiological studies of hospital registry-based data. The program calculates life-table-like information based on Kaplan-Meier's product-limit estimators of the survivorship function S(t) and provides summary measures of average survival times. In addition, two non-parametric tests for the comparison of survival distributions are performed. A report-quality, high resolution plot of the S(t) estimates for all groups being compared complements each set of analyses. KMSURV is not a simple adaptation of a mainframe statistical analysis package and, thus, it utilizes efficiently the interactive environment which is inherent in microcomputing.

Actuarial Analysis

Toward a computer assisted analysis of NOESY spectra: a multivariate data analysis of an RNA NOESY spectrum.

A multivariate data-representation of a portion of the H-NOESY spectrum of an RNA octamer duplex was used to explore the possibility of using Principal Component Analysis and Partial Least Squares Discrimination for pattern recognition. In this case, it is found that the methods can: (i) distinguish slices containing signal from those containing only noise, (ii) locate slices containing overlapping signals, and (iii) in some cases to segregate slices with unique aspects such as those from terminal nucleotides, overlapping signals, purine-H8, pyrimidine-H6 and adenine-H2 containing slices. These properties can easily be included in a scheme to automate spectral analysis. The formulation described here does not distinguish patterns needed to automate sequential assignment of resonances in NOESY spectra of RNA.

Base Sequence

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software

Planning controlled clinical trials on the basis of descriptive data analysis.

In controlled clinical trials the problem of multiplicity of desired inferential statements finds attention at an increasing rate. In this paper the previously proposed concept of Descriptive Data Analysis (DDA), situated between Confirmatory and Exploratory Data Analysis, is applied to the planning aspects of controlled trials for which the problem of multiplicity exists. The (non-Bayesian) DDA planning concept should provide the investigator with tools to draw final conclusions from data of several variables possibly observed at several time points in possibly several groups of subjects by combining his pre-trial medical experience with descriptive inferential statements (confidence intervals and test results) at nominal significance levels. DDA also provides for confirmatory statements concerning individual null hypotheses and partially global hypotheses.

Clinical Trials as Topic

AWGE-ESPCA: An edge sparse PCA model based on adaptive noise elimination regularization and weighted gene network for Hermetia illucens genomic data analysis.

Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.

Animals