Search PubMedSearch

SEARCH · Search PubMed

Results for “data analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

AWGE-ESPCA: An edge sparse PCA model based on adaptive noise elimination regularization and weighted gene network for Hermetia illucens genomic data analysis.

Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.

Animals

Database and search techniques for two-dimensional gel protein data: a comparison of paradigms for exploratory data analysis and prospects for biological modeling.

Two-dimensional (2-D) polyacrylamide gel electrophoresis can detect thousands of polypeptides, separating them by apparent molecular weight (Mr) and isoelectric point (pI). Thus it provides a more realistic and global view of cellular genetic expression than any other technique. This technique has been useful for finding sets of key proteins of biological significance. However, a typical experiment with more than a few gels often results in an unwiedly data management problem. In this paper, the GELLAB-II system is discussed with respect to how data reduction and exploratory data analysis can be aided by computer data management and statistical search techniques. By encoding the gel patterns in a "three-dimensional" (3-D) database, an exploratory data analysis can be carried out in an environment that might be called a "spread sheet for 2-D gel protein data". From such databases, complex parametric network models of protein expression during events such as differentiation might be constructed. For this, 2-D gel databases must be able to include data from other domains external to the gel itself. Because of the increasing complexity of such databases, new tools are required to help manage this complexity. Two such tools, object-oriented databases and expert-system rule-based analysis, are discussed in this context. Comparisons are made between GELLAB and other 2-D gel database analysis systems to illustrate some of the analysis paradigms common to these systems and where this technology may be heading.

Algorithms

Computer analysis of automated Edman degradation and amino acid analysis data.

Computer programs are described that allow facile analysis of data from a protein sequencer and amino acid analyzer. The sequencer program provides automated sequence interpretation while requiring minimal user interaction. The program serves as a powerful aid in deciphering mixture sequences and allows routine monitoring of sequencer performance. The computer program for amino acid analysis data provides the following calculations: mole percent, protein concentration and residues per mole with comparison between theoretical and calculated values. A plot of molecular weight versus deviation from integer values is calculated providing a measure of peptide or protein purity.

Amino Acids

Area normalization of the renal region of interest in radionuclide renography data analysis: a misconception.

Relative renal function is estimated by comparing the area under the second segment of the curve from the renal region of interest in a renographic study. We have examined the problems arising out of area normalization of the renal region of interest in the data analysis for relative renal function evaluation. Error analysis by computer simulation proves that this method of data analysis is highly misleading and erroneous.

Humans

Empirical considerations in orthopaedic research design and data analysis. Part II: The application of data analytic techniques.

To assure that a hypothesis is tested as rigorously as possible, the proper statistical method must be used to analyze the data. But without a strong background in statistics, it may be difficult to determine the efficacy of the data analytic technique used in the study. This paper describes several widely used data analytic techniques and offers examples of their proper application in orthopaedic research design.

Data Interpretation, Statistical

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i) perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii) introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans

Novel data analysis for synchronised spontaneous neuromagnetic activity.

A novel approach to neuromagnetic data analysis is presented. This technique is aimed at studying synchronised spontaneous activity (SSA) and has been used to resolve two different signals from one single evoked response, providing evidence for two possibly distinct sources. The data presented are consistent with a model that permits the generators of spontaneous activity to be synchronised by sensory stimuli.

Brain

A hierarchical, count-based model highlights challenges in scATAC-seq data analysis and points to opportunities to extract finer-resolution information.

BACKGROUND: Data from Single-cell Assay for Transposase Accessible Chromatin with Sequencing (scATAC-seq) is highly sparse. While current computational methods feature a range of transformation procedures to extract meaningful information, major challenges remain. RESULTS: Here, we discuss the major scATAC-seq data analysis challenges such as sequencing depth normalization and region-specific biases. We present a hierarchical count model that is motivated by the data generating process of scATAC-seq data. Our simulations show that current scATAC-seq data, while clearly containing physical single-cell resolution, are too sparse to infer true informational-level single-cell, single-region of chromatin accessibility states. CONCLUSIONS: While the broad utility of scATAC-seq at a cell type level is undeniable, describing it as fully resolving chromatin accessibility at single-cell resolution, particularly at individual locus level, may overstate the level of detail currently achievable. We conclude that chromatin accessibility profiling at true single-cell, single-region resolution is challenging with current data sensitivity, but that it may be achieved with promising developments in optimizing the efficiency of scATAC-seq assays.

Single-Cell Analysis

[Computer-assisted data analysis in a pediatric intensive care unit].

Computer assisted real time data analysis introduces a reasonable method of judgment into patient monitoring systems. From fast changing vital parameters discrete heart and respiration rate samples are immediately evaluated and presented as graphs near the bedside. Thus, statistical routines can increase the better understanding of instable clinical conditions and lend support to the decision making process. The early detection of a pathological trend in a patient whose ability to compensate is still present provides necessary time for diagnostic or preventive countermeasures in case of emergency.

Computers

Histopathological criteria for progressive dementia disorders: clinical-pathological correlation and classification by multivariate data analysis.

Autopsied brains from 55 patients with dementia between 59-95 years of age (mean age 77.9 +/- 8.1 years) and 19 non-demented individuals between 46-91 years of age (mean age 74.3 +/- 10.5 years) were examined to establish histopathological criteria for normal ageing, primary degenerative [Alzheimer's disease (AD)/senile dementia of Alzheimer type (SDAT)] and vascular (multi-infarct) dementia (MID) disorders. Senile/neuritic plaques, neurofibrillary tangles, microscopic infarcts and perivascular serum protein deposits were quantified in the frontal lobe (Brodmann area 10) and in the hippocampus. The demented patients were classified according to the DSM-III criteria into AD/SDAT and MID. Operationally defined histopathological criteria for dementias, based on the degree/amount of the histopathological changes seen in aged non-demented patients, were postulated. The demented patients were clearly separable into three histopathological types, namely AD/SDAT, MID and AD-MID, the dementia type where both the degenerative and the vascular changes are coexistent in greater extent than are seen in the non-demented individuals. Using general clinical, gross neuroanatomical and histopathological data three separate dementia classes, namely AD/SDAT, MID and AD-MID, were visualized in two-dimensional space by multivariate data analysis. This analysis revealed that the pathology in the AD-MID patients was not merely a linear combination of the pathology in AD/SDAT and MID, indicating that AD-MID might represent a dementia type of its own. The clinical diagnosis for AD/SDAT and MID was certain in only half of the AD/SDAT and one third of the MID cases when evaluated histopathologically and by multivariate data analysis. AD/SDAT, MID and AD-MID were histopathologically diagnosed in 49%, 24% and 27%, respectively, of all the dementia cases studied. Opposite correlation between the number of tangles, plaques and the patient age in non-demented and AD/SDAT cases were observed, indicating that the pathogenesis of tangles and plaques in the two groups of patients might be different and that AD/SDAT might not be a form of an exaggerated ageing process.

Aged

[The effect of smoking habit on aortic pulse wave velocity using a new method for data analysis].

We measured aortic pulse wave velocity (PWV) in 168 male adult cases of various arteriosclerotic diseases. In order to evaluate the effects of age, smoking habits, alcohol intake, and blood pressure, we applied the least median of squares (LMS) regression which was considered to be very useful for data analysis. The results showed that PWV level increased with age. Furthermore smoking was associated with increasing PWV level and this effect was also related to age. We concluded that the PWV was valuable as an index of arteriosclerosis, and instead of the classical least squares method, LMS regression was very useful for analysis of medical data.

Adult

Implementing a training resource for large-scale genomic data analysis in the All of Us Researcher Workbench.

A lack of representation in genomic research and limited access to computational training create barriers for many researchers seeking to analyze large-scale genetic datasets. The All of Us Research Program provides an unprecedented opportunity to address these gaps by offering genomic data from a broad range of participants, but its impact depends on equipping researchers with the necessary skills to use it effectively. The All of Us Biomedical Researcher (BR) Scholars Program at Baylor College of Medicine aims to break down these barriers by providing early-career researchers with hands-on training in computational genomics through the All of Us Evenings with Genetics Research Program. The year-long program begins with the faculty summit, an in-person computational boot camp that introduces scholars to foundational skills for using the All of Us dataset via a cloud-based research environment. The genomics tutorials focus on genome-wide association studies (GWASs), utilizing Jupyter Notebooks and the Hail computing framework to provide an accessible and scalable approach to large-scale data analysis. Scholars engage in hands-on exercises covering data preparation, quality control, association testing, and result interpretation. By the end of the summit, participants will have successfully conducted a GWAS, visualized key findings, and gained confidence in computational resource management. This initiative expands access to genomic research by equipping early-career researchers from a variety of backgrounds with the tools and knowledge to analyze All of Us data. By lowering barriers to entry and promoting the study of representative populations, the program fosters innovation in precision medicine and advances equity in genomic research.

Humans

Proficiency of the Tradescantia-micronucleus image analysis system for scoring micronucleus frequencies and data analysis.

The Tradescantia-micronucleus (Trad-MCN) bioassay is an efficient short-term test for genotoxicity of pollutants. In order to increase the efficiency and to standardize the micronucleus (MCN) scoring process, an automated scoring system was developed using the principle of image analysis in computer science. This assemblage is called the Tradescantia-micronucleus image analysis (Trad-MCNIA) system. The MCN frequencies scored by this system were compared with those scored by human observation for its proficiency. A set of low MCN frequency (around 5 MCN/100 tetrads) slides prepared from a control group, a set of medium MCN frequency (around 20 MCN/100 tetrads) slides prepared from sodium azide treated plant cuttings and a set of high MCN frequency (around 50 MCN/100 tetrads) slides prepared from X-ray treated materials were used for this study. In the low MCN frequency slides, the Trad-MCNIA system scored about the same value as human observation. In the medium and high frequency slides, MCN frequencies scored by the system were lower than those scored by human observers. This discrepancy was corrected by increasing the power of the objective of the microscope in the system. The MCN frequencies scored by the system attained 90% congruity with those scored by human observers after the correction. The scoring speed of the system was about 3.5 times as fast as that by human observers, and the data could be statistically analyzed immediately after the data scores were recorded. Further improvements can be made by upgrading the video camera and the computer speed.

Azides

Recruitment in NHLBI population-based studies and randomized clinical trials: data analysis and survey results.

Data on screening and recruitment from current and previous NHLBI population-based studies (PBSs) and randomized clinical trials (RCTs) were examined. In only two of the studies examined was the projected recruitment completed within the planned recruitment period. The shape of the graph of the relation between enrollment of participants and time varies by study. A single summary statistic for measuring the efficiency of recruitment in RCTs and PBSs is proposed and applied to the examined studies. In addition to providing summary data on recruitment for several studies, this article reports the survey results of a questionnaire sent to the coordinating centers of currently and previously funded National Heart, Lung, and Blood Institute and Veteran's Administration studies. The purpose was to ascertain the desirability of recommending that a generic core of information be collected on recruitment and screening in future studies. Most respondents believed that comparing data collected uniformly and prospectively might be helpful in designing further studies. The variables most respondents believed to be potentially useful are described.

Clinical Trials as Topic

A procedure for data analysis of the rodent micronucleus test involving a historical control.

No standard procedure of data analysis for rodent micronucleus tests involving historical controls has been established. In the present paper, under the presumption that the distribution of the historical control is stable and reliable, a procedure with three statistical steps is proposed to analyze the frequency of micronucleated polychromatic erythrocytes (MNPCEs). In the first step, the frequencies of MNPCEs in negative and positive control groups of a current experiment of the micronucleus test are compared with the distribution of historical negative and positive controls to examine the technical validity of the current experiment. In the second step, the frequency of MNPCEs in each treatment group is compared with the distribution of the historical negative control. In the third step, the dose-response relation is tested with the Cochran-Armitage trend test. A Monte Carlo stimulation study shows that the power of this procedure is acceptable and also this procedure is robust. An application of this procedure on real data reveals that it is effective in detecting clastogenic chemicals when the probability of a type I error is nearly .01.

Animals

Secondary data analysis: research method for the clinical nurse specialist.

This article presents a description of secondary data analysis and suggests that this type of research methodology may be helpful in facilitating research by the clinical nurse specialist (CNS). The article discusses the advantages and disadvantages of the use of this method specifically in relation to the CNS and offers suggestions for sources of data.

Data Collection

Comparison of Ehrlich ascites tumour and mouse liver cells by analytical subcellular fractionation combined with a sensitive computational method for data analysis.

A simple method of analytical subcellular fractionation, combined with a sensitive computational method for data analysis and presentation, has been used to reinvestigate the distribution and relative amounts of several enzymes in the cytoplasmic and plasma membranes of two different cell types: one is a neoplastic, transformed cell type (Ehrlich ascites tumour cells), the other an untransformed, highly differentiated cell type (liver hepatocytes plus Kupffer and endothelial cells). In general the distribution of the enzymes in particular membranes is similar in the two cell types, however the relative amounts differ. Ehrlich ascites tumour cells have a higher specific activity of galactosyltransferase and ouabain-sensitive (Na,K)ATPase, while liver cells have higher glucose-6-phosphatase, 5'-nucleotidase and succinate dehydrogenase activity. These differences appear to be correlated with morphological and, in some cases, functional differences between the two cell types.

5'-Nucleotidase