Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

A method for quantification of absolute amounts of nucleic acids by (RT)-PCR and a new mathematical model for data analysis.

Accurate quantification of nucleic acids by competitive (RT)-PCR requires a valid internal standard, a reference for data normalization and an adequate mathematical model for data analysis. We report here an effective procedure for the generation of homologous RNA internal standards and a strategy for synthesizing and using a reference target RNA in quantification of absolute amounts of nucleic acids. Further, a new mathematical model describing the general kinetic features of competitive PCR was developed. The model extends the validity of quantitative competitive (RT)-PCR beyond the exponential phase. The new method eliminates the errors arising from different amplification efficiencies of the co-amplified sequences and from heteroduplex formation in the system. The high accuracy (relative error <2%) is comparable to the recently developed real time detection 5'-nuclease PCR. Also, corresponding computer software has been devised for practical data analysis.

Cell Line↗

Exploratory biochemical data analysis: a comparison of two sample means and diagnostic displays.

The occurrence of acne in women with hyperandrogenemia is well known; a question remains, however, as to whether a further positive relationship can be detected between the intensity of acne and the levels of testosterone, androgen precursors and sex hormone binding globulin (SHBG). A procedure of interactive data analysis extracting relevant information from original data was applied. Exploratory data analysis (EDA) identifies basic statistical features and patterns of data using a variety of diagnostic displays. The need for this step is particularly acute in biochemical and clinical data, the distribution of which is mostly non-Gaussian and often corrupted by the outliers. The omission of EDA can lead to incorrect results and false conclusions. In the EDA (i) several graphical tools for summarizing data are applied, (ii) the peculiarities of a sample distribution are investigated, (iii) a construction of distribution is carried out, (iv) a graphical comparison of the sample distribution with selected theoretical distributions is employed. The proposed procedure is illustrated by typical case study in the evaluation of differences between mean values of serum levels of testosterone, androgen precursors and SHBG in a group of patients with mild and severe forms of acne. A knowledge of the interval estimate of the mean value in both groups enables their comparison at the chosen probability level. As will be apparent from the evaluation of inter-group SHBG differences, an incorrect approach to the determination of group mean values could result in a complete misinterpretation of the data. The results indicate that androgens are not significantly related to the intensity of acne, and that SHBG is higher in patients with more severe forms of acne.

Acne Vulgaris↗

AWGE-ESPCA: An edge sparse PCA model based on adaptive noise elimination regularization and weighted gene network for Hermetia illucens genomic data analysis.

Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.

Animals↗

Database and search techniques for two-dimensional gel protein data: a comparison of paradigms for exploratory data analysis and prospects for biological modeling.

Two-dimensional (2-D) polyacrylamide gel electrophoresis can detect thousands of polypeptides, separating them by apparent molecular weight (Mr) and isoelectric point (pI). Thus it provides a more realistic and global view of cellular genetic expression than any other technique. This technique has been useful for finding sets of key proteins of biological significance. However, a typical experiment with more than a few gels often results in an unwiedly data management problem. In this paper, the GELLAB-II system is discussed with respect to how data reduction and exploratory data analysis can be aided by computer data management and statistical search techniques. By encoding the gel patterns in a "three-dimensional" (3-D) database, an exploratory data analysis can be carried out in an environment that might be called a "spread sheet for 2-D gel protein data". From such databases, complex parametric network models of protein expression during events such as differentiation might be constructed. For this, 2-D gel databases must be able to include data from other domains external to the gel itself. Because of the increasing complexity of such databases, new tools are required to help manage this complexity. Two such tools, object-oriented databases and expert-system rule-based analysis, are discussed in this context. Comparisons are made between GELLAB and other 2-D gel database analysis systems to illustrate some of the analysis paradigms common to these systems and where this technology may be heading.

Algorithms↗

Management of nursing homes using data envelopment analysis.

Data envelopment analysis (DEA) is used to evaluate the relative technical efficiency and assist in the management of a chain of nursing homes. As with any DEA model, variables chosen are particularly important. The study looks at two possibly critical issues. The first is the appropriateness of models that include only financial and economic measures to evaluate administrators when quality care is an expected output. The second issue is the appropriateness of using noncontrollable variables, in this case operating income, to evaluate administrators. We show how efficiency scores differ when quality variables and/or operating income are included. We also demonstrate the usefulness of DEA information to both the home administrator and chain managers for improving operating efficiency.

Costs and Cost Analysis↗

Cost-effectiveness and data envelopment analysis.

Data Envelopment Analysis (DEA) identifies price and technical inefficiencies among decision-making units. With controls for differences in case-mix and standardized outcomes, DEA's "best practice" frontier can be interpreted as a "cost-effectiveness" frontier. This study illustrates the key concepts, identifies the decisions required to use the technique for medical care decision making, and presents an application to a system of nine hospitals that offer obstetric services.

California↗

Spatial data analysis by epidermal Langerhans cells reveals an elegant system.

Langerhans cells are dendritic cells situated in the mammalian epidermis. In human epidermis, the concentration is between 460 and 1000 mm(-2). Langerhans cells fulfill an essential role in skin immune responses. Numerous scientific reports on Langerhans cells have appeared, but with no systematic research on the pattern of the spatial distributions. On the contrary, in certain fields, a spatial distribution is an important theme, and spatial data analysis has a long history. We hypothesized that epidermal Langerhans cells were set in the best formation for their immuno-surveillance by a sophisticated mechanism. To prove this hypothesis, we have imported spatial data analysis into the study of epidermal Langerhans cells. Here, we show that the distribution is completely regular; the pattern of Voronoi divisions fits the territories; the random packing model simulates their bone marrow derivation; a repulsive interaction is demonstrated and a repulsive potential function is estimated. Spatial data analysis-based computer simulation will be a new method of Langerhans cell study. In addition, this procedure shows promise for future distribution research of certain cells.

Animals↗

Gene expression data analysis.

Microarrays are one of the latest breakthroughs in experimental molecular biology, which allow monitoring of gene expression for tens of thousands of genes in parallel and are already producing huge amounts of valuable data. Analysis and handling of such data is becoming one of the major bottlenecks in the utilization of the technology. The raw microarray data are images, which have to be transformed into gene expression matrices--tables where rows represent genes, columns represent various samples such as tissues or experimental conditions, and numbers in each cell characterize the expression level of the particular gene in the particular sample. These matrices have to be analyzed further, if any knowledge about the underlying biological processes is to be extracted. In this paper we concentrate on discussing bioinformatics methods used for such analysis. We briefly discuss supervised and unsupervised data analysis and its applications, such as predicting gene function classes and cancer classification. Then we discuss how the gene expression matrix can be used to predict putative regulatory signals in the genome sequences. In conclusion we discuss some possible future directions.

Animals↗

Computer analysis of automated Edman degradation and amino acid analysis data.

Computer programs are described that allow facile analysis of data from a protein sequencer and amino acid analyzer. The sequencer program provides automated sequence interpretation while requiring minimal user interaction. The program serves as a powerful aid in deciphering mixture sequences and allows routine monitoring of sequencer performance. The computer program for amino acid analysis data provides the following calculations: mole percent, protein concentration and residues per mole with comparison between theoretical and calculated values. A plot of molecular weight versus deviation from integer values is calculated providing a measure of peptide or protein purity.

Amino Acids↗

Area normalization of the renal region of interest in radionuclide renography data analysis: a misconception.

Relative renal function is estimated by comparing the area under the second segment of the curve from the renal region of interest in a renographic study. We have examined the problems arising out of area normalization of the renal region of interest in the data analysis for relative renal function evaluation. Error analysis by computer simulation proves that this method of data analysis is highly misleading and erroneous.

Humans↗

Empirical considerations in orthopaedic research design and data analysis. Part II: The application of data analytic techniques.

To assure that a hypothesis is tested as rigorously as possible, the proper statistical method must be used to analyze the data. But without a strong background in statistics, it may be difficult to determine the efficacy of the data analytic technique used in the study. This paper describes several widely used data analytic techniques and offers examples of their proper application in orthopaedic research design.

Data Interpretation, Statistical↗

Application of the exploratory data analysis for evaluating the toxicity of chlorinated phenol derivatives by various cell models.

Exploratory data analysis based on multivariate statistical analysis techniques was introduced as a new approach to expressing the toxicity of chemical substances at the simultaneous acceptance of various cell models. Using principal component analysis and cluster analysis methods the toxicity of chlorinated phenol derivatives on employing some of the cell models (chlorococcal algae, cyanobacteria, bacteria, micromycetes, plant and animal cells) was characterized. The previous empirical experience that the toxicity of chlorinated phenol derivatives will increase with a growing degree of chlorination and that the presence of the methoxy group will cause a lowering of the toxic effect was demonstrated. The relationship between groups of tests used was presented.

Allium↗

A categorical data analysis of contacts with the Family Health Clinic, Calabar, Nigeria.

The relationships of population, environmental and accessibility variables to registration and attendance by mothers of children under 6 at the Family Health Clinic in Calabar, Nigeria are investigated. The technique used to analyze the data collected is categorical data analysis which proceeds in two stages, variable selection to reduce the variable set and fitting a log-linear model to the reduced set. Details of the statistical procedures used are provided to indicate how categorical data analysis can be used as a valuable tool of analysis in medical geographical studies that employ count or frequency data. It was found that younger mothers and Ibibio women registered more often at the clinic than did their counterparts. However, if the relatively sparse data on fathers is accepted, the association between age and registration is found to be spurious and a model can be substituted which shows younger fathers and fathers who spoke a non-Efik/Ibibio language to be associated with higher clinic registration of mothers. It was further found that for registered mothers the probability of a clinic visit was decreased by mother's age, increased by distance given no travel cost, unaffected by distance given some travel cost, increased by travel cost given a short distance to the clinic and decreased by travel cost given a longer distance from the clinic. These results are discussed in relation to population characteristics such as socio-economic status, clinic procedures such as health worker activities, transportation availability in Calabar, the spatial ecology of the city and local environmental conditions.

Adult↗

Data analysis in behavioral cerebral blood flow activation studies using xenon-133 clearance.

BACKGROUND AND PURPOSE: Three mainstream strategies exist to detect the responses of regional cerebral blood flow to functional activation. We tested the significance of changes in raw regional cerebral blood flow data, regional cerebral blood flow data normalized by division by global cerebral blood flow (dependent model of the regional-to-global cerebral blood flow relation), and regional cerebral blood flow data treating global cerebral blood flow as a covariate (independent model). Both latter models attempt to enhance regional sensitivity by removing global effects. We examined the sensitivity and pitfalls of these three strategies in behavioral activation studies. METHODS: These three strategies of data analysis were applied to changes in regional cerebral blood flow induced by a visuospatial problem-solving task in 38 healthy subjects as measured by the intravenous xenon-133 method with 32 stationary detectors. RESULTS: Mental activation increased blood flow in all regions of interest. Raw data were most sensitive and reliable to detect responses to mental stimulation. Both the independent and dependent models to remove global effects were less sensitive and falsely indicated deactivation in regions that were clearly stimulated. CONCLUSIONS: In behavioral activation paradigms, safe data analysis should be restricted to using raw regional cerebral blood flow increases without normalization or separation of global from regional effects. Studies using complex stimulation tasks should be scrutinized for global cerebral blood flow effects confounding regional responses.

Behavior↗

A visual data analysis system for the medical image processing.

We developed a visual data analysis system that can easily manage a large volume of medical imaging data. This system can analyze sets of imaging data using general image processing methods, so that various kinds of medical imaging data such as ECG charts, X ray image films, and MRI images, can be processed. The system has a graphical user interface (GUI). A physician who is novice at the system can manipulate the imaging data intuitively by pull down menus, pop up menus and buttons within the window system. The system can run on a standard UNIX workstation which is faster and more powerful than most personal computers. The system needs an X window system/Motif and C compiler. These are standard system programs already available on most UNIX workstations. The source code of the system can be retrieved from our anonymous ftp site via Internet.

Computer Graphics↗

Evaluation of replication studies, combined data analysis, and analytical methods in complex diseases.

Due to genetic heterogeneity, phenocopies, incomplete penetrance, misdiagnosis, and unknown mode of inheritance, linkage studies of most complex diseases are unlikely to provide conclusive findings with unambiguously high lod scores. Typically, several marginally significant lod scores or elevated lod scores are observed in a genome-wide screen. However, it is usually difficult to differentiate these findings from false positives (type I errors). Two approaches are commonly used to guard against false positives: replication studies in independent samples and combined data analysis. In the current paper, we evaluated these two common approaches using simulated data where data from multiple groups were available and locations of disease genes were known. We found replication studies and combined data analysis performed similarly in terms of their ability to identify true and false positive linkages. Both approaches confirmed two true linkages and did not confirm any false positive linkages. The results also indicated that it is not appropriate to apply the criteria proposed for confirming significant evidence for linkage to confirm regions with only suggestive evidence for linkage. The current results support previous findings that parametric analysis using an incorrect genetic model can still identify a true linkage.

Environment↗