Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Structured exploratory data analysis (SEDA) for determining mode of inheritance of quantitative traits. I. Simulation studies on the effect of background distributions.

We examine through simulations the effectiveness of a new methodology to help distinguish among monogenic, multifactorial, and sporadic trait transmission from parents to offspring in nuclear family data sets. The major gene index (MGI), which compares the deviation of the offspring from the midparental value with a function of the individual deviations between parents and offspring, aids in the discrimination of multifactorial from sporadic and monogenic models. In contrast with other methodologies, the ability of the MGI to separate multifactorial, monogenic, and sporadic models improves with increased skewness in the trait distribution. The midparental correlation coefficient serves as a further guide for indicating mode of inheritance. A new class of techniques, the offspring between parents function (OBP), is introduced that provides a more sensitive tool to help in assessing mode of transmission through the analysis of the level, shape, and undulation characteristics of the curves. Four data examples are used to illustrate the methodology: erythrocyte catechol-O-methyltransferase (COMT) activity, height, weight, and triglyceride measurements. Height appears largely multifactorial, and weight appears to be mostly sporadic, while COMT and triglyceride measurements suggest the presence of some major gene influences.

Genetics↗

Error entropy in classification problems: a univariate data analysis.

Entropy-based cost functions are enjoying a growing attractiveness in unsupervised and supervised classification tasks. Better performances in terms both of error rate and speed of convergence have been reported. In this letter, we study the principle of error entropy minimization (EEM) from a theoretical point of view. We use Shannon's entropy and study univariate data splitting in two-class problems. In this setting, the error variable is a discrete random variable, leading to a not too complicated mathematical analysis of the error entropy. We start by showing that for uniformly distributed data, there is equivalence between the EEM split and the optimal classifier. In a more general setting, we prove the necessary conditions for this equivalence and show the existence of class configurations where the optimal classifier corresponds to maximum error entropy. The presented theoretical results provide practical guidelines that are illustrated with a set of experiments with both real and simulated data sets, where the effectiveness of EEM is compared with the usual mean square error minimization.

Classification↗

NetAffx Gene Ontology Mining Tool: a visual approach for microarray data analysis.

SUMMARY: The NetAffx Gene Ontology (GO) Mining Tool is a web-based, interactive tool that permits traversal of the GO graph in the context of microarray data. It accepts a list of Affymetrix probe sets and renders a GO graph as a heat map colored according to significance measurements. The rendered graph is interactive, with nodes linked to public web sites and to lists of the relevant probe sets. The GO Mining Tool provides visualization combining biological annotation with expression data, encompassing thousands of genes in one interactive view. AVAILABILITY: GO Mining Tool is freely available at http://www.affymetrix.com/analysis/query/go_analysis.affx

Abstracting and Indexing↗

Estimation of binding parameters by kinetic data analysis: differentiation between one and two binding sites.

A method that enables the discrimination between binding models and the estimation of binding parameters, based solely on kinetic data, is described. Experimental data from association and dissociation experiments were fitted simultaneously to models with mono- or biphasic kinetics with the aid of a non-linear maximum likelihood computer program. Discrimination between two models can be performed statistically. The protocol was used to study the binding of the antitussive [3H]noscapine to guinea pig brain homogenate. Two binding processes could be discriminated by their kinetics, despite the fact that [3H]noscapine apparently binds to one homogeneous population of binding sites in equilibrium binding experiments. This method might find general application when two populations of binding sites are suspected from kinetic data, but when selective ligands are lacking. Since parameter estimates are obtained independent of equilibrium binding data, our approach could also serve as an independent control of such experiments, with respect to both Kd and Bmax.

Animals↗

Using field data analysis for environmental decision making and subsequent remediation at two example sites.

One of the major challenges in remediating contaminated sites is having quick access to quality data on which to base remedial decisions as onsite work progresses. Case studies are presented at two Superfund sites where field screening and field analyses are used to provide these data. Emphasis is placed on the importance of high quality field data, as these data are the basis for remedial decisions prior to receipt of offsite laboratory confirmation. The decision-making processes for remediating contaminated soils and structures are presented in addition to project specifics including data quality objectives, field data collection procedures, quality assurance/quality control procedures, and comparisons of the field data with offsite laboratory results.

Data Collection↗

Psychology without p values. Data analysis at the turn of the 19th century.

Although the fledgling psychology of 100 years ago was assertively empirical, there were no inferential statistics to guide psychologists' data analyses. However, 19th-century developments had left psychology with a rich array of techniques for analyzing and presenting data, some of which remain underutilized today. These include comparisons across replications, within-subject designs, reanalysis of data, analyses of factorial designs, and especially the use of tables and graphs. As the merits of hypothesis-testing statistics are debated at the turn of the 21st century, the history of data-handling practices can remind psychologists that there are many ways to overcome the current uniformity of statistical practice.

Data Interpretation, Statistical↗

Accrual to the Breast Cancer Prevention Trial by participating Community Clinical Oncology Programs: a panel data analysis.

In 1992 patient accrual to the National Cancer Institute-sponsored Breast Cancer Prevention Trial was initiated in the United States and Canada. The Trial will involve 16,000 women who are evaluated to be at high risk of developing breast cancer. Nearly 250 health care organizations are participating in the Trial, including over 40 Community Clinical Oncology Program (CCOP) organizations, which are a component of the NCI's national clinical trials program. A previous NCI-funded evaluation conducted by the University of North Carolina showed the CCOP program to be an effective means of transferring the latest cancer technology, particularly cancer treatments, to the community setting. This paper describes a study designed to evaluate the performance of CCOP organizations in the Breast Cancer Prevention Trial. Using data from the first fifteen months of the Trial, the ability of CCOPs to accrue women is assessed using panel data estimation techniques. An attempt is made to predict accrual by structural, process, and environmental characteristics of participating CCOPs. Factors predictive of accrual include month in which accrual occurred and the extent of competition for trial participants in the CCOP service area. The hypothesized model explains slightly over 19 percent of the variation in accrual performance. The analysis demonstrates the utility of a panel data approach to modeling the dynamics of CCOP participation in a chemoprevention clinical trial.

Breast Neoplasms↗

Mortality data analysis using a multiple-cause approach.

Death certificates are the primary source for information used to define general mortality patterns in the United States. Analyses of mortality data generally are restricted to one of the conditions listed on the certificate--the underlying cause of dealth. We review principles related to the use of mortality data and describe a study using mortality tapes ("multiple-cause tapes") that list all conditions recorded on dealth certificates. Using multiple-cause tapes, we found that the number of deaths associated with seven infectious diseases in 1968, 1969, and 1970 was from 24% (diphtheria) to 81% (rubella) greater than that officially reported. Multiple-cause tapes also permitted a review of the association of deaths attributed to measles and varicella and known complications of these diseases. these observations confirm the usefulness of multiple-cause tapes in analyzing mortality data and emphasize the importance of examining all conditions listed on the death certificate.

Death Certificates↗

DATAC: a multipurpose biological data analysis program based on a mathematical interpreter.

The use of a mathematical command interpreter combined with the structural facility of the C-language allowed us to design a data treatment program having considerable flexibility and being able to handle any types of data (electrophysiological, biochemical and theoretical data). Ensembles of data are treated by the interpreter as if they were simple variables so that an elaborate computation can be performed on the spot by simply writing the appropriate equation on the terminal. These facilities combined with the ability of editing macrocommands at run time provide the user with data treatment possibilities that extend far beyond the possibilities actually implemented in the program. The originality of this program is that the user can easily implement the commands he most often needs, writing them in a language that most scientists will know, algebra.

Biometry↗

The Population Health Information System: data analysis and software.

This article describes the software developed in the process of creating the Population Health Information System. The software can be applied to a range of administrative data and provides standardized data on the health status and health care use of populations by generating population-based rates of discrete events. The standardized approach permits construction of a comprehensive, comparative picture for residents of defined geographic regions. The addition of a user friendly graphic interface will permit regional planners to do their own data analyses and allow out-of-province researchers to adopt the system for their own uses.

Community Health Planning↗

IR spectroscopy together with multivariate data analysis as a process analytical tool for in-line monitoring of crystallization process and solid-state analysis of crystalline product.

Crystalline product should exist in optimal polymorphic form. Robust and reliable method for polymorph characterization is of great importance. In this work, infra red (IR) spectroscopy is applied for monitoring of crystallization process in situ. The results show that attenuated total reflection Fourier transform infra red (ATR-FTIR) spectroscopy provides valuable information on process, which can be utilized for more controlled crystallization processes. Diffuse reflectance Fourier transform infra red (DRIFT-IR) is applied for polymorphic characterization of crystalline product using X-ray powder diffraction (XRPD) as a reference technique. In order to fully utilize DRIFT, the application of multivariate techniques are needed, e.g., multivariate statistical process control (MSPC), principal component analysis (PCA) and partial least squares (PLS). The results demonstrate that multivariate techniques provide the powerful tool for rapid evaluation of spectral data and also enable more reliable quantification of polymorphic composition of samples being mixtures of two or more polymorphs. This opens new perspectives for understanding crystallization processes and increases the level of safety within the manufacture of pharmaceutics.

Algorithms↗

Quality improvement data analysis of a mass casualty event.

Trauma auditing is important for monitoring the process of trauma care and outcome prediction. This pilot study was conducted to evaluate quality improvement (QI) data following a mass casualty event and discuss its impact on the trauma care process and outcome. A pre-designed trauma quality improvement data set was used for all 103 injured patients admitted to Asir Central Hospital, Saudi Arabia, who were involved in a single motor vehicle crash. Most of the trauma management variations from norms occurred during the initial assessment and resuscitation phase of care, and these had the greatest impact on morbidity and mortality. Trauma management variations throughout all phases of care were associated with 10% and 9% incidence of preventable morbidity and mortality, respectively. Efforts including rigorous educational programs should be made to stress the initial assessment and resuscitation phase of care. Successful regionalized trauma care systems involving quality improvement programs report significant reduction in morbidity and mortality rates from trauma.

Accidents, Traffic↗

[Genetic variability and interrelationship of Siberian and Far Eastern larches from RAPD-analysis data].

Genetic diversity of larches from six geographically isolated regions, Tomsk, Irkutsk, Ulan-Ude (Siberia), and Blagoveshchensk, Khabarovsk, Yuzhno-Sakhalinsk (Far East) was examined by means of RAPD analysis. Tree DNA samples were compared using 457 RAPD loci (97% of which were polymorphic), identified with 17 primers of random sequences. In the samples examined, 32 to 49% of the genes were in heterozygous state, mean expected heterozygosity (Hexp) varied from 0.1373 to 0.1891, and the genetic distances (DN) for different sample pairs varied from 0.0361 to 0.1802. The main population parameters were determined for Larix sibirica Ledeb., L. gmelinni (Rupr.) Rupr., and L. kamtschatica (Rupr.) Carr. Analysis of the genetic relationships showed that L. kamtschatica was characterized by highest genetic differentiation from the other larches examined, while larches from Primorskii krai were genetically close to L. sibirica.

DNA, Plant↗