Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

The "iron screen": modification of standard laboratory practice with data analysis.

Multivariate analysis was applied to iron deficiency anemia to generate an efficient sequence of diagnostic laboratory tests. A three step diagnostic system--serum ferritin level and mean corpuscular volume as a screen in all patients, followed by serum iron level and total iron binding capacity in some patients, and by erythrocyte sedimentation rate in a few patients--was constructed using a previously validated data reduction system. When compared to bone marrow iron stores, this system was found to have 96 per cent accuracy. In one year of clinical trial the "iron screen" classified 396 of 416 patients in a hospital setting. This sequential strategy shows how clinical laboratory data can be utilized to render diagnoses of defined probability.

Anemia, Hypochromic↗

Data analysis for detection and localization of multiple abnormalities with application to mammography.

RATIONALE AND OBJECTIVES: In assessing diagnostic accuracy it is often essential to determine the reader's ability both to detect and to correctly locate multiple abnormalities per patient. The authors developed a new approach for the detection and localization of multiple abnormalities and compared it with other approaches. MATERIALS AND METHODS: The new approach involves partitioning the image into multiple regions of interest (ROIs). The reader assigns a confidence score to each ROI. Statistical methods for clustered data are used to assess and compare reader accuracy. The authors applied this new method to a reader-performance study of conventional film images and digitized images used to detect and locate malignant breast cancer lesions. RESULTS: The ROI-based approach, the free-response receiver operating characteristic (FROC) curve, and the patient-based approach handle the estimation of the false-positive rate (FPR) quite differently. These differences affect the measures of the respective areas under the curves. In the ROI-based approach the denominator is the number of ROIs without a malignant lesion. In the FROC approach the average number of false-positive findings per patient is plotted on the x axis of the curve. In contrast, the patient-based approach mishandles the FPR by ignoring multiple detection and/or localization errors in the same patient. The FROC approach does not lend itself easily to statistical evaluations. CONCLUSION: The ROI-based approach appropriately captures both the detection and localization tasks. The interpretation of the ROI-based accuracy measures is simple and clinically relevant. There are statistical methods for estimating and comparing ROI-based estimates of accuracy.

Breast Neoplasms↗

A multivariate laboratory data analysis system: introduction.

In an era severely affected by the advanced stages of technocracy, it should not astound anyone that highly sophisticated technologies have metastasized throughout our hospital system. While simplifying many complex problems, the advantages of modern technology also create many interesting conflicts. One such dilemma surfaces as a consequence of clinical laboratories being able to produce a large number of test results in a relatively short time with a high degree of accuracy. Optimization of laboratory information must precede successful utilization of this extensive and expensive wealth of data.

Computers↗

Non-linearities and data analysis: towards a quantitative investigation of delayed luminescence.

Innovative techniques for the acquisition and analysis of delayed luminescence (DL) signals are proposed and discussed. At a preliminary level, the signals prove to be a useful tool, not only for quantitative analysis, but also for discrimination--on theoretical grounds--among different, possibly competing, mechanisms responsible for DL in photosynthetic organisms. Moreover, DL recordings from non-photosynthetic organisms (S. cerevisiae yeast) with avalanche photodiode (APD) detection will be discussed.

Data Interpretation, Statistical↗

Structure-function inferences based on molecular modeling, sequence-based methods and biological data analysis of snake venom lectins.

Lectins are a structurally and functionally diverse group of proteins from different sources, capable to recognize and bind specifically carbohydrates. Several snake venoms contain calcium-dependent true lectins (SVLs) that recognize galactose. Herein, in order to enlighten some of the structure-function relationships of snake venom lectins (SVLs), we constructed theoretical models for 10 SVLs based on the Crotalus atrox lectin (CaL), the only SVL crystal structure available, and compared with other animal and plant lectins, and C-type lectin-like proteins (CLPs) that do not bind carbohydrates. Although these are theoretical structures, we could identify some SVL features, including: (i) a singular intrachain disulfide bond (Cys(38)-Cys(133)) that is not present in CLPs; (ii) a significant reorientation (39-41A) of the 80's loop position that folds back to the globular domain, assists the carbohydrate recognition domain (CRD), and orients the dimer formation, even in BfL-1 and BfL-2, which did not present the Cys(86) interchain; (iii) a CRD presenting a negative and concave surface that allows the interaction with the specific saccharide hydroxyl groups and calcium ion; (iv) the role of water molecules in some interchain interactions, similar to other animal and plant lectins; and (v) the inability of forming oligomers in contrast to CaL and some CLPs, such as convulxin.

Amino Acid Sequence↗

A statistical approach to data analysis and 3-D geometric description of the human head and face.

Many analytical biomechanical methods require extensive three-dimensional statistical description of anatomical geometry. In particular, to design personal protective items for the human head and face, where good fit is critical, it is inevitable that a three-dimensional statistical description of this complicated structure will be needed. The work here offers an approach to this problem. This approach consists of three steps: (1) osteometric scaling, (2) normative specimen accumulation and (3) statistical testing. Three groups of facial data (24 Asian, 29 Black, and 29 White) were digitized. The effectiveness and accuracy of the statistical approach was tested on these three different experimental specimen sets. The method was found to be very accurate in modelling the most complicated human body parts--head and face. The availability of this detailed geometric information will also open many doors for future research and development of muscle controlled prostheses, repair of ligament damage, and in-vivo bone remodelling.

Asian People↗

A spline function approach for detecting differentially expressed genes in microarray data analysis.

MOTIVATION: A primary objective of microarray studies is to determine genes which are differentially expressed under various conditions. Parametric tests, such as two-sample t-tests, may be used to identify differentially expressed genes, but they require some assumptions that are not realistic for many practical problems. Non-parametric tests, such as empirical Bayes methods and mixture normal approaches, have been proposed, but the inferences are complicated and the tests may not have as much power as parametric models. RESULTS: We propose a weakly parametric method to model the distributions of summary statistics that are used to detect differentially expressed genes. Standard maximum likelihood methods can be employed to make inferences. For illustration purposes the proposed method is applied to the leukemia data (training part) discussed elsewhere. A simulation study is conducted to evaluate the performance of the proposed method.

Algorithms↗

An integrated proteome database for two-dimensional electrophoresis data analysis and laboratory information management system.

We describe an integrated proteome database, termed Yonsei Proteome Research Center Proteome Database (YPRC-PDB) which can store, retrieve and analyze various information including two-dimensional electrophoresis (2-DE) images and associated spot information that were obtained during studies of hepatocellular carcinoma (HCC). YPRC-PDB is also designed to perform as a laboratory information management system that manages sample information, clinical background, conditions of both sample preparation and 2-DE, and entire sets of experimental results. It also features query system and data-mining applications, which are amenable to automatically analyze expression level changes of a specific protein and directly link to clinical information. The user interface is web-based, so that the results from other laboratories can be shared effectively. In particular, the master gel image query is equipped with a graphic tool that can easily identify the relationship between the specific pathological stage of HCC and expression levels of a potential marker protein on the master gel image. Thus, YPRC-PDB is a versatile integrated database suitable for subsequent analyses. The information in YPRC-PDB is updated easily and it is available to authorized users on the World Wide Web (http://yprcpdb.proteomix.org/ approximately damduck/).

Carcinoma, Hepatocellular↗

PowerMV: a software environment for molecular viewing, descriptor generation, data analysis and hit evaluation.

Ideally, a team of biologists, medicinal chemists and information specialists will evaluate the hits from high throughput screening. In practice, it often falls to nonmedicinal chemists to make the initial evaluation of HTS hits. Chemical genetics and high content screening both rely on screening in cells or animals where the biological target may not be known. There is a need to place active compounds into a context to suggest potential biological mechanisms. Our idea is to build an operating environment to help the biologist make the initial evaluation of HTS data. To this end the operating environment provides viewing of compound structure files, computation of basic biologically relevant chemical properties and searching against biologically annotated chemical structure databases. The benefit is to help the nonmedicinal chemist, biologist and statistician put compounds into a potentially informative biological context. Although there are several similar public and private programs used in the pharmaceutical industry to help evaluate hits, these programs are often built for computational chemists. Our program is designed for use by biologists and statisticians.

Computer Simulation↗

The National Leprosy Control Programme of Zimbabwe a data analysis, 1983-1992.

Prevalence and detection rates of leprosy in Zimbabwe as well as patient characteristics were reported by the National Leprosy Control Programme over the 10-year period 1983-1992. The control programme made a new start in 1983 when multidrug therapy was introduced. Prevalence per 10,000 population declined steeply from 3.78 in 1983 to 0.52 in 1987. Prevalence continued to decline to 0.22 in 1992 and was highest in the north-eastern provinces. After an initial increase, the detection rate per 10,000 had declined from 0.19 in 1985 to 0.08 in 1992. The proportion of refugees among new cases had gradually increased since 1988 and amounted to one third in 1991 and 1992. An analysis of records of 802 cases who were newly detected from 1983 to 1992 showed that 51% were of the multibacillary (MB) type, 33% had visible disabilities at detection, 5% were under 15 years of age while the average delay time was 2.6 years. Patients with disabilities reported a longer delay time, were more often men and had more often the MB type of leprosy. The data suggest that transmission of leprosy is low but that cases are not diagnosed early enough to prevent transmission altogether.

Adolescent↗

Gene mapping in the 20th and 21st centuries: statistical methods, data analysis, and experimental design.

In the 20th century geneticists began to unravel some of the simpler aspects of the etiology of inherited diseases in humans. The theory of linkage analysis was developed and applied long before the advent of molecular biology, but only the technological advances of the second half of the 20th century made large-scale gene mapping with a dense genome-spanning set of markers a reality. More recently, the primary topic of interest has shifted from simple Mendelian diseases, for which genotypes of some gene are the cause of disease, to more complex diseases, for which genotypes of some set of genes together with environmental factors merely alter the probability that an individual gets the disease, although individual factors are typically insufficient to cause the disease outright. To this end, a great deal of dogma has evolved about the best way to skin this cat, although to date success has been minimal with any approach. We postulate that the main reason for this is a lack of attention to experimental design. Once the data have been ascertained, the most powerful statistical methods will not be able to salvage an inappropriately designed study (Andersen 1990). Each phenotype and/or population mandates its own individually tailored study design to maximize the chances of successful gene mapping. We suggest that careful consideration of the available data from real genotype-phenotype correlation studies (as opposed to oversimplified theoretically tractable models), and the practical feasibility of different ascertainment schemes dictate how one should proceed. In this review we review the theory and practice of gene mapping at the close of the 20th century, showing that most methods of linkage and linkage disequilibrium analysis are similar in a fundamental sense, with the differences being related more to study design and ascertainment than to technical details of the underlying statistical analysis. To this end, we propose a new focus in the field of statistical genetics that more explicitly highlights the primacy of study design as the means to increase power for gene mapping.

Algorithms↗

Fluorescence data analysis on gel-based biochips.

A series of biochip readers developed for gel-based biochips includes three imaging models and a novel nonimaging biochip scanner. The imaging readers, ranging from a research-grade versatile reader to a simple portable one, use wide-field objectives and 12-bit digital large-coupled device cameras for parallel addressing of multiple array elements. This feature is valuable for monitoring the kinetics of sample interaction with immobilized probes. Depending on the model and the label used, the sensitivity of these readers approaches 0.3 amol of a labeled sample per gel element. In the selective scanner, both the spot size of the excitation laser beam and the detector field of view match the size of the biochip array elements so that the whole row of the array can be read in a single scan. The portable version reads 50-mm long, 150-element, one-dimensional arrays in 5 s. With a dynamic range of 4000:1, a sensitivity of 1-5 amol of a labeled sample per gel element, and a data format facilitating online processing, the scanner is an attractive, inexpensive solution for biomedical diagnostics. Fluorophores for sample labeling were compared experimentally in terms of detection sensitivity, influence on duplex stability, and suitability for multilabel analysis and thermodynamic studies. Texas Red and tetracarboxyphenylporphyn proved to be the best choice for two-wavelength analysis using the imaging readers.

Fluorescent Dyes↗

Iterative partial least squares with right-censored data analysis: a comparison to other dimension reduction techniques.

In the linear model with right-censored responses and many potential explanatory variables, regression parameter estimates may be unstable or, when the covariates outnumber the uncensored observations, not estimable. We propose an iterative algorithm for partial least squares, based on the Buckley-James estimating equation, to estimate the covariate effect and predict the response for a future subject with a given set of covariates. We use a leave-two-out cross-validation method for empirically selecting the number of components in the partial least-squares fit that approximately minimizes the error in estimating the covariate effect of a future observation. Simulation studies compare the methods discussed here with other dimension reduction techniques. Data from the AIDS Clinical Trials Group protocol 333 are used to motivate the methodology.

Anti-HIV Agents↗

Mutational spectra in transgenic animal research: data analysis and study design based upon the mutant or mutation frequency.

Understanding chemically induced changes in mutational spectra can aid in deciphering mechanisms of mutagenesis. In this paper, we propose the use of statistical methods that are based upon the mutation frequency, rather than simple mutant counts which have no relationship to the mutation frequency. These methods have a number of advantages over the current standard analysis: an improved means of identifying those classes/sites of mutation which have treatment-related induction, greater sensitivity to localized differences in spectra (e.g., limited to a single base pair), one-sided tests for induction of mutations, tests of dose-response, and a framework for sample-size estimation in terms of the number of mutants to sequence. As examples, the methods are applied to data from transgenic mutation assays.

Animals↗

Robust algorithm for alignment of liquid chromatography-mass spectrometry analyses in an accurate mass and time tag data analysis pipeline.

Liquid chromatography coupled to mass spectrometry (LC-MS) and tandem mass spectrometry (LC-MS/MS) has become a standard technique for analyzing complex peptide mixtures to determine composition and relative abundance. Several high-throughput proteomics techniques attempt to combine complementary results from multiple LC-MS and LC-MS/MS analyses to provide more comprehensive and accurate results. To effectively collate and use results from these techniques, variations in mass and elution time measurements between related analyses need to be corrected using algorithms designed to align the various types of data: LC-MS/MS versus LC-MS/MS, LC-MS versus LC-MS/MS, and LC-MS versus LC-MS. Described herein are new algorithms referred to collectively as liquid chromatography-based mass spectrometric warping and alignment of retention times of peptides (LCMSWARP), which use a dynamic elution time warping approach similar to traditional algorithms that correct for variations in LC elution times using piecewise linear functions. LCMSWARP is compared to the equivalent approach based upon linear transformation of elution times. LCMSWARP additionally corrects for temporal drift in mass measurement accuracies. We also describe the alignment of LC-MS results and demonstrate their application to the alignment of analyses from different chromatographic systems, showing the suitability of the present approach for more complex transformations.

Algorithms↗

Estimation of Ki in a competitive enzyme-inhibition model: comparisons among three methods of data analysis.

There are a variety of methods available to calculate the inhibition constant (Ki) that characterizes substrate inhibition by a competitive inhibitor. Linearized versions of the Michaelis-Menten equation (e.g., Lineweaver-Burk, Dixon, etc.) are frequently used, but they often produce substantial errors in parameter estimation. This study was conducted to compare three methods of analysis for the estimation of Ki: simultaneous nonlinear regression (SNLR); nonsimultaneous, nonlinear regression, "KM,app" method; and the Dixon method. Metabolite formation rates were simulated for a competitive inhibition model with random error (corresponding to 10% coefficient of variation). These rates were generated for a control (i.e., no inhibitor) and five inhibitor concentrations with six substrate concentrations per inhibitor and control. The KM/Ki ratios ranged from less than 0.1 to greater than 600. A total of 3 data sets for each of three KM/Ki ratios were generated (i.e., 108 rates/data set per KM/Ki ratio). The mean inhibition and control data were fit simultaneously (SNLR method) using the full competitive enzyme-inhibition equation. In the KM,app method, the mean inhibition and control data were fit separately to the Michaelis-Menten equation. The SNLR approach was the most robust, fastest, and easiest to implement. The KM,app method gave good estimates of Ki but was more time consuming. Both methods gave good recoveries of KM and VMAX values. The Dixon method gave widely ranging and inaccurate estimates of Ki. For reliable estimation of Ki values, the SNLR method is preferred.

Algorithms↗

Sulfide as a confounding factor in toxicity tests with the sea urchin Paracentrotus lividus: comparisons with chemical analysis data.

Sperm cell and embryo toxicity tests with the sea urchin Paracentrotus lividus were performed to assess the toxicity of sulfide, which is considered a confounding factor in toxicity tests. For improved information on the sensitivity of these methods to sulfide, experiments were performed in the same aerobic conditions used for testing environmental samples, with sulfide concentrations being monitored at the same time by cathodic stripping voltammetry. New toxicity data for sulfide expressed as median effective concentration (EC50) and no-observed-effect concentration (NOEC) are reported. The EC50 value for the embryo toxicity test (total sulfide at 0.43 mg/L) was three times lower than for the sperm cell test (total sulfide at 1.20 mg/L), and the NOEC values were similar (on the order of total sulfide at 10(-1) mg/L) for both tests. The decrease in sulfide concentration during the bioassay as a consequence of possible oxidation of sulfide by dissolved oxygen was determined by voltammetric analysis, indicating a half-life of about 50 min in the presence of gametes. To check the influence of sulfide concentrations on toxicity effects in real samples, toxicity (with the sperm cell toxicity test) and chemical analyses also were performed in pore-water samples collected with an in situ sampler in sediments of the Lagoon of Venice (Italy). A highly positive correlation between increased acute toxicity and increased sulfide concentration was found. Examination of data revealed that sulfide is a real confounding factor in toxicity testing in anoxic environmental samples containing concentrations above the sensitivity limit of the method.

Animals↗

EQUIL: simulation and data analysis of binding reactions with arbitrary chemical models.

We have developed an algorithm for simulation and analysis of arbitrary chemical systems in equilibrium, with emphasis on ligand binding reactions. The program EQUIL can treat reactions involving multiple ligands, multiple binding sites, ternary complex models, allosteric effectors, competitive and noncompetitive binding, conformational changes, cooperativity, and generally any scheme that can be represented as a set of chemical equations. EQUIL is based on a general thermodynamic model of chemical equilibria; it does not involve nonlinear transformation of experimental data, but it does require the user to define the model of interaction between ligands and receptors by writing down the appropriate chemical reactions. EQUIL contains features of particular importance to ligand binding experiments: variable binding capacities, nonspecific binding, and the ability to simultaneously analyze data from different types of experiments. Furthermore, the simulation feature of EQUIL allows the user to investigate the feasibility of experiments that could possibly distinguish between different reaction models. We illustrate the use of this program on personal computers to analyze and simulate simple and complicated interactions between ligands and receptors.

Algorithms↗