Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical power analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

The relationship between the sibling recurrence-risk ratio and genotype relative risk.

The recurrence-risk ratio of disease in siblings, lambdaS, is a standard parameter used in genetic analysis to estimate the statistical power for detection of a disease locus. However, the relationship between the underlying risk conferred by a disease-susceptibility allele and lambdaS has not been well described. The former is generally quantified as a genotype relative risk, gamma, and measures the ratio of disease risks between those with and those without the susceptibility genotype(s). We demonstrate that lambdaS varies significantly more with respect to gamma and the disease-allele frequency for two-locus multiplicative models than for other two-locus and for single-locus models. For the single- and two-locus dominant-inheritance models that we studied, when a disease-susceptibility allele had a frequency >/=.2, lambdaS had an upper limit of <10. In general, lambdaS values >10 are possible only under recessive inheritance, dominant inheritance with relatively rare (<5%) disease-susceptibility alleles, or when two or more disease loci have alleles acting either epistatically or multiplicatively. We introduce the idea of a restricted sib recurrence-risk ratio (lambda*S) estimated by restriction of sibships to those ascertained through a proband who already has a putative high-risk allele. A lambda*S larger than the lambdaS value estimated from randomly selected probands can serve as an indirect way of testing whether the posited susceptibility allele increases disease risk. Our results demonstrate that a lambdaS of 2-3 may portend successful mapping for a variety of genetic models but that, for some two-locus models, a lambdaS as high as 10 does not guarantee underlying genes easily mapped by linkage.

Alleles↗

Are variants in the CAPN10 gene related to risk of type 2 diabetes? A quantitative assessment of population and family-based association studies.

The calpain-10 gene (CAPN10) on chromosome 2q37.3 was the first candidate gene for type 2 diabetes (T2D) identified through a genomewide screen and positional cloning. One polymorphism (UCSNP-43: G-->A) and a specific haplotype combination defined by three polymorphisms (UCSNP-43, -19, and -63) were linked to an increased risk of T2D in several populations. To quantitatively assess the collective evidence for the effects of CAPN10 on risk of T2D, we conducted a meta-analysis of both population-based and family-based association studies. We retrieved data from the MEDLINE, PubMed, and Online Mendelian Inheritance in Man databases, as well as from other relevant reports and abstracts published up to July 2003. From a total of 26 studies with primary data (21 population-based studies: 5,013 cases and 5,876 controls; 5 family-based studies: 487 parent-offspring trios), we developed a summary database that contains variables of study design, study population/ethnicity, specific polymorphisms and haplotype combinations in CAPN10, and diabetes-related metabolic phenotypes. For population-based studies, we used both fixed-effects and random-effects models to calculate the pooled odds ratio (OR) and 95% confidence interval (CI) for the associations of CAPN10 genotypes with the risk of T2D. We also calculated weighted mean differences for the associations between CAPN10 and diabetes-related quantitative traits. Under either an additive or a dominant effect model, we found no statistically significant relation between CAPN10 genotypes in the UCSNP-43 locus and T2D risk. However, under a recessive model, individuals homozygous for the common G allele had a statistically significant 19% higher risk of T2D than carriers of the A allele (OR 1.19; 95% CI 1.07-1.33). The association between the 112/121 haplotype combination and T2D risk appeared to be overestimated by several initial small studies with positive findings (OR 1.38; 95% CI 1.04-1.84). After we removed these initial studies, this association became nonsignificant (OR 1.11; 95% CI 0.91-1.35). Moreover, we found no evidence for the associations between the UCSNP-43 G/G genotype and the 112/121 haplotype combination and metabolic phenotypes. Our meta-analysis of family-based studies showed only an overtransmission of the rare allele C in UCSNP-44 from heterozygous parents to their affected offspring with T2D. Our analysis indicates that inadequate statistical power, racial/ethnic differences in frequencies of alleles, haplotypes and haplotype combinations, potential gene-gene or gene-environment interactions, publication bias, and multiple hypothesis testing may contribute to the significant heterogeneity in previous studies of CAPN10 and T2D. Our findings also suggest that both large-scale, well-designed association studies and functional studies are warranted to either reliably confirm or conclusively refute the initial hypothesis regarding the role of CAPN10 in T2D risk.

Alleles↗

An efficient test for comparing sequence diversity between two populations.

We address the problem of comparing interindividual genomic sequence diversity between two populations. Although the methods are general, for concreteness we focus on comparing two human immunodeficiency virus (HIV) infected populations. From a viral isolate(s) taken from each individual in a sample of persons from each population, suppose one or multiple measurements are made on the genetic sequence of a coding region of HIV. Given a definition of genetic distance between sequences, the goal is to test if the distribution of interindividual distances differs between populations. If distances between all pairs of sequences within each group are used, then data-dependencies arising from the use of multiple sequences from individuals invalidates the use of a standard two-sample test such as the t-test. Where this problem has been recognized, a typical solution has been to apply a standard test to a reduced dataset comprised of one sequence or a consensus sequence from each patient. Disadvantages of this procedure are that the conclusion of the test depends on the choice of utilized sequences, often an arbitrary decision, and exclusion of replicate sequences from the analysis may needlessly sacrifice statistical power. We present a new test free of these drawbacks, which is based on a statistic that linearly combines all possible standard test statistics calculated from independent sequence subsamples. We describe statistical power advantages of the test and illustrate its use by application to nucleotide sequence distances measured from HIV-1 infected populations in southern Africa (GenBank accession numbers AF110959--AF110981) and North America/Europe. The test makes minimal assumptions, is maximally efficient and objective, and is broadly applicable.

Africa, Southern↗

Protein metabolism in children with edematous malnutrition and acute lower respiratory infection.

This study tested the hypothesis that wholebody protein kinetics remain low in children with edematous malnutrition and acute infection. Thirteen children with edematous malnutrition and acute infection (subjects) were compared with 14 uninfected children with edematous malnutrition early in recovery (control children). Protein kinetics were determined by using a primed, constant intravenous infusion of [13C]leucine and [15N2]urea in the postabsorptive state. Calculations of rates of whole-body protein synthesis and breakdown were based on the rate of leucine appearance; the rate of leucine oxidation was estimated from the rate of urea appearance. Protein synthesis and breakdown rates were lower in subjects than in control children (97 +/- 30 compared with 153 +/- 67, P < 0.01, and 103 +/- 30 compared with 160 +/- 67 mumol leucine.kg-1.h-1, P < 0.01). No difference was found between the two groups in the rate of urea appearance, but this analysis only had a statistical power of 54%. The absence of the expected increase in the rate of protein turnover during acute infection in edematous malnutrition implies that acute phase proteins are made with a corresponding depletion of muscle, hepatic, and other body proteins such as albumin, and that there may also be a blunting of the acute phase response.

Acute Disease↗

Planning research.

The starting point for any research project should be a question. Once this has been defined and the relevant scientific literature reviewed, a protocol should be drawn up. This will be used not only as a guide to the conduct of the study and in the preparation of the final report, but also in seeking any financial support and approvals that are required for the investigation. A protocol is normally arranged in sections covering the background to the study, the question(s) that it will address, the methods that will be used for the collection and analysis of data, the statistical power of the investigation (where relevant), any ethical considerations, and the financial input that will be needed. A pilot study is often helpful where aspects of the study method are untried or of uncertain validity.

Occupational Medicine↗

Recent advances in observer performance methodology: jackknife free-response ROC (JAFROC).

The jackknife free-response receiver operating characteristic (JAFROC) method allows quantitative analysis of observer data such as that observed when radiologists interpret images, which could contain more than one lesion and a location can be reported for each perceived lesion. The method was recently validated with a perception-based simulation model that incorporated the detectability parameter of the standard binormal ROC model, and in addition allowed simultaneous samples from both noise and signal distributions. The total number of noise samples is an important new parameter that measures reader expertise. The new sampling model incorporates search, which is an integral part of lesion detection that has not been possible to model until now. The model was used to generate simulated FROC ratings data, which was used to assess the statistical validity of JAFROC analysis. We found that JAFROC analysis is a statistically valid approach for analysing FROC data and that JAFROC analysis exhibited significantly greater statistical power than the existing ROC approach.

Algorithms↗

Does providing consumer health information affect self-reported medical utilization? Evidence from the Healthwise Communities Project.

OBJECTIVE: To determine whether providing health information to residents of Boise ID had an effect on their self-reported medical utilization. RESEARCH DESIGN: The Healthwise Communities Project (HCP) evaluation followed a quasi-experimental design. SUBJECTS: Random households in metropolitan zip codes were mailed questionnaires before and after the HCP. A total of 5,909 surveys were returned. MEASURES: The dependent variable was self-reported number of visits to the doctor in the past year. A difference-in-differences estimator was used to assess the intervention's community-level effect. We also assessed the intervention's effect on the variance of self-report utilization. RESULTS: Boise residents had a higher adjusted odds of entering care (OR = 1.27, 95% CI 0.88, 1.85) and 0.1 more doctor visits compared with residents in the control cities; however, for both outcomes, the effects were small and not significant. Although the means changed little, the data suggest that the variance of utilization in Boise decreased. CONCLUSIONS: The HCP had a small effect on overall self-reported utilization. Although the findings were not statistically significant, a posthoc power analysis revealed that the study was underpowered to detect effects of this magnitude. It may be possible to achieve larger effects by enrolling motivated people into a clinical trial. However, these data suggest that population-based efforts to provide health information have a small effect on self-reported utilization.

Adult↗

Bayesian spatial analysis and disease mapping: tools to enhance planning and implementation of a schistosomiasis control programme in Tanzania.

OBJECTIVE: To predict the spatial distributions of Schistosoma haematobium and S. mansoni infections to assist planning the implementation of mass distribution of praziquantel as part of an on-going national control programme in Tanzania. METHODS: Bayesian geostatistical models were developed using parasitological data from 143 schools. RESULTS: In the S. haematobium models, although land surface temperature and rainfall were significant predictors of prevalence, they became non-significant when spatial correlation was taken into account. In the S. mansoni models, distance to water bodies and annual minimum temperature were significant predictors, even when adjusting for spatial correlation. Spatial correlation occurred over greater distances for S. haematobium than for S. mansoni. Uncertainties in predictions were examined to identify areas requiring further data collection before programme implementation. CONCLUSION: Bayesian geostatistical analysis is a powerful and statistically robust tool for identifying high prevalence areas in a heterogeneous and imperfectly known environment.

Adolescent↗

Bioinformatic identification of novel early stress response genes in rodent models of lung injury.

Acute lung injury is a complex illness with a high mortality rate (>30%) and often requires the use of mechanical ventilatory support for respiratory failure. Mechanical ventilation can lead to clinical deterioration due to augmented lung injury in certain patients, suggesting the potential existence of genetic susceptibility to mechanical stretch (6, 48), the nature of which remains unclear. To identify genes affected by ventilator-induced lung injury (VILI), we utilized a bioinformatic-intense candidate gene approach and examined gene expression profiles from rodent VILI models (mouse and rat) using the oligonucleotide microarray platform. To increase statistical power of gene expression analysis, 2,769 mouse/rat orthologous genes identified on RG_U34A and MG_U74Av2 arrays were simultaneously analyzed by significance analysis of microarrays (SAM). This combined ortholog/SAM approach identified 41 up- and 7 downregulated VILI-related candidate genes, results validated by comparable expression levels obtained by either real-time or relative RT-PCR for 15 randomly selected genes. K-mean clustering of 48 VILI-related genes clustered several well-known VILI-associated genes (IL-6, plasminogen activator inhibitor type 1, CCL-2, cyclooxygenase-2) with a number of stress-related genes (Myc, Cyr61, Socs3). The only unannotated member of this cluster (n = 14) was RIKEN_1300002F13 EST, an ortholog of the stress-related Gene33/Mig-6 gene. The further evaluation of this candidate strongly suggested its involvement in development of VILI. We speculate that the ortholog-SAM approach is a useful, time- and resource-efficient tool for identification of candidate genes in a variety of complex disease models such as VILI.

Algorithms↗

Methodological guidelines for reading drug-evaluation research.

The present paper discusses methodological issues in psychopharmacological research. The intention is to provide readers of drug-evaluation research with a set of basic guidelines that will assist them to critically evaluate the investigations they encounter in the psychiatric literature. This paper describes the underlying rationale and basic principles associated with applying statistical analyses to drug-evaluation research data, and also addresses 11 additional methodological issues: specification of the research sample with respect to descriptive variables, diagnostic criteria, reliability of the diagnosis, the control group, random assignment of subjects to treatment conditions, blindness, subject attrition, treatment complications and side effects, the power of statistical tests, multivariate statistical analysis, and the reliability of the dependent variable. Evidence is presented to support the premise that there is a need to read drug-evaluation research critically.

Double-Blind Method↗

Antinuclear antibody profile in Italian patients with connective tissue diseases.

In the present work we report data on the specificity of antinuclear antibodies (ANA) in a large series of Italian patients suffering from a broad spectrum of connective tissue diseases (CTD), by using a series of homogeneous and validated techniques. The present study confirms, on the one hand, generally accepted concepts, i.e. that certain autoantibodies are strictly associated to certain disease states (such as anti-PCNA and anti-Sm in systemic lupus erythematosus, Jo 1 in polymyositis, and ACA and Scl-70 in scleroderma); the presence of 'marker' antibodies is, however, restricted to a relative minority of CTD patients. The application of a new methodological approach that considers the entire profile of ANA can greatly augment their diagnostic relevance and may provide useful indications for their interpretation, allowing us to establish for the first time the diagnostic usefulness not only of marker autoantibodies but also of certain associations between non-marker autoantibodies. Finally, the application of a more appropriate and powerful statistical tool (multiple correspondence analysis) has further emphasized the clear relationship existing between antibody specificities and certain disease states.

Adolescent↗

Controlled clinical trials in cancer research.

Knowledge of important aspects of the design and analysis of clinical trials is essential to clinical researchers and readers of medical literature. A brief description of proper trial design, including the contents of a trial protocol, as well as different strategies to avoid bias, is given. The concept of p-values is explained, and some commonly used statistical analysis methods are mentioned. Statistical power is defined, and two useful formulas and examples of estimating sample size are presented. The correct interpretation of trial results is emphasized, and misinterpretations and errors that frequently occur are dealt with. Various issues regarding multiple significance testing, such as interim analyses, multiple endpoints, and subgroup analyses, are addressed.

Controlled Clinical Trials as Topic↗

Toward a new approach in tumor cell heterogeneity studies using the concept of order.

A new methodology was developed to study dynamic processes topographically in biological systems by means of a graph-theoretical method. It is based upon order parameters obtained from a minimal spanning tree analysis coupled with computer simulations. The method was used to analyse the heterogeneous behavior of two neoplastic cell lines after treatment with laminin. The laminin-induced cell detachment was quantitated and shown to be inversely related to cell population density and thus to cellular interactions. Our statistical analysis is a very powerful tool to obtain information from seemingly disorderly heterogeneous biological models.

Animals↗

Improving permutation test power for group analysis of spatially filtered MEG data.

Non-parametric statistical methods, such as permutation, are flexible tools to analyze data when the population distribution is not known. With minimal assumptions and better statistical power compared to the parametric tests, permutation tests have recently been applied to the spatially filtered magnetoencephalography (MEG) data for group analysis. To perform permutation tests on neuroimaging data, an empirical maximal null distribution has to be found, which is free from any activated voxels, to determine the threshold to classify the voxels as active at a given probability level. An iterative procedure is used to determine the distribution by computing the null distribution, which is recomputed when a possible activated voxel is found within the current distributions. Besides the high computational costs associated with this approach, there is no guarantee that all activated voxels are excluded when constructing the maximal null distribution, which may reduce the statistical power. In this study, we propose a novel way to construct the maximal null distribution from the data of the resting period. The approach is tested on the MEG data from a somatosensory experiment, and demonstrated that the approach could improve the power of the permutation test while reducing the computational cost at the same time.

Adult↗

Incorporating prior information via shrinkage: a combined analysis of genome-wide location data and gene expression data.

Transcriptional control is a critical step in regulation of gene expression. Understanding such a control on a genomic level involves deciphering the mechanisms and structures of regulatory programmes and networks. A difficulty arises due to the weak signal and high noise in various sources of data while most current approaches are limited to analysis of a single source of data. A natural alternative is to improve statistical efficiency and power by a combined analysis of multiple sources of data. Here we propose a shrinkage method to combine genome-wide location data and gene expression data to detect the binding sites or target genes of a transcription factor. Specifically, a prior 'non-target' gene list is generated by analysing the expression data, and then this information is incorporated into the subsequent binding data analysis via a shrinkage method. There is a Bayesian justification for this shrinkage method. Both simulated and real data were used to evaluate the proposed method and compare it with analysing binding data alone. In simulation studies, the proposed method gives higher sensitivity and lower false discovery rate (FDR) in detecting the target genes. In real data example, the proposed method can reduce the estimated FDR and increase the power to detect the previously known target genes of a broad transcription regulator, leucine responsive regulatory protein (Lrp) in Escherichia coli. This method can also be used to incorporate other information, such as gene ontology (GO), to microarray data analysis to detect differentially expressed genes.

Bayes Theorem↗

Factors affecting statistical power in the detection of genetic association.

The mapping of disease genes to specific loci has received a great deal of attention in the last decade, and many advances in therapeutics have resulted. Here we review family-based and population-based methods for association analysis. We define the factors that determine statistical power and show how study design and analysis should be designed to maximize the probability of localizing disease genes.

Data Interpretation, Statistical↗

A comparison of bivariate and univariate QTL mapping in livestock populations.

This study presents a multivariate, variance component-based QTL mapping model implemented via restricted maximum likelihood (REML). The method was applied to investigate bivariate and univariate QTL mapping analyses, using simulated data. Specifically, we report results on the statistical power to detect a QTL and on the precision of parameter estimates using univariate and bivariate approaches. The model and methodology were also applied to study the effectiveness of partitioning the overall genetic correlation between two traits into a component due to many genes of small effect, and one due to the QTL. It is shown that when the QTL has a pleiotropic effect on two traits, a bivariate analysis leads to a higher statistical power of detecting the QTL and to a more precise estimate of the QTL's map position, in particular in the case when the QTL has a small effect on the trait. The increase in power is most marked in cases where the contributions of the QTL and of the polygenic components to the genetic correlation have opposite signs. The bivariate REML analysis can successfully partition the two components contributing to the genetic correlation between traits.

Animals↗

Genome scan meta-analysis for hypertension.

BACKGROUND: Genome scans for hypertension have yielded inconsistent results. The non-replication of significant or suggestive linkage might be due to lack of power of individual studies. Here, we conducted a genome scan meta-analysis for hypertension in an attempt to increase statistical power and to enhance evidence of linkage. METHODS: A newly developed Genome Search Meta-analysis (GSMA) method was applied to pool the results obtained from six scans reported in five papers. RESULTS: Our analysis did not find any regions with genome-wide significant linkage to hypertension. We did identify several regions with suggestive linkage, including 2p, 5q, 6q, 8p, 9p, 9q, and 11q. CONCLUSIONS: It seems that no region has a uniformly large impact on hypertension and that susceptibility genes for hypertension may be very difficult to detect.

Blood Pressure↗