Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Penalized Cumulative Probability Model for a Continuous Outcome Subject to Detection Limits.

Mixed-type outcome data occur when the outcome variable's distribution is a mixture of both continuous and discrete ordinal variables. Such mixed-type outcomes are common in biomedical, psychological, and the health sciences, particularly for variables having either a detection or quantitation limit. When interest lies in identifying a combination of genomic features associated with a mixed-type outcome, any method used would require a variable selection strategy for high-dimensional data. Unfortunately, few variable selection methods exist for modeling a mixed-type outcome when the covariate space is high dimensional. This study develops a high-dimensional penalized cumulative probability model (CPM), to allow for the identification of genomic features associated with mixed-type outcome of interest. We demonstrated how such model may be estimated using the iterative penalization procedure-the generalized monotone incremental forward stagewise (GMIFS) algorithm. The Model-X knockoffs procedure was combined with the estimation algorithm to control the false discovery rates (FDR) when performing variable selection. Through extensive simulation studies, our penalized CPM was shown to outperform alternative methods in terms of controlled variable selection performance by achieving high statistical power with the FDR being controlled at the target level. We demonstrate the utility of our method by applying it to predict estimated glomeruli filtration rate (eGFR) in kidney transplant recipients at 24 months post-transplant using baseline gene expression data as predictors. Our CPM model identified five genes associated with this mixed-type outcome which have important links to renal disease, which may provide prognostic guidance for kidney transplantation recipients.

Models, Statistical↗

Usefulness of molecular markers for detecting population bottlenecks via monitoring genetic change.

It is important to detect population bottlenecks in threatened and managed species because bottlenecks can increase the risk of population extinction. Early detection is critical and can be facilitated by statistically powerful monitoring programs for detecting bottleneck-induced genetic change. We used Monte Carlo computer simulations to evaluate the power of the following tests for detecting genetic changes caused by a severe reduction in a population's effective size (Ne): a test for loss of heterozygosity, two tests for loss of alleles, two tests for change in the distribution of allele frequencies, and a test for small Ne based on variance in allele frequencies (the 'variance test'). The variance test was most powerful; it provided an 85% probability of detecting a bottleneck of size Ne = 10 when monitoring five microsatellite loci and sampling 30 individuals both before and one generation after the bottleneck. The variance test was almost 10-times more powerful than a commonly used test for loss of heterozygosity, and it allowed for detection of bottlenecks before 5% of a population's heterozygosity had been lost. The second most powerful tests were generally the tests for loss of alleles. However, these tests had reduced power for detecting genetic bottlenecks caused by skewed sex ratios. We provide guidelines for the number of loci and individuals needed to achieve high-power tests when monitoring via the variance test. We also illustrate how the variance test performs when monitoring loci that have widely different allele frequency distributions as observed in five wild populations of mountain sheep (Ovis canadensis).

Alleles↗

Genetic association studies in mood disorders: issues and promise.

Genetic association is a powerful method for identifying genetic variants that contribute to the molecular basis of complex diseases. There is now a wealth of informative, validated and densely-spaced single nucleotide polymorphism (SNP) markers for use in association studies, and the delineation of the genome-wide haplotype architecture will greatly enhance our ability to conduct whole genome association screens, fine mapping of linkage regions, and systematic screening of functional candidate genes. Single nucleotide polymorphism-based genotyping technology has progressed dramatically to the point of high-throughput methods that can assay up to thousands of SNPs on many samples in one experiment. Genotyping cost remains a limiting factor in complex disease studies, where numerous SNPs and large sample sets are needed to maximize statistical power. Strategies designed to reduce cost include DNA pooling and analysis with tagSNPs. As larger clinical samples become available, it will be increasingly important to test for hidden stratification in case-control studies, as well as transmission distortion in family-based studies, either of which can lead to spurious association findings. As yet, there is no widely-accepted genetic association finding in mood disorders, but functional candidate genes, such as the serotonin transporter, and positional candidates, such as G72/G30 on chromosome 13q, are beginning to be identified in several studies. Relating associated variants to the phenotype represents the next critical step toward establishing the pathogenic role of gene variants in mood disorders.

Gene Pool↗

[Possibilities and limitations of the predictive risk estimates and epidemiological studies following the Chernobyl incident].

The disastrous accident at the nuclear power station at the Chernobyl on 1986 (April 26) has brought attention to the estimation of radiation health effects and many "experts" were attending to the evaluation on oncogenic mortality increase among the Italian population in the next future. On the contrary at that time too few peoples were worried about the possibility of detecting such an increase. Discussion of this topic is notoriously fraught with difficulties arising from differences of opinion how to estimate low-dose risk in humans without data from direct observation. One opinion is to extrapolate from the data points obtained at relatively high doses toward zero dose (zero extrapolation theory). This permit estimates of risk to be made but, in the final analysis, no data from humans exist that show that low-level radiation exposures produce measurable biologic effects. For that this theory is more useful in radio-protection and medico-legal subjects. It is easy on a statistical basis to prove the impossibility to establish an increase in human cancer after low doses of ionizing radiation such as those received environmentally after the Chernobyl's accident. In this condition to observe the numbers of radiation-induced cancer deaths that far exceed the "natural" incidence would require a follow-up a sample more and more greater than the italian population herself. Indeed the statistical power of a hypothetical follow-up study at a suitable confidence level would require a sample size higher than a milliard of persons for the detection of an increase of a generic cancer mortality and higher then seven hundred of millions for the detection of an increase of the specific thyroid cancer mortality. In more detail the following figures for the parameters needed to curring out the evaluation have been used: medium dose equivalent to the thyroid, 2.03 mSv; medium effective dose equivalent up to december '87, 0.6 mSv; thyroid cancer mortality in the italian population, 0.94 10(-5) y-1; total cancer mortality in the italian population, 22.2 10(-2) y-1; risk factor per unit dose equivalent in thyroid, 0.5 10(-6) mSv-1; risk factor per unit effective dose equivalent, 2.0 10(-5) mSv-1. Applying the foregoing values in statistical inference methods it could be achieved that 7.5 18(8) and 1.25 10(9) persons must be followed-up in the next 30 years to detect a significant increase over the "natural" cancer mortality for thyroid and "total body" radioinduced cancers respectively.(ABSTRACT TRUNCATED AT 400 WORDS)

Accidents↗

Susceptibility-induced loss of signal: comparing PET and fMRI on a semantic task.

Functional magnetic resonance imaging (fMRI) has become a popular tool for investigations into the neural correlates of cognitive activity. One limitation of fMRI, however, is that it has difficulty imaging regions near tissue interfaces due to distortions from macroscopic susceptibility effects which become more severe at higher magnetic field strengths. This difficulty can be particularly problematic for language tasks that engage regions of the temporal lobes near the air-filled sinuses. This paper investigates susceptibility-induced signal loss in the temporal lobes and proposes that by defining a priori regions of interest and using the small-volume statistical correction of K. J. Worsley, S. Marrett, P. Neelin, A. C. Vandal, K. J. Friston, and A. C. Evans (1996, Hum. Brain Mapp. 4: 58-83), activations in these areas can sometimes be detected by increasing the statistical power of the analysis. We conducted two experiments, one with PET and the other with fMRI, using almost identical semantic categorization paradigms and comparable methods of analysis. There were areas of overlap as well as differences between the PET and fMRI results. One anticipated difference was a lack of activation in two regions in the temporal lobe on initial analyses in the fMRI data set. With a specific region of interest, however, activation in one of the regions was detected. These experiments demonstrate three points: first, even for almost identical cognitive tasks such as those in this study, PET and fMRI may not produce identical results; second, differences between the two methods due to macroscopic susceptibility artifacts in fMRI can be overcome with appropriate statistical corrections, but only partially; and third, new data acquisition paradigms are necessary to fully deal with susceptibility-induced signal loss if the sensitivity of the fMRI experiment to temporal lobe activations is to be enhanced.

Adult↗

Risk groups in patients with bladder cancer treated with radical cystectomy: statistical and clinical model improving homogeneity.

PURPOSE: In this study we identified homogeneous risk groups, with no survival overlap among the subgroups that make up each risk group, in patients with transitional cell carcinoma of the bladder treated with radical cystectomy alone. MATERIALS AND METHODS: Predictive factors for tumor death were analyzed with univariate and multivariate analysis among a group of 298 patients with transitional cell carcinoma of the bladder treated with radical cystectomy alone. Independent variables were progressively incorporated according to their statistical power in a stepwise process identifying a model with independent subgroups. The risk groups were identified according to different survival cutoff points including subgroups with similar survival. To search a clinical application and to check the strength of this model a new model was also set up using the weight score based on the size of hazard ratio from multivariate analysis. RESULTS: Univariate analysis demonstrated that lymphatic invasion status, pathological stage (P), lymph node status (N) and prostatic stroma status (St) were predictive variables for tumor death, and the latter 3 were independent variables in the multivariate analysis. By taking the most powerful, N, as the reference variable, and progressively incorporating additional variables, a model was found including 7 independent subgroups. In this model only 2 subgroups, N1 and N2-3, included more than 1 category and their survival was also calculated. Three risk groups were identified establishing different survival cutoffs. The 5-year cancer specific survival rate was 86.4% for low risk (P1-2N0St-), 64.4% (range 60.9% to 65.3%) for intermediate risk (P1-2N1St-, P3N0St-, HR = 2.7) and 28.1% (range 0% to 47.7%) for high risk (N2-3, P4, St+, N1P3, HR = 8.7). This model was also reproduced using the weight score based on the size of the hazard ratio from the multivariate analysis CONCLUSIONS: Three homogeneous risk groups were identified with high statistically significant survival differences among them and no survival overlap among subgroups that make up the risk groups.

Adult↗

Attention-deficit/hyperactivity disorder endophenotypes.

Attention-deficit/hyperactivity disorder (ADHD) is a highly heritable disorder with a multifactorial pattern of inheritance. For complex conditions such as this, biologically based phenotypes that lie in the pathway from genes to behavior may provide a more powerful target for molecular genetic studies than the disorder as a whole. Although their use in ADHD is relatively new, such "endophenotypes" have aided the clarification of the etiology and pathophysiology of several other conditions in medicine and psychiatry. In this article, we review existing data on potential endophenotypes for ADHD, emphasizing neuropsychological deficits because assessment tools are cost effective and relatively easy to implement. Neuropsychological impairments, as well as measures from neuroimaging and electrophysiological paradigms, show correlations with ADHD and evidence of heritability, but the familial or genetic overlap between these constructs and ADHD remains unclear. We conclude that these endophenotypes will not be a quick fix for the field but offer potential if careful consideration is given to issues of heterogeneity, measurement and statistical power.

Adoption↗

The interpretation of occupational epidemiologic data in regulation and litigation: studies of auto mechanics and petroleum workers.

Epidemiologic data often serve as scientific basis for policies in regulation and opinions in litigation. The interpretation of epidemiologic data in both regulation and litigation is often challenged and debated. In this commentary, a wide range of issues concerning the interpretation of epidemiologic data in regulation and litigation are discussed. These issues include: case reports, study design, specificity of exposure, interview or recall bias, misclassification of occupation or exposure, confounding multiple exposures, confidence intervals and statistical power, selection of relevant studies, consistency of study results, study cohort definition, cohort membership misclassification, dilution effect, and subcohort or stratified analysis. Epidemiologic studies of auto mechanics and petroleum workers are used as examples to illustrate the importance of relying on sound epidemiologic principles in study interpretation. If these principles are not followed, the interpretation of epidemiologic studies will likely be erroneous and not useful to regulatory policy-makers or to those involved in litigation.

Automobiles↗

Using imprecise probabilities to address the questions of inference and decision in randomized clinical trials.

Randomized controlled clinical trials play an important role in the development of new medical therapies. There is, however, an ethical issue surrounding the use of randomized treatment allocation when the patient is suffering from a life threatening condition and requires immediate treatment. Such patients can only benefit from the treatment they actually receive and not from the alternative therapy, even if it ultimately proves to be superior. We discuss a novel new way to analyse data from such clinical trials based on the use of the recently developed theory of imprecise probabilities. This work draws an explicit distinction between the related but nevertheless distinct questions of inference and decision in clinical trials. The traditional question of scientific interest asks 'Which treatment offers the greater chance of success?' and is the primary reason for conducting the clinical trial. The question of decision concerns the welfare of the patients in the clinical trial, asking whether the accumulated evidence favours one treatment over the other to such an extent that the next patient should decline randomization and instead express a preference for one treatment. Consideration of the decision question within the framework of imprecise probabilities leads to a mathematical definition of equipoise and a method for governing the randomization protocol of a clinical trial. This paper describes in detail the protocol for the conduct of clinical trials based on this new method of analysis, which is illustrated in a retrospective analysis of data from a clinical trial comparing the anti-emetic drugs ondansetron and droperidol in the treatment of postoperative nausea and vomiting. The proposed methodology is compared quantitatively using computer simulation studies with conventional clinical trial designs and is shown to maintain high statistical power with reduced sample sizes, at the expense of a high type I error rate that we argue is irrelevant in some specific circumstances. Particular emphasis is placed on describing the type of medical conditions and treatment comparisons where the new methodology is expected to provide the greatest benefit.

Antiemetics↗

Statistical models of acute mountain sickness.

Acute mountain sickness (AMS) is caused by exposure to altitudes exceeding 2500 m and often resolves by acclimatization without further ascent. Statistical models of AMS score and the probability of an AMS diagnosis were developed to allow the combination of dissimilar exposures for simultaneous analysis. The study population was 302 trekkers from a previous investigation who provided self-reported symptoms upon arrival at 3840 m during hikes through altitudes of 1500 to 6200 m. AMS score (Hackett scale) was estimated by linear regression and the probability of an AMS diagnosis (Lake Louise criteria) by logistic regression. AMS score or probability was significantly associated with exposure day and altitude. Increased altitude over the prior 3 days resulted in higher estimated AMS score or probability and decreased altitude in lower score or probability. The odds ratio (OR) of AMS was 3.6 if not on acetazolamide. Females appeared slightly more susceptible than males (1.5 OR). The approach offers the advantages of (1) improved statistical power by combining exposures, (2) insight into the dose-response relationship of altitude exposure and AMS risk, (3) quantitative tests for the significance of factors that might affect AMS susceptibility, and (4) practical tools to track individual climbers and plan operational ascents.

Acclimatization↗

Effects of pooling mRNA in microarray class comparisons.

MOTIVATION: In microarray experiments investigators sometimes wish to pool RNA samples before labeling and hybridization due to insufficient RNA from each individual sample or to reduce the number of arrays for the purpose of saving cost. The basic assumption of pooling is that the expression of an mRNA molecule in the pool is close to the average expression from individual samples. Recently, a method for studying the effect of pooling mRNA on statistical power in detecting differentially expressed genes between classes has been proposed, but the different sources of variation arising in microarray experiments were not distinguished. Another paper recently did take different sources of variation into account, but did not address power and sample size for class comparison. In this paper, we study the implication of pooling in detecting differential gene expression taking into account different sources of variation and check the basic assumption of pooling using data from both the cDNA and Affymetrix GeneChip microarray experiments. RESULTS: We present formulas for the required number of subjects and arrays to achieve a desired power at a specified significance level. We show that due to the loss of degrees of freedom for a pooled design, a large increase in the number of subjects may be required to achieve a power comparable to that of a non-pooled design. The added expense of additional samples for the pooled design may outweigh the benefit of saving on microarray cost. The microarray data from both platforms show that the major assumption of pooling may not hold. SUPPLEMENTARY INFORMATION: Supplementary material referenced in the text is available at http://linus.nci.nih.gov/brb/TechReport.htm.

Algorithms↗

SNP allele frequency estimation in DNA pools and variance components analysis.

The estimation of single nucleotide polymorphism (SNP) allele frequency in pooled DNA samples has been proposed as a cost-effective approach to whole genome association studies. However, the key issue is the allele frequency window in which a genotyping method operates and provides a statistically reliable answer. We assessed the homogeneous mass extend assay and estimated the variance associated with each experimental stage. We report that a relationship between estimated allele frequency and variance might exist, suggesting that high statistical power can be retained at low, as well as high, allele frequencies. Assuming this relationship, the formation of subpools consisting of 100 samples retains an effective sample size greater than 70% of the true sample size, with a savings of 11-fold the cost of an individual genotyping study, regardless of allele frequency.

Algorithms↗

Ordered subset linkage analysis supports a susceptibility locus for age-related macular degeneration on chromosome 16p12.

BACKGROUND: Age-related macular degeneration (AMD) is a complex disorder that is responsible for the majority of central vision loss in older adults living in developed countries. Phenotypic and genetic heterogeneity complicate the analysis of genome-wide scans for AMD susceptibility loci. The ordered subset analysis (OSA) method is an approach for reducing heterogeneity, increasing statistical power for detecting linkage, and helping to define the most informative data set for follow-up analysis. OSA assesses the linkage evidence in subsets of potentially more homogeneous families by rank-ordering family-specific lod scores with respect to trait-associated covariates or phenotypic features. Here, we present results of incorporating five continuous covariates into our genome-wide linkage analysis of 389 microsatellite markers in 62 multiplex families: Body mass index (BMI), systolic (SBP) and diastolic (DBP) blood pressure, intraocular pressure (IOP), and pack-years of cigarette smoking. Chromosome-wide significance of increases in nonparametric multipoint lod scores in covariate-defined subsets relative to the overall sample was assessed by permutation. RESULTS: Using a correction for testing multiple covariates, statistically significant lod score increases were observed for two chromosomal regions: 14q13 with a lod score of 3.2 in 28 families with average IOP </= 15.5 (p = 0.002), and 6q14 with a lod score of 1.6 in eight families with average BMI >/= 30.1 (p = 0.0004). On chromosome 16p12, nominally significant lod score increases (p </= 0.05), up to a lod score of 2.9 in 32 families, were observed with several covariate orderings. While less significant, this was the only region where linkage evidence was associated with multiple clinically meaningful covariates and the only nominally significant finding when analysis was restricted to advanced forms of AMD. Families with linkage to 16p12 had higher averages of SBP, IOP and BMI and were primarily affected with neovascular AMD. For all three regions, linkage signals at or very near the peak marker have previously been reported. CONCLUSION: Our results suggest that a susceptibility gene on chromosome 16p12 may predispose to AMD, particularly to the neovascular form, and that further research into the previously suggested association of neovascular AMD and systemic hypertension is warranted.

Aged↗

Finding genes influencing susceptibility to complex diseases in the post-genome era.

During the last decade, hundreds of genes that harbor mutations causing simple Mendelian disorders have been identified using a combination of linkage analysis and positional cloning techniques. Traditional approaches to gene mapping have been largely unsuccessful in mapping genes influencing so-called 'complex' genetic diseases, however, because of low power and other factors. Complex genetic diseases do not display simple Mendelian patterns of inheritance, although genes do have an influence and close relatives of probands consequently have an increased risk. These disorders are thought to be due to the combined effects of variation at multiple interacting genes and the environment. Complex diseases have a significant impact on human health because of their high population incidence (unlike simple Mendelian disorders, which tend to be rare). New techniques are being developed aimed specifically at mapping genes conferring susceptibility to complex diseases. A project aimed at mapping genes influencing susceptibility to a complex disease may be undertaken in several stages: establishing a genetic basis for the disease in one or more populations; measuring the distribution of gene effects; studying statistical power using models; carrying out marker-based mapping studies using linkage or association. Quantitative genetic models can be used to estimate the heritability of a complex (polygenic) disease, as well as to predict the distribution of gene effects and to test whether one or more quantitative trait loci (QTLs) exist. Such models can be used to predict the power of different mapping approaches, but are often unrealistic and therefore provide only approximate predictions. Linkage analyses, association studies and family-based association tests are all hindered by low power and other specific problems. Association studies tend to be more powerful but can generate spurious associations due to population admixture. Alternative strategies for association mapping include the use of recent founder populations or unique isolated populations that are genetically homogeneous, and the use of unlinked markers (so-called genomic controls) to assign different regions of the genome of an admixed individual to particular source populations. Linkage disequilibrium observed in a sample of unrelated affected and normal individuals can also be used to fine-map a disease susceptibility locus in a candidate region. New Bayesian strategies make use of an annotated human genome sequence to further refine the position of a candidate disease susceptibility locus.

Genetic Diseases, Inborn↗

Reliability of longitudinal ultrasonographic measurements of carotid intimal-medial thicknesses. Asymptomatic Carotid Artery Progression Study Research Group.

BACKGROUND AND PURPOSE: Serial ultrasonic B-mode measurements of intimal-medial thickness (IMT) of the carotid artery are commonly used as surrogates for describing atherosclerosis progression. This report describes the longitudinal reliability of IMT measurement during a multicenter clinical trial, quantifies the error attributable to differences among readers, and discusses how studies can be efficiently designed. METHODS: Serial B-mode measurements of carotid IMT from the 3-year Asymptomatic Carotid Artery Progression Study (ACAPS; formerly Asymptomatic Carotid Artery Plaque Study) were used to estimate the contributions to longitudinal measurement error of systematic reader effects, nonvisualization, and nonsystematic error and to describe the distribution of "true" progression rates that underlie the observed data. Variance components were estimated from random-effects models fitted to outcome measures formed by averaging IMTs from different sets of carotid artery walls. These were used to contrast the relative efficiency of study designs. RESULTS: Of the total variance of measured IMT, 11% was attributable to systematic differences among readers. Nonvisualization contributed less than 7%. Thus, the predominant source of error was unaccounted for (ie, random error or "noise," which in our analyses included any drift, nonlinearity, and sonographer differences). For studies with measurement protocols similar to ACAPS, follow-up times of 2 years or more are desirable for describing the mean progression rates of cohorts, and of 6 years or more for categorizing progression within individuals. In 3-year studies, sample sizes as low as 237 provide 90% statistical power for detecting risk factors that have correlations with IMT progression of .50 or greater. CONCLUSIONS: The ACAPS measurement protocol provided highly reliable serial IMT data. Moderate-sized multicenter studies using B-mode outcomes are feasible.

Adult↗

Twin studies of eating disorders: a review.

OBJECTIVE: Twin methodology has been used to delineate etiological factors in many medical disorders and behavioral traits including eating disorders. Although twin studies are powerful tools, their methodology can be arcane and their implications easily misinterpreted. METHOD: The goals of this study are to (a) review the theoretical rationale for twin studies; (b) provide a framework for their interpretation and evaluation; (c) review extant twin studies on eating disorders; and (d) explore the implications for understanding etiological issues in eating disorders. DISCUSSION: On the basis of this review, it is not possible to draw firm conclusions regarding the precise contribution of genetic and environmental factors to anorexia nervosa. Twin studies confirm that bulimia nervosa is familial and reveal significant contributions of additive genetic effects and of unique environmental factors in liability to bulimia nervosa. The magnitude of the contribution of shared environment is less clear, but in the studies with the greatest statistical power, it appears to be less prominent than additive genetic factors.

Anorexia Nervosa↗

PCR-based calibration curves for studies of quantitative gene expression in human monocytes: development and evaluation.

BACKGROUND: Quantitative reverse transcription-PCR (RT-PCR) used to detect small changes in specific mRNA concentrations is often associated with poor reproducibility. Thus, there is a need for stringent quality control in each step of the protocol. METHODS: Real-time PCR-based calibration curves for a target gene, tissue factor (TF), and a reference gene, beta-actin, generated from PCR amplicons were evaluated by running cDNA controls. In addition, the reverse transcription step was evaluated by running mRNA controls. Amplification efficiencies of calibrators and targets were determined. Variances within and between runs were estimated, and power statistics were applied to determine the concentration differences that could reliably be detected. RESULTS: Within- and between-run variations (CVs) of cDNA controls (TF and beta-actin), extrapolated from reproducible calibration curves (CVs of slopes, 4.3% and 2.7%, respectively) were 4-10% (within) and 15-38% (between) using both daily and "grand mean" calibration curves. CVs for the beta-actin mRNA controls were 12% (within) and 19-28% (between). Estimates of each step's contribution to the total variation were as follows: CV(RT-PCR), 28%; CV(PCR), 15%; CV(RT), 23% (difference between CV(RT-PCR) and CV(PCR)). PCR efficiencies were as follows: beta-actin calibrator/target, 1.96/1.95; TF calibrator/target, 1.95/1.93. Duplicate measurements could detect a twofold concentration difference (power, 0.8). CONCLUSIONS: Daily PCR calibration curves generated from PCR amplicons were reproducible, allowing the use of a grand mean calibration curve. The reverse transcription step contributes the most to the total variation. By determining a system's total variance, power analysis may be used to disclose differences that can be reliably detected at a specified power.

Actins↗

Predicting reproductive outcome from sperm measurements in Swiss (CD-1) mice.

UNLABELLED: Recent advances in the quantification of sperm characteristics, particularly by techniques of videomicrography, raise the problem of identifying a subset of sperm measurements that can accurately predict reduced reproductive performance in the presence of a reproductive toxicant. This paper discusses and illustrates, with sperm data from Swiss (CD-1) mice, three properties of sperm measurements in addition to the association with reproductive outcome that can be used as objective criteria to distinguish among potentially useful sperm characteristics. Identification of these properties was motivated by the need to single out sperm characteristics that would yield hypothesis tests with good power to distinguish groups with altered sperm characteristics from those with normal sperm characteristics. The list of desirable properties of sperm characteristics includes: a) MEASUREMENT: The most useful sperm characteristics will exhibit low measurement bias and high measurement precision. b) Distribution: Sperm characteristics that follow the class of normal distributions allow easy identification of the most powerful hypothesis testing procedures. c) Variation: Sperm characteristics that exhibit limited variation from individual to individual will make altered values easier to detect. d) Correlation: Sperm characteristics that exhibit a high correlation with reproductive outcome will be most useful. MEASUREMENTs of sperm concentration, sperm motility, and abnormal sperm were examined for each of these properties in Swiss (CD-1) mice. For control animals, evidence of measurement bias between labs and significant variation among studies within labs was found for each sperm characteristic, with measurements of abnormal sperm exhibiting the least bias. Sperm concentration and the natural logarithm of abnormal sperm appeared normally distributed. Individual to individual variation was substantial for sperm motility and sperm concentration, both of which required an approximately 30% decrease in mean to achieve good statistical power. In contrast, abnormal sperm required only a 4% increase in mean to achieve the same power. In spite of the measurement noise associated with these sperm characteristics, data from 25 experiments indicated good agreement between the results of hypothesis tests based on sperm characteristics and reproductive outcome as judged by fertility and the number of pups. The same conclusions reached for reproductive outcome were reached in 19 of 24 experiments for sperm motility, in 19 of 25 experiments for sperm concentration, and in 17 of 24 experiments for abnormal sperm.

Animals↗