Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Comparison of the National Institutes of Health Stroke Scale with disability outcome measures in acute stroke trials.

BACKGROUND AND PURPOSE: Acute stroke trials typically use disability scales as their primary end point. Neurologic impairment scales such as the National Institutes of Health Stroke Scale (NIHSS) are possibly more sensitive to change in patient status. We aimed to compare a range of potential NIHSS end points with modified Rankin Scale (mRS) and Barthel Index (BI) end points. METHODS: We simulated a total of 6000 clinical trials, each with 1400 patients. We estimated statistical power for a range of NIHSS end points, including prognosis-adjusted and fixed dichotomized end points. These end points were compared with the BI and mRS dichotomized at 95 and 1, respectively. RESULTS: The most powerful fixed end point was the NIHSS dichotomized at 1. For prognosis-adjusted outcome, we found greatest power if we defined success as achieving a score of < or =1 or improvement by at least 11 points from baseline. We are more likely to achieve a statistically significant result by using this prognosis-adjusted end point instead of NIHSS < or =1 (odds ratio, 2.8; 95% confidence interval [CI], 2.5 to 3.2). Use of the optimal NIHSS prognosis-adjusted end point rather than BI > or =95 could justify a reduction in sample size of approximately 68% (95% CI, 67% to 69%) without loss of statistical power. CONCLUSIONS: The NIHSS neurologic scale appears more sensitive than the BI or mRS, allowing smaller sample sizes or greater statistical power. The use of an NIHSS prognosis-adjusted end point could allow therapeutic effects from drugs to be more easily identified.

Clinical Trials as Topic↗

A comparison of methods to test mediation and other intervening variable effects.

A Monte Carlo study compared 14 methods to test the statistical significance of the intervening variable effect. An intervening variable (mediator) transmits the effect of an independent variable to a dependent variable. The commonly used R. M. Baron and D. A. Kenny (1986) approach has low statistical power. Two methods based on the distribution of the product and 2 difference-in-coefficients methods have the most accurate Type I error rates and greatest statistical power except in 1 important case in which Type I error rates are too high. The best balance of Type I error and statistical power across all cases is the test of the joint significance of the two effects comprising the intervening variable effect.

Humans↗

A new method in applying power spectral statistics to examine cardio-respiratory interactions in fish.

Power spectral analysis (PSA) provides a powerful tool for determining frequency oscillations in time signals, and it is accepted that mammals can show distinct components in the heart rate (fH) spectrum that are synchronous with ventilatory frequency (fV). Using similar signal processing techniques, these fundamental components at fV are not apparent in the spectrum calculated from fish fH. Here we compare conventional PSA on the R-R interval tachogram generated from ECG traces recorded in rats and fish, with PSA on the raw ECG waveform. The rat R-R tachogram showed a defined sigmoidal component, whereas the fish R-R tachogram was a more chaotic waveform. In agreement with the literature, PSA of these respective waveforms produced a component at the same frequency as ventilation in the rat, but of lower frequency than ventilation for the fish. Applying PSA to the rat ECG produced a spectrum with a fundamental component of similar frequency to that observed in the R-R tachogram spectrum, indicating that the latter adequately contained heart rate variability (HRV) oscillations. However, PSA of the ECG in fish contrasted with that from the R-R tachogram, with components observed in the latter spectrum being absent from the former. This suggests that the frequency components determined by PSA on the fish R-R tachogram were not true components, but were aliased (or folded-back) from higher up in the spectrum. Using established aliasing equations, recalculation of these peaks showed that their true frequency was similar to that of the ventilatory frequency for individual fish. The extent of cardio-respiratory interaction, resulting in fV < f(H/2) in rats but fV > f(H/2) in fish, is suggested to be the origin of the differences observed.

Animals↗

An assessment of recently published gene expression data analyses: reporting experimental design and statistical factors.

BACKGROUND: The analysis of large-scale gene expression data is a fundamental approach to functional genomics and the identification of potential drug targets. Results derived from such studies cannot be trusted unless they are adequately designed and reported. The purpose of this study is to assess current practices on the reporting of experimental design and statistical analyses in gene expression-based studies. METHODS: We reviewed hundreds of MEDLINE-indexed papers involving gene expression data analysis, which were published between 2003 and 2005. These papers were examined on the basis of their reporting of several factors, such as sample size, statistical power and software availability. RESULTS: Among the examined papers, we concentrated on 293 papers consisting of applications and new methodologies. These papers did not report approaches to sample size and statistical power estimation. Explicit statements on data transformation and descriptions of the normalisation techniques applied prior to data analyses (e.g. classification) were not reported in 57 (37.5%) and 104 (68.4%) of the methodology papers respectively. With regard to papers presenting biomedical-relevant applications, 41(29.1 %) of these papers did not report on data normalisation and 83 (58.9%) did not describe the normalisation technique applied. Clustering-based analysis, the t-test and ANOVA represent the most widely applied techniques in microarray data analysis. But remarkably, only 5 (3.5%) of the application papers included statements or references to assumption about variance homogeneity for the application of the t-test and ANOVA. There is still a need to promote the reporting of software packages applied or their availability. CONCLUSION: Recently-published gene expression data analysis studies may lack key information required for properly assessing their design quality and potential impact. There is a need for more rigorous reporting of important experimental factors such as statistical power and sample size, as well as the correct description and justification of statistical methods applied. This paper highlights the importance of defining a minimum set of information required for reporting on statistical design and analysis of expression data. By improving practices of statistical analysis reporting, the scientific community can facilitate quality assurance and peer-review processes, as well as the reproducibility of results.

Analysis of Variance↗

Estimating three-dimensional spinal repositioning error: the impact of range, posture, and number of trials.

STUDY DESIGN: Spinal repositioning sense was tested in normal subjects using a balanced within-subject study design. OBJECTIVES: The study had three objectives: first, to document the number of trials required to derive a representative value of accuracy and precision in spinal repositioning; second, to document the effects of range on spinal repositioning sense; and finally, to document the effect of different lower limb postures on the repositioning performance. SUMMARY OF BACKGROUND DATA: Joint position sense and kinesthesia play an important role in the control of normal movement of the spine. This has important implications for the diagnosis and assessments of specific movement disorders in individuals with spinal pain syndromes. The literature is varied in methods and results in assessing spinal repositioning sense. For some studies, the inability to determine effects for range or differences between patients with low back pain and normal control subjects may be related to the fact that too few trials were performed to detect a statistical difference. METHODS: Twenty-three subjects were tested in standing on a repositioning task for spinal position sense. After a familiarization period, each subject performed 10 matching trials in three ranges (20%, 50%, and 80% of available range) during spinal flexion. The flexion task was performed with knees fully extended, with knees partly flexed, and with the pelvis rotated at 45 degrees to incorporate an asymmetric flexion rotation movement pattern. The three-dimensional coordinates of the repositioning tasks were used to determine accuracy (mean and median of trials) and precision (variable error-standard deviation of trials). The coefficient of variation and statistical power analysis using variables derived from progressively larger numbers of trials were examined. Analysis of variance was used to detect differences for the three ranges and three postures. RESULTS: After the familiarization period, no learning effect was demonstrated across trials. The coefficient of variation and statistical power for the accuracy and precision tended to stabilize after six trials. Using derived variables from six trials, there was a statistically significant range effect. Accuracy in the inner range was worse than that in the outer range (P < 0.05). There was little evidence of a range effect for precision. Posture had little overall impact. Trunk flexion with the knees flexed improved three-dimensional accuracy in the middle range compared with accuracy with the knees extended and during the flexion rotation task. CONCLUSIONS: It was concluded that increasing the number of trials increases the statistical power and stability of the derived variables. In normal subjects, the accuracy of trunk flexion repositioning improves as one moves further into range.

Adult↗

Trick or treat: the effect of placebo on the power of pharmacogenetic association studies.

The genetic mapping of drug-response traits is often characterised by a poor signal-to-noise ratio that is placebo related and which distinguishes pharmacogenetic association studies from classical case-control studies for disease susceptibility. The goal of this study was to evaluate the statistical power of candidate gene association studies under different pharmacogenetic scenarios, with special emphasis on the placebo effect. Genotype/phenotype data were simulated, mimicking samples from clinical trials, and response to the drug was modelled as a binary trait. Association was evaluated by a logistic regression model. Statistical power was estimated as a function of the number of single nucleotide polymorphisms (SNPs) genotyped, the frequency of the placebo 'response', the genotype relative risk (GRR) of the response polymorphism, the strategy for selecting SNPs for genotyping, the number of individuals in the trial and the ratio of placebo-treated to drug-treated patients. We show that: (i) the placebo 'response' strongly affects the statistical power of association studies--even a highly penetrant drug-response allele requires at least a 500-patient trial in order to reach 80 per cent power, several-fold more than the value estimated by standard tools that are not calibrated to pharmacogenetics; (ii) the power of a pharmacogenetic association study depends primarily on the penetrance of the response genotype and, when this penetrance is fixed, power decreases for larger placebo effects; (iii) power is dramatically increased when adding markers; (iv) an optimal study design includes a similar number of placebo- and drug-treated patients; and (v) in this setting, straightforward haplotype analysis does not seem to have an advantage over single marker analysis.

Clinical Trials as Topic↗

Comparing different statistical methods for evaluating diagnostic effectiveness of clinical tests: respiratory distress syndrome as a model.

Phospholipids in specimens of amniotic fluid from 346 patients were quantified and the results evaluated in light of the clinical outcome. Fifty-eight neonates had respiratory distress syndrome. We used this data base to compare different statistical methods for evaluating test effectiveness and diagnostic discrimination. Dichotomizing quantitative tests into binary tests with arbitrary cutoff values was inadequate for comparing test effectiveness. Subgrouping the data into deciles and calculating the incidence of respiratory distress syndrome for each decile avoided the problems of the preceding approach and was easy to calculate and comprehend; however, this method lacked statistical power. Relative operating characteristic curves yielded more statistical power, but results were more difficult to calculate and were not intuitively obvious to most workers in the laboratory. A modified cumulative frequency plot, combining elements of both decile subgrouping and relative operating characteristic curves, was easily calculated and intuitively obvious. These plots, like relative operating characteristic curves, provided an index for quantifying test effectiveness. When used in combination with standard cumulative frequency curves, they also provided direct diagnostic information on disease probability for any value of the clinical assay.

Amniotic Fluid↗

Efficiency of sampling: birthweight and gestational age distributions in two cohorts, < 31 weeks and 500-1499 grams.

We studied the efficiency of two common sampling strategies used to assemble cohorts to study the long-term problems of preterm infants: infants with birthweights of 500-1499 g, and infants with gestational ages (GA) of < 31 weeks. Birthweight, GA and 2-year outcome data from a population based study of infants < 2001 g, the Central New Jersey Brain Hemorrhage Study (NBH), were used to define the birthweight and GA distributions, at enrollment and at the age of 2 years, of overlapping subsets: infants 500-1499 g (n = 599) and infants < 31 weeks of age (n = 522). Using frequencies from the NBH study, we estimated that 1000 infants of 500-1499 g enrolled at birth would produce 712 infants at the age of 2 years, 498 below 31 weeks and 214 above. Enrolling 1000 infants < 31 weeks would produce a cohort of 697 infants at the age of 2, all of whom were < 31 weeks. Neither sampling strategy maximised the statistical power to investigate the pathophysiological determinants of long-term outcomes associated with short GA. Both methods oversampled older GAs. A stratified sampling technique based on GA, designed to produce equal numbers of subjects at each week of GA, would improve statistical power to study long-term outcomes. As we move from descriptive to analytical studies of preterm infants, we need to devise efficient, GA-based, sampling strategies that maximise statistical power to test pathophysiological hypotheses.

Birth Weight↗

Communication disorders: a power analytic assessment of recent research.

This study assessed the relative statistical power of contemporary research in communication disorders. Results of the analysis, based upon an evaluation of two major journals, revealed overall mean power figures of 0.16, 0.44, and 0.73 for small, medium, and large effect sizes, respectively. Interdisciplinary comparisons indicated that low statistical power is not unique to research in communication disorders, but is apparent in other behavioral science areas as well. Several alternatives are offered to the researcher who will want to ensure sufficient power for his investigation on an a priori basis. The implications of this study are discussed in reference to the experimenter/clinician model.

Audiology↗

Power of regression and maximum likelihood methods to map QTL from sib-pair and DZ twin data.

A common study design to map quantitative trait loci (QTL) is to compare the phenotypes and marker genotypes of two or more siblings in a sample of unrelated sib groups, and to test for linkage between chromosome location and quantitative trait values. The simplest case is sib pairs only, in particular dizygotic twin pairs, and a simple and elegant regression method was proposed by Haseman & Elston in 1972 to test for linkage. Since then, several other methods have been proposed to test for linkage. In this study, we derived the statistical power of linear regression and maximum likelihood methods to map QTL from sib pair data analytically, and determined which methods are superior under which set of population parameters. In particular, we considered four regression-based and three maximum likelihood-based approaches, and derived asymptotic approximations of the mean test statistic and statistical power for each method. It was found, both analytically and by computer simulation, that the revisited or new Haseman-Elston method (based upon the mean-corrected crossproduct of the observations on sib-pairs) is less powerful than a full maximum likelihood approach and is also inferior to the Haseman-Elston method under a realistic range of values for the population parameters. We found that a simple regression method, based upon both the squared difference and the mean-corrected squared sum of the observations on sib-pairs, is as powerful as a full maximum likelihood approach. Our derivations of statistical power for regression and maximum likelihood methods provide a simple way to compare alternative methods and obviate the need to perform elaborate computer simulations. DZ twin pairs are likely to be more powerful for linkage analysis than ordinary siblings because they may share more common environmental effects, thereby increasing the proportion of within-family variance that is explained by a QTL.

Chromosome Mapping↗

Power for detecting genetic divergence: differences between statistical methods and marker loci.

Information on statistical power is critical when planning investigations and evaluating empirical data, but actual power estimates are rarely presented in population genetic studies. We used computer simulations to assess and evaluate power when testing for genetic differentiation at multiple loci through combining test statistics or P values obtained by four different statistical approaches, viz. Pearson's chi-square, the log-likelihood ratio G-test, Fisher's exact test, and an F(ST)-based permutation test. Factors considered in the comparisons include the number of samples, their size, and the number and type of genetic marker loci. It is shown that power for detecting divergence may be substantial for frequently used sample sizes and sets of markers, also at quite low levels of differentiation. The choice of statistical method may be critical, though. For multi-allelic loci such as microsatellites, combining exact P values using Fisher's method is robust and generally provides a high resolving power. In contrast, for few-allele loci (e.g. allozymes and single nucleotide polymorphisms) and when making pairwise sample comparisons, this approach may yield a remarkably low power. In such situations chi-square typically represents a better alternative. The G-test without Williams's correction frequently tends to provide an unduly high proportion of false significances, and results from this test should be interpreted with great care. Our results are not confined to population genetic analyses but applicable to contingency testing in general.

Alleles↗

Optimizing detection of QTLs retarding aging: choice of statistical model and animal requirements.

Quantitative trait locus (QTL) analysis makes no assumptions about the identity of genes involved in regulating aging. Moreover, it may be used as the first step in identifying such genes and, thus QTL analysis may be instrumental in formulating new hypotheses about aging. Genetic experiments, however, require hundreds to thousands of animals and are very expensive in mammals. Statistical power to detect longevity genes could be improved by excluding accidental, unrelated to aging mortality. While many early deaths are probably accidental, excluding early mortality altogether eliminates the age-related component, too. We used computer simulations to assess the effect of excluding early age-related, mortality on the statistical power of several common tests, such as t-test, Mann-Whitney and chi(2). Surprisingly, even the age-related, Gompertz component of early mortality reduces the statistical power of the t- and Mann-Whitney tests. For example, in a backcross design, to detect a gene slowing down the rate of aging and increasing mouse life span by 10% (P=0.0001; power=0.8), a regular t-test will require 640 mice, all kept for the entire life span and genotyped. If life spans of only 25% of the longest-lived animals from each of the two groups, carrying a putative longevity allele and not carrying it, are compared, population size can be reduced by two-fold, to about 300, and genotyping by seven-fold, to 90. Confirming simulation results, the significance of the effect of caloric restriction on life span increased from P=3.4x10(-5) to 1.1x10(-7), when life spans of only 40% of the longest-lived mice from each of the two groups, ad libitum fed and calorie restricted, were compared. Finally, finding the optimal combination of statistical test, the number of phenotyped and the number of genotyped animals, which would minimize experimental costs was addressed.

Aging↗

How many patients are needed? Variation and design considerations in bone histomorphometry.

In osteoporosis research, bone histomorphometry plays an important role in documenting the biological effects and possible side-effects of new drug treatments. To ensure that the study is properly scaled, it is important to be concerned with the risk of type II error; that is, the risk of failing to detect a real difference. We therefore calculated the necessary sample size in bone histomorphometric studies according to a specified difference of 15% between two groups. The calculations were based on variance components estimated from three different studies: women with a distal fracture of the forearm (n = 22); patients with pituitary insufficiency (n = 21); and patients with primary hyperparathyroidism (n = 21). Using a significance level of 0.05 and a risk of type II error of 0.20, the statistical power of two different designs was compared: a single biopsy design comparing the responses in two groups after the treatment; and a paired biopsy design in which individual differences (posttreatment minus baseline) were calculated before the comparison of the two groups. We found that the mineral apposition rate, wall thickness, and erosion depth are statistically powerful indices that, in the single biopsy design, require no more than n = 25 in each group to detect differences of 15% between the groups. Bone volume, erosion surface, osteoid surface, mineralizing surface, and activation frequency need group sizes of 100-600 individuals to find a 15% difference to be statistically significant. However, the effect of bisphosphonate treatment, for instance, is large enough to reduce the group size to 20 individuals concerning activation frequency. The remodeling balance reaches extreme group sizes of several thousand for a 15% difference to be statistically significant, but for a 5 microm (approximately 150%) improvement, about 100 individuals are required in the single biopsy design. An analysis of the components of variance showed that the variation between individuals is small and often negligible compared with the variation within individuals, and sample sizes needed for the paired biopsy design are therefore larger than those for the single biopsy design. In conclusion, the most cost-effective histomorphometric study design within a randomized clinical trial appears to be a single biopsy design comparing posttreatment biopsies with scaling performed according to the statistical power of the indices of interest.

Biopsy↗

An empirical investigation into the number of subjects required for an event-related fMRI study.

Optimising the number of subjects required for an event-related functional imaging study is critical for ensuring sufficient statistical power. We report an empirical investigation of this issue by employing a resampling approach to the data of 58 subjects drawn from four previous GO/NOGO studies. Using voxelwise measures and setting the activation map from the complete sample to be a "gold standard", analyses revealed the statistical power to be surprisingly low at typical sample sizes (n = 20). However, voxels that were significantly active from smaller samples tended to be true positives, that is, they were typically active in the gold standard map and correlated well with the gold standard activation measure. The numerous false negatives that resulted from the lower SNR of the smaller samples drove the poor statistical power of those samples. Splitting the sample into two groups provided a test of the reproducibility of activation maps that was assessed using an alternative measure that quantified the distances between centres-of-mass of activated areas. These analyses revealed that although the voxelwise overlap may be poor, the locations of activated areas provide some optimism for studies with typical sample sizes. With n = 20 in each of two groups, it was found that the centres-of-mass for 80% of activated areas fell within 25 mm of each other. The reported analyses, by quantifying the spatial reproducibility for various sample sizes performing a typical event-related cognitive task, thus provide an empirical measure of the disparity to be expected in comparing activation maps.

Adult↗

Two-stage designs in case-control association analysis.

DNA pooling is a cost-effective approach for collecting information on marker allele frequency in genetic studies. It is often suggested as a screening tool to identify a subset of candidate markers from a very large number of markers to be followed up by more accurate and informative individual genotyping. In this article, we investigate several statistical properties and design issues related to this two-stage design, including the selection of the candidate markers for second-stage analysis, statistical power of this design, and the probability that truly disease-associated markers are ranked among the top after second-stage analysis. We have derived analytical results on the proportion of markers to be selected for second-stage analysis. For example, to detect disease-associated markers with an allele frequency difference of 0.05 between the cases and controls through an initial sample of 1000 cases and 1000 controls, our results suggest that when the measurement errors are small (0.005), approximately 3% of the markers should be selected. For the statistical power to identify disease-associated markers, we find that the measurement errors associated with DNA pooling have little effect on its power. This is in contrast to the one-stage pooling scheme where measurement errors may have large effect on statistical power. As for the probability that the disease-associated markers are ranked among the top in the second stage, we show that there is a high probability that at least one disease-associated marker is ranked among the top when the allele frequency differences between the cases and controls are not <0.05 for reasonably large sample sizes, even though the errors associated with DNA pooling in the first stage are not small. Therefore, the two-stage design with DNA pooling as a screening tool offers an efficient strategy in genomewide association studies, even when the measurement errors associated with DNA pooling are nonnegligible. For any disease model, we find that all the statistical results essentially depend on the population allele frequency and the allele frequency differences between the cases and controls at the disease-associated markers. The general conclusions hold whether the second stage uses an entirely independent sample or includes both the samples used in the first stage and an independent set of samples.

Algorithms↗

How would a decline in sperm concentration over time influence the probability of pregnancy?

BACKGROUND: Reports have suggested a decline in sperm concentration during the second half of the 20th century. The effect of this decline on fecundability (the monthly probability of pregnancy) could be detected in principle by a study of time to pregnancy. In practice, the amplitude of this expected effect is not well known and the statistical power of time-to-pregnancy studies to detect it has not been explored. METHODS: We developed a nonparametric model to describe a temporal decline in sperm concentration using data on French semen donors. We then applied this model to 419 Danish couples planning a first pregnancy in 1992, to predict their time to pregnancy as if the pregnancy attempt had begun during earlier decades with higher sperm concentrations. Finally, we used bootstrap simulations to estimate the statistical power of prospective or retrospective studies that compared fecundability (estimated from time to pregnancy) across these time periods. We express the change in fecundability over time as a fecundability ratio (FR), with values less than 1 indicating decreased fecundability. RESULTS: We estimate that the median sperm concentration decreased by 21% from 1977 to 1992 and by 47% from 1947 to 1992. The estimated decline in fecundability with those semen changes was 7% from 1977 to 1992 (FR = 0.93, adjusted) and 15% from 1947 to 1992 (FR = 0.85, adjusted). The total numbers of couples that would be needed in prospective studies of time to pregnancy to detect these changes in fecundability (with a power of 80%) were 12,000 when comparing 1977 to 1992, and 2000 when comparing 1947 to 1992. Retrospective studies of the same size that excluded childless couples had much lower statistical power and were biased toward the null. CONCLUSION: The effect of realistic declines in sperm concentration on time to pregnancy may be observed only with studies that include several thousand couples.

Adult↗

Preoperative magnetic resonance imaging screening for a surgical decision regarding the approach for anterior spine fusion at the cervicothoracic junction.

STUDY DESIGN: In a study investigating the correlation between a set of designed criteria and judgments of surgical experience, 100 cervical magnetic resonance images from different patients were used. OBJECTIVES: To demonstrate reliable and reproducible anatomic measurements that can aid spine surgeons in selecting surgical approaches for anterior spine fusion in the cervicothoracic region. SUMMARY OF BACKGROUND DATA: Surgical approaches to the cervicothoracic junction vary among surgeons. Whereas sternotomy provides maximum exposure, less extensive approaches are preferred to minimize surgical trauma, provided surgical goals are not compromised. No quantitative criteria currently exist to determine before surgery the least invasive surgical approach for sufficiently exposing pertinent anatomy. METHODS: Thirteen geometric variables designed to be clinically practical and to expose important anatomic relations were used to evaluate 100 sagittal scout cervical magnetic resonance image sequences. An experienced spine surgeon independently rated each image for the most appropriate surgical approach to the C5-T2 region. The ratings were tested for interrater reliability using a second spine surgeon. After testing for interrater and intrarater reliability, the geometric measurements were correlated with the surgeon's selected surgical approaches for each intervertebral segment (P < 0.05). RESULTS: Instrument manubrial thoracic distances, reflecting standardized heights of intervertebral discs above or below the superior tip of the manubrium, were the most reliable, reproducible, and correlative with the choice of surgical approach. All the measurements but one, the instrument manubrial thoracic distance for T1/T2, demonstrated interrater and intrarater reliability, with an interclass correlation of at least 0.70. The primary surgeon-investigator indicated the anterior approach with sternotomy (n = 3) or the transverse cervical approach (n = 97) for the C7/T1 exposure, and the anterior approach with sternotomy (n = 43) or the transverse cervical approach (n = 57) for the T1/T2 exposure. The interrater questionnaire reliability results indicated statistical agreement between the primary surgeon-investigator and the second cervical spine surgeon at all vertebral segments evaluated. Instrument manubrial thoracic distances showed the strongest significant correlation with the surgical approach, demonstrating a statistical power of 1. For the C7/T1 exposure, the instrument manubrial thoracic distance for C7/T1 was 1.9 +/- 2 cm (95% confidence interval [CI] = 1.41 to 2.22) for the transverse cervical approach, and -3.3 +/- 1.3 cm (95% CI = -4.79 to -1.75)] for the anterior approach with sternotomy. The instrument manubrial thoracic distance measurements for C5/C6, C6/C7, and T1/T2 also showed nonoverlapping 95% confidence intervals for the transverse cervical versus the anterior approach with sternotomy for the C7/T1 exposure. For the T1/T2 exposure, all four instrument manubrial thoracic distance measurements again showed statistically significant differences between approaches, with nonoverlapping 95% confidence intervals and a statistical power of 1. In addition, the measurements elaborating the anterior-to-posterior distance of the thoracic outlet and the measurements of the angle between the planes of the intervertebral disc and the sternum also showed statistically significant differences between approaches for the T1/T2 segment, with a statistical power of at least 0.9. CONCLUSIONS: Strong correlations exist between objective measurements and the choice of surgical approach for anterior spine fusion. Among investigated anatomic relations, the instrument manubrial thoracic distance correlated most reliably with the surgeons' choice of the anterior approach. Such objective measurements represent tools that cervical spine surgeons can use to determine the surgical approach.

Cervical Vertebrae↗

Metrics comparing simulated early concentration profiles for the determination of bioequivalence.

PURPOSE: To compare the effectiveness of various metrics which evaluate bioequivalence in the early phase of concentration-time profiles. METHODS: Two-period crossover trials were simulated with increasing assumed ratio of the true absorption rate constants of the two formulations, and with various kinetic models. Kinetic sensitivities (KS) and standard errors (SE) of the various metrics were recorded and the percentage of trials accepting bioequivalence (the statistical power) was evaluated. The principal metrics included the partial AUC(AUCP), the intercept obtained by linear extrapolation of the ratios of the lower over higher concentrations (C) measured for the two formulations (I L/H), and the ratios of intercepts extrapolated from logarithmic C/ time values of the two products (MLOG). For comparison, also properties of CMAX and an ideally evaluated measure (Id) were determined. RESULTS: MLOG showed generally the highest statistical power and KS, and also the largest SE, closely followed by I L/H. Partial AUC exhibited lower power and KS, but also smaller SE than the intercept procedures. The three methods had much higher power, KS and SE than CMAX. These comparisons were maintained over various kinetic conditions and experimental designs. The effective evaluation of bioequivalence in the early phase of studies is assured with 3 (or more) measurements until the population average peak of the reference formulation. CONCLUSIONS: The three principal methods assess bioequivalence very effectively in the early phase of a concentration-time profile. MLOG had the highest statistical power, closely followed by I L/H and then by partial AUC.

Area Under Curve↗