Hoehler's adjusted kappa is equivalent to Yule's Y.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to S D Walter.
Explore the source record for details and available documents.
OBJECTIVE: To create gross motor function growth curves for children with Down syndrome (DS) and to estimate the probability that motor functions are achieved by different ages. DESIGN: Nonlinear growth curve analysis by using a 2-parameter (rate, upper limit) model. SETTING: Early intervention programs, schools, and children's homes. PARTICIPANTS: One hundred twenty-one children with DS, ages 1 month to 6 years. MAIN OUTCOME MEASURES: Gross Motor Function Measure (GMFM) and severity of motor impairment. RESULTS: The curves for children with mild (n = 51) and moderate/severe (n = 70) impairment were characterized by a greater increase in GMFM scores during infancy and smaller increases as the children approached the predicted maximum score of 85.9 or 87.9. The estimated probability that a child would roll by 6 months was 51%; sit by 12 months, 78%; crawl by 18 months, 34%; walk by 24 months, 40%; and run, walk up stairs, and jump by 5 years, 45% to 52%. CONCLUSIONS: Children with DS require more time to learn movements as movement complexity increases. Impairment severity affected the rate but not the upper limit of motor function. The results have implications for counseling parents, making decisions about motor interventions, and anticipating the time frame for achievement of motor functions.
BACKGROUND: Approaches to interpretation of quality of life changes in clinical trials have fallen into two camps: those that rely on the distribution of changes and the Effect Size (ES), and those that use some external anchor, such as patient judgments of change, which is then used to compute a Minimally Important Difference (MID), the proportion benefiting from treatment, p(B), and the Number Needed to Treat (NNT). OBJECTIVE: To examine the relationship between the ES and p(B), and the impact of the MID on this relationship. METHODS: Simulation was used based on a normal distribution to compute the proportion of patients benefiting in both parallel group and crossover designs, for various values of the ES and the MID. The agreement of the simulation with empirical data from four studies of asthma and respiratory disease was assessed. The effect of skewness in the distributions of change scores on the relationship between ES and p(B) was also examined. RESULTS: The simulation showed a near-linear relationship between ES and p(B), which was nearly independent of the value of the MID. Agreement of the simulation with the empirical data were excellent. Although the curves differed for crossover and parallel group designs, the general form was similar. Introducing moderate skew into the distributions had minimal impact on the relationship. CONCLUSIONS: The proportion of patients who will benefit from treatment can be directly estimated from the ES, and is nearly independent of the choice of MID. Effect size and anchor based approaches provide equivalent information in this situation.
The absence of standardized assessment protocols with well- defined measurement properties limits comparison of outcomes among those receiving long-term oxygen therapy (LTOT). We describe simple protocols for a hospital test, a simulated home test, and an actual home test, their reliability and relationship to each other. Stable patients with exercise hypoxemia participated. In 74 patients who completed four exercise tests, correlations between tests ranged from 0.85 to 0.78. Of these 27.0% had the same prescription from all four tests. In 46% prescriptions were within 1 L/ min and in 27% within 2 L/min. During exercise the hospital tests suggested slightly higher oxygen prescriptions than did the simulated home tests (2.5 L/min versus 2.0 L/min, p < 0.001). In 23 patients who participated in actual home assessments, the correlations between the home test, the hospital, and the simulated home tests were 0.22 (95% CI -0.24 to 0.67) and 0.27 (95% CI -0.18 to 0.72). In conclusion, standardizing tests for the assessment of LTOT is important. We describe simple hospital and simulated home tests that are reproducible, easy to carry out, and correlate well with each other.
Explore the source record for details and available documents.
The purpose of this article is to describe our clinical experiences in using the Gross Motor Function Measure (GMFM) to evaluate motor development in children with Down syndrome and to provide strategies we found helpful in enhancing a child's adherence to standardized testing. The issues discussed are: (1) strategies for test administration; (2) modifications in administration and scoring; (3) reliability of the GMFM using the modified administration and scoring procedures; and (4) applications of the GMFM for clinical practice. The strategies and recommendations address the particular characteristics of children with Down syndrome and allow for their progress to be monitored relative to other children with Down syndrome rather than to children without motor delays. Future studies validating the use of specific goal areas for the administration and scoring of the GMFM for children with Down syndrome are recommended.
The debate concerning the choice of effect measure for epidemiologic data has been renewed in the literature, and it suggests some continuing disagreement between the pertinent clinical and statistical criteria. In this article, some defining characteristics of the main choices of effect measure [risk difference (RD), relative risk (RR), and odds ratio (OR)] for binary data are presented and compared, with consideration of both the clinical and statistical perspectives. Relationships of these measures to the relative risk reduction (RRR) and number needed to treat (NNT) are also discussed. A numerical comparison of models of constant RD, RR, and OR is made to assess when and by how much they might differ in practice. Typically the models show only small numerical differences, unless extreme extrapolation is involved. The RD and RR models can predict impossible event rates, either less than zero or greater than 100%. Each measure has potential theoretical justification. RD and RR may enjoy some advantages for communication of risk, but OR may be preferred for data analysis. A clear distinction should be maintained between the objectives of data analysis and subsequent risk communication, and different effect measures may be needed for each.
OBJECTIVE: To develop and evaluate alternative methods of adjusting primary medical care capitation payments for variations in relative need for health care among enrolled practice populations. METHODS: We developed alternative needs-based capitation formulae and applied them to a sample of capitation-funded primary care practices to assess each formula's performance against a reference standard of capitation payments based on age, sex and self-assessed health status of the enrolled populations. The alternative formulae were based on: (1) age and sex; (2) age, sex and individually-measured socioeconomic characteristics; (3) age, sex and socioeconomic characteristics imputed from census data for enrollees' neighbourhood of residence; (4) age, sex and standardized mortality ratio for enrollees' neighbourhood of residence. RESULTS: Age/sex-adjusted capitation payments for the six practices studied ranged from 10% higher to 18% lower than the reference standard payments. Capitation formulae based on socioeconomic and mortality data did not perform consistently better than the current age/sex-based formula. CONCLUSIONS: Primary medical care capitation payments adjusted only for age and sex do not reflect the relative health care needs of enrolled practice populations. Our alternative formulae based on socioeconomic and mortality data also failed to reflect relative needs. Methods that use other approaches to adjusting for differences in relative need among enrolled populations should be investigated.
BACKGROUND: It is unclear which of the number or the density of naevi on the skin is the more appropriate measure of risk of melanoma. Furthermore, the relationship between the number of naevi and their density in an individual has not been explored. Thus, for example, it is unknown if larger people tend to have more naevi by virtue of having a larger skin area, or if the density of naevi is similar in people of different body sizes. In this study, we explored the relationship between the number and the density of naevi in a sample of adolescents. SUBJECTS AND METHODS: A sample survey of naevi in 472 grade 9 secondary school students (aged 14-15 years) was conducted in Tasmania, Australia during 1992, and a subset of these individuals was followed up in 1997. Counts of naevi of various sizes were taken on the arm, leg, and back. Naevus density was estimated by using an algorithm to estimate body surface area from the height and weight of an individual. More general relationships of the naevus counts to height and weight were also explored. Finally, we considered whether the relationship between naevus density and the anthropometric variables could be confounded by exposure to ultraviolet radiation. RESULTS: The mean number of naevi was very similar in the two samples. Naevus density was slightly lower in the 1997 sample, mainly because of increasing body size in the cohort. The numbers of naevi were only weakly related to height and weight in males, and there was essentially no relationship in females. Regression analysis showed significant relationships of weight to the back naevus counts in males in 1992 and 1997, and to the arm naevus count in males in 1997; otherwise, none of the regression coefficients for height and weight were statistically significant. This picture did not change following adjustment for potentially confounding variables indicating time spent outdoors or in the sun. Furthermore, there was no evidence that time spent in the sun was related to the body mass index. CONCLUSIONS: It appears that the number and density of naevi in an individual are unrelated. Accordingly, with the present state of knowledge concerning the risk of melanoma, both the number and density of naevi should be considered as equally valid in future studies as markers of the risk of melanoma, and in studies on the natural history of naevi. If the disease mechanism is systemic, and not related to particular naevi, naevus density might form the better marker of risk. However, if the disease mechanism is related to effects on particular naevi, then the risk would vary in proportion to the number of naevi.
OBJECTIVES: To determine the usefulness of endocervical discharge opacity as a risk indicator for chlamydial infection in relation to two acknowledged visual indicators--yellow endocervical discharge and easily induced mucosal bleeding of the cervix. METHODS: Women from two family planning clinics, a therapeutic abortion clinic, and a university student health clinic (n = 1418 total) consented to a pelvic examination and chlamydia testing, and completed a questionnaire on socio-demographics, sexual behaviour, medical history, and symptoms. A case of chlamydia was defined as positive by culture or blocked enzyme immunoassay in an endocervical swab. RESULTS: The prevalence of chlamydial infection in the clinics was 6.3%. All three of the visual indicators--yellow endocervical discharge, easily induced bleeding, and opaque cervical discharge--were statistically significantly and independently associated with chlamydial infection (odds ratios 2.8, 2.3, and 2.9 respectively), independent of clinic type. Adjustment for the other visual indicators made little difference to the odds ratios. CONCLUSION: Opacity of endocervical discharge was at least as important as the other two commonly acknowledged indicators of chlamydial cervicitis--yellow endocervical discharge and easily induced mucosal bleeding of the cervix.
BACKGROUND AND PURPOSE: This study examined the reliability, validity, and responsiveness to change of measurements obtained with a 66-item version of the Gross Motor Function Measure (GMFM-66) developed using Rasch analysis. SUBJECTS AND METHODS: The validity of measurements obtained with the GMFM-66 was assessed by examining the hierarchy of items and the GMFM-66 scores for different groups of children from a stratified random community-based sample of 537 children with cerebral palsy (CP). A subset of 228 children who had been reassessed at 12 months was used to test the hypothesis that children who are young (<5 years of age) and have "mild" CP will demonstrate greater change in GMFM-66 scores than children who are older ((5 years of age) and whose CP is more severe. Data from an additional 19 children with CP who were assessed twice, one week apart, were used to examine test-retest reliability. RESULTS: The overall changes in GMFM-66 scores over 12 months and a time ( severity ( age interaction supported our hypotheses. Test-retest reliability was high (intraclass correlation coefficient=.99). CONCLUSION AND DISCUSSION: This study demonstrated that the GMFM-66 has good psychometric properties. By providing a hierarchical structure and interval scaling, the GMFM-66 can provide a better understanding of motor development for children with CP than the 88 item GMFM and can improve the scoring and interpretation of data obtained with the GMFM.
BACKGROUND AND PURPOSE: Development of gross motor function in children with cerebral palsy (CP) has not been documented. The purposes of this study were to examine a model of gross motor function in children with CP and to apply the model to construct gross motor function curves for each of the 5 levels of the Gross Motor Function Classification System (GMFCS). SUBJECTS: A stratified sample of 586 children with CP, 1 to 12 years of age, who reside in Ontario, Canada, and are known to rehabilitation centers participated. METHODS: Subjects were classified using the GMFCS, and gross motor function was measured with the Gross Motor Function Measure (GMFM). Four models were examined to construct curves that described the nonlinear relationship between age and gross motor function. RESULTS: The model in which both the limit parameter (maximum GMFM score) and the rate parameter (rate at which the maximum GMFM score is approached) vary for each GMFCS level explained 83% of the variation in GMFM scores. The predicted maximum GMFM scores differed among the 5 curves (level I=96.8, level II=89.3, level III=61.3, level IV=36.1, and level V=12.9). The rate at which children at level II approached their maximum GMFM score was slower than the rates for levels I and III. The correlation between GMFCS levels and GMFM scores was (.91. Logistic regression, used to estimate the probability that children with CP are able to achieve gross motor milestones based on their GMFM total scores, suggests that distinctions between GMFCS levels are clinically meaningful. CONCLUSION AND DISCUSSION: Classification of children with CP based on functional abilities and limitations is predictive of gross motor function, whereas age alone is a poor predictor. Evaluation of gross motor function of children with CP by comparison with children of the same age and GMFCS level has implications for decision making and interpretation of intervention outcomes.
Meta-analysis has become a popular technique in many areas of biomedical research. While there have been a number of studies and commentaries on the use of meta-analyses and systematic reviews of clinical trials of therapy, little is known about the use of meta-analysis techniques for the evaluation of screening data. This paper presents the first systematic survey of meta-analyses of screening, with an assessment of their methodologic quality and their statistical methods. Our findings show that meta-analysis has not often been used in this area as yet, and that published meta-analyses of screening are often deficient in their reporting of methodology. A brief examination of the evaluative methods of several policy-making groups reveals that they did not routinely use formal quantitative meta-analyses of screening data, at least until 1997. There is considerable potential for meta-analyses of this kind, but improvement in methodological standards will be required to obtain valid conclusions.
We present a method to estimate the summary receiver operating characteristic (SROC) curve for combining information on a diagnostic test from several different studies. Unlike previous methods that assume the reference standard to be error free, our approach allows for the possibility of errors in the reference standard, through use of a latent class model. The model provides estimates of the sensitivity and specificity of the diagnostic test and the case prevalence in each study; these parameters can then be used in a meta-analysis, for example, using the regression method proposed by Moses et al., of a measure of test discrimination on a measure of the diagnostic threshold, to fit the SROC. The method is illustrated with an example on Pap smears that shows how adjusting for imperfection in the reference standard typically reduces the scatter of data in the SROC plot, and tends to indicate better performance of the test than otherwise.
Five strategies for creating predictive models of lower respiratory tract infection in residents of long-term care facilities were compared. A linear judgment model was derived by administering clinical vignettes to physicians who indicated the risk of infection based on the presence or absence of five predictor variables. A model based on physician consensus was created using the same variables. Three models based on empirical data (logistic regression, proportional hazards, and recursive partitioning) were created from a "derivation" sample of data from a cohort study of lower respiratory tract infections in nursing homes using the five predictor variables. All models were applied to a validation set and compared using receiver operating characteristic (ROC) curves. The data-derived and consensus models showed the highest discriminative ability while the linear judgment model showed inferior performance.
BACKGROUND: Although solar radiation is well established as a risk factor for melanoma, it is less clear how the pattern and timing of exposure to ultraviolet (UV) radiation might be important. The particular objective of this study was to evaluate the association of melanoma risk with various measures of intermittent and chronic exposures to UV radiation, and to assess how these exposures interact with other risk factors such as skin type. METHODS: Data were analysed from a large case-control study (583 cases, 608 controls) of malignant melanoma, carried out in southern Ontario, Canada. RESULTS: Significant risk increases were identified with several measures of intermittent exposure, including beach vacations in adolescence and in the past 5 years, previous sunburn, and use of sunbeds and sunlamps. Chronic exposure, indicated by days of outdoor activity during adolescence and by occupation in recent adult life, was associated with significantly reduced risk. Subgroup analyses showed: no major risk differences by body site of melanoma; stronger association of lentigo maligna melanoma with intermittent exposure; more pronounced effects of beach vacations and sunburn in younger subjects; and consistently higher risks for intermittent exposures among subjects with skin more susceptible to burning. CONCLUSIONS: The data lend limited support to the hypothesis of increased risk associated with intermittent UV exposure. The findings suggest that future studies should take age at diagnosis, host susceptibility and histological subtype into account.
Estimation of sensitivity and specificity for diagnostic or screening tests usually requires independent confirmation of subjects as diseased or nondiseased using a gold standard. In practice, however, application of the confirmatory procedure is usually limited to individuals with one or more positive test results. For situations in which two initial tests are applied, recent literature has shown that one can use the data from confirmed disease cases to estimate the ratio of test sensitivities and the information from confirmed noncases to estimate the ratio of false-positive rates. In this paper, I show that estimates of sensitivity and specificity can be obtained for each test separately, together with an estimate of the disease prevalence. The only additional information required compared with previous methodology is the total number of individuals tested, a quantity that is usually readily available. The assumption that the test errors are independent is required. Although specific patterns of test errors cannot be identified, the overall assumption can be tested using goodness of fit. I illustrate the methods using data on breast cancer screening. Provision of sensitivity and specificity estimates for each test separately provide considerably greater insight into the data than previous methods.
Seven independent assessments of diagnosis were obtained for 92 records of nontrauma emergency department visits in Saint John, New Brunswick, Canada, in 1994. The hospital database was 1.18 times as likely (p < 0.05) as six external physician raters to classify visits as cardiorespiratory, which was consistent for high- and low-pollution days. Kappa was 0.70 (95 percent confidence interval (CI) 0.68-0.73). Kappajs were: asthma, 0.69 (95% CI 0.64-0.73); chronic obstructive pulmonary disease, 0.78 (95% CI 0.74-0.83); respiratory infections, 0.53 (95% CI 0.49-0.57); cardiac, 0.84 (95% CI 0.79-0.88); and other, 0.66 (95% CI 0.62-0.71). Substantial or better interobserver agreement was seen, respiratory infections notwithstanding, and there was no evidence of diagnostic bias in relation to daily air pollution level.