Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Action potential propagation through embryonic dorsal root ganglion cells in culture. II. Decrease of conduction reliability during repetitive stimulation.

1. The reliability of the propagation of action potentials (AP) through dorsal root ganglion (DRG) cells in embryonic slice cultures was investigated during repetitive stimulation at 1-20 Hz. Membrane potentials of DRG cells were recorded intracellularly while the axons were stimulated by an extracellular electrode. 2. In analogy to the double-pulse experiments reported previously, either one or two types of propagation failures were recorded during repetitive stimulation, depending on the cell morphology. In contrast to the double-pulse experiments, the failures appeared at longer interpulse intervals and usually only after several tens of stimuli with reliable propagation. 3. In the period with reliable propagation before the failures, a decrease in the conduction velocity and in the amplitude of the afterhyperpolarization (AHP), an increase in the total membrane conductance, and the disappearance of the action potential "shoulder" were observed. 4. The reliability of conduction during repetitive stimulation was improved by lowering the extracellular calcium concentration or by replacing the extracellular calcium by strontium. The reliability of conduction decreased by the application of cadmium, a calcium channel blocker, 4-amino pyridine, a fast potassium channel blocker, or apamin or muscarine, the blockers of calcium-dependent potassium channels. The reliability of conduction was not effected by blocking the sodium potassium pump with ouabain or by replacing extracellular sodium with lithium. 5. In the period with reliable propagation cadmium, apamin, and muscarine reduced the amplitude of the AHP. The shoulder of the action potential was more pronounced and not sensitive to repetitive stimulation when extracellular calcium was replaced by strontium. It disappeared when cadmium was applied. 6. In DRG somata changes of the intracellular Ca2+ concentration were monitored by measuring the fluorescence of the Ca2+ indicator Fluo-3 with a laser-scanning confocal microscope. During repetitive stimulation, an accumulation of intracellular calcium occurred that recovered very slowly (tens of seconds) after the AP trains. 7. Computer model simulations performed in analogy to the experimental protocols produced conduction failures during repetitive stimulation only when the calcium currents during the APs were reduced. 8. From these findings it is concluded that conduction failures during repetitive stimulation are dependent on an accumulation of intracellular calcium leading to an inactivation of calcium currents, combined with small contributions of an accumulation of extracellular potassium and a summation of slow potassium conductances.

Acetylcholine↗

Measurement of ventricular size: reliability of the frontal and occipital horn ratio compared to subjective assessment.

INTRODUCTION: The frontal and occipital horn ration (FOR) has recently been described as a simple, linear measurement of ventricular size that correlates very well with ventricular volume. This study further characterizes the measurement properties of the FOR by investigating its interobserver reliability and comparing it to a subjective assessment of ventricular size. METHODS: Axial images (CT and MR) of children with hydrocephalus taken before and after third ventriculostomy were reviewed by 4 independent observers. Two observers were blinded to patient identity and clinical status and 2 observers were nonblinded. Each observer independently recorded linear measurements from which the FOR was calculated for each image. Each reviewer also made a separate subjective assessment of the degree of hydrocephalus on a 9-point adjectival scale. Reliability was calculated using a repeated-measures analysis of variance (ANOVA) and an intraclass correlation coefficient (ICC) with random image and observer effects. RESULTS: There were 120 separate observations (4 observers, 30 images). The FOR ranged from 0.33 to 0.75 (mean 0.55, standard deviation 0.11). The reliability coefficient was 0.93 (95% confidence interval, CI 0.80-0.97) between the 2 blinded observers and 0.98 (95% CI, 0.95-0.99) between the 2 nonblinded observer. The overall interobserver reliability for all 4 observers was 0.95 (95% CI 0.92-0.98). The mean FOR for each observer was very similar, regardless of the observer's blinding status. However, the reliability of the observers' subjective assessment of the hydrocephalus was much lower (ICC = 0.77, 95% CI 0. 60-0.88). CONCLUSIONS: The FOR demonstrates excellent interobserver reliability (>0.9) and was superior to subjective assessments of hydrocephalus. In this study, excellent reliability was maintained regardless of the blinding status of the observers. This further demonstrates the properties of the FOR as a simple and reproducible measure of ventricular size. It is suitable for use in clinical studies, possibly even in situations in which observer blinding is not possible.

Cerebral Ventricles↗

A modified National Institutes of Health Stroke Scale for use in stroke clinical trials: preliminary reliability and validity.

BACKGROUND AND PURPOSE: The National Institutes of Health Stroke Scale (NIHSS) is accepted widely for measuring acute stroke deficits in clinical trials, but it contains items that exhibit poor reliability or do not contribute meaningful information. To improve the scale for use in clinical research, we used formal clinimetric analyses to derive a modified version, the mNIHSS. We then sought to demonstrate the validity and reliability of the new mNIHSS. METHODS: The mNIHSS was derived from our prior clinimetric studies of the NIHSS by deleting poorly reproducible or redundant items (level of consciousness, face weakness, ataxia, dysarthria) and collapsing the sensory item into 2 responses. Reliability of the mNIHSS was assessed with the certification data originally collected to assess the reliability of investigators in the National Institute of Neurological Disorders and Stroke (NINDS) rtPA (recombinant tissue plasminogen activator) Stroke TRIAL: Validity of the mNIHSS was assessed with the outcome results of the NINDS rtPA Stroke Trial: RESULTS: Reliability was improved with the mNIHSS: the number of scale items with poor kappa coefficients on either of the certification tapes decreased from 8 (20%) to 3 (14%) with the mNIHSS. With the use of factor analysis, the structure underlying the mNIHSS was found identical to the original scale. On serial use of the scale, goodness of fit coefficients were higher with the mNIHSS. With data from part I of the trial data, the proportion of patients who improved >/=4 points within 24 hours after treatment was statistically significantly increased by tPA (odds ratio, 1.3; 95% confidence limits, 1.0, 1.8; P=0.05). Likewise, the odds ratio for complete/nearly complete resolution of stroke symptoms 3 months after treatment was 1.7 (95% confidence limits, 1.2, 2.6) with the mNIHSS. Other outcomes showed the same agreement when the mNIHSS was compared with the original scale. The mNIHSS showed good responsiveness, ie, was useful in differentiating patients likely to hemorrhage or have a good outcome after stroke. CONCLUSIONS: The mNIHSS appears to be identical clinimetrically to the original NIHSS when the same data are used for validation and reliability. Power appears to be greater with the mNIHSS with the use of 24-hour end points, suggesting the need for fewer patients in trials designed to detect treatment effects comparable to rtPA. The mNIHSS contains fewer items and might be simpler to use in clinical research trials. Prospective analysis of reliability and validity, with the use of an independently collected cohort, must be obtained before the mNIHSS is used in a research setting.

Clinical Trials as Topic↗

An intervention to improve the reliability of manuscript reviews for the Journal of the American Academy of Child and Adolescent Psychiatry.

OBJECTIVE: The effects of methods used to improve the interrater reliability of reviewers' ratings of manuscripts submitted to the Journal of the American Academy of Child and Adolescent Psychiatry were studied. METHOD: Reviewers' ratings of consecutive manuscripts submitted over approximately 1 year were first analyzed; 296 pairs of ratings were studied. Intraclass correlations and confidence intervals for the correlations were computed for the two main ratings by which reviewers quantified the quality of the article: a 1-10 overall quality rating and a recommendation for acceptance or rejection with four possibilities along that continuum. Modifications were then introduced, including a multi-item rating scale and two training manuals to accompany it. Over the next year, 272 more articles were rated, and reliabilities were computed for the new scale and for the scales previously used. RESULTS: The intraclass correlation of the most reliable rating before the intervention was 0.27; the reliability of the new rating procedure was 0.43. The difference between these two was significant. The reliability for the new rating scale was in the fair to good range, and it became even better when the ratings of the two reviewers were averaged and the reliability stepped up by the Spearman-Brown formula. The new rating scale had excellent internal consistency and correlated highly with other quality ratings. CONCLUSIONS: The data confirm that the reliability of ratings of scientific articles may be improved by increasing the number of rating scale points, eliciting ratings of separate, concrete items rather than a global judgment, using training manuals, and averaging the scores of multiple reviewers.

Algorithms↗

Two-unit reliability analysis of questionnaires used in a regulatory system.

An interjudge reliability test was conducted to evaluate the questionnaires used in the surveillance of residential care institutions. Because the reliability test was carried out as part of the routine surveillance program and not as part of a controlled experiment, it was subject to deviations from the optimal reliability test model. However, this nonpure design provided an opportunity to not only examine the reliability of the items in the surveillance tool, but also to gain a better understanding of the use of a reliability test in an "imperfect" field setting. Two different surveyor teams administered the 257 questions on the questionnaires to a representative sample of 32 institutions on two separate occasions. In order to explain the variance in the reliability scores, a multivariate analysis was conducted for two units of analysis: the surveillance questions and the institutions. Based on the results of the reliability test, changes were introduced to improve the questionnaires and their administration.

Aged↗

A reliability study of the universal goniometer, fluid goniometer, and electrogoniometer for the measurement of ankle dorsiflexion.

This study investigated the reliability of three goniometers, the universal, fluid, and electro-goniometers, in the measurement of ankle dorsiflexion. Intra- and interobserver reliability were assessed using 10 healthy volunteers and five observers. A standardized ankle position was used to measure full range of active dorsiflexion. Intraobserver reliability was assessed using one observer over two successive occasions. Interobserver reliability was assessed among five observers over five separate occasions. A one-factor analysis of variance to examine intraobserver reliability demonstrated no significant difference between each of the devices on the two occasions. A multifactorial analysis of variance demonstrated significant differences among observers and again among devices (P < 0.001). Secondary analysis for interdevice reliability demonstrated significant differences among the three devices (p < 0.1). The study suggests that each device cannot be used reliably among observers or be used interchangeably, and clinical judgment based on angular changes of less than 10 degrees are invalid if rigid protocols are not followed.

Adult↗

Reliability of F-scan in-shoe measurements of plantar pressure.

Research by our group and others indicates that many amputations of the lower limb occur after foot ulceration in patients with diabetes. It has been proposed that diabetic foot ulcers are mainly caused by repetitive trauma in areas of high plantar pressure during walking. Recent technology permits in-shoe measurement of plantar pressure. We assessed the reliability of the F-Scan in-shoe system for measurement of plantar pressure (Tekscan Inc., Boston, MA) in 51 subjects from a cohort of 977 diabetic veterans enrolled in a prospective study of risk factors for foot ulceration and amputation (the Seattle Diabetic Foot Study). Subjects were tested twice, wearing their own shoes. We used the coefficient of variation (CV) and the intra-class correlation coefficient (ICC) to estimate the reliability of F-Scan measurements of pressure. Peak pressure over the metatarsal heads proved to have the best indices of reliability, with CVs of 0.150 and 0.155, and ICCs of 0.755 and 0.751. Coefficients of variation for the heel, whole foot, and hallux ranged from 0.148 to 0.240, with ICCs ranging from 0.493 to 0.832. By published standards, peak pressures over the metatarsal heads and right hallux met the criteria for excellent reliability. Our ICCs for high pressures under the foot, heel, metatarsal heads, and hallux, and for peak pressures under the heel and left hallux represented fair-to-good reliability. No F-Scan plantar measurements could be judged by these criteria as having poor reliability. This clinical study found that for elderly patients with diabetes who were wearing their own shoes and were tested on two different days with different insoles, the F-Scan insole system was generally reliable for measurements of high pressure and peak pressure.

Adult↗

Test-retest reliability of psychological and neurobehavioral tests self-administered by computer.

A series of 12 psychological and 7 neurobehavioral performance tests were administered twice to a nonclinical normative sample with 1 week between administrations. The tests were presented in a self-administered computerized format. One week test-retest reliabilities were comparable to conventional administration formats. The results suggest that individual test reliability is not affected when tests are administered as part of an extensive multi-measure battery. Computer administered test reliability coefficients also were compared to a Mixed Format (computer-conventional) administration with mixed format reliabilities generally similar to the reliabilities of published conventional tests but also generally lower than same format testing. Compared to psychological test reliability, neurobehavioral test reliability appeared more vulnerable to decreases with mixed format testing. These conclusions should not be generalized to all computer implemented tests as the qualities of the test implementation will affect the outcome.

Adult↗

Assessing reliability of a measure of self-rated health.

The test-retest reliability of self-rated health is analysed and compared with the reliability of health questions phrased more as well as less precisely. Differences in reliability between men and women and between age groups are also assessed. The study is based on 204 and 409 re-interviews from the 1991 Swedish Level of Living Survey and the 1989 Survey of Living Conditions respectively. The results show that the reliability of self-rated health is as good as or even better than that of most of the more specific questions. Only an indicator of high blood pressure showed significantly higher reliability. The reliability of self-rated health is good in all subgroups studied, and is even excellent among older men. It is concluded that the good overall reliability of self-rated health found in this study is in line with previous results concerning the validity of people's assessments of their general health as well as results concerning the basis upon which they make these judgements.

Adult↗

Validity and reliability of retrospective assessment of disease activity and flare in observational cohorts of lupus patients.

BACKGROUND: The use of validated retrospective tools to assess disease activity in an observational cohort of patients would allow researchers the flexibility to analyze unique exploratory questions. Valid, reliable tools exist for assessing disease activity and flare for Systemic Lupus Erythematosus (SLE) patients. However, these tools have been designed for use in structured settings. Many of these populations under study are subject to strict inclusion, exclusion criteria and disease management protocols. The ability to apply these tools to populations not subject to such control would allow researchers to explore questions about unique populations, new treatment applications or predictors of clinical outcomes. This study sought to establish the reliability and validity of retrospective medical record abstraction of SLE disease activity and flare instruments by two rheumatologists. METHODS: From university rheumatology outpatient clinics, 22 patients were randomly selected to establish intra-rater reliability and 26 patients were selected to establish inter-rater reliability. Two rheumatologists using a Physician Global Assessment (PGA), the Systemic Lupus Erythematosus Disease Activity Index (SLEDAI) and the Safety of Estrogens in Lupus Erythematosus-National Assessment (SELENA) flare tool retrospectively abstracted patient charts. Agreement between the tools was evaluated to assess validity. RESULTS: The mean patient age was 39 y; 96% were female, 54% Caucasian, 27% Hispanic, 19% Asian; median PGA was 1.4 (on a 0-3 scale) and median SLEDAI score was 4 (range 0-27). Intra-rater reliability for PGA, SLEDAI and the SELENA flare tool was 0.88, 0.87 and 0.52, respectively. Inter-rater reliability for PGA, SLEDAI, and SELENA flare was 0.79, 0.75, and 0. 50, respectively. To assess validity, the tools were compared against each other to assess agreement. From the parent study sample (n=54 patients), the disease activity measures, PGA and SLEDAI demonstrated adequate agreement (r=0.60). However, the SELENA flare tool demonstrated poor agreement with either PGA-defined flare or SLEDAI-defined flare (weighted kappa, 0.29 and 0.40 respectively). PGA-defined and SLEDAI-defined flare also demonstrated poor agreement (weighted kappa, 0.35). CONCLUSIONS: These data suggest that investigators can reliably reproduce patient disease activity through retrospective chart abstraction using PGA and SLEDAI. Assessing flare is a more difficult concept. The validity of assessing flare at a specific patient-visit is poor. Retrospective assessment of patient flare risk over a specified time period is conceptually more valid and avoids difficulties assessing timing and duration of flare. As there have been no similar published prospective analyses of validity using the SELENA flare tool, it is not clear if this problem was unique to the method of retrospective chart abstraction, the nature of non-protocol study patient visits, the tool itself or a combination of all three aspects.

Adult↗

An examination of interrater reliability for scoring the Rorschach Comprehensive System in eight data sets.

In this article, we describe interrater reliability for the Comprehensive System (CS; Exner. 1993) in 8 relatively large samples, including (a) students, (b) experienced re- searchers, (c) clinicians, (d) clinicians and then researchers, (e) a composite clinical sample (i.e., a to d), and 3 samples in which randomly generated erroneous scores were substituted for (f) 10%, (g) 20%, or (h) 30% of the original responses. Across samples, 133 to 143 statistically stable CS scores had excellent reliability, with median intraclass correlations of.85, .96, .97, .95, .93, .95, .89, and .82, respectively. We also demonstrate reliability findings from this study closely match the results derived from a synthesis of prior research, CS summary scores are more reliable than scores assigned to individual responses, small samples are more likely to generate unstable and lower reliability estimates, and Meyer's (1997a) procedures for estimating response segment reliability were accurate. The CS can be scored reliably, but because scoring is the result of coder skills clinicians must conscientiously monitor their accuracy.

Adult↗

Interrater reliability of Alzheimer's disease diagnosis.

To determine interrater reliability of dementia diagnosis, 4 physicians experienced in the evaluation of dementia patients applied 3 sets of diagnostic criteria to each of 62 patients, based on a standardized set of medical record information. All patients had undergone similar examinations and follow-up to establish the initial clinical diagnosis (76% had autopsy). Raters were blind to the diagnosis and to follow-up information after the initial evaluation period. This paper presents interrater agreement (kappa values) for a diagnosis of Alzheimer's disease using the American Psychiatric Association diagnostic criteria from the Diagnostic and Statistical Manual (DSM-III), the National Institute of Neurological and Communicative Disorders and Stroke (NINCDS) criteria for the clinical diagnosis of Alzheimer's disease, and the Eisdorfer and Cohen Research Diagnostic Criteria (ECRDC) for primary neuronal degeneration. The NINCDS showed somewhat higher average interrater reliability (kappa = 0.64) than the DSM-III (kappa = 0.55) and considerably higher interrater reliability than the ECRDC (kappa = 0.37). One rater displayed conspicuously lower levels of interrater reliability than the other 3, especially in DSM-III and ECRDC. This study indicates that interrater reliability of DSM-III and NINCDS criteria are comparable. Documentation of interrater reliability and, if necessary, training to improve reliability is an important consideration in research where different observers are diagnosing dementing illnesses.

Aged↗

Reliability of the NINDS Myotatic Reflex Scale.

The assessment of deep tendon reflexes is useful for localization and diagnosis of neurologic disorders, but only a few studies have evaluated their reliability. We assessed the reliability of four neurologists, instructed in two different countries, in using the National Institute of Neurological Disorders and Stroke (NINDS) Myotatic Reflex Scale. To evaluate the role of training in using the scale, the neurologists randomly and blindly evaluated a total of 80 patients, 40 before and 40 after a training session. Inter- and intraobserver reliability were measured with kappa statistics. Our results showed substantial to near-perfect intraobserver reliability, and moderate-to-substantial interobserver reliability of the NINDS Myotatic Reflex Scale. The reproducibility was better for reflexes in the lower than in the upper extremities. Neither educational background nor the training session influenced the reliability of our results. The NINDS Myotatic Reflex Scale has sufficient reliability to be adopted as a universal scale.

Adult↗

Reliability of and correlation between the respiratory therapist written registry and clinical simulation self-assessment examinations.

STUDY OBJECTIVES: The purpose of this study was to determine the reliability of two respiratory therapy self-assessment examinations: the written registry examination (WR), and the clinical simulation examination (CSE). We then used reliability coefficients to test the true correlation between the WR and CSE by employing the Spearman-Brown formula to attenuate for unreliability. DESIGN: This was a nonexperimental correlational study. SETTING: The study was conducted at respiratory therapy education programs located in four states. PARTICIPANTS: Sixty advanced-level respiratory therapy students enrolled in the final semester of their programs. MEASUREMENTS AND RESULTS: Fifty-eight students completed the WR, and 56 students completed the CSE. The reliability coefficient for the WR was 0.79. The reliability coefficient for the CSE when taken as a whole was 0.76. However, the CSE is separated into two sections, information gathering and decision making, which are scored separately. Cronbach alpha computed for the information-gathering section was 0.72, while the alpha coefficient for the decision-making section was only 0.64. The correlation between the WR and CSE was 0.86 after attenuation for reliability. CONCLUSIONS: The estimate of the reliability for the CSE is less than that for the WR, and the two examinations are strongly correlated. This leads us to question whether the CSE adds to the validity or reliability in the testing of respiratory therapists.

Credentialing↗

The Collaborative Longitudinal Personality Disorders Study: reliability of axis I and II diagnoses.

Both the interrater and test-retest-retest reliability of axis I and axis II disorders were assessed using the Structured Clinical Interview for DSM-IV Axis I Disorders (SCID-I) and the Diagnostic Interview for DSM-IV Personality Disorders (DIPD-IV). Fair-good median interrater kappa (.40-.75) were found for all axis II disorders diagnosed five times or more, except antisocial personality disorder (1.0). All of the test-retest kappa for axis II disorders, except for narcissistic personality disorder (1.0) and paranoid personality disorder (.39), were also found to be fair-good. Interrater and test-retest dimensional reliability figures for axis II were generally higher than those for their categorical counterparts; most were in the excellent range (> .75). In terms of axis I, excellent median interrater kappa were found for six of the 10 disorders diagnosed five times or more, whereas fair-good median interrater kappa were found for the other four axis I disorders. In general, test-retest reliability figures for axis I disorders were somewhat lower than the interrater reliability figures. Three test-retest kappa were in the excellent range, six were in the fair-good range, and one (for dysthymia) was in the poor range (.35). Taken together, the results of this study suggest that both axis I and axis II disorders can be diagnosed reliably when using appropriate semistructured interviews. They also suggest that the reliability of axis II disorders is roughly equivalent to that reliability found for most axis I disorders.

Diagnosis, Differential↗

Validity and reliability of self-reported drinking behavior: dealing with the problem of response bias.

This work assesses the validity and reliability of self-reported survey data on drinking behavior. There is evidence to suggest that data are adversely affected by bias from underreporting. This bias affects the validity of measures of consumption of alcohol and can have deleterious effects on the results of some forms of statistical estimation. Data for this study were collected at an isolated military base. The remoteness of this site and the fact that it is a military station made it possible to estimate the actual level of consumption of alcohol for the population by assessing apparent consumption through officially recorded sales of alcohol. The results of eight measures of consumption of alcohol were compared with apparent consumption, as established by documented sales, and the validity and reliability of the various measures were determined using the classical correlational approach. The validity and reliability of the data generated by the self-report survey were also analyzed using LISREL, the measurement model in particular. The results indicate that various instruments used to assess the consumption of alcohol produce very different outcomes in terms of their validity and reliability, some questions being considerably more valid and reliable than others. Two of the more salient characteristics of questions that affect validity and reliability were isolated, namely a question's ability to aid recall and its ability to mitigate the effects of persons providing socially desirable responses. The LISREL results show that these are two underlying factors for the measurement of the consumption of alcohol. It is concluded that questions that produce valid and reliable responses do so for identifiable reasons, and measurement instruments can be improved by incorporating particular features.

Alcohol Drinking↗

Stulberg classification system for evaluation of Legg-Calvé-Perthes disease: intra-rater and inter-rater reliability.

BACKGROUND: Researchers and clinicians commonly use the classification system of Stulberg et al. as a basis for treatment decisions during the active phase of Legg-Calvé-Perthes disease because of its putative utility as a predictor of long-term outcome. It is generally assumed that this system has an acceptable degree of reliability. This assumption, however, is not convincingly supported by the literature. METHODS: The purpose of the present study was to assess the inter-rater and intra-rater reliability of the classification system of Stulberg et al. with use of a pre-test, post-test design. During the pre-test phase, nine raters independently used the system to evaluate the radiographs of skeletally mature patients who had been managed for Legg-Calvé-Perthes disease. The intervention between the pre-test and post-test phases consisted of a consensus-building session during which all raters jointly arrived at standardized definitions of the various joint structures that are assessed with use of the classification system. The effect of these definitions on reliability then was assessed by reevaluating the radiographs during the post-test phase. RESULTS: The pre-test intra-rater reliability coefficients ranged from 0.709 to 0.915, and the post-test coefficients ranged from 0.568 to 0.874. The pre-test inter-rater reliability coefficients ranged from 0.603 to 0.732, and the post-test coefficients ranged from 0.648 to 0.744. Contributing to the variance was a lack of agreement concerning the assessment of joint structures and the way in which the raters translated these evaluations into a classification according to the system of Stulberg et al. CONCLUSIONS: Although intra-rater reliability was marginally acceptable, the degree of variability between the classifications assigned by different raters even after the intervention - calls into question the reliability of the system of Stulberg et al.; consequently, the validity of any treatment decisions, outcome evaluations, or epidemiological studies based on this system is also in question.

Acetabulum↗

Reliability of three classification systems measuring active motion in brachial plexus birth palsy.

BACKGROUND: Several classification systems for the categorization of function in patients with brachial plexus birth palsy have been proposed. The purpose of this investigation was to determine the intraobserver and interobserver reliability of the modified Mallet Classification, Toronto Test Score, and Hospital for Sick Children Active Movement Scale in the evaluation of these patients. METHODS: Eighty children with brachial plexus birth palsy were evaluated by two trained examiners on two different occasions. Intraobserver and interobserver reliability was determined with use of the kappa statistic. RESULTS: On the basis of the kappa statistic, intraobserver reliability was good to excellent for individual elements of the modified Mallet Classification, Toronto Test Score, and Active Movement Scale in all age-groups. Interobserver reliability for individual elements of these three systems ranged from fair to excellent. When aggregate Toronto Test and modified Mallet scores were assessed, positive intraobserver and interobserver correlations were noted (Pearson r = 0.70 to 0.98, p < 0.001). Internal consistency (test-retest reliability) as determined by the Cronbach alpha for the aggregate Toronto Test and modified Mallet scores was excellent for each age-group (alpha > 0.90, p < 0.001). CONCLUSIONS: The modified Mallet Classification, Toronto Test Score, and Active Movement Scale are reliable instruments for assessing upper-extremity function in patients with brachial plexus birth palsy. The natural history and surgical outcomes of these patients can now be conducted with use of these reliable outcomes instruments.

Adolescent↗