Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Development of a new seizure severity questionnaire: initial reliability and validity testing.

PURPOSE: This report describes the initial steps for development of a new scale to assess seizure severity as a treatment response. METHODS: Standard methodology was used to develop the test instrument. Item generation was performed by selecting items from other questionnaires, and asking patients and epileptologists about seizure components. Face and content validity were assessed in a pilot study with patients and observers. The questionnaire was formatted as a structured interview for a reliability study. Construct validity was assessed with three existing questionnaires. Inter-rater reliability and test-retest reliability were performed as two independent ratings on one day, and one re-interview. RESULTS: Based on item generation and pilot testing (33 patients, 28 observers), the Seizure Severity Questionnaire (SSQ) was organized into warning, activity-movement, and recovery (cognitive, emotional, and physical aspects) stages of seizures. Questions reviewed duration, severity, bothersomeness and overall ratings, and the most bothersome aspect of seizures. The mean SSQ Summary Score was 5.78+/-3.24, inter-rater reliability was 0.76 (N=91), and test-retest reliability was 0.74 (N=63). Construct validity showed statistically significant correlations with other scales. CONCLUSION: This study has explored the psychometric properties of the SSQ for face and content validity, inter-rater and test-retest reliability, and construct validity. The Summary Score reliably represents the major components of seizures.

Adult↗

An analysis of the reliability phenomenon in the FitzHugh-Nagumo model.

The reliability of single neurons on realistic stimuli has been experimentally confirmed in a wide variety of animal preparations. We present a theoretical study of the reliability phenomenon in the FitzHugh-Nagumo model on white Gaussian stimulation. The analysis of the model's dynamics is performed in three regimes-the excitable, bistable, and oscillatory ones. We use tools from the random dynamical systems theory, such as the pullbacks and the estimation of the Lyapunov exponents and rotation number. The results show that for most stimulus intensities, trajectories converge to a single stochastic equilibrium point, and the leading Lyapunov exponent is negative. Consequently, in these regimes the discharge times are reliable in the sense that repeated presentation of the same aperiodic input segment evokes similar firing times after some transient time. Surprisingly, for a certain range of stimulus intensities, unreliable firing is observed due to the onset of stochastic chaos, as indicated by the estimated positive leading Lyapunov exponents. For this range of stimulus intensities, stochastic chaos occurs in the bistable regime and also expands in adjacent parts of the excitable and oscillating regimes. The obtained results are valuable in the explanation of experimental observations concerning the reliability of neurons stimulated with broad-band Gaussian inputs. They reveal two distinct neuronal response types. In the regime where the first Lyapunov has negative values, such inputs eventually lead neurons to reliable firing, and this suggests that any observed variance of firing times in reliability experiments is mainly due to internal noise. In the regime with positive Lyapunov exponents, the source of unreliable firing is stochastic chaos, a novel phenomenon in the reliability literature, whose origin and function need further investigation.

Computer Simulation↗

The reliability of assessment of vibration sense.

OBJECTIVES: To assess the reliability of quantitative assessment of vibration sense with a Vibrameter type III. MATERIAL AND METHODS: We examined 111 healthy subjects (21-69 years). For intraobserver reliability, short-term (15 min between measurements) (n=11) and 24-h (n=28) reliability was tested. For interobserver reliability, a second tester performed the second measurement 15 min after the initial test (n=39). We also assessed the independent impact of effects of age, gender and height on vibration thresholds. RESULTS: In our study the intraobserver reliability is good [intraclass correlation coefficients (ICC) ranging from 0.55 to 0.99], whereas the interobserver reliability is moderate (ICC ranging from 0.32 to 0.88). Multiple linear regression analysis showed that age and--to a lesser extent--height was independently associated with the threshold values of the feet, but not with the thresholds of the hands. CONCLUSION: The use of a Vibrameter for measuring vibration thresholds in clinical practice and in multicentre studies is restricted because of the moderate interobserver reliability.

Adult↗

The reliability of observational data: I. Theories and methods for speech-language pathology.

Much research and clinical work in speech-language pathology depends on the validity and reliability of data gathered through the direct observation of human behavior. This paper reviews several definitions of reliability, concluding that behavior observation data are reliable if they, and the experimental conclusions drawn from them, are not affected by differences among observers or by other variations in the recording context. The theoretical bases of several methods commonly used to estimate reliability for observational data are reviewed, with examples of the use of these methods drawn from a recent volume of the Journal of Speech and Hearing Research (35, 1992). Although most recent research publications in speech-language pathology have addressed the issue of reliability for their observational data to some extent, most reliability estimates do not clearly establish that the data or the experimental conclusions were replicable or unaffected by differences among observers. Suggestions are provided for improving the usefulness of the reliability estimates published in speech-language pathology research.

Humans↗

The reliability of the wolf motor function test for assessing upper extremity function after stroke.

OBJECTIVE: To examine the reliability of the Wolf Motor Function Test (WMFT) for assessing upper extremity motor function in adults with hemiplegia. DESIGN: Interrater and test-retest reliability. SETTING: A clinical research laboratory at a university medical center. PATIENTS: A sample of convenience of 24 subjects with chronic hemiplegia (onset >1yr), showing moderate motor impairment. INTERVENTION: The WMFT includes 15 functional tasks. Performances were timed and rated by using a 6-point functional ability scale. The WMFT was administered to subjects twice with a 2-week interval between administrations. All test sessions were videotaped for scoring at a later time by blinded and trained experienced therapists. MAIN OUTCOME MEASURE: Interrater reliability was examined by using intraclass correlation coefficients and internal consistency by using Cronbach's alpha. RESULTS: Interrater reliability was.97 or greater for performance time and.88 or greater for functional ability. Internal consistency for test 1 was.92 for performance time and.92 for functional ability; for test 2, it was.86 for performance time and.92 for functional ability. Test-retest reliability was.90 for performance time and.95 for functional ability. Absolute scores for subjects were stable over the 2 test administrations. CONCLUSION: The WMFT is an instrument with high interrater reliability, internal consistency, test-retest reliability, and adequate stability.

Adolescent↗

Reliability of cervical range of motion using the OSI CA 6000 spine motion analyser on asymptomatic and symptomatic subjects.

Cervical range of motion (ROM) is evaluated in both clinical and research settings. This study's purpose was to determine if ROM data obtained with the OSI CA 6000 Spine Motion Analyser (SMA) from asymptomatic and symptomatic cervical subjects were reliable within and between testers. Cervical ROM was measured in all three planes in 30 adult asymptomatic and 20 adult symptomatic subjects. A standardized protocol was used to fit each subject with the OSI SMA cervical hardware. Subjects were tested in a seated position with the trunk stabilized. Subjects performed four trials of each pain-free cervical motion during testing. The hardware was completely removed and replaced by the same tester and ROM trials in all three planes were repeated for intratester asymptomatic and symptomatic reliability. The same procedure was completed by a second tester for asymptomatic intratester and intertester reliability. Repeated measures analysis of variance and intraclass correlation coefficients (ICC [2,1 and 2 k]) were used to analyse intra- and intertester reliability data. Intratester ICCs were 0.85 or higher (except for flexion 0.76) for asymptomatic subjects and 0. 87 or higher (except for flexion 0.68) for symptomatic subjects for all motions. Intertester ICCs were 0.88 or higher for all motions. Standard error of measurements were less than 3.92 degrees for all motions. Measures of cervical spinal ROM obtained with the OSI SMA showed good intertester reliablity for all motions, and good intratester reliability for all motions with the exception of the motion of flexion for one of the examiners, which showed moderate reliability.

Adult↗

Reliability and within subject variability of VE, VO2, heart rate and blood pressure during submaximum cycle ergometry.

To assess the reliability and within subject variability of steady-rate ventilation (VE), oxygen uptake (VO2), heart rate, systolic and diastolic blood pressure, 4 subjects exercised for 10 minutes at 3 work rates on a bicycle ergometer: 50 W, 125 W and 55% of maximum work rate (55% max). Each testing session included two work rates and only 2 testing sessions were scheduled per week. The order of the work rates was counterbalanced. In 8 to 10 weeks, 3 of the subjects completed 20 trials at 50 W while the fourth subject completed 11 trials, and all the subjects completed 10 trials at 125 W and 55% max. The within subject variability (S2w) was expressed as a percent of the mean steady-rate response. VO2 ranged from 21.2% to 27.5% of VO2max at 50 W, from 37.7% to 49.7% at 125 W and from 42.9% to 63.7% at 55% max. The S2w averaged 6.8% for VE, 4.3% for VO2, 3.2% for heart rate, 7.3% for systolic blood pressure and 10.5% for diastolic blood pressure. Reliability coefficients were calculated for the steady-rate scores by dividing the between subject variation by the total variation. The reliability was similar for VE, VO2 and heart rate and ranged from r = 0.69 to r = 0.97. Systolic and diastolic blood pressure reliabilities were lower and ranged from r = 0.27 and r = 0.80. In summary, the steady-rate ventilation, oxygen uptake and heart rate responses were reliable and consistent. The reliability of blood pressure was low. It is possible that this low reliability may result from variability in stroke volume or total peripheral resistance.

Adult↗

Reliability of chiropractic methods commonly used to detect manipulable lesions in patients with chronic low-back pain.

OBJECTIVE: To assess the intraexaminer and interexaminer reliability of a multidimensional spinal diagnostic method commonly used by chiropractors. DESIGN: An intraexaminer and interexaminer Latin square, repeated measures reliability study. The techniques of diagnosis under investigation included visual postural analysis, pain description by the patient, plain static erect x-ray film of the lumbar spine, leg length discrepancy, neurologic tests, motion palpation, static palpation, and orthopedic tests. PARTICIPANTS: Three experienced chiropractors examined 19 patients, and 2 experienced chiropractors examined 10 and 9 patients, respectively, who were suffering from chronic mechanical low-back pain. RESULTS: Intraexaminer reliability of the decision to manipulate a certain spinal segmental level was moderate (kappa = 0.47). The interexaminer agreement pooled across all spinal joints indicated fair agreement (kappa = 0.27). Interexaminer reliability for individual examiner pairs for the L4/L5 segmental level was slight (kappa = 0.09). At the L5/S1 level, the interexaminer reliability was fair (kappa = 0.25). For the sacroiliac joints, interexaminer reliability was slight (kappa = 0.04 and 0.14). CONCLUSION: This study of commonly used chiropractic diagnostic methods in patients with chronic mechanical low-back pain to detect manipulable lesions in the lower thoracic spine, lumbar spine, and the sacroiliac joints has revealed that the measures are not reproducible. The implementation of these examination techniques alone should not be seen by practitioners to provide reliable information concerning where to direct a manipulative procedure in patients with chronic mechanical low-back pain.

Adult↗

The kinematics of motion palpation and its effect on the reliability for cervical spine rotation.

BACKGROUND: The reliability of a test depends on its standardization. Instrumental measurement of the reproducibility of the test is an effective way to evaluate the level of standardization obtained. Improved standardization is believed to yield greater reliability. OBJECTIVE: The objectives of this study were to measure the technical ability of an examiner to reproduce the kinematics of motion palpation for cervical spine rotation and to evaluate the effect of standardization on the reliability of the test. DESIGN: A study of reproducibility of the kinematics of the test for cervical spine rotation was conducted by means of a computerized system of analysis of movement. The reliability when reproducibility was achieved was compared with reliability when it failed. RESULTS: The data collected enable us to establish a standardized protocol for the execution of the test. The standardized palpation is executed within 6 of inclination from the pure plane of rotation. The successful reproduction of the kinematics of the test raises its reliability to detect the presence of fixations (kappa raising from 0.337 and 0.352 to 0.682). CONCLUSIONS: A greater reliability, arising from a high level of reproducibility, enables us to document the advantages of the standardization of motion palpation in chiropractic.

Adult↗

A reliability study of clinical tooth wear measurements.

STATEMENT OF PROBLEM: Most studies examining tooth wear severity have been performed on dental casts. This indirect approach has limited applicability to dental practice because during the assessment of the casts, the identification of dentin exposure is difficult or even impossible. PURPOSE OF STUDY: The purpose of this study was to assess occlusal and incisal tooth wear clinically to determine the reliability of the assessment procedure and to establish the influence of selected relevant clinical variables (dental quadrant, tooth type, and severity of wear) on the reliability. MATERIAL AND METHODS: Forty-five volunteers (17 men, 28 women; mean age 33.7 +/- 10.7 years), 32 with temporomandibular disorders and 13 free from signs and symptoms of such disorders, were evaluated on 4 occasions. Two trained observers graded tooth wear at 2 different points in time with a 5-point ordinal scale developed for use in this study. The inter-rater and intra-rater reliability of the scale was expressed as Cohen's kappa. The influence of 2 clinical variables, dental quadrant and tooth type, on the values of kappa was tested with 1-way analysis of variance and post hoc Bonferroni tests. Probability levels of P< .05 were considered statistically significant. The influence of the final clinical variable, severity of wear, was assessed qualitatively. RESULTS: The overall values of the inter-rater and intra-rater reliability were substantial (kappa = 0.632 to 0.678). The clinical variable dental quadrant did not influence the kappa values, whereas the inter-rater reliability during the first session was better for incisors and canines than for premolars (1-way analysis of variance: F(3,23)=4.577, P=.012; post hoc Bonferroni tests: P=.030 and.036). Qualitative assessment of severity of wear indicated that the more advanced the tooth wear, the more reliably it could be graded. CONCLUSION: By means of the developed 5-point ordinal scale and within the limitations of this study, it was concluded that tooth wear can be assessed reliably in the clinical dental setting.

Adult↗

Reliability of speech and language therapists using therapy outcome measures.

The Therapy Outcome Measure (TOM) aims to provide Speech and Language Therapists (SLT) with a practical tool to measure outcomes of care by providing a quick and simple measure which can be used over time with patients and clients in a routine clinical setting. The TOM allows therapists to reflect their clinical judgement on the dimensions of impairment, disability/activity, handicap/participation and well-being on an 11-point ordinal scale. The purpose of this paper is to examine the reliability and the influences on reliability of SLT using this measure. Three studies are presented and give information on 73 SLT using the measure with different client groups. Study one assesses the degree of reliability of six SLT using the TOM following training and practice. Reliability was studied on three occasions and the results demonstrate the influence of training and practice. Study two included 56 SLT to examine reliability over a broader range of client groups and to investigate the effect of the SLT specializing on rating patients from within or -out that specialism. Eleven therapists were included in the study, which examined the influence of a specific training approach. The participating SLT achieved a substantial-moderate level of reliability; this was established in all domains. The degree of reliability achieved on the TOM was affected by some, but not extensive, training and experience.

Clinical Protocols↗

Structured interview and uniform assessment improves diagnostic reliability.

OBJECTIVE: To compare a Childhood Uniform Assessment Package (CUAP), including a computerized structured diagnosis, with routine assessment and treatment in public mental health settings. DATA SOURCES/STUDY SETTINGS: Data was collected prospectively on 250 children and adolescents in both public mental health inpatient and outpatient settings in a large metropolitan area and a rural area. STUDY DESIGN: Subjects were randomized to either routine assessment and treatment as usual (ATU) or ATU plus an additional "gold standard" assessment battery Childhood Uniform Assessment Package (CUAP). Outcome measures were taken at admission (baseline), discharge, and again 6 months later. METHODS: The study was conducted at a State Hospital (CUAP, n = 75; ATU, n = 75) and a Community Mental Health center (CUAP, n = 50; ATU, n = 50). The "gold standard" diagnostic process was established at the Children's Medical Center-Dallas. Research focused on a comparison of the CUAP diagnostic process to the existing diagnostic process (ATU) and the service delivery system of an inpatient and outpatient public sector clinical treatment setting. PRINCIPAL FINDINGS: A bachelor's level individual can be trained to administer a highly reliable diagnostic battery to meet a "gold standard," suggesting a possible cost-effective way to assist in diagnostic evaluations. Higher reliability was found between this standardized assessment package (CUAP) and inpatient physicians than for outpatient physicians. The highest interrater reliabilities were found for attention deficit and substance abuse disorders, less so for the other behavior disorders. The use of CUAP results in more reliable diagnoses in public settings than those provided by typical clinical staff by identifying mood and anxiety disorders (disorders with the lowest reliability) with better reliability. The addition of "gold standard" diagnostic assessments (CUAP) did not appear to affect length of stay, number of medication changes, use of seclusion or restraints, and other behavioral interventions in the inpatient setting. Outpatient follow-up services did not differ for CUAP versus ATU either. CONCLUSIONS: A standard uniform assessment package that includes a structured diagnostic instrument can improve overall diagnostic reliability but may not have a significant overall impact in clinical treatment strategies or outcomes without additional intervention to assure proper use of the information. A well-trained bachelor's level assistant can administer such a battery.

Adolescent↗

How valid and reliable are patient satisfaction data? An analysis of 195 studies.

OBJECTIVE: To assess the properties of validity and reliability of instruments used to assess satisfaction in a broad sample of health service user satisfaction studies, and to assess the level of awareness of these issues among study authors. DESIGN: Examination and analysis of 195 papers published in 1994 in 139 journals. The following databases were searched: British Nursing Index, CINAHL, EMBASE, MedLine, Popline, and PsycLIT. MAIN MEASURES: Number and types of strategies used for content, criterion, and construct validity, and for stability and internal consistency. Associations between validity/reliability and other study characteristics. RESULTS: Eighty-nine (46%) of the 195 studies reported some validity or reliability data; 76 reported some element of content validity; 14 reported criterion validity, with patient's intent to return the most commonly used criterion; four reported construct validity. Thirty-four studies reported internal consistency reliability, 31 of which used Cronbach's coefficient alpha; eight studies reported test-retest reliability. Only 11 studies (6% of the 181 quantitative studies) reported content validity and criterion or construct validity and reliability. 'New' instruments designed specifically for the reported study demonstrated significantly less evidence for reliability/validity than did 'old' instruments. CONCLUSION: With few exceptions, the study instruments in this sample demonstrated little evidence of reliability or validity. Moreover, study authors exhibited a poor understanding of the importance of these properties in the assessment of satisfaction. Researchers must be aware that this is poor research practice, and that lack of a reliable and valid assessment instrument casts doubt on the credibility of satisfaction findings.

Data Collection↗

Reliability, dependability, and precision of anthropometric measurements. The Second National Health and Nutrition Examination Survey 1976-1980.

The components of reliability for eight anthropometric measures were studied in 95 male and 134 female subjects from the Second National Health and Nutrition Examination Survey (NHANES II). The contributions to unreliability variance (Sr2) that occur as a result of measuring errors (Sp2, imprecision variance) and of intrasubject fluctuations in a measurement due to physiologic factors (Sd2, undependability) were estimated (Sr2 = Sp2 + Sd2). Unreliability was then related to the between-subject variance (S2) to estimate the reliability (R = 1 - (Sr2/S2)) of the measurement. Four of the anthropometric measurements (weight, height, sitting height, and arm circumference) had reliabilities in excess of R = 0.97. In the first three of these, measurement imprecision made up two thirds or less of unreliability, and undependability (Sd2) was stable by two weeks. Lesser but still acceptable reliabilities were obtained for triceps and subscapular skinfolds, bitrochanteric breadth, and elbow breadth (R = 0.81-0.95). For these variables imprecision (Sp2) was the major source of error. Furthermore, the unreliability (Sr2) between observers was twice as high or more than the unreliability within observers for these variables, evidence that imprecision (Sp2) is the single most important source of unreliability in these anthropometric measurements. Unreliability standard deviations of skinfolds increased in a linear manner with skinfold thickness corresponding to an unreliability coefficient of variation of 13-19 per cent. None of the other measurements showed such scale effects. Analyses of the kind suggested will help epidemiologists decide whether reliability can be increased by improving precision, and whether there is a need to improve reliability in the first place. Reliability appears to be adequate for all anthropometry in the NHANES II.

Adult↗

Reliability of personal interview data in a hospital-based case-control study.

Responses to interview questions were compared for concordance among 492 individuals interviewed more than once in a hospital-based case-control surveillance system in the United States, Canada, and Israel between 1976 and 1982. Reliability of the data was determined using the Kappa statistic and the intraclass correlation coefficient. Reliability was good to excellent for demographic factors, such as birthplace, and for medical conditions/procedures that require hospitalization or continuing medical care, such as hysterectomy. Reliability was fair to good for less serious or less well-defined medical conditions/procedures, such as cystic breast disease, and for current habits, such as daily coffee consumption. Regarding medication use, reliability was poor to fair for drugs taken intermittently, such as aspirin and penicillin, and good to excellent for drugs taken on a regular basis, such as oral contraceptives. As expected, medications were reported more consistently when duration of use was prolonged. The data were also analyzed according to two intervals between interviews (less than 1 year and greater than or equal to 1 year). For most factors, reliability was not materially affected by interval. Where differences were observed, reliability tended to be better when the second interview followed the first by less than 1 year. These results suggest that structured interviews administered to hospital patients by trained personnel can elicit reliable data on demographic and medical history factors.

Adult↗

Reliability and interrelations among serum sex hormones in postmenopausal women.

Serum sex hormones may be related to the risk of several diseases in postmenopausal women including osteoporosis, heart disease, and breast and endometrial cancer. For assessment of the relation of sex hormones to disease, the measurements should be reliable, valid, and practical. In this paper, the authors evaluated the short-term (4-week) and long-term (2-year) reliability of serum sex hormones and interrelations among serum sex hormones in white postmenopausal women recruited in Pittsburgh, Pennsylvania, 1981-1986. For comparison, the authors simultaneously evaluated the short- and long-term reliability of other commonly measured risk factors, i.e., lipids, lipoproteins, and blood pressure. Serum concentrations of estrone, estradiol, testosterone, and androstenedione were measured by extraction, column chromatography, and radioimmunoassay. Reliability was estimated by calculating the intraclass correlation coefficients (R) and their 95% confidence interval. About 50% of the estradiol levels were below the sensitivity of the assay and, therefore, these results should be interpreted with some caution. The intraclass correlation coefficient for testosterone was 0.92 (95% confidence interval 1.0-0.82), suggesting that a single measure may be reliable in characterizing women for epidemiologic research. Over 4 weeks, estrone could be measured more reliably (R = 0.72) than over 2 years (R = 0.56), but the variability over the long term was similar to that observed for other biologic variables, suggesting that, in situations where the relation between estrone and disease is fairly substantial, a single measure may be used. For estradiol and androstenedione, the intraclass correlations were small, indicating poor reproducibility and the need for more measurements. Estrone concentrations were 11 pg/ml or 46% higher in women with measurable estradiol. Estrone was also positively related to androstenedione concentrations (r = 0.33, p less than 0.001). Concentrations of estradiol are extremely low in postmenopausal women, and accordingly, there is a greater possibility of laboratory error. Since the data suggest that estrone levels can be more reliably measured and are, in fact, related to estradiol levels, it is possible that estrone levels may be used to indicate the total estrogen status of postmenopausal women.

Androstenedione↗

Assessment of the reliability of pediatric screening: a tool for occupational or physical therapists.

This study was conducted to determine interrater and test-retest reliability characteristics of the instrument, Pediatric Screening: A Tool for Occupational and Physical Therapists. This protocol was developed by two public school therapists to be used as a decision-making mechanism for systematically assessing the students' relative need for therapy services. The subjects were 75 children, aged 3 to 16 years, with various types and degrees of disability. Each was scored on the screening tool by three different school therapists within one week to determine interrater reliability. Each of the therapists also tested two or three of the children again several weeks later to determine test-retest reliability. Analysis of interrater reliability using the Spearman-Brown prediction formula showed total scores on the screening tool to be reliable at the .90 level. Test-retest reliability measurements using the Pearson product-moment correlation coefficients showed that total scores were highly correlated (r = .96; p less than .001). These measures indicated that the Pediatric Screening tool is a highly reliable instrument in terms of scoring between therapists and by individual therapists across time.

Adolescent↗

Item reliability of the Milani-Comparetti Motor Development Screening Test.

The purpose of this study was to determine the level of interobserver and test-retest reliability of the Milani-Comparetti Motor Development Screening Test. Sixty healthy children, aged 1 through 16 months, were videotaped during administration of the Milani-Comparetti test. Four pediatric physical therapists independently viewed each videotape and scored the responses. Interobserver reliability was determined by calculation of percentage of agreement and the G statistic between a primary observer and each therapist. Forty-three children were retested within one week by the initial tester to examine test-retest reliability. Test-retest reliability was determined by percentage of agreement of items between the two test sessions and using the Kappa statistic. Interobserver percentage of agreement for the individual items on the Milani-Comparetti test ranged from 79% to 98%. The G statistic was significant for all items indicating the high percentage-of-agreement values were not due merely to chance agreement. Test-retest agreement ranged from 80% to 100%. Using Kappa statistic guidelines, excellent test-retest reliability (K greater than .75) was found for 82% of the test items, with good reliability of the remaining items. Acceptable interobserver and test-retest reliability was found for all items on the Milani-Comparetti test. Use of the Milani-Comparetti test as a clinical screening tool for prediction or follow-up of motor development in children at risk for developmental delays requires further evaluation.

Child Development↗