Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Structured interview and uniform assessment improves diagnostic reliability.

OBJECTIVE: To compare a Childhood Uniform Assessment Package (CUAP), including a computerized structured diagnosis, with routine assessment and treatment in public mental health settings. DATA SOURCES/STUDY SETTINGS: Data was collected prospectively on 250 children and adolescents in both public mental health inpatient and outpatient settings in a large metropolitan area and a rural area. STUDY DESIGN: Subjects were randomized to either routine assessment and treatment as usual (ATU) or ATU plus an additional "gold standard" assessment battery Childhood Uniform Assessment Package (CUAP). Outcome measures were taken at admission (baseline), discharge, and again 6 months later. METHODS: The study was conducted at a State Hospital (CUAP, n = 75; ATU, n = 75) and a Community Mental Health center (CUAP, n = 50; ATU, n = 50). The "gold standard" diagnostic process was established at the Children's Medical Center-Dallas. Research focused on a comparison of the CUAP diagnostic process to the existing diagnostic process (ATU) and the service delivery system of an inpatient and outpatient public sector clinical treatment setting. PRINCIPAL FINDINGS: A bachelor's level individual can be trained to administer a highly reliable diagnostic battery to meet a "gold standard," suggesting a possible cost-effective way to assist in diagnostic evaluations. Higher reliability was found between this standardized assessment package (CUAP) and inpatient physicians than for outpatient physicians. The highest interrater reliabilities were found for attention deficit and substance abuse disorders, less so for the other behavior disorders. The use of CUAP results in more reliable diagnoses in public settings than those provided by typical clinical staff by identifying mood and anxiety disorders (disorders with the lowest reliability) with better reliability. The addition of "gold standard" diagnostic assessments (CUAP) did not appear to affect length of stay, number of medication changes, use of seclusion or restraints, and other behavioral interventions in the inpatient setting. Outpatient follow-up services did not differ for CUAP versus ATU either. CONCLUSIONS: A standard uniform assessment package that includes a structured diagnostic instrument can improve overall diagnostic reliability but may not have a significant overall impact in clinical treatment strategies or outcomes without additional intervention to assure proper use of the information. A well-trained bachelor's level assistant can administer such a battery.

Adolescent↗

How valid and reliable are patient satisfaction data? An analysis of 195 studies.

OBJECTIVE: To assess the properties of validity and reliability of instruments used to assess satisfaction in a broad sample of health service user satisfaction studies, and to assess the level of awareness of these issues among study authors. DESIGN: Examination and analysis of 195 papers published in 1994 in 139 journals. The following databases were searched: British Nursing Index, CINAHL, EMBASE, MedLine, Popline, and PsycLIT. MAIN MEASURES: Number and types of strategies used for content, criterion, and construct validity, and for stability and internal consistency. Associations between validity/reliability and other study characteristics. RESULTS: Eighty-nine (46%) of the 195 studies reported some validity or reliability data; 76 reported some element of content validity; 14 reported criterion validity, with patient's intent to return the most commonly used criterion; four reported construct validity. Thirty-four studies reported internal consistency reliability, 31 of which used Cronbach's coefficient alpha; eight studies reported test-retest reliability. Only 11 studies (6% of the 181 quantitative studies) reported content validity and criterion or construct validity and reliability. 'New' instruments designed specifically for the reported study demonstrated significantly less evidence for reliability/validity than did 'old' instruments. CONCLUSION: With few exceptions, the study instruments in this sample demonstrated little evidence of reliability or validity. Moreover, study authors exhibited a poor understanding of the importance of these properties in the assessment of satisfaction. Researchers must be aware that this is poor research practice, and that lack of a reliable and valid assessment instrument casts doubt on the credibility of satisfaction findings.

Data Collection↗

Reliability, dependability, and precision of anthropometric measurements. The Second National Health and Nutrition Examination Survey 1976-1980.

The components of reliability for eight anthropometric measures were studied in 95 male and 134 female subjects from the Second National Health and Nutrition Examination Survey (NHANES II). The contributions to unreliability variance (Sr2) that occur as a result of measuring errors (Sp2, imprecision variance) and of intrasubject fluctuations in a measurement due to physiologic factors (Sd2, undependability) were estimated (Sr2 = Sp2 + Sd2). Unreliability was then related to the between-subject variance (S2) to estimate the reliability (R = 1 - (Sr2/S2)) of the measurement. Four of the anthropometric measurements (weight, height, sitting height, and arm circumference) had reliabilities in excess of R = 0.97. In the first three of these, measurement imprecision made up two thirds or less of unreliability, and undependability (Sd2) was stable by two weeks. Lesser but still acceptable reliabilities were obtained for triceps and subscapular skinfolds, bitrochanteric breadth, and elbow breadth (R = 0.81-0.95). For these variables imprecision (Sp2) was the major source of error. Furthermore, the unreliability (Sr2) between observers was twice as high or more than the unreliability within observers for these variables, evidence that imprecision (Sp2) is the single most important source of unreliability in these anthropometric measurements. Unreliability standard deviations of skinfolds increased in a linear manner with skinfold thickness corresponding to an unreliability coefficient of variation of 13-19 per cent. None of the other measurements showed such scale effects. Analyses of the kind suggested will help epidemiologists decide whether reliability can be increased by improving precision, and whether there is a need to improve reliability in the first place. Reliability appears to be adequate for all anthropometry in the NHANES II.

Adult↗

Reliability of personal interview data in a hospital-based case-control study.

Responses to interview questions were compared for concordance among 492 individuals interviewed more than once in a hospital-based case-control surveillance system in the United States, Canada, and Israel between 1976 and 1982. Reliability of the data was determined using the Kappa statistic and the intraclass correlation coefficient. Reliability was good to excellent for demographic factors, such as birthplace, and for medical conditions/procedures that require hospitalization or continuing medical care, such as hysterectomy. Reliability was fair to good for less serious or less well-defined medical conditions/procedures, such as cystic breast disease, and for current habits, such as daily coffee consumption. Regarding medication use, reliability was poor to fair for drugs taken intermittently, such as aspirin and penicillin, and good to excellent for drugs taken on a regular basis, such as oral contraceptives. As expected, medications were reported more consistently when duration of use was prolonged. The data were also analyzed according to two intervals between interviews (less than 1 year and greater than or equal to 1 year). For most factors, reliability was not materially affected by interval. Where differences were observed, reliability tended to be better when the second interview followed the first by less than 1 year. These results suggest that structured interviews administered to hospital patients by trained personnel can elicit reliable data on demographic and medical history factors.

Adult↗

Reliability and interrelations among serum sex hormones in postmenopausal women.

Serum sex hormones may be related to the risk of several diseases in postmenopausal women including osteoporosis, heart disease, and breast and endometrial cancer. For assessment of the relation of sex hormones to disease, the measurements should be reliable, valid, and practical. In this paper, the authors evaluated the short-term (4-week) and long-term (2-year) reliability of serum sex hormones and interrelations among serum sex hormones in white postmenopausal women recruited in Pittsburgh, Pennsylvania, 1981-1986. For comparison, the authors simultaneously evaluated the short- and long-term reliability of other commonly measured risk factors, i.e., lipids, lipoproteins, and blood pressure. Serum concentrations of estrone, estradiol, testosterone, and androstenedione were measured by extraction, column chromatography, and radioimmunoassay. Reliability was estimated by calculating the intraclass correlation coefficients (R) and their 95% confidence interval. About 50% of the estradiol levels were below the sensitivity of the assay and, therefore, these results should be interpreted with some caution. The intraclass correlation coefficient for testosterone was 0.92 (95% confidence interval 1.0-0.82), suggesting that a single measure may be reliable in characterizing women for epidemiologic research. Over 4 weeks, estrone could be measured more reliably (R = 0.72) than over 2 years (R = 0.56), but the variability over the long term was similar to that observed for other biologic variables, suggesting that, in situations where the relation between estrone and disease is fairly substantial, a single measure may be used. For estradiol and androstenedione, the intraclass correlations were small, indicating poor reproducibility and the need for more measurements. Estrone concentrations were 11 pg/ml or 46% higher in women with measurable estradiol. Estrone was also positively related to androstenedione concentrations (r = 0.33, p less than 0.001). Concentrations of estradiol are extremely low in postmenopausal women, and accordingly, there is a greater possibility of laboratory error. Since the data suggest that estrone levels can be more reliably measured and are, in fact, related to estradiol levels, it is possible that estrone levels may be used to indicate the total estrogen status of postmenopausal women.

Androstenedione↗

Assessment of the reliability of pediatric screening: a tool for occupational or physical therapists.

This study was conducted to determine interrater and test-retest reliability characteristics of the instrument, Pediatric Screening: A Tool for Occupational and Physical Therapists. This protocol was developed by two public school therapists to be used as a decision-making mechanism for systematically assessing the students' relative need for therapy services. The subjects were 75 children, aged 3 to 16 years, with various types and degrees of disability. Each was scored on the screening tool by three different school therapists within one week to determine interrater reliability. Each of the therapists also tested two or three of the children again several weeks later to determine test-retest reliability. Analysis of interrater reliability using the Spearman-Brown prediction formula showed total scores on the screening tool to be reliable at the .90 level. Test-retest reliability measurements using the Pearson product-moment correlation coefficients showed that total scores were highly correlated (r = .96; p less than .001). These measures indicated that the Pediatric Screening tool is a highly reliable instrument in terms of scoring between therapists and by individual therapists across time.

Adolescent↗

Item reliability of the Milani-Comparetti Motor Development Screening Test.

The purpose of this study was to determine the level of interobserver and test-retest reliability of the Milani-Comparetti Motor Development Screening Test. Sixty healthy children, aged 1 through 16 months, were videotaped during administration of the Milani-Comparetti test. Four pediatric physical therapists independently viewed each videotape and scored the responses. Interobserver reliability was determined by calculation of percentage of agreement and the G statistic between a primary observer and each therapist. Forty-three children were retested within one week by the initial tester to examine test-retest reliability. Test-retest reliability was determined by percentage of agreement of items between the two test sessions and using the Kappa statistic. Interobserver percentage of agreement for the individual items on the Milani-Comparetti test ranged from 79% to 98%. The G statistic was significant for all items indicating the high percentage-of-agreement values were not due merely to chance agreement. Test-retest agreement ranged from 80% to 100%. Using Kappa statistic guidelines, excellent test-retest reliability (K greater than .75) was found for 82% of the test items, with good reliability of the remaining items. Acceptable interobserver and test-retest reliability was found for all items on the Milani-Comparetti test. Use of the Milani-Comparetti test as a clinical screening tool for prediction or follow-up of motor development in children at risk for developmental delays requires further evaluation.

Child Development↗

Reliability of goniometric measurements and visual estimates of knee range of motion obtained in a clinical setting.

The purpose of this study was to examine the intratester and intertester reliability for goniometric measurements of knee flexion and extension passive range of motion (PROM). In addition, parallel-forms reliability for PROM measurements of the knee obtained by use of a goniometer and by visual estimation was examined. The intertester reliability for visual estimates of the PROM of the knee was also examined. Repeated measurements were obtained on 43 patients in a clinical setting. The intraclass correlation coefficients (ICCs) for intratester reliability of measurements obtained with a goniometer were .99 for flexion and .98 for extension. Intertester reliability for measurements obtained with a goniometer was .90 for flexion and .86 for extension. The ICCs for parallel-forms reliability for measurements obtained with a goniometer and by visual estimation ranged from .82 to .94. The intertester reliability for measurements obtained by visual estimation was .83 for flexion and .82 for extension. Results suggest clinicians should use a goniometer to take repeated PROM measurements of a patient's knee to minimize the error associated with these measurements.

Adult↗

Facioscapulohumeral dystrophy natural history study: standardization of testing procedures and reliability of measurements. The FSH DY Group.

BACKGROUND AND PURPOSE: The natural history of facioscapulohumeral muscular dystrophy (FSHD) has not been studied prospectively. Knowledge of the natural progression of any disease provides essential information for the design of clinical trials. We present a protocol for the study of the natural history of FSHD using quantitative muscle testing (QMT), manual muscle testing (MMT), and functional testing. SUBJECTS: Thirty-two persons with FSHD (mean age = 36.1 years, SD = 9.6, range = 17-49) and 32 age- and gender-matched volunteer controls (mean age = 35.8 years, SD = 8.0, range = 23-50) served as subjects. METHODS: Using standardized testing procedures, we examined intrarater reliability of the MMT, QMT, and functional testing measurements in both groups. We also examined interrater reliability in 7 subjects with FSHD. Eighteen muscle groups were tested for each subject using QMT and MMT. RESULTS: Intraclass correlation coefficient (ICC) values ranged from .86 to .99 for intrarater reliability and from .86 to .99 for interrater reliability of QMT measurements. Weighted kappa values of .81 to .98 for intrarater reliability and .50 to 1.00 for interrater reliability were obtained for MMT measurements. Intrarater ICCs for various functional testing measures ranged from .60 to .97. In addition, the comparability of the two QMT machines used in the study was demonstrated by testing the same set of volunteer controls on each machine's linear force transducer (ICC = .89-.98). CONCLUSION AND DISCUSSION: We conclude that this standardized testing protocol produces reliable measurements of muscle strength and functional ability in subjects with FSHD.

Adolescent↗

Measurement of scapular asymetry and assessment of shoulder dysfunction using the Lateral Scapular Slide Test: a reliability and validity study.

BACKGROUND AND PURPOSE: The Lateral Scapular Slide Test (LSST) is used to determine scapular position with the arm abducted 0, 45, and 90 degrees in the coronal plane. Assessment of scapular position is based on the derived difference measurement of bilateral scapular distances. The purpose of this study was to assess the reliability of measurements obtained using the LSST and whether they could be used to identify people with and without shoulder impairments. Subjects. Forty-six subjects ranging in age from 18 to 65 years (X=30.0, SD=11.1) participated in this study. One group consisted of 20 subjects being treated for shoulder impairments, and one group consisted of 26 subjects without shoulder impairments. METHODS: Two measurements in each test position were obtained bilaterally. From the bilateral measurements, we derived the difference measurement. Intraclass correlation coefficients (ICC [1,1]) and the standard error of measurement (SEM) were calculated for intrarater and interrater reliability of the difference in side-to-side measures of scapular distance. Sensitivity and specificity of the LSST for classifying subjects with and without shoulder impairments were also determined. RESULTS: The ICCs for intrarater reliability were .75, .77, and .80 and .52, .66, and .62, respectively, for subjects without and with shoulder impairments in 0, 45, and 90 degrees of abduction. The ICCs for interrater reliability were .67, .43, and .74 and .79, .45, and .57, respectively, for subjects without and with shoulder impairments in 0,45 and 90 degrees of abduction. The SEMs ranged from 0.57 to 0.86 cm for intrarater reliability and from 0.79 to 1.20 cm for interrater reliability. Using the criterion of greater than 1.0 cm difference, sensitivity and specificity were 35% and 48%, 41% and 54%, and 43% and 56%, respectively, for 0, 45, and 90 degrees of abduction. Sensitivity and specificity based on the criterion of greater than 1.5 cm difference were 28% and 53%, 50% and 58%, and 34% and 52%, respectively, for the 3 scapular positions. CONCLUSION AND DISCUSSION: Our results suggest that measurements of scapular positioning based on the difference in side-to-side scapular distance measures are not reliable. Furthermore, the results suggest that sensitivity and specificity of the LSST measurements are poor and that the LSST should not be used to identify people with and without shoulder dysfunction.

Adult↗

A comparison of five low back disability questionnaires: reliability and responsiveness.

BACKGROUND AND PURPOSE: The aim of this study was to examine 5 commonly used questionnaires for assessing disability in people with low back pain. The modified Oswestry Disability Questionnaire, the Quebec Back Pain Disability Scale, the Roland-Morris Disability Questionnaire, the Waddell Disability Index, and the physical health scales of the Medical Outcomes Study 36-Item Short-Form Health Survey (SF-36) were compared in patients undergoing physical therapy for low back pain. SUBJECTS AND METHODS: Patients with low back pain completed the questionnaires during initial consultation with a physical therapist and again 6 weeks later (n=106). Test-retest reliability was examined for a group of 47 subjects who were classified as "unchanged" and a subgroup of 16 subjects who were self-rated as "about the same." Responsiveness was compared using standardized response means, receiver operating characteristic curves, and the proportions of subjects who changed by at least as much as the minimum detectable change (MDC) (90% confidence interval [CI] of the standard error for repeated measures). Scale width was judged as adequate if no more than 15% of the subjects had initial scores at the upper or lower end of the scale that were insufficient to allow change to be reliably detected. RESULTS: Intraclass correlation coefficients (2,1) calculated to measure reliability for the subjects who were classified as "unchanged" and those who were self-rated as "about the same" were greater than.80 for the Oswestry and Quebec questionnaires and the SF-36 Physical Functioning scale and less than.80 for the Waddell and Roland-Morris questionnaires and the SF-36 Role Limitations-Physical and Bodily Pain scales. None of the scales were more responsive than any other. DISCUSSION AND CONCLUSION: Measurements obtained with the modified Oswestry Disability Questionnaire, the SF-36 Physical Functioning scale, and the Quebec Back Pain Disability Scale were the most reliable and had sufficient width scale to reliably detect improvement or worsening in most subjects. The reliability of measurements obtained with the Waddell Disability Index was moderate, but the scale appeared to be insufficient to recommend it for clinical application. The Roland-Morris Disability Questionnaire and the Role Limitations-Physical and Bodily Pain scales of the SF-36 appeared to lack sufficient reliability and scale width for clinical application.

Adult↗

The neurologic and adaptive capacity score is not a reliable method of newborn evaluation.

BACKGROUND: The Neurologic and Adaptive Capacity Score (NACS) is a multi-item scale that was published in 1982 to measure the effects of intrapartum drugs on the neonate. Although this scoring system has been widely used in obstetric anesthesia research, studies confirming its reliability have not been published. The purpose of this study was to assess the reliability of the NACS. METHODS: Two teams of observers were trained to perform the NACS on healthy, term neonates born in the vertex presentation. Two examinations were performed on each neonate within the first 2.5 h of life. Simultaneous (or "split-half") reliability was assessed using the alpha coefficient. Test-retest reliability was assessed using the intraclass correlation coefficient. The test was considered to be reliable if a was greater than 0.7 and the intraclass correlation coefficient was greater than 0.6. RESULTS: Two hundred babies were studied. The a was 0.47 and the intraclass correlation coefficient was 0.38 (95% confidence interval, 0.24-0.52). CONCLUSIONS: The NACS had poor reliability both on simultaneous testing and in the test-retest situation when used to evaluate term, healthy neonates. The authors suggest that other measures need to be developed to evaluate the effect of intrapartum drug administration in the neonate. Health measurement scales should undergo rigorous assessment for reliability and validity before they are used in clinical practice or for research purposes.

Adult↗

Description and interobserver reliability of the Tufts Assessment of Motor Performance.

This paper describes the conceptual basis for the development of a new clinical evaluation instrument, the Tufts Assessment of Motor Performance (TAMP). The TAMP is a 32-item, diagnosis-independent, criterion-referenced test that samples physical performance items in the areas of mobility, activities of daily living and physical aspects of communication. The administrative and scoring criteria of the TAMP are presented, and the multiple measurement dimensions are described. The documentation of patient status and progress, as described in the functional and performance profiles, is outlined. The paper also reports initial interobserver reliability on the intraitem tasks and the summary indexes of the two profiles. Forty individuals (20 adults and 20 children) with neurologic and musculoskeletal disorders comprised the reliability sample. Kappa and intraclass correlations were used to estimate the reliability of three independent raters on individual tasks and aggregate scores, respectively. Task reliability for the assistance and approach measurement dimensions were generally higher than for the more qualitative pattern and proficiency dimensions. Yet over 90% of all the tasks had acceptable reliability, while all the summary indexes had high interobserver reliability. Determination of interobserver reliability data is the initial phase of defining the most appropriate and technically valuable items, and will serve as a basis for item revision and reduction to enhance the clinical utility of the test.

Activities of Daily Living↗

Reliability model from the in vitro durability tests of a left ventricular assist system.

A reliability test of the Novacor N100PC left ventricular assist system (LVAS) with valved conduits, including a pump/drive unit with compact controller and LVAS monitor was performed. The initial test objective was to demonstrate sufficient reliability for clinical use as a long-term circulatory support system. The subsequent objective, a test to failure, was intended to provide an assessment of the durability of the design and to determine the LVAS wearout modes. Testing began in April 1993 and was performed with 12 systems on gravity-feed mock circulatory loops. The pump/ drive units were submersed in body temperature saline for the duration of the test. Each of the LVAS units was operated at nominal afterloads of 75, 90, and 105 mm Hg, with test conditions varied to yield nominal pump outputs of 5.6, 7.1, and 8.3 L/min. Failure was defined as the inability of the LVAS to maintain an average pump output of 4 L/min or an average output pressure of 60 mm Hg. After 3 years, all systems remained on test, with durations of 2.3 to 3.0 years. Analysis of the testing to that date, using a constant hazard rate model, indicated a minimum demonstrated reliability of 94% at a 60% confidence level, or 86% at a 90% confidence level, for a 2 year mission time. This greatly surpasses the reliability level included in the STS-ASAIO Long-Term Mechanical Circulatory Support System Reliability Recommendation (80% reliability, 60% confidence level for a 1 year mission time). In the subsequent test-to-failure protocol, all systems ran failure-free for at least 3 years. System failures occurring at longer durations were caused by a single common cause: wear of the energy converter's armature support bearings and shafts. The wearout mode was gradual and could be diagnosed noninvasively before failure. An analysis using a Weibull model was performed, using the test durations of those devices that failed, those that were electively removed from test for analysis of the wear mode, and those that continued on test. As of April 1998, the test results showed a reliability, at a 60% confidence level, of >99.9% for a 1 year mission time, 99.5% for a 2 year mission, and 92.0% for a 3 year mission (>99.9%, 98.3%, and 85.9% for equivalent mission times, at a 90% confidence level). Systems continue on test after as long as 4.9 years.

Equipment Failure Analysis↗

Reliability of transient-evoked otoacoustic emissions.

OBJECTIVE: This investigation addressed four factors affecting transient-evoked otoacoustic emission (TEOAE) reliability: 1) The effect of evoking-stimulus level, 2) the effect of analyzing bandwidth, 3) the effect of slight-mild hearing loss, and 4) the effect of variability in the stimulus spectrum. DESIGN: TEOAEs at 80, 74, 68, and 62 dB pSPL evoking-stimulus levels were measured in 25 ears spanning a range of hearing levels from normal to mild hearing loss for a minimum of 10 test sessions. Reliability was assessed for 1/6-, 1/3-, 1/2-, and 1-octave analyzing bandwidths. RESULTS: Evoking-stimulus level, hearing loss, and center frequency did not significantly affect reliability. With decreasing analyzing bandwidth, reliability decreased. Intrasubject test-retest standard deviations were 1.2 dB for a broadband analyzing bandwidth and 1.4, 1.5, 1.6, and 1.8 dB for 1-, 1/2-, 1/3-, and 1/6-octave analyzing bandwidths, respectively. Stimulus variability within narrower bandwidths was of sufficient magnitude to influence test-retest reliability, and attempts to correct for the variations in stimulus spectrum were unsuccessful. Slopes of the input-output functions differed across frequencies, with shallower slopes at higher frequencies. CONCLUSIONS: In general, TEOAE amplitude is highly reliable. For those individuals in this study who were more variable, the variability was at low frequencies or across the entire frequency spectrum. For clinical applications, the choice of analyzing bandwidth should be based on consideration of both frequency specificity (where narrow analyzing bandwidths are optimal) and reliability (where wide analyzing bandwidths are optimal).

Acoustic Impedance Tests↗

Strategies to manipulate reliability: impact on statistical associations.

OBJECTIVE: To examine the effects of improving measurement reliability on associations between risk factors and childhood psychiatric disorder. METHOD: Data were from a general population sample of parents (N = 211) with children aged 6 to 16 years. Reliability of measurement was improved in three ways: by increasing the number of items in a scale (internal-consistency reliability), by averaging assessments of the same variables collected on two different occasions, and by constructing latent variable measures. To assess the effects of improving reliability, selected risk factors were regressed on parental assessments of childhood oppositional defiant disorder (ODD) and overanxious disorder (OAD). RESULTS: Improving reliability led to systematic increases in the magnitude of standardized regression coefficients between family dysfunction and ODD (beta = .30-.51) and between family dysfunction and OAD (beta = .24-.48). In multiple regression, improving reliability served to strengthen the specificity of associations between ODD, OAD, and family dysfunction and maternal depressed mood. Although latent variable methods produced the largest associations, the standard errors of these estimates were also larger, resulting in wider confidence intervals and slightly larger significance values. CONCLUSIONS: Improving reliability of measurement results in larger associations between risk factors and childhood disorder and may increase the opportunity of revealing differential associations between variables.

Adolescent↗

Patient rating of wrist pain and disability: a reliable and valid measurement tool.

OBJECTIVE: The goal of this study was to develop a reliable and valid tool for quantifying patient-rated wrist pain and disability. DESIGN: Survey, tool development, reliability, and validity study. SETTING: Upper extremity unit. PARTICIPANTS: One hundred members of the International Wrist Investigators were surveyed by mail to assist in development of the scale. Patients with distal radius (n = 64) or scaphoid (n = 35) fractures were enrolled in a reliability study, and 101 patients with distal radius fractures were enrolled in a validity study. INTERVENTION: Information from the expert survey, biomechanical literature, and patient interviews was used as a basis for item generation and definition of structural limitations for a scale that would be practical in the clinic. Patients with distal radius or scaphoid fractures completed the Patient-Rated Wrist Evaluation (PRWE) on two occasions to determine test-retest reliability. Patients with distal radius fractures (n = 101) completed the PRWE and the SF-36 and were tested with traditional impairment measures at baseline and at two, three, and six months after fracture to determine construct and criterion validity. MAIN OUTCOME MEASURES: Reliability coefficients (ICCs) and validity correlations (Pearson product moment correlations). RESULTS: Patient opinions on pain and on ability to do activities of daily living and work were thought to be the most important dimensions to include in subjective outcome tools. Brevity and simplicity were seen as essential in the clinic environment. A fifteen-item questionnaire (the PRWE) was designed to measure wrist pain and disability. Test-retest reliability was excellent (ICCs > 0.90). Validity assessment demonstrated that the instrument detected significant differences over time (p < 0.01) and was appropriately correlated with alternate forms of assessing parameters of pain and disability. CONCLUSIONS: The PRWE provides a brief, reliable, and valid measure of patient-rated pain and disability.

Activities of Daily Living↗

Radiographic fracture assessments: which ones can we reliably make?

OBJECTIVE: To identify the fracture characteristics that can be reliably assessed by analysis of plain radiographs of tibial plateau fractures. DESIGN: Radiographic review study. PARTICIPANTS: Five orthopaedic traumatologists served as observers. INTERVENTION: Observers made assessments based on the radiographs of fifty-six tibial plateau fractures. Precise definitions of the assessments to be made were agreed on by all observers. The tested assessments included raters' abilities to identify and locate fracture lines, identify the presence of fracture displacement and comminution, make quantitative measurements of displacement, and characterize qualitative features of fractures. For thirty-eight of the fractures that had a computed tomography (CT) scan available, assessments were repeated using both radiographs and CT scans. MAIN OUTCOME MEASURES: To characterize interobserver reliability, percentage agreement and kappa statistics were calculated for categorical variables, and intraclass correlation coefficients (ICC) were calculated for noncategorical variables. RESULTS: Reliability of the assessments varied widely. Determining the location of fracture lines had the greatest reliability, whereas the subjective assessments of fracture stability and energy showed the poorest reliability. Although the ICCs for quantitative measurements approached acceptable levels, the tolerance limits were extremely wide. The addition of a CT scan improved the reliability of most assessments, but not to a statistically significant degree. CONCLUSIONS: Many basic radiographic interpretations relied on in making treatment decisions are made variably by observers. Using experienced raters and precise definitions of fracture assessments does not guarantee a high level of agreement. Discrete assessments have higher interrater agreements than do more qualitative assessments. Quantitative measures have wide tolerance limits and, therefore, probably cannot be used reproducibly to classify fractures or make treatment decisions. We conclude the reliability of fracture classification is limited by raters' abilities to agree on basic radiographic assessments.

Humans↗