Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Estimating sleep patterns with activity monitoring in children and adolescents: how many nights are necessary for reliable measures?

STUDY OBJECTIVES: This study provides estimates of reliability for aggregated values from 1 to 7 recording nights for five commonly used actigraphic measures of sleep patterns, reliability as a function of night type (weeknight or weekend night), and stability of measures over several months. DESIGN AND SETTING: Data are from three studies that obtained 7 nights of actigraph data (using Mini Motionlogger actigraphs and associated validated algorithms [ASA]) on children and adolescents living at home on self-selected sleep-wake schedules. PARTICIPANTS: Participants were 169 children aged 12-60 months, and 55 adolescents aged 11-16 years. MEASUREMENTS AND RESULTS: Up to 28% of weekly recordings may be unacceptable for analysis in young participants because of illness, technical problems, and participant noncompliance; studies aiming to collect 5 nights of actigraph data should record for at least 1 full week. Reliability estimates for values aggregated over any 5 nights were adequate (> or = .70) for sleep start time, wake minutes, and sleep efficiency. Measures of sleep minutes and sleep period were less reliable and may require 7 or more nights for estimates of stable individual differences. Reliability for 1- or 2-night aggregates were poor for all measures. We found significant and high correlations between summer and fall session measures for all five variables when weekend nights were included. CONCLUSIONS: Five or more nights of usable recordings are required to obtain reliable actigraph measures of sleep for children and adolescents.

Adolescent↗

ICSD diagnostic criteria for narcolepsy: interobserver reliability. Intemational Classification of Sleep Disorders.

STUDY OBJECTIVES: To estimate the reliability of the diagnosis of narcolepsy after clinical interview and polysomnographic evaluation among sleep medicine doctors, before and after training in application of the International Classification of Sleep Disorders (ICSD). SETTING: Videotaped semi-structured interviews of 10 patients complaining of daytime sleepiness of different etiologies. Questions referred to ICSD criteria for narcolepsy. A further series of 10 cases of narcolepsy without cataplexy were simulated, with at least a random one to three of the ICSD polysomnographic criteria at pathological levels. PARTICIPANTS AND DESIGN: Seventeen doctors were required to classify each videotaped case as "ascertained," "possible," or "excluded" narcolepsy, in two sessions: one before and one after discussion of ICSD criteria. The observers were invited to confirm or exclude the diagnosis of narcolepsy in the 10 simulated cases, according to the given polysomnographic findings, before and after an agreed proposal of the interpretation of ICSD polysomnographic criteria. Interobserver reliability was calculated using Kappa statistics. INTERVENTIONS: N/A. MEASUREMENTS AND RESULTS: Interobserver reliability of clinical judgement improved from "substantial" at baseline (Kappa 0.61) to "almost perfect" after training (Kappa 0.95). Interobserver reliability of polysomnographic findings was "fair" at baseline (Kappa 0.24), unanimous after the proposed interpretation of ICSD polysomnographic criteria. CONCLUSIONS: Baseline reliability of diagnostic judgement in suspected narcolepsy was found satisfactory among Italian sleep medicine doctors. Educational training, based on discussion of ICSD criteria, further improved agreement. Diagnosis based on polysomnographic findings, not reliable at baseline, needed a strict interpretation of ICSD criteria to attain standardization.

Adult↗

Reliability of ratings on the Glasgow Outcome Scales from in-person and telephone structured interviews.

OBJECTIVE: To determine test-retest reliability and interrater reliability for structured interviews for the Glasgow Outcome Scale (GOS) using in-person and telephone contact. METHODS: Study 1: Thirty head-injured patients were interviewed face-to-face and then reinterviewed by telephone a few days later by the same rater. Study 2: Fifty-six head-injured patients were interviewed by telephone and then face-to-face interviews were carried out by a different person up to 1 month later. Agreement between ratings on the GOS and the extended GOS (GOSE) in each of the studies was assessed using the kappa statistic weighted with quadratic weights. RESULTS: Values of kappa(w) for the test-retest reliability study were.92 for both GOS and GOSE, and for interrater reliability study were.85 for the GOS and.84 for the GOSE. CONCLUSIONS: The findings indicate good test-retest and interrater reliability for the structured interviews. In most circumstances a structured interview over the telephone can provide a reliable assessment of the GOS, and can safely be substituted for in person contact.

Adult↗

Reliability of isokinetic trunk muscle strength measurement.

OBJECTIVE: To determine the intrarater and interrater reliability of reciprocal concentric trunk flexion and extension peak torque values at different angular velocities using the Cybex NORM isokinetic dynamometer. DESIGN: Trunk flexor and extensor muscles of 15 healthy subjects were assessed at 60 degrees/sec and 90 degrees/sec angular velocities. Each subject was tested by two physicians three times, with at least 48 hr between test sessions. Intraclass correlation coefficients were used to determine the intrarater and interrater reliability of the reciprocal concentric trunk flexion and extension peak torques at 60 degrees/sec and 90 degrees/sec angular velocities. RESULTS: For intrarater reliability intraclass correlation coefficients for trunk flexion, peak torque values ranged from 0.89 to 0.95; for the extensors, values ranged from 0.80 to 0.92. The intraclass correlation coefficient values for interrater reliability were also found to be highly reliable, and peak torque values demonstrated intraclass correlation coefficients values that ranged from 0.95 to 0.98. CONCLUSIONS: Reciprocal concentric trunk flexion and extension peak torque measurements at 60 degrees/sec and 90 degrees/sec angular velocities had a high intrarater and interrater reliability with the Cybex NORM isokinetic dynamometer in healthy subjects.

Adult↗

The Diagnostic Interview Schedule for Children-Revised Version (DISC-R): II. Test-retest reliability.

OBJECTIVE: Test-retest reliability and internal consistency of the Diagnostic Interview Schedule for Children-Revised Version (DISC-R) were examined to evaluate the effects of a recent revision process. METHOD: A sample of outpatients 11 to 17 years old and their parents were administered the DISC-R in a test-retest design. RESULTS: For the parent interview, four diagnoses had sufficient cases to examine reliability; agreement was good to excellent for three of these. For the child interview, two of the four diagnoses with adequate cases showed good or excellent reliability. Using a combined parent-child algorithm, three of five diagnoses showed good or excellent reliability. Test-retest reliability for symptom scales was excellent for the parent DISC-R and good for the child version, except for oppositional defiant disorder. Maternal depressive symptoms did not affect reliability of reporting about the child's symptomatology. Internal consistency was satisfactory for symptom items comprising most diagnoses. CONCLUSIONS: Although some revision of the instrument will be necessary, the findings from this study suggest that the DISC-R is a promising research instrument for the diagnosis of psychopathology in older children and adolescents.

Adolescent↗

Test-retest reliability of anxiety symptoms and diagnoses with the Anxiety Disorders Interview Schedule for DSM-IV: child and parent versions.

OBJECTIVE: To examine the test-retest reliability of the DSM-IV anxiety symptoms and disorders in children with the Anxiety Disorders Interview Schedule for DSM-IV: Child and Parent Versions (ADIS for DSM-IV:C/P). METHOD: Sixty-two children (aged 7-16 years) and their parents underwent two administrations of the ADIS for DSM-IV:C/P with a test-retest interval of 7 to 14 days. RESULTS AND CONCLUSIONS: Results revealed that the ADIS for DSM-IV:C/P is a reliable instrument for deriving DSM-IV anxiety disorder symptoms and diagnoses in children. The ADIS for DSM-IV:C/P was found to have excellent reliability in symptom scale scores for separation anxiety disorder, social phobia, specific phobia, and generalized anxiety disorder and good to excellent reliability for deriving combined diagnoses of these disorders, as well as using child-only and parent-only interview information. Reliability coefficients were generally similar and, in most instances, superior to those found in previous ADIS-C/P reliability studies. Limitations and directions for future research are discussed.

Adolescent↗

Spanish-language services assessment for children and adolescents (SACA): reliability of parent and adolescent reports.

OBJECTIVE: To assess test-retest reliability of the service utilization screening section of the Services Assessment for Children and Adolescents (SACA) interview among Spanish-speaking parents and adolescents, correspondence between parent and adolescent reports, and the correlation between reliability and participants' demographic and service use characteristics. METHOD: The English SACA was translated and administered from September 1999 through January 2000 in Los Angeles County, California, on two separate occasions to eligible parents with a child (4-17 years old) who was a client of a local public mental health authority. Adolescents of these parents (12-17 years old) were also interviewed. Reliability was measured by the kappa statistic. RESULTS: Adult and adolescent reports about lifetime and previous year service setting use exhibited good reliability, but concordance of parents and adolescents did not. Children's service utilization appears to be correlated with reliability of parent reports, and child gender appears to be correlated with reliability of adolescent reports. CONCLUSION: The SACA appears to be a useful tool for screening Spanish-speaking families about child and adolescent mental health service use. These findings must be considered preliminary until replicated in a larger sample of culturally diverse Spanish-speaking families.

Adolescent↗

Reliability of photographic analysis in determining change in scar appearance.

Photographs frequently are used to document change in the management of hypertrophic scars. The purpose of this study was to design a scale for the analysis of photographs of hypertrophic scars and to test its reliability. The subjects were four occupational and physical therapists, (two novices and two experts), in scar management. Existing scales were modified to produce a new scale. The subjects twice rated four slides from each of ten patients' scars, in random order. They used a Latin Square design. Interrater and test-retest reliabilities were calculated using a weighted kappa statistic. The newly developed scale demonstrated interrater reliability, which ranged between 0.66 and 0.90 for all items. The test-retest reliability ranged between 0.73 and 0.89 for all items. The new scale had substantial reliability (using a single rater) and was at least as reliable when used by novice therapists. This indicated that training had no effect.

Adult↗

On the methods and theory of reliability.

This paper reviews the most frequently used and misused reliability measures appearing in the mental health literature. We illustrate the various types of data sets on which reliability is assessed (i.e., two raters, more than two raters, and varying numbers of raters with dichotomous, polychotomous, and quantitative data). Reliability statistics appropriate for each data format are presented, and their pros and cons illustrated. Inadequancies of some methods are highlighted. The meaning of different levels of reliability obtained with various statistics is discussed. This critique is intended for the reading professional and the investigator who has an occasional need for reliability assessment. Statistical expertise is not required and theoretical material is referenced for the interested reader. Necessary formulas for computations are presented in the appendices. A summary table of some suitable reliability measures is presented.

Humans↗

Can severely mentally ill adults reliably report their needs?

Proponents of psychiatric rehabilitation believe that active patient participation in rehabilitation planning is essential. Participation assumes, however, that severely mentally ill patients can reliably report their needs. Unfortunately, reliably responding to a needs assessment may be diminished by cognitive deficits that are associated with psychosis and low intelligence. To address this issue, 35 severely mentally ill outpatients completed the Needs and Resources Assessment twice, with 1 week intervening. Their reliability on this task was compared with that of 16 control subjects. Results showed that subjects in both groups performed more reliably (measured as raw agreement) on the more highly structured, standardized items. Patients were less reliable than controls on individualized and standardized items. Further analyses suggested that the patients' raw agreement was found to be significantly associated with thought disturbance. Assessment methods must be sufficiently structured to help cognitively disordered patients provide reliable information for rehabilitation planning.

Adult↗

Histologic feature reliability in childhood neural tumors. Childhood Brain Tumor Consortium.

We studied intraobserver reproducibility in recognizing the presence or absence of 57 histologic feature or patterns in a random subset of tumors (822) from the Childhood Brain Tumor Consortium database. The study protocol maximized consistency of the observer. We found that only six histologic features had high (> or = 0.75) reliability estimates while a large number had intermediate estimates of 0.50-0.74. Supratentorial or infratentorial tumor location sometimes altered reliability. Reliability estimates were unacceptable for certain histologic features often used as diagnostic criteria, descriptors of tumor characteristics, or markers of anaplasia. We hypothesize that low reliability reflects, in part, the need for more specific operational definitions, particularly those with subjective boundaries (e.g. granular bodies) may also contribute to low reliability. We also show that the kappa statistic, a commonly used measure of reliability, is inappropriate for very common or uncommon histologic features (e.g. features at the extremes of prevalence in the study cases) and we offer a simple empiric method for determining when an alternative measure, the Jaccard statistic, is appropriate.

Brain Neoplasms↗

The Revised Behavior and Symptom Identification Scale (BASIS-R): reliability and validity.

BACKGROUND: To assess outcomes of health services, providers need brief, responsive, reliable, and valid measures that can be implemented in clinical settings with minimal cost and burden. The Behavior and Symptom Identification Scale (BASIS-32) is a self-report measure developed in 1984 to assess mental health treatment outcomes. During the past 3 years, multiple methods were used to revise the instrument to improve reliability, validity, and applicability to diverse groups of mental health service recipients. OBJECTIVE: The objective of this study was to field test the revised instrument, make further changes based on analysis of the field test data, and assess reliability and validity of the final version (BASIS-24). METHODS: A field test was implemented at 27 treatment sites across the United States. A total of 2656 inpatients and 3222 outpatients participated. Factor analytic methods, classic test theory, and item response theory modeling were used to select the most discriminating, nonredundant items for inclusion in the final version of the instrument and to assess its reliability and validity. Item response theory modeling was used to score the instrument. RESULTS: The final instrument includes 24 items assessing 6 domains: depression/ functioning, interpersonal relationships, self-harm, emotional lability, psychosis, and substance abuse. Test-retest and internal consistency reliability were acceptable. Tests of construct and discriminant validity supported the instrument's ability to discriminate groups expected to differ in mental health status, and its correlation with other measures of mental health. CONCLUSIONS: Analyses of the BASIS-24 supported its reliability and validity for assessing mental health status from the patient's perspective.

Adolescent↗

Test-retest reliability of symptom-limited cycle ergometer tests in patients with chronic obstructive pulmonary disease.

BACKGROUND: Symptom-limited exercise tests are widely used to evaluate the effects of pulmonary rehabilitation in patients with chronic obstructive pulmonary disease (COPD), but the reliability of these tests is not well established in COPD patients. OBJECTIVES: We compared test-retest reliability of two repeated symptom-limited exercise tests between COPD patients and healthy elderly subjects and between male and female patients. METHOD: Fifty-six COPD patients (40 men, 16 women) and 16 healthy subjects (6 men, 10 women) performed two symptom-limited exercise tests approximately 2 weeks apart. Measures of oxygen uptake (VO2), minute ventilation (VE), heart rate, and ratings of breathlessness and leg fatigue were obtained at peak exercise at each symptom-limited exercise test. RESULTS: Repeated measures of peak exercise responses were stable for patients and healthy subjects and for male and female patients. Although mean percent error (absolute difference/mean) for peak exercise responses was low, some individuals' values exceeded 10%. There was no difference in the percent error between COPD patients and healthy subjects or between men and women with COPD. Test-retest reliability was lower for breathlessness ratings than for other peak exercise responses for all groups. CONCLUSIONS: Repeated symptom-limited exercise tests are reliable in COPD patients and healthy subjects. However, some individuals are less reliable, and these patients may require more than one exercise test to establish reliable performance.

Aged↗

The reliability of selected pain provocation tests for sacroiliac joint pathology.

OBJECTIVE: To assess the inter-rater reliability of seven pain provocation tests for pain of sacroiliac origin in low back pain patients. SUMMARY OF BACKGROUND DATA: Previous studies on the reliability of such tests have produced inconclusive and conflicting results. METHODS: Fifty-one patients with low back pain, with or without radiation into the lower limb, were assessed by one examiner and another drawn from a pool of five. Percent agreement and the Kappa statistic were used to evaluate the reliability of the seven tests. RESULTS: Percent agreement and the Kappa statistic ranged in value from 78% and 0.52 (P < 0.001) to 94% and 0.88 (P < 0.001), respectively, when results for all examiner pairs were pooled. However, two tests demonstrated only marginal reliability when performed by one pair of assessors that examined 43% of the patients. CONCLUSIONS: Five of seven tests employed in this study were reliable, the other two were potentially reliable. These tests may be used to detect a sacroiliac source of low back pain, although sensitivity and specificity studies are needed to determine their diagnostic power.

Adult↗

A test to measure lift capacity of physically impaired adults. Part 1--Development and reliability testing.

STUDY DESIGN: Two laboratory studies and one field study evaluated the safety and test-retest reliability of a new test of lift capacity. The first two studies were conducted in a carefully controlled laboratory setting. The first study investigated the safety and intra-rater reliability of the EPIC Lift Capacity test protocol with healthy adult subjects. The second study assessed the safety and inter-rater reliability of the test with disabled subjects. The third study was conducted in the field with 65 evaluators and investigated the safety and intra-rater reliability of the test with healthy adult subjects. OBJECTIVE: To assess the safety and reliability of a new test of lift capacity. SUMMARY OF BACKGROUND DATA: A new test of lift capacity has been developed. Test development occurred within the context of ergonomic standards and guide-lines of the major professional associations and public agencies that govern test development in the United States. METHODS: In study no. 1, 26 healthy subjects participated. In study no. 2, 14 disabled subjects participated. In study no. 3, 318 healthy subjects participated. After subjects underwent basic screening and warm-up, the EPIC Lift Capacity test was administered. One to 2 weeks later, the test was administered again. Correlations between the times of testing were calculated. RESULTS: No subjects were injured. Hamstring soreness the next day that resolved without complication was reported by some healthy subjects. None of the disabled subjects reported new symptoms. CONCLUSION: The safety and reliability of the EPIC Lift Capacity test was adequately demonstrated in a laboratory setting and across multiple field sites with evaluators who have varying types and degrees of professional preparation.

Adult↗

Compensatory spinopelvic balance over the hip axis and better reliability in measuring lordosis to the pelvic radius on standing lateral radiographs of adult volunteers and patients.

STUDY DESIGN: Sagittal alignments, including lumbar lordosis and spinopelvic balance (measured from C7, S1, and hip axis reference points for the relative positions of the spine and sacropelvis over the hips), were studied on standing 36-in. lateral radiographs of adult volunteers (control subjects) and patients who had specific spinal disorders. OBJECTIVES: To determine the most reliable methods for measuring lumbopelvic lordosis and to define significant spinopelvic compensations for sagittal balance. SUMMARY OF BACKGROUND DATA: Measurements for standing sagittal balance, obtained using a C7 plumb line, and segmental angulations of the spinal vertebrae, including lordosis to the sacrum, have been reported. Absolute values, even for normative data, have had wide variation and limited clinical usefulness. Correlations of sagittal balance with the reported spinopelvic angulations (spinal vertebral and sacropelvic angulations) have not been well defined. In addition, determinates of balance (spinal and pelvic) have not been studied for reliability, and compensatory mechanisms for maintenance of balance have not been carefully evaluated. Better recognition of the correlations and more reliable methods to measure lordosis and balance and the spinopelvic compensations for its maintenance may be beneficial in treating patients who have spinal disorders. METHODS: Measurements on standing 36-in. lateral radiographs were made for sagittal alignments in adult volunteers (n = 50) and in adult patients who had symptomatic degenerative lumbar disc disease (n = 50), low grade L5-S1 isthmic (lytic) spondylolisthesis (n = 30), and idiopathic or degenerative scoliosis (n = 30). All participants exhibited clinical compensation for balance. Data were analyzed for significant correlations within each group to determine compensatory correlations of spinopelvic balance with the other sagittal alignments. Intraobserver and interobserver reliability for the parameters evaluated were calculated. This included two methods for determining lordosis (S1 end-plate and pelvic radius techniques). RESULTS: Plumb line measurements for balance from the S1 and hip axis reference points, as defined, were similar in all four groups. However, the groups appeared to adjust for balance by using common and distinctive spinopelvic compensations that resulted in significantly and characteristically different angular alignments among the four groups. Lordosis and balance measurements were closely correlated, and the correlation was characterized by pelvic rotation and translation around the hip axis. The subjects with less lordosis typically stood with the C7 plumb line anterior to and at a longer distance from the sacral reference point. This was primarily because of posterior sacropelvic translation around the hip axis and not because the sagittal plumb line initially moved anteriorly away from the sacrum. This was true in all four groups and gave the appearance that the sacropelvis was less well balanced over the hips in the subjects with less lordosis. Even small differences in lordosis appeared to be associated with considerable adjustments in the other spinopelvic alignments. Therefore, it was important to determine that lordosis was lumbopelvic more reliably measured by the pelvic radius technique. CONCLUSIONS: Lower lumbar lordosis, by the pelvic radius technique, and compensatory sacropelvic translation around a hip axis, in addition to measurements from this axis to the C7 plumb line, were the primary determinates and most reliable radiographic assessments for sagittal balance. Understanding the common and characteristically different compensations that occur with balance in these patients who had specific spinal disorders may help to improve their care.

Adult↗

The reliability and validity of the Biering-Sorensen test in asymptomatic subjects and subjects reporting current or previous nonspecific low back pain.

STUDY DESIGN: A reliability study and case-control study were conducted. OBJECTIVES: To determine the reliability and discriminative validity of the Biering-Sorensen test. SUMMARY OF BACKGROUND DATA: A low Biering-Sorensen score has been found to predict who will have nonspecific low back pain. However, the reliability of the test remains controversial, implying that some studies may have produced results that underestimated the magnitude of the predictive validity of this test. METHODS: Two raters measured the time holding a specific position (holding time) of 63 subjects (23 currently experiencing nonspecific low back pain, 20 who had had an episode, and 20 who were asymptomatic) while they performed the Biering-Sorensen test twice, 15 minutes apart. A standardized protocol was followed. Test-retest reliability was evaluated by calculating intra-class correlation coefficients (ICC 1,1), 95% confidence intervals (CI), and standard errors of the measurement (SEM) for the total group and for the subgroups. A three-way analysis of variance was used to determine whether test order, subject gender, or symptom status affected holding time. RESULTS: High reliability indices were obtained for the Biering-Sorensen test in subjects with current nonspecific low back pain (ICC [1,1], 0.88; 95% CI, 0.73-0.95; SEM, 11.6 seconds), in subjects who had had nonspecific low back pain (ICC [1,1], 0.77; 95% CI, 0.52-0.90; SEM, 17.5 seconds), and in asymptomatic subjects (ICC [1,1], 0.83; 95% CI, 0.62-0.93; SEM, 17.4 seconds). Results of an analysis of variance showed that subjects asymptomatic for low back pain had a significantly longer holding time than the other two groups (P < 0.05). CONCLUSIONS: The Biering-Sorensen test provides reliable measures of position-holding time and can discriminate between subjects with and without nonspecific low back pain.

Adult↗

Reliability and validity of the back performance scale: observing activity limitation in patients with back pain.

STUDY DESIGN: A single group design to examine reliability and validity of the Back Performance Scale. OBJECTIVES: To examine intertester reliability, test-retest reliability, and concurrent validity of the Back Performance Scale. SUMMARY OF BACKGROUND DATA: The Back Performance Scale is a condition-specific performance measure of activity limitation in patients with back pain. It includes five tests of daily activities requiring mobility of the trunk: sock test, pick-up test, roll-up test, fingertip-to-floor test, and lift test. Discriminative ability and responsiveness to important change have previously been demonstrated. METHODS: A total of 41 patients with back pain participated in the study. Two physiotherapists examined test performances concurrently, but independently. The patients filled in three questionnaires, two reflecting perceived disability (Der Funktionsfragenbogen Hannover, Roland-Morris Disability Questionnaire) as well as one for fear avoidance of daily activities and work (Fear Avoidance Belief Questionnaire). One physiotherapist retested the patients after 2 to 3 days. RESULTS: Intertester agreement of the Back Performance Scale sum score was very high (intraclass correlation coefficient 2.1): 0.996. Within-patient standard deviation (sw) on the 16-point Back Performance Scale was very low: 0.25. Test-retest reliability was high (intraclass correlation coefficient = 0.91, sw = 1.3). Intertester agreement of the separate tests was also very high, ranging from kappa= 0.90-1.00. Test-retest reliability was moderate to high (kappa= 0.55-0.83). A high correlation was demonstrated between the Back Performance Scale and the Der Funktionsfragenbogen Hannover: Spearman rho (rho) = 0.825, P < 0.01. Correlation between the Back Performance Scale and Roland-Morris Disability Questionnaire was moderate: rho = 0.454, P < 0.01. No correlation was demonstrated between the Back Performance Scale and the Fear Avoidance Belief Questionnaire. CONCLUSION: The Back Performance Scale appears to be a reliable and valid outcome measure of activity limitation.

Activities of Daily Living↗