Search PubMedSearch

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The test-retest reliability of standardized instruments among homeless persons with substance use disorders.

OBJECTIVE: Standardized instruments are widely used to assess homeless persons, but basic data on their reliability and validity in these populations have not been available. The purpose of this study was to examine the reliability of standardized instruments used in a cooperative agreement on homeless persons with substance use disorder. METHOD: This study examined the 1-week test-retest reliability of the Alcohol Dependence Scale, the Addiction Severity Index and the Personal History Form, using 189 randomly selected subjects participating in a multisite study of services for homeless persons with alcohol and other drug abuse problems. In addition to scales and items, factors hypothesized to influence reliability related so subject, interviewer and setting were examined. RESULTS: Results showed substantial reliability for scale scores (> .60) but mixed reliability for individual items. Reliability was greater when items were factual and based on a recent time interval, and when subjects were interviewed in a protected setting. Higher reliability was also related to younger age, female gender, a first episode of homelessness and lower severity of psychiatric problems. CONCLUSIONS: Reliability should be examined in individual studies of homeless persons, and efforts should be made to minimize controllable sources of unreliability.

Adult

Reliability of radiographic grading of osteoarthritis of the hip and knee.

We review studies on the reliability of radiographic assessment of osteoarthritis of the hip and knee. Reliability studies were reported for 10 among 24 identified scores. In general, moderate to good agreement was found for overall scores and for separate grading of joint space narrowing of the hip and osteophytes of the knee in the majority of studies, while reliability tended to be lower for other radiographic features. Overall scores of the knee were more reliable than overall scores of the hip, and intra-rater-reliability was considerably higher than inter-rater-reliability in most instances. Comparison of reliability between scores can only be made with caution, given the difference in the design of reliability studies, particularly the different qualification of involved observers. The limits of existing knowledge on reliability of commonly used radiologic scores are outlined, and, proposals are made to overcome those limits in future studies.

Adult

The reliability of goniometric measurements of active and passive wrist motions.

A reliability study was conducted to determine (a) the intrarater and interrater reliability of goniometric measurement of active and passive wrist motions under clinical conditions and (b) the effect of a therapist's specialization on the reliability of measurement. Randomly paired therapists performed repeated measurements of active and passive wrist motions in 48 subjects who had been referred to one of four occupational therapy or hand management clinics for evaluation and treatment. The data were analyzed with an intraclass correlation coefficient. A posteriori data analyses were performed to determine the effects of identified sources of error on the reliability of measurement. The results indicated that measurement of wrist motion by individual therapists is highly reliable and that intrarater reliability is higher than interrater reliability for all active and passive motions. Interrater reliability was generally higher among specialized therapists for reasons not immediately apparent from this study. With the exception of pain, identified sources of error were found to have surprisingly little effect on the reliability of measurement.

Adult

Reliability of clinical findings in temporomandibular disorders.

The aim of the present investigation was to study the interexaminer reliability of orthopedic tests and palpation techniques routinely used in the clinical diagnosis of disorders of the masticatory system. The tests were performed by a dentist and a physiotherapist, who both used the tests routinely when examining patients with temporomandibular disorders. Seventy-nine patients participated in this study. In the analysis, percentage agreement, intraclass correlation, and Cohen's kappa were used. The interexaminer reliability of the tests measuring maximal active mouth opening and registration of clicking during active mouth opening was high. The interexaminer reliability was fair for the tests measuring the intensity of pain during active movements and moderate for tests recording joint sounds (kappa = 0.47 to 0.59). There was high interobserver agreement on several items of the traction and translation tests, although the kappa values were low. The interexaminer reliability of the multitest scores for compression was substantial for joint sounds (kappa = 0.66) and fair for pain (kappa = 0.40). The interexaminer reliability of the multitest scores for muscle palpation and joint palpation was moderate (kappa = 0.51) and fair (kappa = 0.33), respectively. It can be concluded that most variables determined during active movements can be measured with satisfactory reliability, whereas variables for other tests are not measured with the same reliability on the basis of the kappa scores. The main symptoms of temporomandibular disorders can be evaluated reliably with multitest scores. It is recommended that clinicians calibrate their techniques regularly to improve the reliability of results in daily practice.

Adolescent

Reliability, Device Agreement and Validity of Load-Velocity Profiles: A Systematic Review with Meta-analysis.

BACKGROUND: For a valid one-repetition maximum (1RM) prediction via load-velocity (LV) relationships, high reliability and accuracy must be assumed. OBJECTIVE: Since individual study results indicate ambivalent prediction, this systematic review and meta-analysis was designed to provide a updated and comprehensive overview, extending knowledge about the validity and reliability of commercially available velocity sensors in Part I and the validity and reliability of velocity-based 1RM prediction models in Part II. METHODS: A systematic literature search was conducted in PubMed/MEDLINE, Web of Science, and Scopus. Validity and/or reliability studies or velocity-based 1RM prediction evaluations were included. Methodological quality was assessed using adapted COSMIN. The analysis was performed for intraclass correlation coefficient (ICC), Lin's concordance correlation coefficient (CCC), and Pearson's correlation coefficient (r). The review was preregistered in PROSPERO (CRD42025634595). RESULTS: Sixty-three studies were included for sensor validity and reliability and 38 for 1RM prediction models. Part I: Velocity sensors demonstrated good-to-excellent pooled validity and device agreement (ICC = 0.91-0.92 [0.83-0.97]; k = 55 and 439, respectively); intra- and inter-day reliability were classified as good to excellent with ICC = 0.90-0.91 [0.85-0.95] (k = 228 and 608, respectively), with sensor technology moderating the results. However, substantial heterogeneity and wide ranges of study-level estimates indicated considerable variability across moderators, linear position transducer (LPT) generally showing more consistent performance than inertial measurement units (IMU). Part II: Velocity-based 1RM prediction showed ICCs = 0.90 [0.83-0.94] (k = 124) and ICC = 0.91 [0.72-0.98] (k = 9); for reliability and validity, respectively. DISCUSSION: Commercial velocity sensors generally provide high relative validity and reliability. Results varied depending on exercise complexity, intensity, sensor technology, and modeling approach. While velocity-based 1RM prediction demonstrated high average validity, large heterogeneity in lower body exercises significantly biased the results. Furthermore, the dearth of measurement error and agreement analyses prohibits final conclusions. CONCLUSION: Therefore, velocity-based monitoring and 1RM prediction require cautious interpretation, as sensor- and exercise-specific evidence remains limited.

Load–velocity relationship

The reliability of distance run tests for children in grades K-4.

The purpose of this study was to determine test-retest reliability for the 1-mile, 3/4-mile, and 1/2-mile distance run/alk tests for children in Grades K-4. Fifty-one intact physical education classes were randomly assigned to one of the three distance run conditions. A total of 1,229 (621 boys, 608 girls) completed the test-retests in the fall (October), with 1,050 of these students (543 boys, 507 girls) repeating the tests in the spring (May). Results indicated that the 1-mile run/walk distance, as recommended for young children in most national test batteries, has acceptable intraclass reliability (.83 less than R less than .90) for both boys and girls in Grades 3 and 4, has minimal (fall) to acceptable (spring) reliability for Grade 2 students (.70 less than R less than .83), but is not reliable for children in Grades K and 1 (.34 less than R less than .56). The 1/2 mile was the only distance meeting minimal reliability standards for boys and girls in Grades K and 1 (.73 less than R less than .82). Results also indicated that reliability estimates remained fairly stable across gender and age groups from the fall to spring testing periods, with the exception of the noticeably improved values for Grade 2 students on the 1-mile run/walk test. Criterion-referenced reliability (P, percent agreement) was also estimated relative to Physical Best and Fitnessgram run/walk standards. Reliability coefficients for all age group standards were acceptable to high (.70 less than P less than .95), except for Fitnessgram standards for 5-year-old girls on the 1-mile test for both fall and spring and for 6-year-old boys and girls on the 1-mile test administered in the spring.

Child

Getting the story straight: evaluating the test-retest reliability of a university health history questionnaire.

This study was designed to establish the reliability of a health history questionnaire used as a screening tool for incoming university students. The authors used a test-retest design, with a test interval of 6 months, on a sample of medical and nursing students. The analysis focused on overall reliability of the questionnaire and reproducibility of specific items, based on question format. Questionnaire items of specific interest were those with dichotomous yes/no response options versus open-ended format questions, those using the words frequently or recently, or those that asked multiple questions. Demographic characteristics of the subjects were considered in the evaluation of reliability. Overall reliability of the questionnaire (93.6%) was above the anticipated level of 90%, and subject sex or program of study did not show any significant differences in reproducibility of responses. Although wording of questions did not affect item reliability, dichotomous format questions demonstrated a higher degree of reliability (96.4%) than the overall reliability of the questionnaire. Recommendations for enhancing the reliability of the questionnaire are based on item analysis and information gathered from interviews with subjects.

Adult

Goniometric reliability in a clinical setting. Subtalar and ankle joint measurements.

Measurements of the subtalar joint neutral (STJN) position and passive range of motion (PROM) of the ankle joint and the subtalar joint (STJ) are often part of a physical therapy evaluation. These measurements may be used in treatment planning, such as in the prescription of specialized shoes or orthoses. Therefore, reliability of these measurements, as they are obtained clinically, must be determined. The purpose of this study was to examine the reliability of measurements of the STJN position and of ankle and STJ PROM. To determine reliability, repeated measurements of the STJN position and of STJ PROM were taken on the involved feet of 43 patients with neurologic orthopedic disorders (including both feet of 7 patients), and measurements of ankle PROM (dorsiflexion and plantar flexion) were taken on 42 of these patients (including both feet of 7 patients). Intraclass correlation coefficients (ICCs) for intratester reliability ranged from .74 to .90 for ankle and STJ measurements. The ICCs for intertester reliability were .25 for measuring the STJN position, .32 for STJ inversion, and .17 for SJJ eversion. The ICCs for intertester reliability were .50 for ankle dorsiflexion and .72 for ankle plantar flexion. Goniometric measurements of the STJN position and of PROM of the ankle and STJ appear to be moderately reliable if taken by the same therapist over a short period of time. With the exception of ankle plantar flexion, these measurements cannot be considered to be reliable between therapists.

Adolescent

The reliability and functional validity of visual and semiautomatic sleep/wake scoring in the Møll-Wistar rat.

The present paper has three major objectives: first, to document the reliability of a published criteria set for sleep/wake scoring in the rat; second, to develop a computer algorithm implementation of the criteria set; and third, to document the reliability and functional validity of the computer algorithm for sleep/wake scoring. The reliability of the visual criteria was assessed by letting two raters separately score 8 hours of polygraph records from the light period from five rats (14,040 10-second scoring epochs). Scored stages were waking, slow-wave sleep-1, slow-wave sleep-2, transition type sleep and rapid eye movement (REM) sleep. The visual criteria had good interrater reliability [Cohen's kappa (kappa) = 0.68], with 92.6% agreement on the waking/nonrapid eye movement (NREM) sleep/REM sleep distinction (kappa = 0.89). This indicated that the criteria allow separate raters to independently classify sleep/wake stages with very good agreement. An independent group of 10 rats was used for development of an algorithm for semiautomatic computer scoring. A close implementation of the visual criteria was chosen. The algorithm was based on power spectral densities from two electroencephalogram (EEG) leads and on electromyogram (EMG) activity. Five 2-second fast Fourier transform (FFT) epochs from each EEG/EMG lead per 10-second sleep/wake scoring epoch were used to take the spatial and temporal context into account. The same group of five rats used in visual scoring was used to appraise reliability of computerized scoring. The computer score was compared with the visual score for each rater. There was a lower agreement (kappa = 0.57 and 0.62 for the two raters) than in interrater visual scoring [percent agreement 87.7 and 89.1% (kappa = 0.82 and 0.84) in the waking/NREM sleep/REM sleep distinction]. Subsequently, the computer scores of the raters were compared. The interrater reliability was better than the interrater reliability for visual scoring (kappa = 0.75), with 92.4% agreement for the waking/NREM sleep/REM sleep distinction (kappa = 0.89). The computer scoring algorithm was applied to data from a third independent group of rats (n = 6) from an acoustical stimulus arousal threshold experiment, to assess the functional validity of the scoring directly with respect to arousal threshold. The computer algorithm scoring performed as well as the original visual sleep/wake stage scoring. This indicated that the lower intrarater reliability did not have a significant negative influence on the functional validity of the sleep/wake score.

Algorithms

Improvement of reliability of an oral examination by a structured evaluation instrument.

The main purposes of this study were to estimate the reliability of oral examinations administered to medical students during a clinical clerkship and to improve the reliability of this evaluation technique. In the first part of the study, the reliability of oral examinations as traditionally administered was estimated. The average intraclass reliability coefficient for these examinations was .48. Cassette recordings of these oral examinations were also rated by the faculty members. The average intraclass reliability coefficient of the ratings of the taped performances was .82. In the second part of the study, the reliability of oral examinations was investigated with the raters using a newly developed evaluation form. The average intraclass reliability of the oral examination using the evaluation form was .67, a noticeable increase over the .48 obtained without the form. The average intraclass reliability of ratings made from tape recordings of these oral examinations was .62.

Clinical Clerkship

Reliability of the Glasgow Coma Scale when used by emergency physicians and paramedics.

We sought to determine the reliability of the Glasgow Coma Scale (GCS) when used by emergency physicians and paramedics. We performed a prospective sequential trial in a classroom setting, with subjects blinded to others' scoring. Nineteen university-affiliated emergency physicians and 41 professional paramedics from an urban EMS system voluntarily participated. Participants viewed four videotaped scenes in which a patient is assessed by a paramedic. The first three scenes represented severe, intermediate, and no/mild alteration in level of consciousness (LOC). The findings in the fourth scene were identical to the first, allowing determination of intrarater reliability. The Kappa statistic was used to determine interrater reliability; the reliability coefficient determined intrarater reliability. Kappa was significant (p < 0.0001) for severe (kappa = 0.48), intermediate (kappa = 0.34), and no/mild (kappa = 0.85) conditions. Intrarater reliability (r1,2) for emergency physicians was 0.66 (p < 0.01) and for paramedics was 0.63 (p < 0.01). The GCS shows statistically significant reliability (i.e., significant agreement) between emergency physicians and emergency medical technician-paramedics. It also has a significant level of intrarater reliability.

Allied Health Personnel

Shoulder abduction strength measurement in football players: reliability and validity of two field tests.

Musculoskeletal and neurologic injuries affecting shoulder strength are common in contact sports. Full-strength recovery is desired before resumption of competition. On-field assessment of shoulder strength is usually done by manual muscle testing, which lacks sensitivity and reliability. Our objective was to determine the reliability and validity of two field instruments capable of quantifying shoulder abduction strength. Twenty junior football players underwent bilateral isokinetic (60 degrees/s) and isometric shoulder abduction strength measurements using a Cybex 340 isokinetic dynamometer. Test-retest measurements of both shoulders of each player were made using strain gauge (SG) and handheld dynamometer (HHD) instruments. Players were tested during rested and competition conditions. Within and between session reliabilities were calculated using the intraclass coefficient, and validity was assessed using Pearson's correlation coefficient. Overall reliability for each device was calculated using Lisrel analysis. SG was found to be superior to HHD in overall reliability and validity. Within-session reliability in the rested and competition states was 0.75 and 0.78, respectively, for SG and 0.60 and 0.81, respectively, for HHD. Between-session reliability in the rested and competition states dropped to 0.51 and 0.63, respectively, for SG and 0.55 and 0.70, respectively, for HHD. Validity was 0.41 and 0.70 for SG when correlated with Cybex at 0 degree and 60 degrees/s respectively. Validity for HHD was 0.28 and 0.42 for Cybex speeds of 0 degree and 60 degrees/s, respectively. SG reliability and validity were similar when testing was done one shoulder at a time or both shoulders concurrently.(ABSTRACT TRUNCATED AT 250 WORDS)

Adolescent

Lifetime DSM-IV diagnosis of alcohol, cannabis, cocaine and opiate dependence: six-month reliability in a multi-site clinical sample.

Psychiatric research increasingly emphasizes the diagnosis of symptoms and syndromes on a longitudinal basis. This study tests the reliability of lifetime DSM-IV diagnoses of alcohol, cannabis, cocaine and opiate dependence. The CIDI-SAM was administered at intervals not less than six months apart to a multi-site sample of 201 clinical respondents. The reliability of lifetime diagnosis of the syndromes, of the criteria which constitute the syndromes, and of the ages of onset reported for the criteria and for the dependence syndromes as a whole, were studied and the effects of patient characteristics suspected to degrade reliability were examined. There was generally good agreement, statistically, at both the syndrome and criterion level between the two interviews. Lifetime diagnoses for three of the drugs--alcohol, cannabis and opiates--were made at or near levels of agreement generally considered excellent under less strict testing conditions, and cocaine dependence was only marginally below this level. Most criteria showed good reliability and all delivered about equal results when averaged across the four substances, although a relationship between reliability and centrality of the symptom to the individual drug abuse pattern was found. Age of onset was almost uniformly highly reliable. Most patient characteristics bore no detectable relationship to reliability, although patients with multiple drug use patterns may warrant more careful probing by interviewers. Overall, these data indicate that lifetime symptoms and diagnoses can be queried reliably, although they must be reported with less confidence than current state diagnoses.

Age of Onset

Determinants of reliability in psychiatric surveys of children aged 6-12.

The reliability of young children's self reports of psychiatric information is a concern of epidemiologists and clinicians alike. This paper explores the determinants of test-retest reliability in a sample of children from the general population using reliability coefficients constructed from a kappa statistic. Age, cognitive ability, and gender are related to consistency of reports in a test-retest paradigm. Controlling for age, cognitive ability and gender, children report more reliably on observable behaviors, and less reliably on questions involving unspecified time, reflections of one's own thoughts, and comparison of themselves with others. The reliability of reports of emotions lies between these two extremes. Surprisingly, sentence length of up to 40 words and psychiatric impairment of the child as measured by the Child Global Assessment Scale did not influence reliability. As might be expected, parents' reports of their children are more reliable than their children's reports.

Affective Symptoms

Effects of cognitive impairment on the reliability of geriatric assessments in nursing homes.

OBJECTIVE: To explore the relationship between an elderly subject's cognitive status and the reliability of multidimensional assessment data. DESIGN: Survey, with cognitive status as the independent variable and interrater reliability as dependent variable. SETTING: Medicare/Medicaid-certified nursing homes. PARTICIPANTS: 147 residents age 65 or older. MEASUREMENTS: Dual assessments of elderly nursing home residents were performed by nurse assessors using the Health Care Financing Administration's new Minimum Data Set for Nursing Home Resident Assessment and Care Screening (MDS). Assessments were classified on the basis of residents' cognitive status, and levels of disagreement between assessors were analyzed. MAIN RESULTS: Overall assessment reliability, agreement concerning a resident's activities of daily living status, and the reliability of estimates of his or her communication skills and sensory abilities were significantly affected by a resident's cognitive status. The presence of cognitive impairment made these measurements less reliable--especially those related to communication skills, vision, and hearing. CONCLUSIONS: Assessments of residents suffering from cognitive impairment were significantly less reliable than assessments of cognitively intact residents. However, these differences in reliability were not uniform across all assessment domains. When treating the cognitively impaired elderly, clinicians must exercise caution in their reliance on standardized measurements that may be less reliable for this population.

Activities of Daily Living

Reliability of a standardized and expanded Brief Psychiatric Rating Scale: a replication study.

This study aimed to determine the replicability of the interrater reliability coefficients obtained with a standardized and expanded Brief Psychiatric Rating Scale (BPRS-E) in a 1991 psychometric evaluation. Furthermore, intrarater reliability was assessed. At item level, interrater concordance turned out to be satisfactory for most of the BPRS-E items. However, only a few of the items reached acceptable chance-corrected coefficients. In contrast to the previous study, the anxiety-depression subscale met the standard of acceptable interrater reliability in the present study. As in the 1991 study, the 10-item psychotic disintegration scale as well as BPRS-18 global scores met (or closely approximated) this standard. The 6 additional items of BPRS-E did not contribute to the scale's reliability. Joining the samples of the 1991 and replication studies (to cover the range of symptoms' severity and heterogeneity more fully) did not improve interrater reliability. Intrarater reliability coefficients were globally comparable to interrater reliability coefficients. In all, the results of this replication study suggest that only the anxiety-depression subscale, the 10-item psychotic disintegration scale and the BPRS-18 global scale can be used reliably in unselected groups of psychiatric inpatients in acute distress.

Adolescent

Factors affecting reliability coefficients of health attitude scales.

This study determined the minimum number of health attitude items and minimum sample size required to achieve maximum scale reliability coefficients, using different methods of estimating reliability. A 54-item alcohol attitude scale was administered to 700 participants. The scale produced .96 and .91 reliability coefficients, using the Cronbach Alpha (CA) and the Split-half (S-B) methods, respectively. A computer program randomly selected groups of participants and items from the pool of participants and items using different increments. A matrix of coefficients of reliability for both methods was calculated for different groups of items and sample size. To replicate the study, a 30-item cancer attitude scale was administered to more than 1,000 representative participants and produced reliability coefficients of .94 (using CA) and .82 (using S-B). The same computer and statistical procedures were repeated for the second data set. Results from both analyses consistently demonstrated that sample size has an insignificant effect on the coefficient values of reliability. Reliability increased as the number of items reached 18. Adding more items only negligibly increased the coefficients. Overall, the CA method consistently produced higher coefficient values of reliability compared to the S-B method.

Attitude to Health

Reliability of the National Institutes of Health Stroke Scale. Extension to non-neurologists in the context of a clinical trial.

BACKGROUND AND PURPOSE: The reliability of the National Institutes of Health Stroke Scale (NIHSS) has been established through testing its use in live and videotaped patients. This reliability testing has primarily focused on the use of the scale by neurologists. We sought to determine the reliability of the NIHSS as used by non-neurologists in the context of a clinical trial. METHODS: In anticipation of the initiation of a randomized trial of a new therapy for patients with acute ischemic stroke, 30 physician investigators (30% of whom were not neurologists) and 29 non-physician study coordinators were trained in the use of the NIHSS at an informational and training conference using standardized videotaped patient examinations. A series of 4 patients were rated initially. After 3 months, the same 4 patients were rerated, providing a measure of intraobserver reliability. An additional series of 4 new patients were also rated after 3 months and, with the initial 4 ratings, provided data for assessment of interobserver reliability. RESULTS: Overall, 28% of the raters had previous experience with the NIHSS, and 22% had previously used the videotapes as used in the present trial. The coefficients of determination (r2) were each greater than .95 when the means of the two ratings of the same 4 cases were compared between (1) neurologists and other types of physicians, (2) physicians and study coordinators, (3) raters who had prior experience with the NIHSS and those without prior experience, and (4) raters who had used the videotapes in the past and those who had never viewed the tapes. The calculated r2s were greater than .98 for the initial rating of the first 4 cases and for the later rating of the 4 new cases. The slopes of the regression lines were all near 1, indicating that the raters were similarly calibrated. The intraclass correlation coefficients were .93 and .95, reflecting high levels of intraobserver and interobserver reliability. CONCLUSIONS: These data extend the previously demonstrated reliability of the NIHSS to non-neurologists and show that both a variety of physician investigators and nurse study coordinators can be rapidly trained to reliably apply the scale in the context of an actual clinical trial.

Cerebrovascular Disorders