Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Inter-rater and test-retest reliability of the Spanish language version of the eating disorder examination interview: clinical and research implications.

BACKGROUND: To examine the inter-rater and test-retest reliability of the Spanish Language version of the Eating Disorder Examination (S-EDE) in monolingual Latina women. Established measures are needed to study Latino groups, and short-term test-retest reliability findings are needed to provide context for clinical treatment and outcome studies. METHODS: Inter-rater reliability (IRR) and short-term (5-14 days) test-retest reliability (TRR) of the S-EDE (using intraclass correlation coefficients [ICCs]) were examined in a non-clinical study group of 60 monolingual Latina women. RESULTS: IRR was excellent for objective bulimic episodes (ICC = 0.99) but was modest for subjective bulimic episodes (ICC = 0.55). TRR was good for objective bulimic episodes (ICC = 0.79) but was unacceptable for subjective bulimic episodes (ICC = 0.22). IRR and TRR kappa coefficients (0.56 and 0.37, respectively) for identifying the presence or absence of objective bulimic episodes were modest. For the S-EDE subscales, both IRR (ICCs ranged from 0.80 to 0.98) and TRR (ICCs ranged from 0.67 to 0.90) were good to excellent. CONCLUSIONS: These findings provide preliminary support for the reliability of the S-EDE for use with Latina women. The constructs of eating disorder psychopathology measured by the S-EDE subscales (restraint, eating concern, weight concern, and shape concern) and the core feature of binge eating (objective bulimic episodes) show high short-term consistency. The results for subjective bulimic episodes are consistent with previous studies that have questioned whether these eating behaviors are reliable indicators of eating disorders. Additional evaluation is needed with clinical groups.

Adult↗

The reliability and validity of the self-reported drinking measures in the Army's Health Risk Appraisal survey.

BACKGROUND: The reliability and validity of self-reported drinking behaviors from the Army Health Risk Appraisal (HRA) survey are unknown. METHODS: We compared demographics and health experiences of those who completed the HRA with those who did not (1991-1998). We also evaluated the reliability and validity of eight HRA alcohol-related items, including the CAGE, weekly drinking quantity, and drinking and driving measures. We used Cohen's kappa and Pearson's r to assess reliability and convergent validity. To assess criterion (predictive) validity, we used proportional hazards and logistical regression models predicting alcohol-related hospitalizations and alcohol-related separations from the Army, respectively. RESULTS: A total of 404,966 soldiers completed an HRA. No particular demographic group seems to be over- or underrepresented. Although few respondents skipped alcohol items, those who did tended to be older and of minority race. The alcohol items demonstrate a reasonable degree of reliability, with Cronbach's alpha = 0.69 and test-retest reliability associations in the 0.75-0.80 range for most items over 2- to 30-day interims between surveys. The alcohol measures showed good criterion-related validity: those consuming more than 21 drinks per week were at 6 times the risk for subsequent alcohol-related hospitalization versus those who abstained from drinking (hazard ratio, 6.36; 95% confidence interval=5.79, 6.99). Those who said their friends worried about their drinking were almost 5 times more likely to be discharged due to alcoholism (risk ratio, 4.9; 95% confidence interval=4.00, 6.04) and 6 times more likely to experience an alcohol-related hospitalization (hazard ratio, 6.24; 95% confidence interval=5.74, 6.77). CONCLUSIONS: The Army's HRA alcohol items seem to elicit reliable and valid responses. Because HRAs contain identifiers, alcohol use can be linked with subsequent health and occupational outcomes, making the HRA a useful epidemiological research tool. Associations between perceived peer opinions of drinking and subsequent problems deserve further exploration.

Adult↗

Radiographic assessment of pediatric proximal radius fractures: interrater and intrarater reliability.

Three methods of measuring pediatric proximal radius fracture radiographs were compared using injury films of 32 patients. Angulation and displacement were independently measured by four physicians. One physician measured the films by each method a second time 2 months later. Values for interrater and intrarater reliability were determined using inter- and intra-class coefficients (ICC). Interrater reliability was poor for methods using the axis of the proximal radial fragment or the proximal radial physis as a reference (ICC = 0.47 and 0.42, respectively). Measurement of the angle between a line parallel to the proximal radius articular surface and the radial shaft had the highest interrater reliability (0.76); measurement of displacement had the lowest interrater reliability (0.09). The intrarater reliability was excellent for all methods (0.93-0.99) and was also highest when the proximal articular surface reference was used. Of described methods, use of the proximal radius articular surface and the radial shaft as references had the highest interrater and intrarater reliability.

Adolescent↗

Agreement between orthopedic surgeons and neurosurgeons regarding a new algorithm for the treatment of thoracolumbar injuries: a multicenter reliability study.

INTRODUCTION: Considerable variability exists in the management of thoracolumbar (TL) spine injuries. Although there are many influences, one significant factor may be the treating surgeon's specialty and training (ie, orthopedic surgery vs. neurosurgery). Our objective was to assess the agreement between spinal orthopedic and neurologic surgeons in rating the severity of TL spine injuries with a new treatment algorithm. This information could be important in establishing consensus-based protocols for managing these challenging injuries. METHODS: Twenty-eight spinal surgeons (8 neurosurgeons and 20 orthopedic surgeons) reviewed 56 TL injury case histories. Each case was classified and scored according to the TL injury severity score (TLISS). The case histories were reordered and the physicians repeated the exercise 3 months later. At both intervals the surgeons were asked if they agreed with the final treatment recommendation of the TLISS algorithm. The reliability and decision validity of the TLISS was compared. RESULTS: Between-group interrater reliability was similar to within group reliabilities. Intrarater reliability was also similar between groups. The between speciality interrater reliability of the TLISS management recommendation was moderate (74% agreement, kappa=0.532). Orthopedic and neurosurgeons agreed with the TLISS management recommendation 91.4% and 94.4% of the time, respectively. CONCLUSIONS: The TLISS demonstrated good reliability in terms of intraobserver and interobserver agreement on the algorithmic treatment recommendations. The recommendation for operation seems to be consistent between fellowship-trained orthopedic and neurosurgical spine surgeons. This type of classification system may reduce the existing variability and initial management decision for treatment of TL injuries.

Algorithms↗

Assessing past treatment history: test-retest reliability of the Treatment Response to Antidepressant Questionnaire.

A reliable and valid instrument has yet to be developed that elicits antidepressant treatment history via patient interview. The goal of the present study was to establish the test-retest reliability of the Treatment Response to Antidepressant Questionnaire (TRAQ). The TRAQ is a semistructured interview that was designed to collect systematically information regarding previous antidepressant treatment, adequacy of trials, and nature of response. Fifty subjects who sought outpatient treatment as part of the Rhode Island Methods to Improve Diagnostic Assessment and Services (MIDAS) project participated in the study. Patients were interviewed initially by a psychologist, who administered the TRAQ. An average of 5 to 6 days later, a psychiatrist who was blind to the results of the initial evaluation readministered the TRAQ to each of these patients. Reliability of recall of antidepressant trials, trial adequacy, and nature of response were evaluated using the kappa statistic. The mean duration of the TRAQ interviews was 3.30 minutes (SD=2.03 minutes). The reliability of recall of antidepressant trials ranged from 0.81 to 0.95, with an overall kappa of 0.91. The kappa for trial adequacy, depending on the definition used, ranged from 0.72 to 0.84. The kappa for determining positive versus negative response was 0.72. Thus, the test-retest reliability of the TRAQ was found to be in the good to excellent range for each of the principal outcome measures. The TRAQ can be administered by non-MDs as a reliable measure for collecting standardized information regarding antidepressant treatment history via patient interview.

Adult↗

Interrater and intrarater reliability of the Modified Ashworth Scale in children with hypertonia.

PURPOSE: The Modified Ashworth Scale (MAS) is a six-point scale used to assess spasticity. This study assessed interrater and intrarater reliability of the MAS for children with hypertonia. METHODS: Five raters participated in this examination of interrater and intrarater reliability. The study included 17 children who showed hypertonus. Elbow flexor, hip adductor, quadriceps, hamstring, gastrocnemius, and soleus muscles were tested bilaterally. RESULTS: Results demonstrated good interrater reliability (intraclass correlation coefficient [ICC] >0.75) for elbow flexors and hamstrings and poor interrater reliability (ICC <0.50) for other muscles. Intrarater scores were good (ICC >0.75) for hamstrings and moderate (ICC = 0.50 to 0.75) for other muscles. CONCLUSION: Interrater reliability of the MAS may be lower than desired for clinical use for muscles other than hamstrings and elbow flexors, and intrarater reliability may also be lower than desired for muscles other than the hamstrings.

Child↗

Reliability of classification of fractures of the tibial plafond according to a rank-order method.

BACKGROUND: Many orthopedic classification systems, including those for tibial plafond fractures, are either unvalidated or have demonstrated problems with interobserver reliability. Classification of tibial plafond fractures according to a rank-order method has shown excellent interobserver reliability with several observers. The purpose of this study is to determine the reliability of a rank order classification of plafond fractures with a large number of observers. METHODS: A radiographic review study was completed by 69 orthopedists of varying training levels. Observers ranked 10 fractures of the tibial plafond based on anteroposterior and lateral ankle radiographs. Fractures were ranked in increasing severity from 1 to 10. No instructions were given regarding determination of severity. Agreement between rankings was analyzed by the intraclass correlation coefficient (ICC). RESULTS: Rankings were performed by viewing prints at the annual Orthopaedic Trauma Association meeting and through the Orthopaedic Trauma Association website using digital images. The overall ICC was 0.62. There was no difference in the ICC between traumatologists and general orthopedists (p > 0.5). Eleven observers commented that the radiographs did not represent the full spectrum of injury severity. CONCLUSIONS: The interobserver reliability of the rank-order classification in this study was fair to good, which is better than previously reported for plafond fracture classification systems. It remains to identify and validate a series of tibial plafond fractures that represent a full spectrum of injury and can be ranked with excellent interobserver reliability. A series of cases such as this may then serve a measurement standard for severity of bony injury against which individual cases may be reliably compared.

Classification↗

Reliability of radiological classifications used in Legg-Calve-Perthes disease.

Radiological assessment is a valuable tool in the assessment, management and prognostication of Perthes disease. Radiological assessment, however, is not an easy task and all classification systems used in Perthes disease have some degree of interrater and intrarater variabilities. In the past, there were some isolated studies to find the reliability of the classifications used in Perthes disease. In this study, we comprehensively studied three most commonly used radiological classifications (Salter-Thompson, lateral pillar and Catterall). We had 44 patients' radiographs (anteroposterior and lateral) taken in the fragmentation stage, and two experienced observers assessed and classified the radiographs on two separate occasions. In this study, we found that the average interrater reliability of the Salter-Thompson, lateral pillar and Catterall classifications was 0.163 (0.08-0.236), 0.722 (0.581-0.824) and 0.433 (0.280-0.546), respectively. The intrarater reliability was 0.313 and 0.699 for the Salter-Thompson, 0.707 and 0.658 for the lateral pillar and 0.38 and 0.577 for the Catterall classifications. Further, we tried to determine the possible reason for the low reliability associated with the Catterall classification. We think that the quantitative method of lateral pillar has better intrarater and interrater reliabilities than other classification systems, and the reliability of the Catterall classification can be significantly improved if some radiological parameters such as metaphyseal reaction and identification of the junction of involved to uninvolved region can be optimized.

Humans↗

Reliability and the adaptive utility of discrimination among alarm callers.

Unlike individually distinctive contact calls, or calls that aid in the recognition of young by their parents, the function or functions of individually distinctive alarm calls is less obvious. We conducted three experiments to study the importance of caller reliability in explaining individual-discriminative abilities in the alarm calls of yellow-bellied marmots (Marmota flaviventris). In our first two experiments, we found that calls from less reliable individuals and calls from individuals calling from a greater simulated distance were more evocative than calls from reliable individuals or nearby callers. These results are consistent with the hypothesis that marmots assess the reliability of callers to help them decide how much time to allocate to independent vigilance. The third experiment demonstrated that the number of callers influenced responsiveness, probably because situations where more than a single caller calls, are those when there is certain to be a predator present. Taken together, the results from all three experiments demonstrate the importance of reliability in explaining individual discrimination abilities in yellow-bellied marmots. Marmots' assessment of reliability acts by influencing the time allocated to individual assessment and thus the time not allocated to other activities.

Acoustic Stimulation↗

Reliability analysis of event-related brain potentials to olfactory stimuli.

Olfactory event-related potentials (OERP) have been used to investigate olfactory processing in health and disease. However, the reliability of the OERP has yet to be established statistically. The present study examined test-retest reliability of the OERP over a 4-week interval. EEG was recorded from Fz, Cz, and Pz, using a single-stimulus paradigm with amyl acetate. Reliabilities for ERP component latencies and interpeak amplitudes were assessed as intraclass and Pearson product-moment correlation coefficients. Reliabilities were higher for latency than for amplitude. Highest correlation coefficients were observed for P2 latency, specifically at Cz and Pz P3 amplitude and latency exhibited high reliability at Cz and Pz. Fz demonstrated weakest correlation coefficients. The data suggest that OERP reliability is comparable to that of auditory and visual ERPs, supporting the use of OERPs in both basic research and clinical assessment.

Adult↗

A study examining inter- and intrarater reliability of three scales for measuring severity of psoriasis: Psoriasis Area and Severity Index, Physician's Global Assessment and Lattice System Physician's Global Assessment.

BACKGROUND: There is a lack of consensus as to the best way of monitoring psoriasis severity in clinical trials. The Psoriasis Area and Severity Index (PASI) is the most frequently used system and the Physician's Global Assessment (PGA) is also often used. However, both instruments have some drawbacks and neither has been fully evaluated in terms of 'validity' and 'reliability' as a psoriasis rating scale. The Lattice System Physician's Global Assessment (LS-PGA) scale has recently been developed to address some disadvantages of the PASI and PGA. OBJECTIVES: To evaluate the inter-rater and intrarater reliability of the PASI, PGA and LS-PGA. METHODS: On the day before the study, 14 dermatologists (raters), with varied experience of assessing psoriasis, received detailed training (2.5 h) on use of the scales. On the study day, each rater evaluated 16 adults with chronic plaque psoriasis in the morning and again in the afternoon. Raters were randomly assigned to assess subjects using the scales in a specific sequence, either PGA, LS-PGA, PASI or PGA, PASI, LS-PGA. Each rater used one sequence in the morning and the other in the afternoon. The primary endpoint was the inter-rater and intrarater reliability as determined by intraclass correlation coefficients (ICCs). RESULTS: All three scales demonstrated 'substantial' (a priori defined as ICC > 80%) intrarater reliability. The inter-rater reliability for each of the PASI and LS-PGA was also 'substantial' and for the PGA was 'moderate' (ICC 75%). CONCLUSIONS: Each one of the three scales provided reproducible psoriasis severity assessments. In terms of both intrarater and inter-rater reliability values, the three scales can be ranked from highest to lowest as follows: PASI, LS-PGA and PGA.

Adult↗

The diagnosis of hypersensitivity to ingested foods. Reliability of skin prick testing and the radioallergosorbent test with different materials.

The diagnostic reliability in food allergy of skin prick tests (SPT) and the radio-allergosorbent test (RAST) was investigated in paediatric patients with respiratory and skin allergies. SPT and RAST were found to be reliable for the diagnosis of allergy to codfish, peas, nuts, peanuts and egg white. Positive SPT and RAST to cereals were common, but were most often without clinical significance or were correlated with respiratory allergy to the inhalation of flour dust. SPT and RAST were only partly reliable with regard to allergy to cow's milk, and were mostly reliable when used together and showing corresponding results. Experimental allergosorbents for RAST with soy beans and white beans were not reliable. The study shows the need to improve the diagnostic materials and to establish the diagnostic reliability of the material and tests used for each food item in question.

Adolescent↗

Interexaminer reliability of six orthopaedic tests in diagnostic subgroups of craniomandibular disorders.

Interexaminer reliability is defined as the degree of consistency among examiners when making observations of the same clinical variable. In the present study, the interexaminer reliability of six orthopaedic tests was determined in a group of 79 patients with signs and/or symptoms of craniomadibular disorders (CMD), subdivided into three subgroups of patients with a mainly myogenous, a mainly arthrogenous, and a combined myogenous and arthrogenous disorder. Multi-test scores were composed for each test and combinations of tests for the three main symptoms of CMD, viz. pain, joint noises and restriction of movement. Although the orthopaedic tests showed different reliability scores, overall reliability of the determination of these three main symptoms of CMD was satisfactory. In the subgroups, arthrogenous signs and symptoms could be determined reliably with the set of six tests, whereas the reliability of the tests in determining pain and joint noises in the myogenous group was rather low. It may be concluded that the tests are well suited to evaluate arthrogenous signs and symptoms, but that the clinician should be aware of erroneous results of the tests in evaluating pain of a myogenous origin.

Adult↗

Reliability: on the reproducibility of assessment data.

CONTEXT: All assessment data, like other scientific experimental data, must be reproducible in order to be meaningfully interpreted. PURPOSE: The purpose of this paper is to discuss applications of reliability to the most common assessment methods in medical education. Typical methods of estimating reliability are discussed intuitively and non-mathematically. SUMMARY: Reliability refers to the consistency of assessment outcomes. The exact type of consistency of greatest interest depends on the type of assessment, its purpose and the consequential use of the data. Written tests of cognitive achievement look to internal test consistency, using estimation methods derived from the test-retest design. Rater-based assessment data, such as ratings of clinical performance on the wards, require interrater consistency or agreement. Objective structured clinical examinations, simulated patient examinations and other performance-type assessments generally require generalisability theory analysis to account for various sources of measurement error in complex designs and to estimate the consistency of the generalisations to a universe or domain of skills. CONCLUSIONS: Reliability is a major source of validity evidence for assessments. Low reliability indicates that large variations in scores can be expected upon retesting. Inconsistent assessment scores are difficult or impossible to interpret meaningfully and thus reduce validity evidence. Reliability coefficients allow the quantification and estimation of the random errors of measurement in assessments, such that overall assessment can be improved.

Bias↗

The development, validity and reliability of a multimodality objective structured clinical examination in psychiatry.

OBJECTIVES: To evaluate the development, validity and reliability of a multimodality objective structured clinical examination (OSCE) in undergraduate psychiatry, integrating interactive face-to-face and telephone history taking and communication skills stations, videotape mental state examinations and problem-oriented written stations. METHODS: The development of the OSCE on a restricted budget is described. This study evaluates the validity and reliability of 4 15-18-station OSCEs for 128 students over 1 year. Face and content validity were assessed by a panel of clinicians and from feedback from OSCE participants. Correlations with consultant clinical 'firm grades' were performed. Interrater reliability and internal consistency (interstation reliability) were assessed using generalisability theory. RESULTS: The OSCE was feasible to conduct and had a high level of high perceived face and content validity. Consultant firm grades correlated moderately with scores on interactive stations and poorly with written and video stations. Overall reliability was moderate to good, with G-coefficients in the range 0.55-0.68 for the 4 OSCEs. CONCLUSIONS: Integrating a range of modalities into an OSCE in psychiatry appears to represent a feasible, generally valid and reliable method of examination on a restricted budget. Different types of stations appear to have different advantages and disadvantages, supporting the integration of both interactive and written components into the OSCE format.

Clinical Competence↗

Pretreatment blood pressure reliably predicts progression of chronic nephropathies. GISEN Group.

BACKGROUND: Random, nontimed blood pressure (BP) measurements in the outpatient clinic may fail to provide reliable information on actual daily BP control in renal patients on chronic antihypertensive therapy. METHODS: In a cohort of 163 patients with proteinuric chronic nephropathies followed prospectively with repeated BP and glomerular filtration rate (GFR) measurements, we compared baseline and follow-up pretreatment, morning ("trough," measured by standard procedures, and "0 minutes," measured by an automatic device) and post-treatment (120 minutes) measurements, with BP monitored up to 600 minutes after treatment administration. We then evaluated which BP value most reliably predicted GFR decline (delta GFR) and progression to end-stage renal failure (ESRF) over a median (interquartile range) follow-up of 20 (9 to 25) months. RESULTS: GFR decline was more reliably predicted by systolic as compared with diastolic BP and by pretreatment as compared to post-treatment BP, regardless of the timing and method of measurement, respectively. In particular, at the 120-minute baseline and follow-up measurements, systolic BP had no predictive value in patients with less severe renal insufficiency and baseline diastolic BP, regardless of the level of renal dysfunction. The BP predictive value was remarkably higher in ramipril than in conventionally treated patients. All follow-up-but no baseline-measurements reliably predicted the risk of ESRF in the entire study group. CONCLUSIONS: In patients with progressive chronic nephropathies, systolic BP and pretreatment morning BP measurements are the most reliable predictors of disease outcome and may serve to guide antihypertensive therapy in routine clinical activities and in prospective controlled trials, particularly in patients on angiotensin-converting enzyme inhibitor therapy. Reliability and relevance of single measurements taken at different times after treatment administration are questionable.

Adult↗

Retrospective self-report of alcohol consumption: test-retest reliability by telephone.

The Timeline Follow-Back (TLFB) is an interview technique for obtaining detailed retrospective self-reports of alcohol consumption with excellent reliability for various composite variables when both administrations are in person. Because the telephone offers practical advantages over face-to-face interviewing for follow-up assessments in longitudinal studies of problem drinkers, this study was undertaken to compare the test-retest reliability of a 12-week TLFB interview when the second administration was by telephone to that when the second interview was in person. In addition, because the reliability of the TLFB has been previously assessed using composite variables, we examined the reliability of the TLFB at the item level. Research participants were 30 adult medical patients who drank frequently, and 75 college students who were problem drinkers. Test-retest reliability as measured by intraclass correlation was generally high, 0.79 or greater for the number of days of drinking > 6 standard drinks, 0.90 or greater for the number of abstinent days, and 0.80 or greater for the greatest number of drinks consumed on any 1 day, in both the most recent 4-week interval and in the entire 12-week interval. Test-retest correlation coefficients for composite variables derived from the interview data were not systematically affected by whether the second interview was in person or by telephone. Furthermore, item-level correlations were also substantial. Findings support the use of the telephone for follow-up interviews, potentially reducing costs of longitudinal studies and facilitating multisite studies with centralized data collection, and lend further general support to the reliability of the TLFB.

Adolescent↗

Long-term reliability and validity of alcoholism diagnoses and symptoms in a large national telephone interview survey.

The long-term reliability and validity of telephone lay interview assessments of alcoholism were examined in the context of a large national community-based survey of over 8,000 male Vietnam era veterans. A subsample of 146 men was interviewed twice by telephone using the same structured interview an average of 15 months apart to evaluate the long-term reliability of alcoholism symptoms and diagnoses. In addition, a search of Department of Veterans Affairs patient treatment files of inpatient hospitalizations between 1970 and 1993 yielded a subsample of 89 interviewed men with a past discharge diagnosis of alcohol dependence. The test-retest reliability of alcohol abuse and alcohol dependence diagnoses was good, with kappa coefficients of 0.74 and 0.61, respectively. The reliability of individual alcoholism symptoms was fair to good, with kappas of 0.46 to 0.67. Ninety-six percent of individuals identified by Department of Veterans Affairs patient treatment files as having an alcohol dependence diagnosis were correctly diagnosed by the telephone interview. The results of the present study provide additional evidence for the long-term reliability and validity of lifetime alcoholism diagnoses, and suggest that the reliability and validity of telephone interview assessments of alcoholism are as good as that of an in-person interview. Telephone administration of structured psychiatric interviews appears to be an attractive alternative to in-person interviewing for gathering information about alcoholism and alcohol-related problems.

Adult↗