Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Reliability of substance use disorder diagnoses among African-Americans and Caucasians.

Cross cultural research on substance use disorders (SUD) demands diagnostic measures and criteria that apply equally well to persons of different ethnic backgrounds. To evaluate the reliability of SUD in different ethnic groups, comparisons were made of the one week test/retest agreement on DSM-IV lifetime dependence disorders for 196 African-American (AA) and 107 Caucasian (C) respondents using the Composite International Diagnostic Interview-Substance Abuse Module (CIDI-SAM). Overall we found excellent reliability, using kappa (k) statistics, in diagnosing both AA and C respondents with alcohol dependence (AA k = 0.78; C k = 0.80) and opiate dependence (AA k = 0.77; C k = 0.71), good reliability for diagnosing both AA respondents (k = 0.63) and C respondents with cocaine dependence (k = 0.67), and to good reliability for both AA and C respondents with cannabis dependence (AA k = 0.50; C k = 0.69). Reliability of the dependence/abuse criteria was consistent with the overall diagnostic reliability but some variation was noted. No significant differences in the kappas were found between the two ethnic groups for any of the substance dependence diagnoses, and only one dependence or abuse criterion (continued use of cocaine despite physical/psychological problems) differed significantly between AA and C respondents. These initial results indicate that DSM-IV dependence diagnoses as measured by the CIDI-SAM apply equally well to AA and C respondents.

Adolescent↗

Reliability of echocardiographic assessment of left ventricular structure and function: the PRESERVE study. Prospective Randomized Study Evaluating Regression of Ventricular Enlargement.

OBJECTIVES: The study was done to evaluate reliability of echocardiographic left ventricular (LV) mass. BACKGROUND: Echocardiographic estimation of LV mass is affected by several sources of variability. METHODS: We assessed intrapatient reliability of LV mass measurements in 183 hypertensive patients (68% men, 65 +/- 9 years) enrolled in the Prospective Randomized Enalapril Study Evaluating Regression of Ventricular Enlargement (PRESERVE) trial after a screening echocardiogram (ECHO) showed LV hypertrophy. A second ECHO was repeated at randomization (45 +/- 25 days later). Two-dimensional (2D)-guided M-mode or 2D linear measurements of LV cavity and wall dimensions were verified by one experienced reader. RESULTS: Mean LV mass was similar at first and second ECHO (243 +/- 53 vs. 241 +/- 54 g) and showed high reliability as estimated by intraclass correlation coefficient (RHO) = 0.93. Within-patient 5th, 10th, 90th and 95th percentiles of between-study difference in LV mass were -32 g, -28 g, +25 g and +35 g. Mean LV mass fell less from the first to the second ECHO than expected from a formula to predict regression to the mean (2 +/- 19 vs. 17 +/- 12 g, p < 0.001). Reliability was also high for LV internal diameter (RHO = 0.87), septal (RHO = 0.85) and posterior wall thickness (RHO = 0.83). Substantial or moderate reliability was observed for measures of LV systolic function and diastolic filling (RHO from 0.71 to 0.57). CONCLUSIONS: Left ventricular mass had high reliability and little regression to the mean; between-study LV mass change of +/-35 g or +/-17 g had > or = 95% or > or = 80% likelihood of being true change.

Aged↗

Development, test reliability and validation of a classification for revision hip arthroplasty.

The objective of the study was to develop a valid and reliable classification system for failed hip arthroplasties. The study uses research principals derived from multi-attribute utility theory and consensus group techniques. The development of the severity measure was done in two phases. Phase I of the study included: (a) questionnaire development, (b) submission of the questionnaire to the respondents, (c) data synthesis of the responses and item reduction, and (d) classification development and inter-observer reliability testing. Phase II included: (a) resubmission of the instrument to the respondents for suggestions/feedback, (b) instrument revision by the co-investigators based on the respondents' second feedback, and (c) inter-observer reliability testing and intraoperative validity testing of the instrument. The questionnaires sought to capture expert opinion as to what clinical determinants obtained preoperatively (during patient interview, physical exam and review of plain radiographs - AP pelvis and hip lateral) that would in their clinical experience reveal intraoperative severity. There was an 80% (16/20) response rate from the outside experts invited to participate in the study. Based on item reduction and test retest analysis, a five-grade radiographic classification for the acetabulum as well as the femur was developed. This system was then reviewed by 13 of the initial outside experts (16, 80%) who participated in the first round. Inter-rater reliability testing of the final format of the classification revealed a weighted kappa statistic value of 0.88 between the two-blinded raters (inter-rater reliability) and 0.87 between the blinded raters and the reference standard (intraoperative validity). We conclude that the study developed a reliable and valid radiographic classification system for failed hip arthroplasty.

Arthroplasty, Replacement, Hip↗

A comparison of within- and between-day reliability of discrete 3D lower extremity variables in runners.

It is important to understand the day-to-day variability that is attributed to repositioning of markers especially when assessing a treatment effect or response over time. While previous studies have reported reliability of waveform patterns, none have assessed the repeatability of discrete points such as peak angles, velocities and angular excursions which are often used when making statistical and clinical comparisons. The purpose of this study was to compare the within- and between-day variability of discrete kinematic, kinetic, and ground reaction force (GRF) data collected during running. Comparisons for 20 recreational runners were evaluated for within- and between-day reliability of discrete 3D kinematic, kinetic, and GRF variables. The results indicated that within-day comparisons were more reliable than between-day. Joint angular velocity and angular excursion values were more reliable between-days as compared to absolute peak angle measures and may be more useful in interpreting changes in treatment over time. Between-day kinematic and kinetic sagittal plane values were more reliable than secondary plane values. Reliability of GRF data was greater than kinematic and kinetic data for between-day comparisons.

Adult↗

Reliability of the Child Behavior Checklist for the assessment of behavioral problems of children and youth with mild mental retardation.

The assessment of psychopathology in persons with mental retardation requires reliable and valid instruments. In the present study, the reliability of the Child Behavior Checklist was determined, using data of 42 children and youth with mild mental retardation, with ages from 10 to 18 years. Kappa coefficients and intra-class correlations were computed to determine the reliability at item level and syndrome level. At item level, mean kappa's for inter-rater and test-retest reliability were 0.267 and 0.52, respectively. At syndrome level, mean intra-class correlations for inter-rater and test-retest reliability were 0.493 and 0.775, respectively. These results suggest that the Child Behavior Checklist may not always represent a reliable checklist for the assessment of psychopathology among children and youth with mild mental retardation.

Adolescent↗

The reliability and validity of the Tekdyne hand dynamometer: Part I.

The purpose of this study was to evaluate the reliability of Tekdyne hand dynamometer in a controlled loading environment using the Instron 1331. Test-retest reliability, the intertool reliability of three different Tekdyne hand dynamometers, and the effects of surface area and the configuration of the forces applied to the Tekdyne hand dynamometer were studied. In addition, intertool reliability between a Jamar dynamometer modified with foam padding and the same Jamar dynamometer without padding was calculated to explore this device as an alternative measurement device. Intratool and intertool reliabilities of the three Tekdyne tools were high (ICC = 0.993-10.998, SEM = 0.106-0.045 psi and ICC = 0.995, SEM = 0.080 psi, respectively) when tested on the Instron. Both the surface area and the configuration of the force applied to the Tekdyne dynamometer appeared to influence the output reading of this device. The measurements obtained on the Tekdyne hand dynamometer correlated well with those obtained with the Jamar hand dynamometer when a controlled load was applied (r = 0.988). The Tekdyne hand dynamometer is a reliable tool in a controlled loading environment; however, further study is needed to determine its validity with respect to the Jamar dynamometer for testing human grip strength.

Equipment Design↗

Reliability and sensitivity to change assessed for a summary measure of lower body function: results from the Women's Health and Aging Study.

A summary performance measure comprised of a hierarchical balance task, a 4-meter walk, and five repetitive chair stands is increasingly being used as a predictor of independent living for older persons. The reliability and sensitivity to change of this summary performance measure have not been investigated, however. Because a measure can be reliable while being unresponsive to change, this study presents information on both the reliability and sensitivity to change for the summary performance measure. This is a 3-year prospective cohort study of 1,002 moderately to severely disabled older women. Short- and long-term reliability was assessed by intraclass correlation coefficients (ICC). Sensitivity to change was assessed by slope differences for three age categories (65-74, 75-84, and >or=85) over six 6-month follow-up periods. Sensitivity to change was also assessed by summary performance change scores for those who did and did not suffer from one of four medical events [myocardial infarction (MI), stroke, hip fracture, or congestive heart failure (CHF)] at follow-up. The summary performance measure showed excellent reliability. Intraclass correlation coefficients ranged from 0.88 to 0.92 for measures made 1 week apart. The 6-month average intraclass correlation coefficient was 0.77 (range 0.72-0.79). The summary performance measure was also highly responsive to change. Subjects who suffered an incident MI, stroke, hip fracture, or CHF at follow-up were significantly more likely to have poorer summary performance change scores (-2.25) compared with those who did not have one of these medical events (-0.24). Additionally, subjects who suffered one of these events improved their summary performance scores in the following assessment period by 0.72. With increasing utilization of the summary performance measure by researchers and clinicians it is important that the measurement properties of this instrument are known. Our results show that the summary performance measure has excellent reliability and is highly sensitive to change.

Activities of Daily Living↗

The test-retest reliability of the frequency of multiple drug use in young drug users entering treatment.

Assessment of multiple drug use relies primarily on self-report. Several studies support the reliability of client self-reports of drug use but these studies have not involved assessment of the actual frequency of drug use. This test-retest reliability study assessed the frequency of drug use in a clinical sample of 103 multiple drug users, aged 16-25 years. At initial assessment, all participants completed the Drug Use History Form (DUHF) that inquired about the number of drug-using days and the daily frequency of use for 13 drug classes during four time intervals. The DUHF was readministered 2-4 weeks later. Reliability was assessed using Intra-class correlations (ICC's). The results indicated that clients do, in general, reliably report both the number of days of use and daily frequencies. The two frequency measures were not highly correlated. Reliability estimates declined over time but most markedly after 90 days, suggesting that assessments of drug use can be reliably extended beyond 30 days. Frequency estimates based solely on the number of days of use of a substance may be unreliable estimates of actual drug consumption, indicating limitations to this commonly used outcome measure.

Adolescent↗

Reliability of the life chart schedule for assessment of the long-term course of schizophrenia.

We report on the inter-rater reliability of the Life Chart Schedule (LCS). The LCS is designed to assess the long-term course of schizophrenia in four key domains (symptoms, treatment, residence, and work) over two time periods (past two years, entire period of illness). The subjects were 27 consecutive admissions to a schizophrenia research unit. The LCS was filled out by pairs of raters, blinded to each others' ratings, using the same data (interview with subject and chart). Reliability was examined for 45 LCS ratings selected from all four domains and both time periods. Selected ratings pertained to the duration of specified experiences, the quality of these experiences, and the long-term time trend. The kappa statistic and the intra-class correlation coefficient (ICC) were used to determine inter-rater reliability for continuous and categorical ratings, respectively. LCS ratings proved reliable in all four key domains and both time periods. The reliability was fair to excellent for ratings of duration of experience (ICC ranged from 0.53 to 0.99), quality of experience (kappa ranged from 0.46 to 0. 92) and long-term time trends (kappa ranged from 0.66 to 0.94). The LCS can be used to obtain reliable ratings of the long-term course of schizophrenia in multiple domains.

Adolescent↗

Lateral pelvic displacement during walking: retest reliability of a new method of measurement.

This study examined the retest reliability of a new method of measuring the amplitude and symmetry of lateral pelvic displacement (LPD) during walking. Previous methods of quantifying LPD did not consider variations in walking direction or enable quantification of side-to-side symmetry of LPD. Therefore, the new method was developed to measure the amplitude and symmetry of LPD relative to the base of support provided by the feet. Three-dimensional motion analysis was used to collect the coordinates of markers placed on the sacrum and heels of 20 unimpaired adults. Subjects performed three 10 m walking trials at their preferred speed. Between trial 1 and 2 the markers were removed and then reattached. Between trial 2 and 3 the markers remained attached. Retest reliability was high for consecutive measurements of amplitude of LPD (ICC (2,1)=0.91), and marker reattachment did not affect the degree of reliability (ICC (2,1)=0.94). The reliability was also high for consecutive measurements of symmetry of LPD (ICC (2,1)=0.80). Marker reattachment did however reduce the reliability of symmetry measures (ICC (2,1)=0.54). This study provides data on the retest reliability of a new method of measuring LPD in unimpaired subjects against which the performance of patients can be compared. Copyright 1998 Elsevier Science B.V. All rights reserved

Journal Article↗

Inter- and intrarater reliability of retrospective drug utilization reviewers.

OBJECTIVE: To assess inter- and intrarater reliability among 23 pharmacist and physician retrospective drug utilization reviewers and to assess interrater reliability after a reviewer training session. DESIGN: Exploratory study. SETTING: Maryland Medicaid's retrospective drug utilization review (DUR) program. PARTICIPANTS: 23 physician and pharmacist retrospective drug utilization reviewers. INTERVENTIONS: None. MAIN OUTCOME MEASURES: Profiles rated as "intervention indicated" or "intervention not indicated." Cochran's Q test, overall percent agreement, and the unweighted kappa statistic were used in the analysis of review consistency. RESULTS: Intrarater reliability showed substantial consistency among the 23 reviewers; the percent agreement was 82.9% with kappa = 0.66. Interrater reliability, however, was poor, with an overall agreement of 69.6% and kappa = 0.16. Interrater reliability was also poor after a one-hour reviewer training session (agreement 81.8%, kappa = -0.19). CONCLUSION: The implicit review process used in the retrospective DUR program that we evaluated was unreliable. Since reliability is a necessary but not sufficient condition for validity of an indicator of inappropriate drug use, the validity of the DUR implicit review process is in question.

Data Interpretation, Statistical↗

Test-retest reliability of cognitive EEG.

OBJECTIVE: Task-related EEG is sensitive to changes in cognitive state produced by increased task difficulty and by transient impairment. If task-related EEG has high test-retest reliability, it could be used as part of a clinical test to assess changes in cognitive function. The aim of this study was to determine the reliability of the EEG recorded during the performance of a working memory (WM) task and a psychomotor vigilance task (PVT). METHODS: EEG was recorded while subjects rested quietly and while they performed the tasks. Within session (test-retest interval of approximately 1 h) and between session (test-retest interval of approximately 7 days) reliability was calculated for four EEG components: frontal midline theta at Fz, posterior theta at Pz, and slow and fast alpha at Pz. RESULTS: Task-related EEG was highly reliable within and between sessions (r0.9 for all components in WM task, and r0.8 for all components in the PVT). Resting EEG also showed high reliability, although the magnitude of the correlation was somewhat smaller than that of the task-related EEG (r0.7 for all 4 components). CONCLUSIONS: These results suggest that under appropriate conditions, task-related EEG has sufficient retest reliability for use in assessing clinical changes in cognitive status.

Adolescent↗

Test-retest reliability of the Saladin card.

BACKGROUND: Test-retest reliability is a measure of the confidence that results will be identical when the same patient is measured with an instrument in the same manner on more than one occasion. METHOD: Using the Saladin Near Point Balance Card--an instrument designed to test various near visual findings, including visual acuity, phorias, AC/A ratios, dynamic retinoscopy, fixation disparity, associated phorias, fixation disparity curves, as well as accommodative and vergence facilities--28 first- and second-year optometry students were evaluated on two occasions by one clinician, with the tests separated by approximately two weeks. Subjects were required to demonstrate 20/20 distance visual acuity and at least 100 sec of arc stereopsis. A total of 38 findings were compared, which included near horizontal and vertical phorias, associated phoria, and fixation disparity curves. Twenty-two findings were performed through the subjects' habitual prescription and the remaining findings were taken through the lens indicated by MEM retinoscopy. PURPOSE: The purpose of this study was to investigate whether the Saladin Near Point Balance Card had acceptable test-retest reliability. If reliability can be demonstrated, this instrument could be used in clinical situations to diagnose visual departures from normal. RESULTS: All but two of the 38 tests performed demonstrated acceptable test-retest reliability. Coefficients of repeatability and 95% limits of agreement were calculated. The Saladin Near Point Balance Card demonstrated acceptable test-retest reliability. CONCLUSION: Since it is light-weight, portable, easily and quickly administered, and reliable, the Saladin card can be used by clinicians who are performing screenings or examinations in non-clinical situations, such as nursing homes or schools.

Adult↗

Reporting of instrument validity and reliability in selected clinical nursing journals, 1989.

Before research findings are applied to practice, the quality of the research must be assessed so that flawed research does not lead inadvertently to flawed practice. Two critical indicators of research quality are the validity and reliability of the data collection instruments. This article summarizes the principles of instrument validity and reliability and identifies deviations from these principles in a random sample of 55 research studies published in 1989 in five refereed nursing journals targeted toward practicing clinicians. Using a valid and reliable instrument, the investigators found that even with a policy of giving authors "the benefit of the doubt," 47% of the research studies contained no evidence of validity for any data collection instruments and 36% had no evidence of reliability; 29% had no evidence of either validity or reliability. Content validity, a basic requirement for all research instruments, was addressed in only 27% of the studies. This article provides documentation, justification, and suggestions for nursing educators, journal editors, and researchers to take action to improve the reporting of instrument validity and reliability to help ensure the quality of the research on which nursing practice is based.

Clinical Nursing Research↗

The reliability of reporting adverse experiences.

The reliability of reporting of life-events was examined in 52 subjects attending clinics. An inventory of events and longer-standing difficulties was administered on 2 occasions, 7-14 days apart. High levels of reliability were found for the number of events, the mean score for distress or change over all events, and for the single event with highest score. The reporting of individual events was less reliable: only 70% of those events reported at either interview were reported under the same heading at both interviews. Subjective reactions to events differ in reliability according to the type of response, and they are less reliable for single events than overall. Lastly, the reliability of highly distressing events in fact lower than for the less distressing. These findings point to some of the shortcomings of inventory methods in life-event research.

Female↗

Reliability, validity and ease of use of a portable point-of-care coagulation device in a pharmacist-managed anticoagulation clinic.

UNLABELLED: In a pharmacist-managed anticoagulation clinic, portable point-of-care coagulation devices may facilitate patient monitoring by providing rapid INR measurement. Few studies, however, have validated this type of device. OBJECTIVE: To evaluate the reliability, validity and ease of use of the CoaguChek S, a new portable coagulation device. METHODS: A total of 100 patients followed at a pharmacist-managed anticoagulation clinic attended two study visits. INRs were measured using the CoaguChek S and the standard laboratory technique. RESULTS: Reliability: The test-retest reliability (precision) of the CoaguChek S, estimated by the intraclass correlation coefficient (ICC) and a 95% confidence interval (95% CI), was high (0.98 (0.98-0.99)) and comparable to the standard laboratory technique (0.99 (0.98-0.99)). Interrater reliability was also high (0.97 (0.95-0.98)). Reliability coefficients did not vary with the test-strip lot number nor the CoaguChek S operator. VALIDITY: When compared with standard laboratory procedure, the ICC (95% CI) was equal to 0.93 (0.91-0.95). The mean difference (95% CI) between INR measured by the laboratory and the CoaguChek S was equal to -0.02 units (-0.06-0.03). The mean absolute and relative absolute differences (95% CI) were equal to 0.24 units (0.21-0.27) and 9% (8%-10%), respectively. Differences tended to increase for INRs greater than 3 units as seen by a mean difference (95% CI) of -0.17 units (-0.35-0.02). This represented a mean absolute difference (95% CI) of 0.44 units (0.33-0.55) and a mean relative absolute difference of 12% (9%-15%). Concordance between therapeutic decisions based on CoaguChek S and laboratory results was high (Kappa = 0.68). In 34 cases (18%), the therapeutic decision would have been different. However, in 15 of these discordant observations, the difference between the CoaguCheck S and laboratory INR was <or=0.25 units. Ease of use: In 3% of cases, no INR could be measured by the CoaguChek S. The percentage of extra finger pricks and extra test-strips were equal to 25.8% and 23.7%, respectively. CONCLUSIONS: When used by health professionals in a pharmacist-managed anticoagulation clinic, the CoaguChek S is reliable, valid and easy to use. However, its validity tends to decrease as the INR increases, possibly due to the low sensitivity of the thromboplastin. If the CoaguChek S INR is supratherapeutic, we would therefore recommend confirming the results with a standard laboratory measurement.

Adult↗

Interrater reliability of diagnosing complex regional pain syndrome type I.

BACKGROUND: Diagnosis of complex regional pain syndrome type I (CRPS I) is based on clinical observation of symptoms. As little information is available on the reliability of CRPS I diagnosis, we evaluated the agreement between therapists with regard to the presence and severity of CRPS I and its symptoms. METHODS: The interrater reliability was evaluated in 37 presumed CRPS I patients by three observers; one consultant anesthesiologist and two resident anesthesiologists. Patients were assessed on the basis of Veldman's CRPS criteria. RESULTS: The interrater reliability for diagnosing CRPS I was good for the majority of observer combinations. The percentage of agreement for the absence or presence of CRPS I was good (88%-100%). Cohen's Kappa's ranged from 0.60 to 0.86. The agreement for the mean symptom score ranged from 70.2% to 88.6%; Kappa's were lower and showed more variation. Interrater reliability for assessment of the severity of CRPS I and its symptoms was poor. Factors influencing the interrater reliability were symptom type, individual observers and sample population. CONCLUSION: Diagnosing CRPS I can be performed on the basis of clinical observation. Further assessment of severity of CRPS I and its symptoms should be performed with reliable and valid measurement instruments.

Adult↗

Beyond the Spearman-Brown: a structural approach to maximal reliability.

The requirement of parallel parts has long been the cornerstone of classic reliability theory. By recasting reliability in a structural equation framework, items, raters, or judges no longer need to be treated as equivalent entities. Instead, unique reliability estimates can be determined for each and collectively used to assess the maximal reliability of a weighted composite, with the composite reliability submitted to inferential test. Procedures are shown to generalize from single to multifactor applications. Ramifications of a structural approach to reliability determination are probed, and the dilemma posed by possible falsification of the true score hypothesis presented for individual researcher consideration.

Humans↗