Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

The validity and reliability of pain measures in adults with cancer.

To be most useful, clinical trials of cancer pain treatments should use pain measures that are both reliable and valid. A great variety of measures are now available that may be used to assess cancer pain. However, there are not yet any clear guidelines for selecting one or more measures over the others. The purpose of this article is to summarize the evidence concerning the validity and reliability of cancer pain measures. One hundred sixty-four articles were identified that provided psychometric data of pain measures among patients with cancer. The results indicate that commonly used single-item ratings of pain intensity are all valid and adequately reliable as measures of pain intensity, although some scales appear to be easier for patients with cancer to understand and to use than others. Multiple-item measures of pain intensity are reliable, but evidence concerning their validity is lacking. There is a paucity of research examining the psychometric properties of measures of cancer pain interference, pain relief, pain site, the temporal aspects of pain, and pain quality. This lack of evidence limits the conclusions that may be drawn concerning the reliability and validity of these other pain measures. Composite measures that combine ratings of pain intensity and pain interference into a single score appear to be both valid and reliable for describing patient populations, although their usefulness in clinical trials may be limited because they can obscure the contributions of intensity and interference to the total score. Proxy measures of cancer pain (pain ratings made by someone other than the patient) may be useful when patients are not able to provide pain ratings, but they should not be used as replacements for patient ratings when patient self-report measures are available. The discussion includes specific recommendations for selecting from among the available pain measures, as well as recommendations for future research into the assessment of cancer pain.

Adolescent↗

[The reliability of the german version of the barthel-index and the development of a postal and telephone version for the application on stroke patients].

BACKGROUND AND PURPOSES: Aim of the study was to translate the original version of the Barthel-Index (BI) into German and to investigate the reliability of the German version. In addition, a German version of the BI for postal and telephone use was developed. METHODS: Data were collected in four neurological hospitals in Germany. The translation of the BI followed the protocol of the Medical Outcomes Trust. The interrater reliability of the German version of the BI was investigated in 72 patients after acute stroke. The reliability of the postal and telephone version of the BI was compared with face-to-face interview in 147 patients three months after stroke. Reliability was assessed using simple weighted kappa-statistics. RESULTS: The interrater reliability of the German version of the BI was excellent (mean kappa 0.93). The mean kappa coefficient was 0.79 for the postal version of the BI and 0.80 for the telephone version. Thus, the agreement between the postal and the telephone administration of the BI compared to the face-to-face interview was substantial to excellent. CONCLUSIONS: Our study published the first German version of the BI which was investigated for interrater reliability in a standardized way. The development of a postal and a telephone version allows the widespread use of the German BI for the follow up of stroke patients in different access paths.

Acute Disease↗

Chiropractic biophysics digitized radiographic mensuration analysis of the anteroposterior cervicothoracic view: a reliability study.

OBJECTIVE: To investigate the reliability of a radiographic measurement procedure that uses a computer and sonic digitizer to determine projected spinal displacements from an ideal, normal position. DESIGN: A blind, repeated-measure design was used. Anteroposterior cervicothoracic spine radiographs were presented in random order to each of 3 examiners. Each film was digitized, and the films were randomized for a second examination. SETTING: Private, primary care chiropractic clinic. MAIN OUTCOME MEASURES: Intraclass correlation coefficients for intraexaminer and interexaminer reliability for measures on radiographs comparing the perpendicular distance (T(x)) from a vertical axis line drawn through the center of T4 and the center of C2, the linear distance (vertebra(apex)) from the center of the vertebra most displaced from a line connecting the centers of C2 and T4, the angle (Rz) formed by the intersection of the vertical axis line and the upper thoracic line, and the angle of intersection (CDA) between the upper thoracic line and the cervical line. RESULTS: Intraexaminer reliability for T(x) distance was 0.99 to 1.00, with confidence intervals from 0.98-1.00; for vertebra(apex) was 0.96 to 0.97, with confidence intervals from 0.92-0.98; for Rz was 0.94 to 0.98, with confidence intervals from 0. 89-0.99; and for CDA was 0.92 to 0.95, with confidence intervals from 0.84-0.97. Interexaminer reliabilities for the 3 examiners ranged from 0.97 to 0.99. CONCLUSIONS: Measures similar to those described in this study are commonly used to quantify and categorize spinal displacements from true vertical alignment (i.e., scoliosis measurements). Intraclass correlation coefficient values >0.70 are considered accurate enough for use in clinical and research applications. The measures tested here would fit within these guidelines of reliability. Establishing reliability is an important first step in evaluating these measures so that future studies of validity may be undertaken.

Biophysics↗

Reliability of 3 methods for assessing shoulder strength.

The reliability of tests for isometric strength of the shoulder joint in symptomatic subjects has yet to be established. For this purpose, interrater and intrarater agreement trials were undertaken to ascertain the reliability of manual muscle tests, a handheld dynamometer, and a spring-scale dynamometer for 5 different shoulder movements in symptomatic subjects. Intraclass correlation coefficients were calculated from a random-effects model. All movements tested with the handheld dynamometer demonstrated excellent reliability for the interrater trial (rho = 0.79-0.92). Excellent reliability was also demonstrated for elevation, external rotation, and internal rotation for the intrarater trial (rho = 0.79-0.96). For the interrater trial, measurement of the lift-off maneuver with the handheld dynamometer was significantly more reliable than with manual muscle tests (P =.002). In summary, the handheld dynamometer was the most reliable and discriminatory means for assessing strength of the rotator cuff in symptomatic subjects.

Adult↗

Intrarater and interrater reliability of a manual technique to assess anterior humeral head translation of the glenohumeral joint.

The purpose of this study was to determine the intrarater and interrater reliability of a manual anterior humeral head translation test. Fifteen subjects were positioned lying in a supine position with their identity shielded from examiners. A standard manual anterior humeral head translation test was performed and repeated with the glenohumeral joint in 90 degrees of elevation in the scapular plane, with use of the grading method proposed by Altchek and Dines in 1993. Reliability was assessed with the coefficient of agreement and kappa statistic. Intrarater reliability was 81.4% comparing grade I and II translation. This decreased to 54% when examiners distinguished between grades I, I+, II, and II+. Interrater reliability for the same comparisons was 70.4%, decreasing to 37.3%. On the basis of these data, the technique of manually assessing anterior humeral head translation studied has poor overall interrater reliability and only fair intrarater reliability. The test-retest accuracy of humeral head translation is enhanced when examiners only determine the relationship of the humeral head relative to the glenoid rim.

Female↗

A questionnaire assessment of nutrition knowledge--validity and reliability issues.

OBJECTIVE: This study describes an evaluation of validity and reliability measures in a questionnaire designed to assess knowledge of applied nutrition in children participating in an after-school care dietary intervention programme being undertaken in an area of high social disadvantage. DESIGN: Three domains were assessed: Knowledge of Applied Nutrition (KN), Knowledge of Food Preparation (KP) and Perceived Confidence in Cooking Skills (PC). Four pilot studies were undertaken to determine item reliability, test-retest reliability, discrimination and difficulty indices, and content, cognitive and face validity. SETTING: Primary schools in Dundee, Scotland and Newcastle upon Tyne, England. SUBJECTS: Ninety-eight children aged 11 years. RESULTS: The final instrument comprised 36 questions (18 KN items, 9 KP items and 9 PC items) presented on four sides of paper, which could be self-completed in less than 15 minutes. Question formatting included open and closed structures (KP) and multiple choice (KN and PC) items. All knowledge questions could be answered correctly by 5 to 95% of the target population, with discrimination scores ranging from 0.06 to 0.83. Retest reliability scores were significant (KN 0.458, KP 0.577, PC 0.381, ) and internal reliability (Cronbach's alpha) of each component was also significant. CONCLUSION: The test meets basic psychometric criteria for reliability and validity and forms a suitable instrument for measuring changes associated with intervention work aimed at improving food and dietary knowledge.

Child↗

Reliability and validity of a modified gait scoring system and its use in assessing tibial dyschondroplasia in broilers.

1. The gait scoring system for broilers developed by Kestin et al. (Veterinary Record, 131: 190-194, 1992) has been widely used to evaluate leg problems. The many factors and measures associated with this scale have empirically established its external (biological) validity. However, published test-retest (within-observer) reliabilities are poor, and inter-observer reliabilities are unknown. We evaluated several modifications to this scale aimed at improving its objectivity and reliability. 2. Eighteen naïve observers scored a standardised video of birds exhibiting varying degrees of lameness, either using Kestin et al.'s system, or our modified system. 3. Test-retest reliability (0.906) for Kestin et al.'s system was higher than previously reported. Inter-rater reliability was also good (0.892). The modified system offered significantly better test-retest (0.948) and inter-rater reliabilities (0.943), without incurring costs in terms of time taken or difficulty of use. The systems were consistent, assigning individual birds the same score on average. 4. It is concluded that the modified system offers the advantages of reduced error within and between studies. 5. In a second experiment, we used our modified scoring system to examine the relationship between tibial dyschondroplasia (TD) and gait score in 267 selected broilers. 6. Neither the presence nor severity of TD affected gait score, suggesting that, at least in this strain of broilers, other leg problems like slipped tendons or torsional deformities had more influence on gait impairment than did TD.

Animals↗

Reliability testing of the Danish version of the Kidney Disease Quality of Life Short Form.

OBJECTIVE: The questionnaire Kidney Disease Quality of Life Short Form version 1.3 (KDQOL-SF) is valuable for assessing the health-related quality of life in patients treated with chronic dialysis. The aim of this study was to translate and test the reliability of the KDQOL-SF for use in Denmark. MATERIAL AND METHODS: Translation into Danish and back-translation into English were performed. Pilot, field and internal consistency reliability tests were performed. RESULTS: Cronbach's alpha coefficients for the internal reliability test ranged from 0.77 to 0.93 for the eight generic scales. In a test involving all patients, two of the disease-specific scales had Cronbach's alpha coefficients of <0.70 ("social support" = 0.67; and "quality of social interaction" = 0.43). After removing one item from the scale "quality of social interaction", Cronbach's alpha reached 0.63. A test of the scores of peritoneal dialysis (PD) patients discovered low reliability for three disease-specific scales. The KDQOL-SF manual and the Danish manual for the Short Form 36 (SF36) differed in the scoring of four generic scales: "role limitation-physical", "bodily pain", "general health" and "social function". CONCLUSIONS: With the exception of the scale "quality of social interaction" the Danish translation of the KDQOL-SF achieved values in the internal consistency reliability test of the same level as the original U.S. version. When data were stratified according to dialysis treatment, the reliability of PD patients scores was lower. Generic data from the questionnaire SF36 should be scored according to the Danish SF36 manual.

Aged↗

Small group learning in the final year of a medical degree: a quantitative and qualitative evaluation.

The new undergraduate medical curriculum in Manchester uses problem-based learning (PBL) throughout the course. However, the major difference from other PBL schools is that in years 3 & 4 (phase 2) the students can use clinical experience when discussing the paper cases. The process is then developed further in year 5 (phase 3), in which there are no set PBL 'triggers' and students bring their own cases to the groups for discussion. In this study, we have explored what happens in the phase 3 (year 5) group sessions and how the students view them. A questionnaire and focus groups were used to generate data, from which a model was developed of what happens in a 'good' group session. The data suggest that most groups run on a case-presentation and discussion format, most commonly about clinical management and diagnosis. Students want tutors to act as an expert resource and to be flexible in allowing students to direct the discussions. University guidance about the group sessions was not generally used.

Journal Article↗

Ensuring reliability in UK written tests of general practice: the MRCGP examination 1998-2003.

Reliability in written examinations is taken very seriously by examination boards and candidates alike. Within general education many factors influence reliability including variations between markers, within markers, within candidates and within teachers. Mechanisms designed to overcome, or at least minimize, the impact of such variables are detailed. Methods of establishing reliability are also explored in the context of a range of assessment situations. In written tests of general practice within the Membership of the Royal College of General Practitioner (MRCGP) examination considerable effort has put been put into achieving acceptable levels of reliability. Current mechanisms designed to ensure high reliability are described and related to the evolution of the written component of the examination. In addition to description of marker selection and training, question development including construct a detailed example of specific and generic marking schedules is provided. Examination results for the Written Paper of the MRCGP from 1998 to 2003 are reported including Cronbach's alpha coefficients and standard error of measurements, mean scores (and SD) and pass rates. In addition individual discrimination scores for each question in the October 2002 paper are shown. Consistent high reliability of the written component of the MRCGP examination provides valuable lessons in terms of selection, training and monitoring of markers as well as practical methods of moderating factors affecting candidate variability. The challenge for examination developers is to carry these important lessons forward into a modernized assessment structure of UK general practice.

Education, Medical, Graduate↗

Reliability of the Rey-Osterrieth Complex Figure in use with memory-impaired patients.

Rater reliability was evaluated for the system most widely used to assess copy and recall of the Rey Complex Figure: the Osterrieth (1944) 18-item scoring system. The study sample consisted of 95 subjects (49 males, 46 females), most of whom were elderly individuals (M = 59.83, SD = 15.21 years) suffering from memory impairment. Four raters rated copy and delayed-recall protocols, and three raters re-rated the protocols after an interval of 3 months. Results revealed excellent inter- and intra-rater reliability coefficients (.85-.97) for total scores. However, reliabilities for the 18 individual items ranged from poor (.14) to excellent (.96). Differences in both reliability and level of subject performance were observed as a function of item and conditions of copy versus recall. It is concluded that the Osterrieth scoring system supports excellent reliability in use with memory-impaired patients using total scores. Nevertheless, individual-item reliability would benefit from enhancement, for example, via amplified delineation of relevant decision criteria.

Adult↗

The COOP/WONCA charts in an acute psychiatric ward. Validity and reliability of patients' self-report of functioning.

We wanted to examine whether persons needing acute psychiatric admittance can give us reliable information about their functional status with the COOP/WONCA charts (Dartmouth Primary Care Cooperative Information Project-World Organization of National Colleges, Academies and Academic Associations of General Practitioners/Family Physicians), and if this information parallel that of therapists and ward staff. To examine internal consistency, external reliability and test-retest reliability, all patients consecutively admitted to an acute psychiatric department were asked to fill in the COOP/WONCA, Short-Form-36 Health Survey (SF-36) and Symptom Checklist 90--Revised (SCL-90-R) 1-2 weeks after admittance. Their therapists scored a special version of the GAF, and the ward staff scored an observer version of the COOP/WONCA. The first 59 patients repeated the COOP/WONCA scoring after a few days to estimate of test-retest reliability. Of a total of 267 persons admitted in the study period, 102 were included. Non-inclusion was largely due to early discharge. The internal consistency of the COOP/WONCA corresponded to a Cronbach's alpha=0.85. External reliability versus the SF-36 resulted in a correlation of r=-0.82 (P<0.001), and versus the observer version of COOP/WONCA we found r=0.46 (P<0.001). Correlation with GAF was low and not significant. The COOP/WONCA correlated significantly with the SCL-90-R scores (r=0.59, P<0.001). The test-retest correlation was r=0.85 (P<0.001). Correlations varied between sub-groups of patients, but all had consistently high correlations between the various self-scored measures. The COOP/WONCA gives a consistent and reliable report of acutely admitted psychiatric inpatients' own views of their functional capacity. The therapists' views and to some degree the ward staff's views diverge from the patients' opinions, especially for the more seriously disturbed patients.

Acute Disease↗

Evaluation of cardio-respiratory fitness and perceived exertion for patients with depressive and anxiety disorders: a study on reliability.

PURPOSE: The implementation of a physical reconditioning programme for patients with depressive and/or anxiety disorders requires a thorough evaluation of the physical fitness and the perceived exertion during exercise. This implies the use of reliable and clinically useful instruments. The present study examined the reliability of the Franz ergocycle test, as measure for cardio-respiratory fitness, and the Borg Category Ratio 10 Scale, as measure for subject-perceived exertion. METHOD: Sixty-eight hospitalized patients performed test and re-test of the Franz ergocycle test and the Borg CR 10 Scale with a between interval of 1 week. RESULTS: The Physical Work Capacity 130 and the Physical Work Capacity 150, determined by the Franz ergocycle test, had a proper to good test-re-test reliability (r ranged from 0.74 to 0.90). The Borg Category Ratio 10 Scale had a moderate reliability (r ranged from 0.42 to 0.82). CONCLUSIONS: The Franz ergocycle test seems to be a reasonable reliable instrument for measuring physical work capacity of these patients. Possible explanations for the simply moderate reliability of the Borg Category Ratio 10 Scale could be the low level of physical activity prior to hospitalization, and the depressive and anxiety symptoms that might influence the perceived exertion.

Adult↗

The Orpington Prognostic Scale for patients with stroke: reliability and pilot predictive data for discharge destination and therapeutic services.

PURPOSE: To determine the inter-rater and test-retest reliability of the Orpington Prognostic Scale (OPS) in patients with stroke. Pilot data were gathered to evaluate its predictive validity for discharge destination and therapeutic services required on discharge. METHOD: Ninety-four consecutive patients, admitted to hospital due to stroke participated. Pairs of physiotherapists (PT) and occupational therapists (OT) assessed patients using the OPS on days 7 and 14 post stroke. For inter-rater reliability, one rater performed the OPS while the other observed, each scoring the scale independently. For test-retest reliability, two different raters tested the subjects separately within the same day. Data were gathered on the discharge destination and the number of follow-up services prescribed. RESULTS: The inter-rater reliability as measured by the intraclass correlation coefficient (ICC) was 0.99 (95% CI 0.97 - 0.99). For test-retest reliability, the ICC was 0.95 (95% CI 0.90 - 0.98). The accuracy for predicting discharge to home using OPS 5.0 was 65% (95% CI 0.52 - 0.76). OPS scores were not related to number of follow-up services prescribed. CONCLUSIONS: Despite high inter-rater and test-retest reliability, the OPS has limited predictive accuracy for discharge destination and is a poor predictor of follow-up services.

Aged↗

Test-retest reliability, internal item consistency, and concurrent validity of the wheelchair seating discomfort assessment tool.

Discomfort is a common problem for wheelchair users. Few researchers have investigated discomfort among wheelchair users or potential solutions for this problem. One of the impediments to quantitative research on wheelchair seating discomfort has been the lack of a reliable method for quantifying seat discomfort. The purpose of this study was to establish the test-retest reliability, internal item consistency, and concurrent validity of a newly developed Wheelchair Seating Discomfort Assessment Tool (WcS-DAT). Thirty full-time, active wheelchair users with intact sensation were asked to use this and other tools in order to rate their levels of discomfort in a test-retest reliability study format. Data from these measures were analyzed in SPSS using an intraclass correlation coefficient (ICC) model (2,k) to measure the test-retest reliability. Cronbach's alpha was used to examine the internal consistency of the items within the WcS-DAT. Concurrent validity with similar measures was analyzed using Pearson product-moment correlations. ICC scores for all analyses were above the established lower bound of .80, indicating a highly stable and reliable tool. In addition, alpha scores indicated good consistency of all items without redundancy. Finally, correlations with similar tools, such as the Chair Evaluation Checklist and the Short Form of the McGill Pain Questionnaire, were significant at the .05 level, and many were significant at the .001 level. These results support the use of the WcS-DAT as a reliable and stable tool for quantifying wheelchair seating discomfort. Its application will enhance the ability to assess and to research this important problem and will provide a means to validate the outcomes of specialized seating interventions for the study population of wheelchairs users.

Persons with Disabilities↗

Reliability of biomarkers of iron status, blood lipids, oxidative stress, vitamin D, C-reactive protein and fructosamine in two Dutch cohorts.

Biomarkers are widely used in epidemiology, yet there are few reliability studies to assess the appropriateness of using these biomarkers for the assessment of exposure-disease relationships. The aim of the study was to assess the reliability of 20 biomarkers in serum collected from two Dutch centres (Utrecht and Bilthoven) participating in the European Prospective Investigation into Cancer and Nutrition (EPIC) at two points several years apart. Blood samples were collected from 30 men from Bilthoven and 35 women from Utrecht. Ferritin, total iron, total iron-binding capacity, transferrin saturation, transferrin, C-reactive protein, bilirubin, cholesterol, triglycerides, apo lipoprotein-A, apo lipoprotein-B, high-density lipoproteins, low-density lipoproteins, uric acid, creatinine, reactive oxygen metabolites, the ferric-reducing ability of plasma, protein thiol oxidation, fructosamine, and vitamin D biomarkers in serum were analysed from the blood samples at the two points of time. For all biomarkers, except C-reactive protein, there were no substantial changes in the mean levels over time. Uric acid, ferritin, creatinine, HDL, and apo lipoprotein-B levels consistently showed the highest reliability for men and women (intra-class correlation = 0.69-0.86). Among women, the ferric-reducing ability of plasma, and protein thiol oxidation had poor reliability; and among men iron-related biomarkers (except serum ferritin) had poor reliability. With the exception of a few gender-specific differences, most of the 20 biomarkers performed well and can be considered to have sufficient reliability to be used in future cohort studies.

Biomarkers↗

Putney Auditory Single Word Yes/No Assessment (PASWORD). Development of a reliable test of yes/no at a single word level in patients unable to participate in assessments requiring a specific motor response: an exploratory study.

BACKGROUND: There are very few formal language assessments aimed at the very severely neurologically impaired individual. These individuals often have multiple deficits on top of their communication impairment that demand a novel approach to assessment. The authors set out to devise a tool (PASWORD) to enable professionals in this field to screen their clients' ability to understand and respond to very simple, closed questions by using their preferred modality. AIMS: PASWORD assessment was examined for its reliability as a tool for determining whether profoundly neurologically impaired individuals could indicate yes or no reliably to simple closed questions. METHODS & PROCEDURES: The assessment comprises 20 high-frequency objects that can be presented visually, auditorily or in the tactile modality, depending on the subject's abilities. Subjects are then asked a one-word level closed question related to the object and were asked to give a yes or no response. Seventy control subjects with no neurological impairment first underwent the PASWORD test to assess for ambiguities in the test items (group 1, n = 74). Neurologically impaired subjects at the Royal Hospital for Neuro-disability in the two experimental groups were allocated according to the opinion of the experienced multidisciplinary team, which observed the subjects closely over a period of time, in different contexts. Those that the team agreed were reliably able to answer simple yes/no questions on observation (group 2, n = 9) and those the team agreed could not indicate yes/no reliably to closed questions (group 3, n = 7). The results of the PASWORD were then compared with the original opinion of the team. Groups 2 and 3 were studied on two occasions to examine test-retest reliability. OUTCOMES & RESULTS: There were no real ambiguities in the test items and the results of the PASWORD reflected the opinion of the experienced multidisciplinary team. As there are no other tests of single-word closed questions, which can be presented in any of three modalities and during which the subject can respond in any augmentative or alternative fashion, the opinion of the team at the Royal Hospital, a centre of excellence for neurorehabilitation, was used against which to measure reliability. CONCLUSIONS: The findings imply that PASWORD is a useful screening tool when assessing the basic language skills of neurologically impaired individuals, whose extensive physical, linguistic and cognitive deficits often render the administration of more traditional language assessments impossible.

Adult↗

Reliability of electric response audiometry using 80 Hz auditory steady-state responses.

The reliability of the Auditory Steady State Response (ASSR) has not been thoroughly evaluated despite its recent application as a clinical tool for threshold estimation. The purpose of this study was to examine test-retest (TR) reliability of ASSR threshold estimates in an empirical research design. The ASSR, tested using modulation frequencies approximately 80 Hz and above, was evaluated against pure tone audiometry (PTA), and the slow vertex potential (SVP, N1-P2). Sixteen normal-hearing young female adults were tested twice, one week apart. Varying degrees of sensorineural hearing loss of a notched configuration were simulated with filtered masking noise. Test-retest reliability was assessed using Pearson-product moment correlation analysis, supplemented by other post-hoc analyses. Results demonstrated moderately strong TR reliability for ASSR at 1000, 2000 and 4000 Hz (r = 0.83-0.93); however, the reliability of ASSR at 500 Hz was weaker (r = 0.75). Results suggest that ASSR-ERA is a reliable test at mid-high frequencies, at least with the configuration and degrees of simulated sensorineural hearing loss examined in this study.

Adolescent↗