Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Reliability of self-reports: data from the Canadian Multi-Centre Osteoporosis Study (CaMos).

Reliable questions enhance study design. We assessed the reliability of questions that gather demographic, sun exposure, reproductive history, and physical activity information. Subjects were participants in the Canadian Multicentre Osteoporosis Study (CaMos), a cohort study of Canadian adults recruited January 1996 to September 1997 in nine cities, stratified by sex, age, and location. Following personal interviews, 367 subjects were re-administered part of the questionnaire by telephone. Reliability was assessed using kappa and intra-class correlation. Reliability was excellent for employment status, reproductive history, weight and height (0.91 to 0.97), not differing greatly when stratified by age group or sex. Physical activity and sun exposure were reported with fair to good reliability (0.44 to 0.58), except for moderate activity (kappa = 0.30, 95% confidence interval 0.23, 0.37). Stratification by body mass index did not show significant differences. Many items can be reported reliably, especially those of height, weight, employment status and reproductive history, and, to a lesser extent, physical activity and sun exposure. Similar questions might be used reliably in future studies.

Aged↗

Validity and reliability of lupus activity measures in the routine clinic setting.

As part of a cohort study of 150 patients with systemic lupus erythematosus (SLE), we investigated the validity and reliability of several indices of lupus activity, including the UCSF/JHU Lupus Activity Index (LAI), the SLE Disease Activity Index (SLEDAI), and a simple Core Index combining common elements. Validity was assessed by measuring correlations of these indices at the first cohort visit with the physician's global assessment (PGA) of SLE activity. The correlation of M-LAI (LAI modified so as not to contain PGA) and SLEDAI with PGA was 0.64 (95% CI 0.50, 0.70) and 0.55 (95% CI 0.42, 0.64), respectively. Reliability was assessed in a study of 6 patients seen twice, one week apart, by 9 physicians. The interrater reliability and test-retest reliability was greater for LAI (or M-LAI) than for SLEDAI. The Core Index performed better in its correlation with PGA (R = 0.78), although it contained no treatment data or serologic tests. Its interrater reliability and test-retest reliability were comparable with LAI. We conclude that (1) all indices have high validity; (2) LAI and the Core Index have higher reliability; and (3) these indices can be readily assimilated into routine clinic practice.

Adult↗

Is there a difference in the reliable measurement of temporomandibular disorder signs between experienced and inexperienced examiners?

AIMS: To determine whether there is a difference in terms of reliability between experienced examiners and inexperienced examiners in the measurement of signs of temporomandibular disorders (TMD). METHODS: A total of 27 patients seen for treatment of TMD were rated blindly and in random sequence by 2 experienced and 2 inexperienced examiners. The examiners participated in a 4-hour calibration session on the day preceding the reliability study. Both experienced and inexperienced examiners participated in the calibration session to reduce the effect of examiner subjectivity and allow the study focus to be on the effect of experience. The rating followed the Research Diagnostic Criteria for Temporomandibular Disorders and included mandibular movements, joint sounds, and digital palpation of muscles and joints. Intraclass correlation coefficients and kappa statistics were calculated to estimate interrater reliability. The Wilcoxon signed rank test was performed to test for differences between experienced and inexperienced examiners' results, and the Friedman test was used for differences between all 6 examiner combinations. RESULTS: Excellent overall reliability was found for vertical mandibular motions, acceptable reliability was found for the summed muscle palpation pain sites, and moderate to poor reliability was found for excursive movements, joint sounds, and single muscle palpation pain sites. No significant differences in the measurement results could be found between the experienced examiners and the inexperienced examiners. CONCLUSION: Examiner calibration rather than professional experience seems to be the most important factor for reliable measurement of TMD symptoms.

Adolescent↗

Reliability of a measurement of neck flexor muscle endurance.

BACKGROUND AND PURPOSE: Neck flexor muscle endurance has been negatively correlated with cervical pain and dysfunction. The purposes of this study were to determine rater reliability in subjects both with and without neck pain and to determine whether there was a difference in neck flexor muscle endurance between the 2 groups. SUBJECTS: Forty-one subjects with and without neck pain were enrolled in this repeated-measures reliability study. METHODS: Two raters used an isometric neck retraction test to assess neck flexor muscle endurance for all subjects during an initial session, and subjects without neck pain returned for testing 1 week later. RESULTS: For the group without neck pain, intrarater reliability was good to excellent (intraclass correlation coefficient [ICC(3,1)]=.82-.91), and interrater reliability was moderate to good (ICC[2,1]=.67-.78). The associated standard error of measurement (SEM) ranged from 8.0 to 11.0 seconds and from 12.6 to 15.3 seconds, respectively. For the group with neck pain, interrater reliability was moderate (ICC[2,1]=.67, SEM=11.5). Neck flexor muscle endurance test results for the group without neck pain (mean=38.95 seconds, SD=26.4) and the group with neck pain (mean=24.1 seconds, SD=12.8) were significantly different. DISCUSSION AND CONCLUSION: Reliability coefficients differed between the 2 groups and ranged from moderate to excellent and improved after the first test session. The interrater reliability of data obtained with the neck flexor muscle endurance test in people with neck pain must be improved in order for clinicians to distinguish a clinically meaningful change from measurement error. Neck flexor muscle endurance was both statistically and clinically greater for subjects without neck pain than for those with neck pain.

Adult↗

Intra-rater and inter-rater reliability of the 10-point Manual Muscle Test (MMT) of strength in children with juvenile idiopathic inflammatory myopathies (JIIM).

OBJECTIVE: Children with juvenile idiopathic inflammatory myopathies (JIIM) present with muscle inflammation and decreased strength that may affect their functional abilities. The purpose of this study was to determine the intra-rater and inter-rater reliability of the 0 to 10-point manual muscle testing method for children with JIIM. METHODS: For the intra-rater and inter-rater reliability studies, 10 and 9 children with JIIM participated, respectively. For intra-rater reliability, one pediatric therapist completed two assessments in one day with a one-hour break. For inter-rater reliability, four therapists assessed the same child within a single morning. RESULTS: Spearman correlations for intra-rater reliability ranged from 0.70 to 1.00. Kendall's W coefficient for inter-rater reliability of groups of muscles (total, proximal, distal, and peripheral) ranged from 0.51 to 0.76. CONCLUSIONS: The total, proximal, and peripheral Manual Muscle Test (MMT) score, using the 0-10 point scale, has acceptable reliability in JIIM patients.

Adolescent↗

Measurement precision and reliability in craniofacial anthropometry: implications and suggestions for clinical applications.

Craniofacial anthropometry has become an important tool used by both clinical geneticists and reconstructive surgeons. Yet little attention has been paid to the potentially serious problem of measurement error. This paper examines intra-observer measurement error and precision (also called repeatability or reliability) for 52 commonly used anthropometric variables of the head and face. Two factors proved critical to reliability: magnitude of the measurement in question and the degree to which its constituant landmarks could be readily identified. Thus, all of the measurement variables with means above 10 cm proved to have good or excellent reliability. In contrast measurement variables with means below 10 cm were more likely to have poor reliability. This trend was especially evident in variables with means of 6 cm or less where 18 of the 20 variables in this range had poor reliability. The least reliable variables were those like philtrum breadth, columella breadth, and nasal root breadth that combine small magnitude with difficult to define landmarks. While these results suggest that it may be prudent to avoid using craniofacial variables with small dimensions this may be neither practical nor desirable. In such cases repeat measurements may be the best means for optimizing reliability.

Adult↗

Interexaminer reliability of the electromagnetic radiation receiver for determining lumbar spinal joint dysfunction in subjects with low back pain.

Twenty subjects (6 male, 14 female) with low back pain were examined by two experienced and licensed chiropractic doctors (E1 and E2). Both examiners examined the patients using a Toftness Electromagnetic Radiation Receiver (EMRR) and by manual palpation (MP) of the spinous processes. Interexaminer reliability was calculated at three sites (L3, L4, L5) for the following combinations: a) E1,MP--E2,MP; b) E1,EMRR--E2,EMRR; c) E1,MP--E2,EMRR; and) d) E2,MP--E1,EMRR, and intraexaminer reliability was calculated for the following variables: e) E1,MP--E1,EMRR; and f) E2,MP--E2,EMRR. Results of a Kappa coefficient analysis for interexaminer reliability of the stated combinations and at the specific sites were: a) -0.071, 0.400, 0.200; b) -0.013, 0.100, -0.120; c) 0.286, 0.300, 0.200; d) -0.081, 0.000, 0.048. These results predominantly indicate a poor to fair interexaminer reliability. The results of a Kappa coefficient analysis for intraexaminer reliability of the stated combinations were: e) 0.111, 0.400, 0.737; f) 0.000, 0.100, 0.368. These results indicate a poor to fair reliability. It was concluded that in subjects with low back pain the EMRR may not be a reliable indicator of spinal joint dysfunction.

Adult↗

Effects of specific criteria and calibration on examiner reliability.

The purpose of this pilot study was to investigate the use of specific criteria and examiner calibration on the reliability of inexperienced examiners on dental sealant evaluations. Dental (N = 8) and dental hygiene (N = 8) students participated as examiners. The study objectives were to identify differences in calibrated and non-calibrated examiners, examiners calibrated by an expert or non-expert, and reliability between dental and dental hygiene student examiners. A criterion-referenced evaluation form was used to evaluate dental sealant end product on 20 teeth, twice by each examiner. Eight of 16 examiners participated in a one-hour calibration session between evaluations. The session consisted of a discussion of operational definitions, the evaluation procedure for dental sealants, and use of the criterion-referenced form. Intra- and interexaminer reliabilities were measured. There were no statistically significant differences (p less than .05) in intraexaminer reliability. Although calibration produced no significant increase in interexaminer reliability, the post-training reliability scores for the group calibrated by an expert decreased, and scores for the group calibrated by a non-expert increased. No significant difference was found in reliability between dental and dental hygiene student examiners.

Humans↗

Reliability in perimetry.

As perimetric instrumentation becomes more sophisticated, patient reliability emerges as an important limiting factor in testing. Modern instrumentation for threshold and suprathreshold perimetry incorporate up to five separate indicators of patient reliability. For these perimetric methods, patient reliability is enhanced with specific techniques such as refractive correction, control of pupil size, and actively monitoring patient responses. With the manual (Goldmann) perimeter and the tangent screen, special statokinetic techniques help in both assessment and enhancement of patient reliability. In screening perimetry, reliability is assessed by analyzing the relative number, relative location, and repeatability of misses. Reliability in confrontation perimetry is both assessed and enhanced by using finger-counting and color-naming techniques. Review of the ophthalmic literature on perimetry shows how the various methods of patient reliability assessment and enhancement can be applied in the clinic.

Humans↗

A study of the reliability of carcinoembryonic antigen blood levels in following the course of colorectal cancer.

Twenty-three patients were studied to assess the reliability of carcinoembryonic antigen (CEA) levels in following the course of colorectal cancer. CEA estimations were made prior to surgery and again postoperatively. The resected specimens were allocated a Dukes' Stage and histological grading (well, moderate or poorly differentiated). In addition, sections were stained for the presence of CEA by an immunoperoxidase method. Of the 23 patients, twelve had either disseminated disease at initial surgery or subsequently developed metastasis/recurrence. Eleven remain disease-free at a minimum follow-up of one year. In all of these the reliability of plasma CEA values in reflecting the disease status has been assessed. No false positive elevations of CEA were found. Three factors emerge as positive predictors of CEA estimation reliability: pre-operative CEA elevation; tumour grading as well differentiated; dark staining for the presence of CEA. These factors identified 15 of the 18 patients (83%) in whom CEA appeared reliable and were not present in any of the five patients where CEA was not reliable. This reliability achieves statistical significance (X2 = 8.5, p less than 0.02). Histological demonstration of CEA may contribute to the reliability placed on plasma CEA estimations and should be considered if serial estimations are to be performed.

Carcinoembryonic Antigen↗

Reliability of measuring isometric and isokinetic peak torque, rate of torque development, integrated electromyography, and tibial nerve conduction velocity.

To determine the reliability of measures used in neuromuscular diagnosis and rehabilitation, 23 adults underwent identical testing on two occasions. Intraclass correlation coefficients (ICC) showed the reliability of peak torque measurement to depend both on the movement tested and velocity of contraction (leg extension ICC = 0.64-0.94, plantar flexion ICC = 0.55-0.76, leg press ICC = 0.72-0.91). Peak rate of torque development (RTD) and the percentage of peak torque at peak RTD were not reliable for any movement (ICC = 0.02-0.28). Mean RTD between 30% and 60% of peak torque was unreliable for leg press (ICC = 0.46), yet fairly reliable for both knee extension (ICC = 0.61) and plantar flexion (ICC = 0.63). Mean integrated electromyography (IEMG) showed fair to good reliability for isometric and 1.05 rad.s-1 leg press (ICC = 0.66, 0.90, respectively), and plantar flexion and leg extension (ICC = 0.75-0.89). Tibial nerve conduction velocity was highly reliable (ICC = 0.89). A range of reliabilities can be expected when measuring these variables, and must be considered when interpreting neuromuscular data.

Adult↗

Forearm pronation and supination: reliability of absolute torques and nondominant/dominant ratios.

This study examined the reliability of pronation and supination measurements expressed in absolute units (newton-meters [Nm]) and nondominant/dominant ratios (%), and determined isometrically using the BTE (WS20) and the Cybex (340) dynamometers. Twenty-one healthy men and 22 healthy women were tested twice on each machine, within 14 days. Twelve of 16 reliability coefficients for absolute torques were considered acceptable (> 0.75) when determined as the reliability of two repetitions on one occasion, while 14 of 16 coefficients were acceptable when determined as the reliability of two repetitions on each of two occasions. However, reliability coefficients for the nondominant/dominant ratios were not acceptable on one (0 of 8 coefficients > 0.75) or two (2 of 8 coefficients > 0.75) occasions. Ratio data were also characterized by larger standard errors of measurement (SEMs) and wider 95% confidence intervals, relative to the sample mean and standard deviation. Overall, reliability coefficients, SEMs, and 95% confidence intervals were similar for men and women, pronation and supination movements, and the BTE and Cybex dynamometers. The authors suggest that absolute scores be used when possible, as these data tend to provide more reliable measurements--especially when the clinician has only limited test occasions to establish baseline scores.

Bias↗

Inter- and intraexaminer reliability of a single, digital inclinometric range of motion measurement technique in the assessment of lumbar range of motion.

OBJECTIVE: The between and within examiner reliability of a range of motion digital inclinometer was evaluated for lumbar flexion, extension and right and left lateral flexion. DESIGN: Blinded, lumbar range of motion instrumentation reliability. SETTING: Private college research and ambulatory patient care facility. PARTICIPANTS: Twenty-eight asymptomatic persons recruited from a private college. This included students, staff and faculty that ranged from 23-36 yr, with no history of back pain or surgery or back injury 6 wk prior to entry into the study. INTERVENTION: Lumbar range of motion examination, twice by each examiner, or four times in all per subject. MAIN OUTCOME MEASURE: Lumbar range of motion, measured in degrees. RESULTS: Intraclass correlation (ICC) revealed lack of reliability for this device except flexion, but intrinsic limitations of the instrument suggests that flexion as well may not be reliable. The p value of .05 was used for statistical significance and Burdock's recommended value of .75 represented the minimum R value for reliability. CONCLUSION: Most of the reliability values did not meet Burdock's recommended minimum R value, and the R values for flexion may have met minimum criteria due to intrinsic limitations of the instrument itself. Due to the findings of this study, we conclude that the Orthoranger II digital inclinometer is not reliable, between and/or within examiners, for measuring lumbar flexion, extension or lateral flexion. Because there was evidence to suggest other variables which were not accounted for, and which could have affected final results, the development of a streamlined protocol may result in more consistent findings. Further research is needed to either support or dispute these results before this instrument can be recommended as an assessment tool in clinical practice or in clinical trials.

Adult↗

ECT seizure duration: reliability of manual and computer-automated determinations.

Reliable monitoring of electroencephalographic (EEG) and electromyographic electroconvulsive therapy (ECT) seizure duration has become important as these assessments have become a routine part of the clinical practice of ECT. In this regard, accurate automated seizure duration determinations would be particularly valuable. As a result, the present study was performed to assess the reliability of available computer-automated determinations of seizure duration (Thymatron Model DGx ECT machine; Somatics, Inc.) and to explore the factors upon which such reliability as well as the determinations of experienced raters depend. We found that the experienced human raters had very high interrater reliability, significantly higher than either did with the automated Thymatron DGx ratings. In general, the reliability of all ratings declined in the context of artifact, poor postictal suppression, or an EEG seizure end point that was reached gradually. Reliability was also greater for continuation ECT as compared with the index course. The reliability of Thymatron DGx versus experienced human ratings was particularly sensitive to these factors, ranging from 0.68 when several of these factors were simultaneously present to 0.999 when all these factors were absent.

Adult↗

Reliability and concurrent validity of the BROM II for measuring lumbar mobility.

OBJECTIVE: To establish the reliability of the BROM II device for measuring lumbar mobility in the sagittal, coronal and transverse planes and its validity against the double inclinometer method. DESIGN: Blind intra- and interexaminer reliability and concurrent validity. Interexaminer reliability was determined between two examiners. SETTING: Chiropractice teaching college. SUBJECTS: Forty-seven asymptomatic chiropractic students (27 men and 20 women, age range 18 to 38 yr). MAIN OUTCOME MEASURE: Lumbar mobility measurement in degrees using the BROM II and double inclinometer techniques. RESULTS: Intraclass correlation coefficients (ICC) showed good intra- and interexaminer reliability of the BROM II for flexion (0.91 and 0.77, respectively) and lateral flexion (0.91 and 0.85 respectively). Less support was given to the reliability of the instrument in extension (0.63 and 0.35, respectively) and rotation (0.57 and 0.36 respectively). Concurrent validity of the BROM II and double inclinometer methods was partially supported (ICC in all planes range from 0.27 to 0.75). CONCLUSION: The BROM II was found to be a reliable instrument in the measurement of lumbar mobility in the sagittal (flexion) and coronal planes. However, before this device can be recommended as an assessment tool in clinical practice or clinical trials, further investigation into its reliability in a symptomatic group of patients is required. (J Manipulative Physiol Ther 1995; 18:497-502).

Adolescent↗

[Reliability and factorial structure of a rating scale for amyotrophic lateral sclerosis].

The Modified Norris Scale is a rating scale for amyotrophic lateral sclerosis (ALS), which consists of two parts, the Limb Norris Scale and the Norris Bulbar Scale. The Limb Scale has 21 items to evaluate extremity function and the Bulbar Scale has 13 items to evaluate bulbar function. Each item is rated in 4 ordinal categories. Considering the habitual difference, we translated the English scale into Japanese one with minor modification, and added more detailed explanations for all categories of each item. Then we examined reliability and factorial structure of the translated scale. The subjects were 23 patients with motor disturbance and each subject was rated twice by 2-4 neurologists. As a measure of reliability, the Kappa coefficient proposed by Cohen (1960) and Kraemer (1980) was calculated for each item and the intraclass correlation coefficient (ICC) was evaluated for total scores of each of two scales. To analyze the factorial structure, the factor analysis was carried out. The minimum and the maximum Kappa values were .70 and .97 for intra-rater reliability of the Limb Scale's items, .60 and .83 for inter-rater reliability of the Limb Scale's items, .41 and 1.00 for intra-rater reliability of the Bulbar Scale's items and .26 and .81 for inter-rater reliability of the Bulbar Scale's items, respectively. Concerning the factorial structure, the contribution of the first factor was 83.6% for the Limb Scale and that for the Bulbar Scale was 66.7%. This indicates unidimensionality of both Scales. The ICCs for the total scores were .97 (95%C.I. .95-.99) for the Limb Scale and .86 (.73-.93) for the Bulbar Scale, respectively. On the basis of these results, the Scale has unidimensionality and high reliability enough for practical use.

Adult↗

Reliability of Spanish translations of select urological quality of life instruments.

PURPOSE: Many patients with urological disease do not speak English. In medical studies restricting patients to those who speak only English undermines efforts to understand disease because restrictions decrease efficiency of patient recruitment, and because language and culture are associated with variable outcomes. In Spanish speaking locations, such as South Florida, studies would suffer severe selection bias if patients were required to speak English. To allow grouping in future studies of English and Spanish speaking patients we examined the English-Spanish reliability of select instruments that measure health related quality of life in patients with urological disease. MATERIALS AND METHODS: We assembled available Spanish versions and translated English versions of questions regarding satisfaction, the American Urological Association symptom index, the University of California, Los Angeles Prostate Cancer Index and a pain inventory. We then examined English-Spanish reliability by asking bilingual men 50 years old or older to complete English and Spanish versions at the same sitting. A convenience sample was recruited from outpatients and volunteers at the Miami Veterans Affairs Medical Center and population based subjects living in largely Hispanic Hialeah, Florida. Reliability estimates were calculated with kappa coefficients for categorical data and intraclass correlation coefficients for quantitative data. RESULTS: A total of 100 subjects a median of 59 years old completed the questionnaire, including 55 born in Puerto Rico or Cuba, while the remainder were born at various sites throughout the Americas and Spain. Reliability estimates showed that kappa = > 0.81 for almost all items. For 2 items relating to health and social interactions reliability was poor, and stratification showed that poor reliability was primarily a feature of subjects in good health who are theoretically socially active. CONCLUSIONS: Almost all items tested have excellent English-Spanish reliability in a mixed sample of bilingual men. Nonreliability of 2 items relating to health and social interactions probably originates from the effect of language on perception, and invalidates English and Spanish grouping of these items. Because the sample represents many dialects of Spanish, the translations tested may be transported to other cities. In studies that use these instruments investigators can reasonably group answers from English and Spanish speaking study subjects or study the effects of acculturation on quality of life.

Hispanic or Latino↗

The reliability of hand-written and computerised records of birth data collected at Baragwanath hospital in Soweto.

This study examined the reliability of hand written and computerised records of birth data collected during the Birth to Ten study at Baragwanath Hospital in Soweto. The reliability of record-keeping in hand-written obstetric and neonatal files was assessed by comparing duplicate records of six different variables abstracted from six different sections in these files. The reliability of computerised record keeping was assessed by comparing the original hand-written record of each variable with records contained in the hospital's computerised database. These data sets displayed similar levels of reliability which suggests that similar errors occurred when data were transcribed from one section of the files to the next, and from these files to the computerised database. In both sets of records reliability was highest for the categorical variable infant sex, and for those continuous variables (such as maternal age and gravidity) recorded with unambiguous units. Reliability was lower for continuous variables that could be recorded with different levels of precision (such as birth weight), those that were occasionally measured more than once, and those that could be measured using more than one measurement technique (such as gestational age). Reducing the number of times records are transcribed, categorising continuous variables, and standardising the techniques used for measuring and recording variables would improve the reliability of both hand-written and computerised data sets.

Data Collection↗