Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

The reliability of multitest regimens with sacroiliac pain provocation tests.

BACKGROUND: Studies concerning the reliability of individual sacroiliac tests have inconsistent results. It has been suggested that the use of a test regimen is a more reliable form of diagnosis than individually performed tests. OBJECTIVE: To assess the interrater reliability of multitest scores by using a regimen of 5 commonly used sacroiliac pain provocation tests. METHODS: Two examiners examined 78 subjects. The threshold for a positive selection was set at 3 positive tests out of 5 tests performed. The test order and the order in which the subjects were examined were randomized per patient, and the examiners were blinded from all information regarding the subjects tested. Fifty-nine of the subjects were symptomatic for low back pain, and 19 of the subjects were asymptomatic. Weighted kappa statistic, bias-adjusted kappa, prevalence-adjusted kappa, and 95% CI intervals were used to evaluate the interrater reliability of the test regimen. RESULTS: Weighted kappa was found to be 0.70 (95% CI = 0.45-0.95). CONCLUSIONS: A multitest regimen of 5 sacroiliac joint pain provocation tests is a reliable method to evaluate sacroiliac joint dysfunction, although further study is needed to assess the validity of this test method.

Adult↗

Further reliability analysis of the Harrison radiographic line-drawing methods: crossed ICCs for lateral posterior tangents and modified Risser-Ferguson method on AP views.

OBJECTIVE: To determine whether the newly derived interclass and intraclass correlation coefficients (ICCs)would overstate or understate the results from 2 previously published studies, which used better known ICCs that assume nested factors, and to determine mean absolute differences of observers' measurements for 3 previous studies. STUDY DESIGN: Retrospective analysis of data from 2 blind studies with repeated-measure design. Two newly derived ICCs, appropriate to situations with 3 random factors (patients, examiners, and occasions) that bear a crossed (as opposed to nested) interrelationship, were applied to data from an experiment with random crossed factors. MAIN OUTCOME MEASURES: Observer reliability is determined with ICCs, 95% CIs, and observer error analysis (mean absolute differences of observers' measurements) for angles and distances derived from Harrison's modified Risser-Ferguson line-drawing method on anteroposterior (AP) lumbar and AP cervical radiographic views. Observer error analysis for angles and distances derived from Harrison's posterior tangent method on lateral cervical views was also determined. RESULTS: The majority of ICCs for reliability of line drawing on both AP cervical and AP lumbar radiographs were in the high range; 13 of 16 ICCs were greater than 0.88. The other 3 ICC values (0.61, 0.76, 0.78) concerned determining the sacral base on AP lumbar views. The new ICCs underestimated observer reliability compared with previously published results (intraclass ICCs lower by 0.01-0.02 and interclass ICCs lower by 0.03-0.10). For an error analysis on data from both AP views, the mean absolute differences of observers' measurements were 1.1 degrees to 1.8 degrees for angles and 1.2 mm to 2.3 mm for distances. For the lateral cervical analysis, the observer error was in the interval 0.8 degrees to 3.2 degrees for angles and <1 mm for distances. CONCLUSIONS: The ICCs assuming random crossed factors understate reliability compared with previously published ICC results assuming nested factors. Reliability of the Harrison modified Risser-Ferguson method of line-drawing analysis on AP views is in the high range, with the majority of ICCs >0.88. For both the Harrison modified Risser-Ferguson method on AP views and posterior tangent method on lateral cervical views, the mean absolute differences of observers' measurements are small.

Cervical Vertebrae↗

Accuracy and reliability of a new, protractor-based neck goniometer.

OBJECTIVES: To assess the reliability of the SpinT, a new protractor-based device, for measuring active cervical spine ranges of motion. In addition, to compare the accuracy of the Cervical Ranges of Motion (CROM) instrument and SpinT measurements of rotation about the Y axis with and without tilt, the former motion occurring during natural rotation of the head. STUDY DESIGN: Interexaminer reliability, intraexaminer reliability, and accuracy trials were conducted. METHODS: Two examiners made 2 individual measurements of each of the individual cervical ranges of motion of 23 patients (15 men, 8 women; aged 21 to 42 years) with no cervical symptoms. The patients were asked to move their necks to end range while they sat upright. The accuracy of the CROM instrument and SpinT goniometers was assessed with a testing instrument capable of rotating and/or tilting to preset angles and upon which either device could be positioned. RESULTS: There was excellent agreement between the SpinT measurements of rotation about the Y axis compared with the readings from the testing platform regardless of the angle of tilt, whereas the CROM instrument displayed poor concordance when the tilt exceeded 5 degrees. The reliability trials generally yielded close agreement between the examiners, especially regarding measurements of rotation left and right and extension and revealed higher concordance regarding intraexaminer results. CONCLUSION: This study indicates that SpinT measurements of active cervical ranges of motion are reliable and that the SpinT goniometer accurately measures rotation with associated tilt.

Adult↗

Testing surgical skills of obstetric and gynecologic residents in a bench laboratory setting: validity and reliability.

OBJECTIVE: Resident surgical skills are acquired mainly through observing and later performing procedures in the operating room. Evaluation of surgical skills has traditionally been done through subjective faculty evaluation, a technique that has poor reliability and unknown validity. Our goal was to develop specific surgical tasks, both laparoscopic and open abdominal, that could be objectively and reliably evaluated in a bench laboratory setting. STUDY DESIGN: The prospective development of a reliable and valid resident surgical skills test in a bench laboratory setting was our goal. A written test of surgical knowledge and 12 skills tests were administered to 36 residents. Laparoscopic bench tasks were simulated with the use of a box and camera with a video display. Six laparoscopic tasks were assessed, including placing pegs on a board, running the bowel simulation, and other tasks that involve hand-eye coordination and manual dexterity. Open abdominal skills simulated incision closure, suturing a vaginal cuff, knot tying, and using a tie on a passer. Residents were timed at each given station and were given a rating score by 2 examiners. RESULTS: Knowledge scores showed a significant improvement by residency level. Assessment of construct validity (the ability to discriminate among residency levels) demonstrated significant differences on the rating of overall performance and individual tasks by level (determined by 1-way analysis of variance). Interrater reliability (agreement between 2 raters) with the use of intraclass correlation was 0.79 for the total score. The cost to administer the bench laboratory test was less than $50 and required 30 hours of faculty time. CONCLUSION: The results of this study suggest that surgical bench laboratory tasks can assess residents' surgical skills with good reliability and validity on most tasks. Our previous study, which used an animal laboratory, was expensive, and the bench laboratory model may provide an alternative means to assess surgical skills.

Clinical Competence↗

Interobserver and intraobserver reliability of the measurement of shoulder internal rotation by vertebral level.

Internal rotation is commonly measured as the vertebral level reached by the fully extended thumb. The purpose of this study was to evaluate interobserver and intraobserver reliability with the use of this method. Three male subjects were used for internal rotation measurement. Eleven orthopaedic surgeons and 2 physical therapists served as examiners. Each subject had a radiographic marker placed at a random vertebral level, and the subject's extended thumb was placed at this marker. All examiners then independently measured internal rotation based on vertebral level. To assess intraobserver reliability, this process was repeated twice. After all measurements were completed, an anterior-posterior radiograph of each subject was obtained to define the vertebral level of the marker. This process was repeated 2 additional times with the marker and subject's thumb positioned at different levels than in the previous examination. Intraclass correlation coefficients were calculated to determine reliability. Results demonstrated poor interobserver reliability and reasonable intraobserver reliability. The mean clinical measurement deviated from the mean actual measurement by 1 vertebral level. Despite being the standard method in which shoulder internal rotation is measured, measurement of internal rotation by vertebral level is not readily reproducible between observers.

Adult↗

Reliability and validity of a food-frequency questionnaire for Chinese postmenopausal women.

OBJECTIVES: (1). To determine the reliability and validity of a food-frequency questionnaire (FFQ) for use in epidemiological research in postmenopausal women; and (2). to compare the volume estimation (VE) and weight estimation (WE) method of administration of this questionnaire. DESIGN: An initial list of foods was derived and modified after pre-testing in 22 subjects. Test-retest reliability was assessed in 21 subjects who had repeat administrations of the questionnaire 14 days apart (FFQ1, FFQ2). The validity of the FFQ was assessed by comparing nutrient intakes with those from a 4-day food record. SETTING: Chengdu, People's Republic of China. SUBJECTS: Twenty-two postmenopausal women (50-70 years) were recruited from The Second University Hospital, West China University of Medical Sciences, Chengdu and participated in the pre-test. Another 21 women (50-70 years) were randomly selected from the general population of all five districts of Chengdu and participated in the reliability and validity sub-studies. RESULTS: Energy, protein, carbohydrate, magnesium and sodium intakes in this sample were less than the Recommended Dietary Allowances (RDAs) for 45-70-year-old women in China. Intake of non-cooking fat was higher than the Chinese RDA. Pearson correlation coefficients and intra-class correlation coefficients (ICCs) for reliability of the VE FFQ ranged from 0.51 to 0.85 and from 0.51 to 0.81, respectively; for the WE FFQ, they ranged from 0.22 to 0.86 and from 0.21 to 0.81. Correlation coefficients and ICCs for validity of the WE FFQ ranged from 0.36 to 0.69 and from 0.34 to 0.57, respectively; corresponding values for the VE FFQ were -0.30 to 0.65 and -0.14 to 0.65. CONCLUSIONS: Both the VE and WE FFQs were reliable and valid except for sodium intake. The VE FFQ provided more valid estimates of nutrient intakes than did the WE FFQ.

Aged↗

Test-retest reliability of tone-burst-evoked otoacoustic emissions.

In this study, the short- and long-term test-retest reliabilities of tone-burst-evoked otoacoustic emissions (TBOAEs) with 12 different tone-burst stimuli (4 frequencies [1, 1.5, 2 and 3 kHz] at 3 stimulus levels [approximately 76, approximately 67 and approximately 55 dB pcSPL]) were examined in 30 normal hearing subjects. Click-evoked and spontaneous OAEs were recorded in parallel with TBOAEs to facilitate cross-comparisons and the generalization of results. Findings for click-evoked and spontaneous OAEs were comparable with most literature data. High reliability for TBOAEs was established for high and mid stimulus levels at all frequencies tested with reference to test-retest prevalence rate, test retest occurrence, intra-subject test retest difference and correlation coefficient. Derived half-octave band analysis at the frequency corresponding to the stimulus was found to reflect real TBOAE performance more reliably than broadband analysis. No significant difference between short- and long-term reliabilities was noted from all results. Similar test-retest reliabilities for high-level TBOAEs and click-evoked OAEs was obtained, suggesting that TBOAEs could potentially contribute to clinical assessment.

Acoustic Stimulation↗

Reliable exposure assessment strategies for physical ergonomics stressors in construction and other non-routinized work.

The objective of this research was to provide guidelines for the reliable assessment of ergonomics exposures in non-routinized work. Using a discrete-interval observational sampling approach, two or three observers collected a total of 5852 observations on tasks performed by three construction trades (iron workers, carpenters and labourers) for periods of several weeks. For each observation, nine exposure variables associated with awkward body postures, tool use and load handling were recorded. The frequency of exposure to each variable was calculated for each worker during each of the tasks on each of the days. ANOVA was used to assess the importance of task in explaining between-worker and within-worker variability in exposures across days. A statistical re-sampling method (bootstrap) was used to evaluate the reliability of exposure estimates for groups of workers performing the same task for different sampling periods. Most exposures were found to vary significantly across construction tasks within trade, and between-worker exposure variability was generally smaller than within-worker exposure variability within task. Bootstrapping showed that the reliability of the group estimates exposure for the most variable exposures within task tended to improve as the assessment periods approached 5-6 d, with marginal improvements for longer assessment periods. Reliable group estimates of exposure for the least variable exposures within task were obtained with 1 or 2 d of observation. The results of this study demonstrate that an initial estimate of the important environmental or task sources of exposure variability can be used to develop an efficient sampling strategy that provides reliable estimates of ergonomics exposures during non-routinized work.

Adult↗

Test-retest reliability of isometric and isoinertial testing in symmetric and asymmetric lifting.

The aim of this study was to evaluate test-retest reliability of a dynamometer in measuring lifting strength or force parameters under several combinations of ergonomic factors. Thirteen healthy participants were tested on peak force (PF), related variables and isometric strength (IS) twice, at intervals of 3 months. Correlation coefficients for all parameters in the sagittal plane were 0.60-0.85. Coefficients of variations (CVs) of methodology error for PF in the sagittal plane were 6.2-6.9%. Correlation coefficients and CVs for IS at 90 degrees to the lateral plane were 0.51-0.54 and 16.6-17.9%, respectively. In paired t-tests of the parameters under all conditions, there was a significant difference between test and retest. In the test and retest, ratings of perceived exertions for the low back and the right arm in isometric lifting were significantly higher than those in dynamic lifting. It was concluded that the test-retest reliability of dynamic forces in the dynamometer was high. The peak force in the sagittal plane was considered reliable. In isometric lifting, isometric strength in the sagittal plane seemed reliable, while that at right angles to the lateral plane was considered to be less reliable.

Adult↗

Test-retest reliability of the emotional stroop task: examining the paradox of measurement change.

The Emotional Stroop (ES) task (I. H. Gotlib & C. D. McCann, 1984) has been proposed as an experimental measure to assess the processing of emotion or the bias in attention of emotion-laden information. However, study results have not been consistent. To further examine its reliability for empirical research, the authors of this study administered the ES task to 33 participants on 2 separate occasions separated by 1 week. Results indicated that retest reliabilities for reaction times (RTs) derived from the 3 separate emotion conditions (manic, neutral, and depressive) across the 1 week interval were very high. However, consistent with previous research, the reliabilities were very low for the interference indices (manic and depressive). These low reliabilities reflect the very high intercorrelation between the RTs derived from the 3 conditions. The authors concluded that a better indicator of the reliability for this task is the individual RTs from each emotion condition.

Adult↗

Reliability of assessment of urgency and other symptoms indicating anal sphincter, large bowel or urinary dysfunction.

OBJECTIVE: The Radiumhemmets Scale of Disease-Specific Symptom Assessment-Prostate Cancer has been used in several studies. However, no test-retest reliability study of it has been conducted concerning the assessment of urinary, anal sphincter or large bowel function. The aim of this study was to evaluate the reliability of items assessing these functions. MATERIAL AND METHODS: We investigated 89 prostate cancer patients randomly selected from a group of patients diagnosed in Stockholm. The patients answered 24 questions assessing anal sphincter, large bowel and urinary function twice, with a 3-week interval in-between, to assess reliability. RESULTS: Most of the questions assessing bowel and urinary symptoms showed substantial or near-perfect agreement. The kappa value for bowel symptom items was > or = 0.60 for all items, except for defecation urgency (0.40-0.55). The kappa value for urinary symptom items varied between 0.43 and 1.0, except for urinary urgency (0.30-0.39). CONCLUSIONS: When comparing the impact of different symptoms of anal sphincter, large bowel or urinary tract dysfunction, it may be important to consider that defecation urgency and urinary urgency have the highest measuring error (low reliability). This error dilutes assessed associations with, for example, decreased quality of life. Nevertheless, the test-retest reliability for anal sphincter, large bowel and urinary symptoms indicates that surveys yield meaningful information.

Aged↗

Is real-time ultrasonic bladder volume estimation reliable and valid? A systematic overview.

To assess the reliability and validity of real-time ultrasonic estimation of bladder volume we conducted an overview of the published literature identified using MEDLINE search (1966-96) and scanning of the bibliographies of known primary and review articles. Short-listed papers were classified into reliability (observer agreement) and validity (comparison of ultrasound estimation with actual bladder volume) studies. Study selection and data extraction were performed independently in duplicate. There were 81 subjects enrolled in 3 reliability studies and 504 subjects in 16 validity studies. Where reported, the index of concordance for reliability ranged from 0.923 to 1.00, while for validity it ranged from 0.914 to 0.983. However, there were several inadequacies in the design, conduct and analysis of these studies, leaving some doubt about the trustworthiness of the high levels of reliability and validity reported in the literature.

Female↗

Long-term test-retest reliability of category loudness scaling in normal-hearing subjects using pure-tone stimuli.

The present study investigates the test-retest reliability of category loudness scaling with pure tones for each of the scaling categories: 'very soft', 'soft', 'OK', 'loud', 'very loud' and 'too loud' at the audiometric frequencies 0.5, 1 kHz, 2 kHz and 4 kHz. Category loudness scaling at two sessions separated by between 1 and 4 weeks was obtained from 16 normal-hearing subjects who all had normal otoscopy, present acoustic reflexes at audiometric frequencies 0.5-4 kHz and middle ear pressure within +/-50 daPa. Intra-subject between-session reliability was found not to be frequency dependent, and comparison with other studies revealed that reliability is not dependent on the applied stimulus signal. Test-retest reliability varied between the different categories: In the categories 'very soft', 'loud', 'very loud' and 'too loud' the test reliability is in the same range as found for hearing thresholds determination, whereas for the 'soft' and 'OK' categories it is poorer. The greater uncertainty for intermediate levels should be considered when using category loudness scaling, e.g. for calculating hearing aid parameters.

Adult↗

Enhancing reliability in portfolio assessment: 'shaping' the portfolio.

This paper reports a follow-on project that assessed a series of portfolios assembled by a cohort of participants attending a course for prospective general practice trainers. In an attempt to enhance reliability, a framework for defining and addressing problems using a reflective practice model was offered to participants. The reliability of the judgements made by a panel of assessors about individual 'components', together with an overall global judgement about performance were studied. The reliability of individual assessors' judgements (i.e. their consistency) was moderate, but inter-rater reliability did not reach a level that could support making a safe summative judgement. Despite offering a possible structure for demonstrating reflective processes, the levels of reliability reached were similar to the earlier work and other subjective assessments generally, and perhaps reflected individuality of personal agendas of both the assessed and the assessors, and variations in portfolio structure and content; even agreement among the assessors about evidence of the framework being used was poor. Suggestions for approaches in the future are made. The conclusion remains that while portfolios might be valuable as resources for learning, as assessment tools they should be treated as problematic.

Journal Article↗

Reliability of exophthalmos measurement and the exophthalmometry value distribution in a healthy Dutch population and in Graves' patients. An exploratory study.

AIM: The purpose of this study was to test the reliability of an exophthalmometer commonly used in the Netherlands; to determine the exophthalmometry value distribution with this instrument and to assess the upper exophthalmometry limits of normal in a healthy, adult, Caucasian, Dutch population. Furthermore, to assess the effects of gender and age on exophthalmometry readings in this group and in a group of Graves' patients by comparing healthy, adult, Caucasian, Dutch individuals with adult, Caucasian, Dutch Graves' patients. METHODS: To test the reliability of our Hertel exophthalmometer, we determined the interobserver variation between two observers by measuring 160 eyes in healthy, adult, Caucasian, Dutch females and males (10 females and 10 males in each decade between 20 and 60 years of age). These data were also used for the assessment of the Hertel value distribution and for defining the upper limits of normal in these individuals by logistic regression analysis. The effects of disease, age and gender were established using these data plus data of a retrospective study of 393 adult, Caucasian, Dutch females (n=294) and males (n=99) with Graves' orbitopathy in whom Hertel values were measured with the same exophthalmometer. RESULTS: Exophthalmometry using an Hertel exophthalmometer appeared reliable (Pearson correlation coefficient for interobserver variation 0.89; 96% of the Hertel values measured by two observers were within the limits (of 2 mm) of agreement). Hertel values usually show a normal distribution in healthy individuals and in Graves' patients and are sex- and age-dependent, but there was no dependence on age in this small series in adults. Logistic regression analysis revealed an upper limit of normal of 16 mm in females and 20 mm in males in our group, using the exophthalmometer described. CONCLUSIONS: Exophthalmometry is reliable and absolute measurement of proptosis is feasible. International standardization of Hertel exophthalmometry is required in order to compare exophthalmometry data in the literature reliably.

Adult↗

WAIS-R test-retest reliability in a normal elderly sample.

We examined the 1-year test-retest reliability of WAIS-R Verbal, Performance, and Full-Scale IQs in a sample of 101 older normal individuals (mean age = 67.1). The respective Pearson rs were .86, .85, and .90. The median retest reliability coefficient for the WAIS-R subtests was .71. The test-retest reliability for the Verbal-Performance Discrepancy was .69. These data indicate that IQ scores are reliable in older normal individuals for this retest interval, but less confidence can be placed in the reliability of subtest scores and the Verbal-Performance Discrepancy.

Aged↗

Reliability of hand preference items and factors.

The present study examined the test-retest reliability of a 32-item version of the Waterloo Handedness Questionnaire (Steenhuis & Bryden, 1987, 1988, 1989) on 500 subjects. The questionnaire was shown to be reliable in terms of basic factor structure. High test-retest reliability was also found within subjects' responses and across items for both right-handers and left-handers, although left-handers were less consistent than right-handers, particularly with regard to direction of hand preference on individual questionnaire items. Furthermore, the direction of hand preference was more reliable than was the degree of hand preference. These data support a multidimensional view of hand preference in which both direction and degree can be reliably assessed.

Adult↗

Assessing the reliability of clinical scales when the data have both nominal and ordinal features: proposed guidelines for neuropsychological assessments.

The purpose of this article is to present, for the first time, a comprehensive methodology for assessing the reliability of a clinical scale that is frequently utilized in neuropsychological research and in biomedical studies, more generally. The dichotomous-ordinal scale is characterized by a single category of "absence" and two or more ordinalized categories of "presence" of a symptom trait, state, or behavior, and it also has special properties that need to be understood in order for its reliability to be appropriately assessed. Using the Brief Psychiatric Rating Scale (BPRS) as a clinical example, we cover the principles of expressing scale reliability in terms of a dichotomy ("absence" - "presence" of a given BPRS symptom); as a trichotomy ("none"; "mild to moderate" symptomatology; and "severe" symptomatology); and as the full 7-category dichotomous-ordinal scale: "none," "very mild," "mild," "moderate," "moderately severe," "severe," and "extremely severe." Criteria are presented that can be used to evaluate which of these three formats produces the most reliable results. Finally, we address, with a second sample, the important issue of replication, or whether the original reliability findings generalize to other independent populations.

Adult↗