Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Development of a reliable measure of walking within and outside the local neighborhood: RESIDE's Neighborhood Physical Activity Questionnaire.

BACKGROUND: The RESIDential Environment project (RESIDE) is a longitudinal study evaluating the impact of a new residential design code on walking. OBJECTIVE: To develop a reliable measure of walking--undertaken within and outside the neighborhood--and overall physical activity. METHODS: A test-retest reliability study was undertaken (n = 82, mean age 39 years). The instrument was based on the International Physical Activity Questionnaire (IPAQ-short version) and Active Australia Survey. It measured usual frequency and duration of (1) recreational- and transport-related walking within and outside the neighborhood and (2) other vigorous and moderate physical activities. RESULTS: Reliability of recall of whether participants had walked within (k = 0.84) and outside (0.73) the neighborhood was acceptable. Similarly, recall of frequency and duration of transport and recreational-related walking within the neighborhood was excellent (ICC > or = 0.82), as was recall of transport-related walking trips outside the neighborhood (ICC > or = 0.84). Reliability for duration of recreational walking outside the neighborhood was fair to good (ICC = 0.55). The reliability of indices of total physical activity based on MET min/week (ICC = 0.82) and MET min/week dichotomized to 'sufficient' physical activity for health (kappa = 0.67) were both acceptable. CONCLUSIONS: The Neighborhood Physical Activity Questionnaire (NPAQ) is sufficiently reliable for studies examining environmental correlates of walking within the neighborhood.

Adult↗

Myotonometer intra- and interrater reliabilities.

OBJECTIVES: To assess the intra- and interrater reliabilities of the Myotonometer, a hand-held, computerized, electronic device that quantifies muscle stiffness (tone/compliance). DESIGN: Reliability study. SETTING: Research laboratory. PARTICIPANTS: Thirty-five healthy, nondisabled adults (age range, 22-42 y). INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Two raters used the Myotonometer to evaluate subjects' lateral gastrocnemius and biceps brachii muscles. Muscles were measured in a relaxed state and during a voluntary isometric contraction. Coefficients were calculated for each muscle and each condition (relaxed, contracted). Results were analyzed by using Design II intraclass correlation coefficients. RESULTS: Reliability coefficients were highest when the instrument exerted moderate to strong forces against the muscle (range, 0.50-2.00 kg; intrarater reliability R range, .84 - .99; interrater reliability R range, .75 - .96). CONCLUSIONS: Myotonometer measurements had high to very high intra- and interrater reliabilities for measurements of the lateral gastrocnemius and biceps brachii muscles.

Adult↗

Test-retest reliability of hand-held dynamometric strength testing in young people with cerebral palsy.

OBJECTIVE: To evaluate the test-retest reliability of measuring lower-limb strength with a hand-held dynamometer in young people with cerebral palsy (CP). DESIGN: One rater measured the isometric strength of the lower limbs in 10 participants with CP on 2 occasions separated by 6 weeks. SETTING: University movement rehabilitation laboratory in Australia. PARTICIPANTS: Ten young people (mean age +/- standard deviation, 13.5+/-3.4 y) with spastic diplegic CP. Eight of the participants walked independently and 2 walked with assistive devices. INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Retest reliability of lower-limb strength, expressed in the units of measurement for the interpretation of group mean and individual scores and as intraclass correlation coefficients (ICC(2,1)). RESULTS: For groups, mean lower-limb strength increases of 7 kg (30%) could be interpreted as real change using 95% confidence intervals (CIs). For individuals, for strength gains to be interpreted as real change using 95% CIs, strength increases would need to be greater than 16.8 kg (70%) for the measurement of knee extension and to be greater than 4.3 kg (25%) for ankle plantarflexion. Measurement of hip extension strength was not reliable for group mean or individual scores. All reliability coefficients were greater than.80. CONCLUSION: A hand-held dynamometer can reliably measure changes in lower-limb strength for groups of young people with CP. It is uncertain whether this method is useful for evaluating change in individuals. Relying only on a coefficient of reliability to decide the usefulness of a measurement can be misleading.

Adolescent↗

Reliability of electromyographic power spectral analysis of back muscle endurance in healthy subjects.

OBJECTIVE AND DESIGN: To investigate the reliability (within-day and between-days) of measurements of electromyographic (EMG) power spectral values in measurement of the fatigue rate of the back muscles. METHODS: Twelve healthy male subjects were tested in the unsupported trunk holding position for 60 seconds. Two trials were performed on each of two separate sessions 3 days apart. Surface recording electrodes were placed over the iliocostalis lumborum and multifidus and a branched electrode technique was used to decrease cross-talk. OUTCOME MEASURES: Median frequency (MF) was extracted from the EMG signals by fast Fourier transform. Initial MF and the MF slope over time were computed from linear regression analysis. The reliability of the initial MF and the MF slope of the iliocostalis lumborum and multifidus was examined by Pearson's product moment correlation coefficients (Pearson's r), paired t tests, and a one-way analysis of variance (ANOVA) from which intrasubject coefficients of variation (CVintra) and intraclass correlation coefficients (ICC) were derived. RESULTS: For the initial MF, within-day and between-days reliability of the iliocostalis lumborum and multifidus were good (Pearson's r = .79-.94 nonsignificant paired t test, CVintra = 6.5% to 8.5%, ICC = .79-.93). The MF slope showed moderate variability for the iliocostalis lumborum (Pearson's r = .39-.55, nonsignificant paired t test, CVintra = 33.0% to 48.7%, ICC = .37-.56) while better reliability was found for the multifidus (Pearson's r = .77-.87, nonsignificant paired t test, CVintra = 25.8% to 27.5%, ICC = .78-.82). CONCLUSION: The present study indicated that the trunk holding test with the use of EMG power spectral analysis can be a reliable method to measure the fatigue rate of the back muscles if adequate measures are employed to minimize cross-talk. The better reliability of monitoring the fatigue rate of the multifidus may lead to its future use as a clinical measure.

Adult↗

Reliability of procedures used in the physical examination of non-specific low back pain: a systematic review.

The purpose of this systematic review was to determine the quality of the research and to assess the reliability of different types of physical examination procedures used in the assessment of patients with non-specific low back pain. A search of electronic databases (MEDLINE, PEDro, AMED, EMBASE, Cochrane, and CINAHL) up to August 2005 identified 48 relevant studies which were analysed for quality and reliability. Pre-established criteria were used to judge the quality of the studies and satisfactory reliability, and conclusions emphasised high quality studies (> or = 60% methods score). The mean quality score of the studies was 52% (range 0 to 88%), indicating weak to moderate methodology. Based on the upper threshold used (kappa/ICC > 0.85) most procedures demonstrated either conflicting evidence or moderate to strong evidence of low reliability. When the lower threshold was used (kappa/ICC > 0.70) evidence about pain response to repeated movements changed from contradictory to moderate evidence for high reliability. Most procedures commonly used by clinicians in the examination of patients with back pain demonstrate low reliability.

Humans↗

Measurement of functional ability following traumatic brain injury using the Clinical Outcomes Variable Scale: a reliability study.

This study determined the inter-tester and intra-tester reliability of physiotherapists measuring functional motor ability of traumatic brain injury clients using the Clinical Outcomes Variable Scale (COVS). To test inter-tester reliability, 14 physiotherapists scored the ability of 16 videotaped patients to execute the items that comprise the COVS. Intra-tester reliability was determined by four physiotherapists repeating their assessments after one week, and three months later. The intra-class correlation coefficients (ICC) were very high for both inter-tester reliability (ICC > 0.97 for total COVS scores, ICC > 0.93 for individual COVS items) and intra-tester reliability (ICC > 0 97). This study demonstrates that physiotherapists are reliable in the administration of the COVS.

Brain Injuries↗

The reliability of distinguishing primary versus secondary negative symptoms.

The objective appearance of negative symptoms in schizophrenia and other psychotic disorders may be a direct reflection of a primary neural abnormality or may be secondary to a variety of factors such as neuroleptic side effects, depression, positive symptoms, or environmental understimulation. Although there is a consensus that it is important to be able to disentangle "primary" versus "secondary" negative symptoms, optimal strategies for doing so remain unclear. Concerns have been raised about making this distinction based on clinical judgment because of potential low reliability in the absence of extensive training and/or highly specialized rating scales. This is particularly important in terms of the application of DSM-IV criteria for schizophrenia, in which negative symptoms play a prominent role. In the context of the DSM-IV schizophrenia field trial project, we examined the reliability of making the primary versus secondary distinction in a multicenter sample of 462 subjects with nonorganic psychotic disorders. Each subject was assessed by two raters, half in an interrater design (i.e., conjoint interviews) and half in a test-retest design (i.e., independent interviews by two raters conducted 1 day apart). All raters used the same semistructured interview instrument, which included an abbreviated version of the Scale for the Assessment of Negative Symptoms (SANS). In addition to the usual SANS ratings, raters were asked to indicate their judgment as to whether the symptom was primary, secondary, or unknown (inadequate information to assess). No formal training was provided. Reliability, as quantified by kapp, indicated only a fair degree of agreement ranging from 0.48 to 0.68 for interrater reliability (median, 0.50) and 0.34 to 0.66 for test-retest reliability (median, 0.38). Negative symptoms were rated as primary approximately twice as often as secondary, and raters believed they had adequate information to make this distinction based only on cross-sectional evaluation in all but 10% of the cases. These data suggest that the primary versus secondary distinction should not be incorporated into the application of operationalized diagnostic criteria. Implications are discussed in terms of balancing reliability and validity in the assessment of negative symptoms.

Adult↗

Reliability of in vivo neutron activation analysis for measuring body composition: comparisons with tracer dilution and dual-energy x-ray absorptiometry.

In vivo neutron activation (IVNA) analysis has the capacity to measure several total body elements in human subjects. Although it has been considered a criterion method for the past 3 decades, the reliability of IVNA analysis has been tested only in phantom calibrations. In 5 male weight-stable patients with AIDS, total body N, Ca, Cl, Na, P, and C were measured three times in 16 weeks at Brookhaven National Laboratory. With tracer dilution methods for total body water (TBW) by 3H2O and for extracellular water (ECW) by 35SO4 and NaBr, and dual-energy x-ray absorptiometry (DXA), total body calcium (TBCa) and fat percentage were measured within 2 weeks of IVNA measurements. For comparison, tracer dilution for TBW by D2O and ECW by NaBr, plus DXA measurements, were performed three times in 5 weight-stable healthy volunteers. The reliability of the IVNA technique was very high in patients with AIDS; it ranged from 0.99 for total body chloride (TBCI) to 0.84 for total body phosphorus (TBP), and it agreed with phantom calibration results in the literature. The reliability for measuring fat percentage and TBCa by DXA was similar in patients with AIDS and in healthy volunteers. Tracer dilution for measuring TBW by 3H2O in patients with AIDS and by D2O in healthy volunteers had a reliability score similar to those found with IVNA and DXA. The reliability scores for measuring ECW in patients with AIDS by 35SO4 and NaBr, 0.66 and 0.68, respectively, were the lowest among all measurements, whereas the reliability score for NaBr in healthy volunteers was 0.96, as with the other measurements.

Absorptiometry, Photon↗

Palpation for muscular tenderness in the anterior chest wall: an observer reliability study.

OBJECTIVE: To asses the interobserver and intraobserver reliability (in terms of day-to-day and hour-to-hour reliability) of palpation for muscular tenderness in the anterior chest wall. DESIGN: A repeated measures designs was used. SETTING: Department of Nuclear Medicine, Odense University Hospital, Denmark. PARTICIPANTS: Two experienced chiropractors examined 29 patients and 27 subjects in the interobserver part, and 1 of the 2 chiropractors examined 14 patients and 15 subjects in the intraobserver studies. INTERVENTION: Palpation for muscular tenderness was done in 14 predetermined areas of the anterior chest wall with all subjects sitting. Each dimension was rated as absent or present for tenderness or pain for each location. All examinations were carried out according to a standard written procedure. RESULTS: Based on a pooled analysis of data from palpation of the anterior chest wall, we found kappa values of 0.22 to 0.31 for the interobserver reliability. For the intraobserver reliability, we found kappa values of 0.21 to 0.28 for the day-to-day reliability and 0.44 to 0.49 for the hour-to-hour reliability. CONCLUSION: Our results indicated great variations between experienced chiropractors palpating for intercostal tenderness or tenderness in the minor and major pectoral muscles in a population of patients with and without chest pain. This may hamper the ability of clinicians to diagnose and classify the musculoskeletal component of chest pain if based exclusively on palpation of the anterior chest wall.

Case-Control Studies↗

Chiropractic biophysics digitized radiographic mensuration analysis of the anteroposterior lumbopelvic view: a reliability study.

OBJECTIVE: To investigate the reliability of a radiographic measurement procedure that uses a computer and sonic digitizer to determine projected spinal displacements from an ideal normal position. DESIGN: A blind, repeated-measure design was used. Anteroposterior lumbopelvic radiographs were presented to each of 3 examiners in random order. Each film was digitized, and the films were randomized for a second run. SETTING: Private, primary-care chiropractic clinic. MAIN OUTCOME MEASURES: The angle of the sacral base in comparison to a true horizontal line (horizontal base angle), lumbodorsal angle, lumbosacral angle, and the thoracic translational displacement from true vertical determined as the perpendicular distance from the center of T12 to a vertical axis line drawn from the center of the S1 spinous process cephalad and parallel to the lateral edge of the x-ray film. RESULTS: Intraexaminer reliability for the (a) horizontal base angle was .72 to .94, with confidence intervals included in the range of .52 to .97; (b) lumbodorsal angle was .90 to .96, with confidence intervals in the range of .82 to .98; (c) lumbosacral angle was .84 to .96, with confidence intervals in the range of .72 to .98, and (d) thoracic translational displacement from vertical was .95 to.97, with confidence intervals included in the range of .91 to .99. Interexaminer reliability for the three examiners ranged from .71 to .97. CONCLUSIONS: Measures similar to those described in this study are commonly used to measure and categorize spinal displacements from true vertical alignment (ie, scoliosis measurements). Most patient assessment methods used in chiropractic have poor or unknown reliability. The one possible exception to this rule is spinal displacement analysis performed on radiographs. In chiropractic, intraclass correlation coefficients values greater than .70 are considered accurate enough for use in clinical and research applications. The measures tested here would fit within these guidelines of reliability. Establishing reliability is an important first step in evaluating these measures so that future studies of validity may be undertaken.

Analysis of Variance↗

Assessing the utility of reliability indices for automated visual fields. Testing ocular hypertensives.

Monocular (right eye) visual fields were recorded with the Humphrey Visual Field Analyzer (30-2 Program) at baseline as well as 6 and 12 months later in 120 patients with established ocular hypertension. Indices of field reliability (fixation loss, less than 20%; false-positives and false-negatives, less than 33%) and field sensitivity (mean deviation [MD] and pattern standard deviation [PSD]) were examined. At baseline, 35% of patients exhibited low reliability (LR) fields, a figure which decreased to approximately 25% at 6 and 12 months, respectively. During this period, over 50% of patients produced at least one LR field, whereas 8.3% were unable to produce even one reliable field. Exhibition of a LR field appeared to be independent of patient age. Fixation errors, the major cause of LR fields, decreased by approximately 10% over the 12-month period; most patients had between 20 and 32% fixation errors. The incidence of significant defects identified by PSD was greater than that for MD; this was true for both reliable and LR fields. It is suggested that increasing the fixation loss criteria for assessing patient reliability to a 33% cutoff might substantially increase the percentage of fields graded reliable with minimal effect on the sensitivity or specificity of the test.

Adult↗

Comparison of reliability indices in conventional and high-pass resolution perimetry.

PURPOSE: The purpose of this study is to compare reliability indices in conventional (Humphrey) and high-pass resolution (Ring) perimetry in healthy subjects followed prospectively at 6-month intervals. METHODS: Of the 146 healthy subjects (mean age, 50.24 years; range, 30-84 years) enrolled in the study, 102 have been tested twice and 71 three times. The authors compared the reliability indices, fixation losses, false-positive rate, and false-negative rate between the two techniques, both cross-sectionally and serially. RESULTS: Fixation losses were slightly higher with high-pass resolution perimetry, whereas false-positive errors were higher with conventional perimetry. False-negative errors were uncommon with either technique. Of 319 fields, 30 (9.4%) conventional and 39 (12.2%) high-pass resolution perimetry fields were unreliable using the current suggested reliability criteria. Nearly all unreliable fields were due to high fixation errors. Using alternative criteria derived from baseline 95th percentile values, unreliable fields were attributed more equally to all three reliability parameters. In subjects tested three times, the reliability indices remained constant. CONCLUSION: The results of this study showed that healthy subjects have comparable reliability indices when tested with conventional and high-pass resolution perimetry.

Adult↗

A novel approach to assess inter-rater reliability in the use of the Overt Aggression Scale-Modified.

The ability of raters to apply measures of efficacy accurately and reproducibly in clinical trials of psychotropic medications is vital to the validity of study data. In a large, multi-center trial using the Overt Aggression Scale-Modified (OAS-M), inter-rater reliability was evaluated using standardized patients (SP) in a live 'mock' interview. Thirty raters experienced in the OAS-M each interviewed two actors trained to portray patients with intermittent explosive disorder, borderline and antisocial personality disorder, and/or post-traumatic stress disorder. Inter-rater reliability of OAS-M scores was evaluated using the intra-class correlation coefficient (ICC). The joint reliability of the raters in the study with an expert rater was also assessed by an ICC. The ICC was 0.96 between the raters, and 0.98 between the raters and expert rater. Results suggest that SP can be utilized in rater assessment and that raters administering the OAS-M in this model of rater reliability assessment demonstrate a high level of consistency and reliability. The method described here may be useful in future assessments of inter-rater reliability.

Adult↗

One-year test-retest reliability of auditory ERPs in young and old adults.

The reliability of ERP measures was investigated in a sample covering the adult life span (n = 59, age 21-92). This sample was divided into a young and an old group. ERPs to an auditory two-stimuli oddball task were recorded in the sample at two occasions separated by 12-14 months (T1 and T2). The recordings of T1 were split in half, to assess within session reliability. Correlations were calculated for N1, P2 and P3 peak latency and amplitude, for average amplitude during 50 ms epochs within the defined P3-window 250-550, and for average amplitude in successive 15 ms epochs from 1 to 705 ms. The results show that amplitude measures were more reliable than the latency measures at all electrodes. Time window/epoch amplitude measures yielded reliabilities in the same range as peak amplitude. Reliabilities peaked around the conventionally studied N1, P2 and P3, and this is seen as a validation of the components. In general, the old group exhibited weaker P3 peak latency reliabilities than the young group. However, many of these differences did not reach statistical significance. Implications of the findings are discussed.

Adult↗

[Reliability and validity of self-reported drug use among secondary school students].

OBJECTIVE: To assess the reliability and validity of self-reported use among secondary school students. METHOD: Validity was assessed in a representative sample of nearly 1,300 students by analyzing: a) the proportion of questions on drug use left unanswered compared with that for other questions; b) the proportion of inconsistencies between related questions; c) the proportion of questions wrongly completed; d) admision of ficticious drug use; e) the relationship between self-reported drug use and that of friends', and f) willingness to admit cannabis and ecstasy use. Reliability was analyzed using the kappa index, the proportion of specific agreement, and the intra-class correlation coefficient in a test-retest procedure in a randomized subsample of 349 students. RESULTS: The response rate to questions on drug use was high and was similar to that for more neutral questions. Only 0,3% of the secondary school students reported having used a fictitious drug. Except in the case of heroin, individual drug use was directly related to friends' perception of consumption. A very low proportion of students would not be willing to admit to use of cannabis (2%) or ecstasy (3,4%). Questions referring to drug use at some time during students' lives showed greater reliability than those referring to more recent drug use (last 12 months). The kappa indexes for drug consumption at some time during students lives ranged from 0.65 to 0.87, except for ecstasy and LSD (0.51 and 0.52 respectively). The age of first drug consumption was highly reliable (intra-class correlation coefficients ranged from 0.71 to 1). CONCLUSIONS: Except for ectasy, amphetamines and LSD, indexes of reliability and validity were generally good and similar to those obtained in other studies. These findings support the idea that information about self-reported drug use obtained through a questionnaire is reliable and valid, although the absolute prevalences of the use of some drugs should be interpreted with caution. The indicator of consumption at some time during students' lives is especially useful in studies monitoring drug use among this population.

Adolescent↗

The reliability of physical examination for carpal tunnel syndrome.

The goal of this study was to determine the interobserver and intraobserver reliability of static and moving two-point discrimination, Semmes-Weinstein monofilament testing, Tinel's test, manual motor testing of abductor pollicis brevis, vibration and Phalen's test in the diagnosis of carpal tunnel syndrome. Twelve patients with suspected carpal tunnel syndrome were examined in an outpatient setting. The interobserver reliability was satisfactory for all tests except for Semmes-Weinstein monofilament testing. Intraobserver reliability was also satisfactory for all tests. Static two point discrimination had higher reliability than moving two-point discrimination. Seven tests for the diagnosis of carpal tunnel syndrome were reliable in the hands of skilled health care professionals. Hand surgeons and hand therapists examined patients more reliably than occupational health workers.

Adult↗

Reliability of peroneal reaction time measurements.

OBJECTIVE: The purpose of this study was to demonstrate the reliability of reaction time-measurements on a tilting platform under consideration of various influencing factors. DESIGN: The peroneal reaction time of 30 healthy subjects was examined in an experimental study. BACKGROUND: Peroneal reaction time measurements have been used to objectively evaluate functional instability of the ankle joint, but the reliability of the method has not been proven yet. METHODS: The reaction time after sudden inversion of the ankle were determined by surface EMG. RESULTS: The median latency of the peroneus brevis was 66 ms and that of the peroneus longus was 63 ms. No differences between male and female subjects and between left and right legs could be found. An increase of reaction time was caused by neuromuscular fatigue (P=0.033, for both the peroneus brevis and the peroneus longus). A decrease in reaction time resulted if the foot was held in 15 degrees of plantar flexion (P=0.0004 for the peroneus brevis, P=0.002 for the peroneus longus). The reliability was examined by circadian and by day-to-day measurements. The coefficient of correlation (Spearman's rho) between the peroneus brevis and days 1-5 was 0.67 (P=0.177) and for the peroneus longus 0. 00 (P0.999). The same results were obtained after the circadian measurements. CONCLUSION: Determination of peroneal reaction time was proven as a reliable measurement method. RELEVANCE: Reliability and validity are basic preconditions of a test to become accepted as a clinical measurement method. This paper demonstrates the reliability of measuring the peroneal reaction time. Thus, assuming validity, the peroneal reaction time measurement is justified as a clinical test.

Adult↗

Intra- and inter-observer variability and reliability of prostate volume measurement via two-dimensional and three-dimensional ultrasound imaging.

We describe the results of a study to evaluate the intra- and inter-observer variability and reliability of prostate volume measurements made from transrectal ultrasound (TRUS) images, using either the (optimal) height-width-length (HWL) method (V = pi/6 HWL) with two-dimensional (2D) TRUS images (obtained as cross-sections of three-dimensional [3D] TRUS images) or manual planimetry of 3D TRUS images (the 3D US method). In this study, eight observers measured 15 prostate images, twice via each method, and an analysis of variance (ANOVA) was performed. This analysis shows that, with the 3D US method, intra-observer prostate volume estimates have 5.1% variability and 99% reliability, and inter-observer estimates have 11.4% variability and 96% reliability. With the HWL method, intra-observer estimates have 15.5% variability and 93% reliability, and inter-observer estimates have 21.9% variability and 87% reliability. Thus, in vivo prostate volume estimates from manual planimetry of 3D TRUS images have much lower variability and higher reliability than HWL estimates from 2D TRUS images.

Analysis of Variance↗