Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Reliability of retrospective self-reports of alcohol consumption among women: data from a U.S. national sample.

Drinking histories (retrospective self-reports) can be a valuable resource for time-ordered analyses of causes and consequences of drinking. However, there is a scarcity of data on the reliability of drinking histories from general population samples. We report here on the reliability and consistency of reported ages of onset and typical drinking frequencies, quantities and volume, from drinking histories provided in 1981 and 1986 by national samples of women drinkers with and without drinking problems. Statistical reliability was generally modest, yet large percentages of women gave exactly the same reports 5 years apart. Reliability was apparently reduced by limited response options, and was lower among younger drinkers, whose drinking was more changeable between 1981 and 1986. We discuss ways to improve reliability and to make best use of drinking histories.

Adult↗

[Test-retest reliability of measures of social network in the "Pró -Saúde" Study].

OBJECTIVE: To evaluate test-retest reliability of social network-related information of the" Pr -Sa de" study. METHODS: A test-retest reliability study was conducted using a multidimensional questionnaire applied to a cohort of university employees. The same questionnaire was filled out twice by 192 non-permanent employees with two weeks apart. Agreement was estimated using kappa statistics (categorical variables), weighted kappa statistics, log-linear models (ordinal variables), and intraclass correlation coefficient (discrete variables). RESULTS: Estimates of reliability were higher than 0.70 for most variables. Stratified analyses revealed no consistently varying patterns of reliability according to gender, age or schooling strata. Log-linear modelling showed that, for the study ordinal variables, the model of best fit was "diagonal agreement plus linear by linear association". CONCLUSIONS: The high level of reliability estimated in this study suggests that the process of measurement of social network-related aspects was adequate. Validation studies, which are currently being conducted, will complete the quality assessment of this information.

Adult↗

Translation into Brazilian Portuguese, cultural adaptation and evaluation of the reliability of the Disabilities of the Arm, Shoulder and Hand Questionnaire.

The objective of the present study was to translate, adapt and validate a Brazilian Portuguese version of the Disabilities of the Arm, Shoulder and Hand (DASH) Questionnaire. The study was carried out in two steps. The first was to translate the DASH into Portuguese and to perform cultural adaptation and the second involved the determination of the reliability and validity of the DASH for the Brazilian population. For this purpose, 65 rheumatoid arthritis patients of either sex (according to the classification criteria of the American College of Rheumatology), ranging in age from 18 to 60 years and presenting no other diseases involving the upper limbs, were interviewed. The patients were selected consecutively at the rheumatology outpatient clinic of UNIFESP. The following results were obtained: in the first step (translation and cultural adaptation), all patients answered the questions. In the second step, Spearman's correlation coefficients for interobserver evaluation ranged from 0.762 to 0.995, values considered to be highly reliable. In addition, intraclass correlation coefficients ranged from 0.97 to 0.99, also highly reliable values. Spearman's correlation coefficients and the intraclass correlation coefficients obtained during intra-observer evaluation ranged from 0.731 to 0.937 and from 0.90 to 0.96, respectively, being highly reliable values. The Ritchie Index showed a weak correlation with Brazilian DASH scores, while the visual analog scale of pain showed a good correlation with DASH score. We conclude that the Portuguese version of the DASH is a reliable instrument.

Adolescent↗

[Reliability of the Portuguese-language version of the Social Phobia Inventory (SPIN) among adolescent students in the city of Rio de Janeiro].

It is believed that social phobia has its onset during adolescence and precedes other mental disorders; it is thus important to investigate the condition among young people. To date there is no self-reported scale validated for the Brazilian population. The present study investigated the reliability of the Portuguese-language version of the Social Phobia Inventory (SPIN) among adolescent students from public schools in the city of Rio de Janeiro. After SPIN was translated into Portuguese, a test-retest reliability study was carried out with 190 students. Intra-class correlation coefficient (ICC) and weighted kappa (kappaw2) were estimated, log-linear models were fitted, and Bland & Altman graphs were built. The Portuguese version showed good internal consistency (Cronbach's alpha=0.88) and good reliability for total score (ICC=0.78). Reliability for single items was not very good (kappaw2 between 0.32 and 0.65). The "semi-association" model fit most of the items best. Based on these findings we concluded that the Portuguese-language version of SPIN showed good reliability results, similar to those obtained with the original English-language version.

Adolescent↗

Reliability of surface electromyographic measurements from subjects with spinal cord injury during voluntary motor tasks.

In this study, the reliability of surface electromyographic data (root-mean-square) for volitional motor tasks drawn from a standardized protocol was assessed. For each motor task, 5 s epochs of data were analyzed with a new method to generate a measure called the voluntary response index (VRI). The VRI consists of two components, magnitude and similarity index (SI), that were separately analyzed for repeatability. We examined three repetitions of each of 10 volitional motor tasks in 69 subjects with spinal cord injury (American Spinal Injury Association [ASIA] Impairment Scale [AIS], classifications C and D: 34 AIS-C and 35 AIS-D) for short-term (within-day) reliability. In 6 of the 69 subjects (3 each, AIS-C and AIS-D), the entire study was repeated after 1 week and results were assessed for intermediate-term (1 week apart) reliability. The reliability of the method for voluntary motor tasks was assessed by intraclass correlation coefficient (ICC), analysis of variance, coefficient of variance, and Pearson's correlation. Good reliability was found for magnitude (ICC = 0.71-0.99, Pearson's r = 0.77-0.99) and for SI (ICC = 0.65-0.96, Pearson's r = 0.72-0.93) for three repeated tests (within-day). Significant difference was found for studies completed 1 week apart for magnitude (p = 0.02) but not for SI (p = 0.57). In addition, SI showed less variation than magnitude (p < 0.001). No significant difference of magnitude and SI between tasks was observed.

Algorithms↗

COMFORT scale: a reliable and valid method to measure the amount of stress of ventilated preterm infants.

OBJECTIVE: Assessment of clinimetric properties and diagnostic quality of a stress measurement scale (COMFORT scale). DESIGN: Sample of an open population. SETTING: Neonatology department (Neonatal Intensive Care Unit), Academic Medical Centre/Emma Children's Hospital, Amsterdam, The Netherlands. METHOD: One clinical expert and 9 observers observed ventilated premature born babies simultaneously. Criterion validity was assessed by correlating the COMFORT scale with the clinical judgment regarding the amount of stress. Interobserver reliability was assessed on the clinical judgment as well as on the COMFORT scale. Diagnostic qualities were evaluated with a ROC curve. RESULTS: On 19 ventilated prematurely born babies (mean gestational age 30 weeks, mean birth weight 1385 gm), one clinical expert and 9 observers made 30 paired observations. The criterion validity of the COMFORT scale was good (Pearson's r of 0.84). The interobserver reliability of the clinical judgment was very good (weighted Kappa 0.84). The interobserver reliability of each item varied from good to almost perfect (weighted Kappa of 0.64 for muscle tone to 1.00 on heart rate). The reliability of the total COMFORT scale score was satisfying (intra-class correlation coefficient of 0.94). The diagnostic quality of the COMFORT scale was excellent, at a cut-off point of 20 the sensitivity was 100 percent, the specificity was 77 percent, and the area under the curve (AUC) of 0.95. CONCLUSION: In this first evaluation, the COMFORT scale appears to be a valid and reliable measurement tool to assess the stress of ventilated prematurely born babies.

Female↗

Measuring maternal sensitivity in teen mothers: reliability and feasibility of two instruments.

This study's purpose was to compare the reliability and feasibility of two instruments measuring maternal sensitivity in adolescent mothers-the Maternal Sensitivity Q-Sort (MBQS) and the Ainsworth Maternal Sensitivity Scale (AMSS). After 2, 4, and 6 hours of home observations, 2 raters completed the MBQS and the AMSS on 10 adolescent mothers. The intraclass correlations for interrater reliability of the MBQS and AMSS were .80 and .81, respectively. There was no significant difference in raters' scores among times 1, 2, or 3 for either instrument, implying that perhaps only 2 hours of observation is required. Training times for the MBQS and AMSS were 15 hours and 13 hours, respectively. Completion time for the MBQS averaged 59 minutes compared to 5 minutes for the AMSS. At a pay rate of $12/hour for one rater, completion of 30 visits costs $354 for the MBQS compared to $30 for the AMSS. There was magnitude bias in both instruments such that the lower the sensitivity score, the greater was the difference in ratings. Results suggest that while both tools are equally reliable, the AMSS is the most cost-effective and time-efficient. Agreement among raters leads only to reliability, not validity. Testing needs to be done on a larger sample of adolescents to further evaluate reliability as well as the relative validity of the measures.

Adolescent↗

Considerations in the choice of interobserver reliability estimates.

Two types of interobserver reliability values may be needed in treatment studies in which observers constitute the primary data-acquisition system: trial reliability and the reliability of the composite unit or score which is subsequently analyzed, e.g., daily or weekly session totals. Two approaches to determining interobserver reliability are described: percentage agreement and "correlational" measures of reliability. The interpretation of these estimates, factors affecting their magnitude, and the advantages and limitations of each approach are presented.

Journal Article↗

Evaluating interobserver reliability of interval data.

Previous recommendations to employ occurrence, nonoccurrence, and overall estimates of interobserver reliability for interval data are reviewed. A rationale for comparing obtained reliability to reliability that would result from a random-chance model is explained. Formulae and graphic functions are presented to allow for the determination of chance agreement for each of the three indices, given any obtained per cent of intervals in which a response is recorded to occur. All indices are interpretable throughout the range of possible obtained values for the per cent of intervals in which a response is recorded. The level of chance agreement simply changes with changing values. Statistical procedures that could be used to determine whether obtained reliability is significantly superior to chance reliability are reviewed. These procedures are rejected because they yield significance levels that are partly a function of sample sizes and because there are no general rules to govern acceptable significance levels depending on the sizes of samples employed.

Journal Article↗

Measuring the environment for friendliness toward physical activity: a comparison of the reliability of 3 questionnaires.

OBJECTIVES: We tested the reliability of 3 instruments that assessed social and physical environments. METHODS: We conducted a test-retest study among US adults (n = 289). We used telephone survey methods to measure suitableness of the perceived (vs objective) environment for recreational physical activity and nonmotorized transportation. RESULTS: Most questions in our surveys that attempted to measure specific characteristics of the built environment showed moderate to high reliability. Questions about the social environment showed lower reliability than those that assessed the physical environment. Certain blocks of questions appeared to be selectively more reliable for urban or rural respondents. CONCLUSIONS: Despite differences in content and in response formats, all 3 surveys showed evidence of reliability, and most items are now ready for use in research and in public health surveillance.

Adult↗

The Neer classification system for proximal humeral fractures. An assessment of interobserver reliability and intraobserver reproducibility.

The radiographs of fifty fractures of the proximal part of the humerus were used to assess the interobserver reliability and intraobserver reproducibility of the Neer classification system. A trauma series consisting of scapular anteroposterior, scapular lateral, and axillary radiographs was available for each fracture. The radiographs were reviewed by an orthopaedic shoulder specialist, an orthopaedic traumatologist, a skeletal radiologist, and two orthopaedic residents, in their fifth and second years of postgraduate training. The radiographs were reviewed on two different occasions, six months apart. Interobserver reliability was assessed by comparison of the fracture classifications determined by the five observers. Intraobserver reproducibility was evaluated by comparison of the classifications determined by each observer on the first and second viewings. Kappa (kappa) reliability coefficients were used. All five observers agreed on the final classification for 32 and 30 per cent of the fractures on the first and second viewings, respectively. Paired comparisons between the five observers showed a mean reliability coefficient of 0.48 (range, 0.43 to 0.58) for the first viewing and 0.52 (range, 0.37 to 0.62) for the second viewing. The attending physicians obtained a slightly higher kappa value than the orthopaedic residents (0.52 compared with 0.48). Reproducibility ranged from 0.83 (the shoulder specialist) to 0.50 (the skeletal radiologist), with a mean of 0.66. Simplification of the Neer classification system, from sixteen categories to six more general categories based on fracture type, did not significantly improve either interobserver reliability or intraobserver reproducibility.

Evaluation Studies as Topic↗

Interobserver reliability and intraobserver reproducibility of the modified Ficat classification system of osteonecrosis of the femoral head.

Anteroposterior and lateral plain radiographs of 116 osteonecrotic femoral heads were reviewed to assess the interobserver reliability and intraobserver reproducibility of the modified Ficat classification system. The radiographs were reviewed initially and then again six months later by three adult reconstructive surgeons, two general orthopaedic surgeons, two orthopaedic residents, and one musculoskeletal radiologist. All eight observers agreed on the classification of twenty hips (17 per cent) at both the first and the second review of the radiographs. Paired comparisons revealed a mean interobserver kappa reliability coefficient of 0.46 (range, 0.30 to 0.67) for the first review and 0.45 (range, 0.30 to 0.66) for the second. For all observers, the mean rate of perfect agreement between the first and the second review was 68 per cent (range, 56 to 80 per cent). The mean kappa value for intraobserver reproducibility was 0.59 (range, 0.44 [one of the residents] to 0.73 [one of the general orthopaedic surgeons]). No observer or pair of observers had excellent reproducibility or reliability (kappa > 0.75). The poor interobserver reliability and fair intraobserver reproducibility diminishes any meaningful comparison of studies in which the modified Ficat classification system has been used and illuminates the need for a more reliable and reproducible classification system.

Femur Head Necrosis↗

Interobserver reliability and intraobserver reproducibility of the system of King et al. for the classification of adolescent idiopathic scoliosis.

The classification of adolescent idiopathic scoliosis with use of the system of King et al. has become widely accepted since its introduction. The purpose of the present study was to establish the interobserver reliability and intraobserver reproducibility of this classification system. The preoperative radiographs of sixty-three patients who were managed operatively for adolescent idiopathic scoliosis were classified by five observers with the system of King et al. Interobserver reliability was assessed by comparison of the classification of the curves among the observers, and intraobserver reproducibility was evaluated by comparison of the classifications of each set of radiographs by each observer on two occasions three weeks apart. The median interobserver reliability kappa coefficient for the classification system of King et al. was 0.44 (range, 0.28 to 0.50), and the median intraobserver reproducibility kappa coefficient was 0.64 (range, 0.44 to 0.72). According to the definition of Landis and Koch, the classification system of King et al. is substantially reproducible but is only moderately reliable. However, according to the stricter definition of Svanholm et al., its reproducibility is only fair and its reliability is poor.

Adolescent↗

The reliability and validity of the self-reported patient-specific index for total hip arthroplasty.

BACKGROUND: The Patient-Specific Index is unique in that it reflects how individual patients weigh concerns in rating the outcome of total hip arthroplasty. The Patient-Specific Index was originally administered by an interviewer, which is not always feasible and can be costly. The purposes of the present study were (1) to create a self-reported version of the Patient-Specific Index, (2) to determine the reliability of this new self-reported version, and (3) to determine the relationship between the scores on the new self-reported version and those on the original interviewer-administered version. METHODS: A self-reported version of the Patient-Specific Index was developed, and a pilot test was performed on ten patients. Patients who were scheduled for a total hip arthroplasty or who had recently had a total hip arthroplasty were eligible for the reliability and validity testing. A copy of the new self-reported Patient-Specific Index was mailed to the patients, and they completed it independently. The patients' ratings of the importance and severity of twenty-four concerns prior to total hip arthroplasty were added together to create a summary Patient-Specific Index score. To determine test-retest reliability, patients completed the self-reported Patient-Specific Index a second time, two weeks later. To determine criterion validity, participants also completed the interviewer-administered Patient-Specific Index. RESULTS: Fifty-five patients completed the study. The random-effects intraclass correlation test-retest coefficient was 0.79 (greater than 0.75 represents excellent reliability). The mean Patient-Specific Index scores on the self-reported version and on the interviewer-administered version were 173 and 165 points, respectively (Student t test, p = 0.45). The self-reported Patient-Specific Index was concordant with the interviewer-administered Patient-Specific Index (intraclass correlation coefficient, 0.78). CONCLUSIONS: We concluded that a self-reported version of the Patient-Specific Index, which focuses on the concerns of individuals, is reliable and has criterion validity compared with an interviewer-administered version.

Activities of Daily Living↗

Reliability and intraoperative validity of preoperative assessment of standardized plain radiographs in predicting bone loss at revision hip surgery.

BACKGROUND: The most challenging aspect of revision hip surgery is the management of bone loss. A reliable and valid measure of bone loss is important since it will aid in future studies of hip revisions and in preoperative planning. We developed a measure of femoral and acetabular bone loss associated with failed total hip arthroplasty. The purpose of the present study was to measure the reliability and the intraoperative validity of this measure and to determine how it may be useful in preoperative planning. METHODS: From July 1997 to December 1998, forty-five consecutive patients with a failed hip prosthesis in need of revision surgery were prospectively followed. Three general orthopaedic surgeons were taught the radiographic classification system, and two of them classified standardized preoperative anteroposterior and lateral hip radiographs with use of the system. Interobserver testing was carried out in a blinded fashion. These results were then compared with the intraoperative findings of the third surgeon, who was blinded to the preoperative ratings. Kappa statistics (unweighted and weighted) were used to assess correlation. Interobserver reliability was assessed by examining the agreement between the two preoperative raters. Prognostic validity was assessed by examining the agreement between the assessment by either Rater 1 or Rater 2 and the intraoperative assessment (reference standard). RESULTS: With regard to the assessments of both the femur and the acetabulum, there was significant agreement (p < 0.0001) between the preoperative raters (reliability), with weighted kappa values of >0.75. There was also significant agreement (p < 0.0001) between each rater's assessment and the intraoperative assessment (validity) of both the femur and the acetabulum, with weighted kappa values of >0.75. CONCLUSIONS: With use of the newly developed classification system, preoperative radiographs are reliable and valid for assessment of the severity of bone loss that will be found intraoperatively.

Acetabulum↗

Assessment of the reliability of a technique to measure postural sway in horses.

OBJECTIVE: To assess the reliability of the center-of-pressure (COP) values obtained from a force platform for analysis of postural sway in horses. ANIMALS: Six 2-year-old horses that were free from lameness and neurologic disease. PROCEDURE: Horses stood stationary with all 4 hooves on a force platform; COP data were collected at 1,000 Hz and 3-dimensional kinematics collected at 60 Hz for 10 seconds. Five trials were recorded at each of 3 time periods (15-minute intervals) or at 1 time period on 3 separate days. Mean values for each set of 5 trials and actual, normalized, and relative COP variables were calculated. The reliability was quantified by use of agreement boundary. RESULTS: The COP results within and across days were similar and provided small agreement boundary limits (eg, across days, in order of least relative reliability: area, +/- 62 mm2; mediolateral range, +/- 8 mm; radius, +/- 2 mm; craniocaudal range, +/- 4 mm; and velocity, +/- 3 mm/s). Head height possessed the greatest relative intraday reliability (12%) but a high agreement boundary limit (+/- 0.15 m). CONCLUSIONS AND CLINICAL RELEVANCE; The use of a force platform to analyze postural sway in a group of young healthy horses was found to produce reliable results and may provide a simple and sensitive measure for assessing balance deficiencies in horses. Agreement boundaries provide 95% confidence intervals for use as limits of error and variability in measurements that, if exceeded, may signify meaningful effects.

Animals↗

Reliability of the WAIS-R with a mixed patient sample.

Reliability of the WAIS-R for a mixed sample of psychiatric and neurological patients was determined. Split-half reliabilities and standard errors of measurement were computed for all subtests (exept Digit Span and Digit Symbol) and the Verbal, Performance, and Full Scale IQs. Only the Arithmetic subtest had a reliability estimate which was significantly lower than that of the 35-44-yr.-old standardization group. All other reliabilities and standard errors of measurement approximated those reported for the normative group. It was concluded that the WAIS-R is a reliable instrument for evaluating patients in a clinical setting.

Adult↗

An assessment of the validity and reliability of two perceived exertion rating scales among Hong Kong children.

This study evaluated the validity and reliability of the Chinese-translated (Cantonese) versions of the Borg 6-20 Rating of Perceived Exertion (RPE) scale and the Children's Effort Rating Table (CERT) during continuous incremental cycle ergometry with 10- to 11-yr.-old Hong Kong school children. A total of 69 children were randomly assigned, with the restriction of groups being approximately equal, to two groups using the two scales, CERT (n = 35) and RPE (n = 34). Both groups performed two trials of identical incremental continuous cycling exercise (Trials 1 and 2) 1 wk, apart for the reliability test. Objective measures of exercise intensity (heart rate, absolute power output, and relative oxygen consumption) and the two subjective measures of effort were obtained during the exercise. For both groups, significant Pearson correlations were found for perceived effort ratings correlated with heart rate (rs > or = .69), power output (rs > or = .75), and oxygen consumption (rs > or = .69). In addition, correlations for CERT were consistently higher than those for RPE. High test-retest intra-class correlations were found for both the effort (R = .96) and perceived exertion (R = .89) groups, indicating that the scales were reliable. In conclusion, the CERT and RPE scales, when translated into Cantonese, are valid and reliable measures of exercise intensity during controlled exercise by children. The Effort rating may be better than the Perceived Exertion scale as a measure of perceived exertion that can be more validly and reliably used with Hong Kong children.

Attitude to Health↗