Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

The emergency severity index triage algorithm version 2 is reliable and valid.

OBJECTIVES: Initial studies have shown improved reliability and validity of a new triage tool, the Emergency Severity Index (ESI), over conventional three-level scales at two university medical centers. After pilot implementation and validation, the ESI was revised to include pediatric and updated vital signs criteria. The goal of this study was to assess ESI version (v.) 2 reliability and validity at seven emergency departments (EDs) in three states. METHODS: In part 1, interrater reliability was assessed using weighted kappa analysis of written training cases and postimplementation by a random sampling of actual patient triages. In part 2, validity was analyzed using a prospective cohort with stratified random sampling at each site. The ESI was compared with outcomes including resource consumption, inpatient admission, ED length of stay, and 60-day all-cause mortality. RESULTS: Weighted kappa analysis of interrater reliability ranged from 0.70 to 0.80 for the written scenarios (n = 3289) and 0.69 to 0.87 for patient triages (n = 386). Outcomes for the validity cohort (n = 1042) included hospitalization rates by ESI triage level: level 1, 83%; 2, 67%; 3, 42%; 4, 8%; level 5, 4%. Sixty-day all-cause mortality by triage level was as follows: level 1, 25%; 2, 4%; 3, 2%; 4, 1%; and 5, 0%. CONCLUSIONS: ESI v. 2 triage produced reliable, valid stratification of patients across seven sites. ESI triage should be evaluated as an ED casemix identification system for uniform data collection in the United States and compared with other major ED triage methods.

Algorithms↗

Interrater reliability of plaque morphology classification in patients with severe carotid artery stenosis.

OBJECTIVE: Ultrasonographic assessment of carotid artery plaque morphology is widely used to identify patients at high risk for stroke. However, the reliability of plaque analysis in high-grade stenosis is uncertain. We determined the interrater reliability of sonographic plaque morphology analysis in patients with severe carotid artery stenosis. MATERIALS AND METHODS: Duplex Doppler was performed on 114 patients with 80-99% stenosis of the internal carotid artery using a Siemens Quantum 2000 D with a handheld 7.5 MHz transducer. B-mode pictures with and without color coding were printed on a Sony color video printer UP-5000 W. Three raters independently evaluated plaque echolucency, heterogeneity, calcification, and surface structure. Interrater agreement was calculated by a jackknife procedure generating kappa values and two-sided 95% confidence intervals. RESULTS: Kappa values and 95% confidence intervals were 0.05 (-0.07 to 0.16) for plaque surface structure, 0.15 (0.02 to 0.28) for plaque heterogeneity, 0.18 (0.09 to 0.29) for plaque echogenicity, and 0.29 (0.19 to 0.39) for plaque calcification. The upper bounds of all of the confidence intervals were below the 0.40 level suggested for minimal reliability. CONCLUSION: The low interrater agreement indicated that unaided visual assessment of static B-mode pictures to assess plaque morphology in patients with severe carotid artery stenosis is not reliable. Other evaluation procedures and standardized criteria, as yet undeveloped, are needed to improve reliability.

Arteriosclerosis↗

The reliability of child psychiatric diagnosis. A comparison among Danish child psychiatrists of traditional diagnoses and a multiaxial diagnostic system.

The study was conducted to compare an experimental multiaxial diagnostic system (MAS) with traditional multicategorical diagnoses in child psychiatric work. Sixteen written case histories were circulated to 21 child psychiatrists, who made diagnoses independently of one another, using two different diagnostic systems. Diagnostic reliability was measured as percentage of interrater agreement. The highest diagnostic reliability was obtained in psychotic disorders, the lowest in personality disorders. The MAS implied improved diagnostic reliability of mental retardation, somatic disorders and developmental disorders. Adjustment reaction (reactio maladaptiva) was the diagnosis most commonly used, but with varying reliability in both systems. The reliability of the socio-economic and psychosocial axes were generally high.

Adjustment Disorders↗

The interrater reliability of a Dutch version of the Structured Clinical Interview for DSM-III-R Personality Disorders.

This study presents data on the interrater reliability of a Dutch version of the Structured Clinical Interview for DSM-III-R Personality Disorders (SCID-II). Seventy outpatients were interviewed before the start of their treatment by one rater, while a second rater observed. Both raters were instructed to make independent ratings and the second rater was not allowed to participate in the discussion. On criterion level, interrater reliabilities appear to be very good, with a few exceptions (most reliabilities are higher than 0.75). However, all 5 observation criteria had poor interrater reliabilities. Agreement on personality disorder, on the whole, was excellent (overall kappa = 0.80). The possible reasons why relatively lower reliabilities are found with some criteria are discussed. Finally, problems encountered during the interviews are addressed and possible adjustments of the SCID-II are suggested.

Adolescent↗

Quantification of dental plaque on lingual tooth surfaces using image analysis: reliability and validation.

AIM: The aim of this study was to increase the versatility and further validate the method reported by Smith et al. (2001) by testing the reliability of plaque measurement against two well-known dental plaque quantification methodologies using image analysis in a clinical trial. METHOD: The teeth of 40 subjects were disclosed before digital images of the labial and lingual surfaces of their upper and lower incisors were acquired. The amount of plaque present was quantified using a modification of the method described by Smith et al. (2001). The method was modified for obtaining images of the lingual surfaces by incorporating the use of orthodontic occlusal mirrors and 5-mm pieces of moistened blue articulating paper used to enable calibration. Plaque measurements were made from 320 upper and lower anterior teeth from the 40 subjects by two operators. Fliess' coefficient of reliability was used to assess intra- and inter-operator reliability and the independent sample t test was used to assess statistical significance between test and control groups after checking the data for normality. For validation, measurements were recorded using the Turesky et al. (1970) (modification of the Quigley & Hein (1962) plaque index and the Addy et al. (1983) plaque area index. The results were compared with the image analysis method using Pearson's correlation coefficient. RESULTS: The results for reliability were within Fliess' range of "excellent" for both intra-operator repeatability and inter-operator reproducibility. Pearson's correlation coefficients showed highly significant values indicating the close similarity between all three methods. CONCLUSIONS: This method for the measurement of dental plaque on lingual surfaces of anterior teeth proved reliable. The combined results from the labial and lingual surfaces of anterior teeth using image analysis produced trial conclusions comparable with the alternate plaque quantification methods used, with less clinician time and further producing a permanent database of images for future use.

Adult↗

Reliability of caries data in three clinical trials.

The reliability of caries data obtained during three clinical trials is presented. Each of the three trials, conducted by different examiners, lasted 3 years and involved children aged 12 to 15 years. Reliability was quantified in terms of reliability coefficient and error variance, which allowed the effect of error upon the efficiency of each study to be measured. Reliability tended to be reduced when precaviation (initial) lesions were included in the DMFS count, but there was little difference between the reliability coefficients for each of four different surface-types. Error was more important in 3-year incremental rather than prevalence data, but nevertheless had a rather small influence upon sample size estimation or confidence limits of percent caries reductions.

Adolescent↗

Reliability in estimating taurodontism of permanent molars from orthopantomograms.

Taurodontism is a morphologic dental trait showing continuous expressivity, and criteria of the degree of pulp chamber elongation vary in different investigations. The aim of this investigation was to test a simple method of assessing taurodontism in the developing dentition from orthopantomograms in order to determine its reliability for later use in epidemiologic investigations. The method was also compared with other methods. Forty-three children 10-16.9 yr of age with one or more taurodontic permanent first or second molars were selected for the study. A subgroup of 16 children with two longitudinally obtained radiographs was used in a follow-up study of the same tooth in two different formational stages. The follow-up time averaged 2.2 yr. The distance between the baseline connecting the mesial and distal points of the cementoenamel junction and the highest point of the floor of the pulp chamber was measured (Measure 3). A tooth was classified as taurodontic when Measure 3 reached or exceeded 3.5 mm. This distance remained unchanged during the course of tooth development. Intraexaminer reliability of two examiners in reproducing the same classification was, on average, 96.2%, and the interexaminer reliability was 93.2%. The reliability was greater for the first than for the second molars. Results obtained by this method agreed well with those obtained by other methods. Measure 3 proved to be reliable in assessing taurodontism in the developing dentition from orthopantomograms in epidemiologic investigations.

Adolescent↗

Inter-examiner reliability in the clinical examination of temporomandibular disorders: influence of age.

OBJECTIVE: The aim of this study was to investigate the influence of the age of the subject on inter-examiner reliability of the clinical signs of temporomandibular disorder (TMD). METHODS: Forty-three elderly (ES) and 44 younger adults (YS) were selected. The female/male distribution was almost the same in the two groups. All participants underwent clinical examination according to the Research Diagnostic Criteria for TMD, performed successively by two clinicians. RESULTS: For metric measurements - with the exception of unassisted opening - the ES gave a significantly lower range of motion with both examiners and significantly worse percentage agreement between the examiners. A remarkable inter-examiner disagreement in the elderly was found with laterotrusion and protrusion movements. The prevalence of joint sounds was rated inconsistently by the examiners. The reliability of detection was not different in the two groups. The prevalence of tender muscle sites was also inconsistent. The overall percentage agreement for subjects with at least one tender muscle point was not age dependent. Because of the very low prevalence in ES, further statistical assessment of reliability is not possible. CONCLUSIONS: The age-dependent lower range of motion and the inferior reliability of metric measurements in the elderly could lead to wrong diagnoses. The reliability of detecting joint sounds and tender muscles was not age dependent within the limitations of the study.

Adult↗

High-frequency audiometry: test reliability and procedural considerations.

This study compared the reliability of a recently developed high-frequency audiometer (HFA) [Stevens et al., J. Acoust. Soc. Am. 81, 470-484 (1987)] with a less complicated system that uses supraaural earphones (Koss system). The new approach permits calibration on an individual basis, making it possible to express thresholds at high frequencies in dB SPL. Data obtained from 50 normal-hearing subjects, ranging in age from 10-60 years, were used to evaluate the effects on reliability of threshold variance, earpiece/earphone fitting variance, and the variance associated with the HFA calibration process. Without earpiece/earphone replacement, the reliability of thresholds for the two systems is similar. With replacement, the HFA showed poorer reliability than the Koss system above 11 kHz, largely due to errors in estimating the calibration function. HFA reliability is greater for subjects with valid calibration functions over the entire frequency range. When average correction factors are applied to the Koss data in an effort to convert threshold estimates to dB SPL, individual transfer functions are not represented accurately. Thus the benefit of being able to express thresholds at high frequencies in dB SPL must be weighed against the additional source of variability introduced by the HFA calibration process.

Adolescent↗

Reliability of isokinetic and isometric knee-extensor force in older women.

Because of the need for efficient, consistent strength measurements, the test-retest reliability of concentric, isometric, and eccentric strength; concentric work; and concentric power was determined in older women without a familiarization session. The reliability of measures derived from a single peak score were compared with those derived from an averaged score. On 2 occasions 25 older women with a mean age of 72 +/- 6 years performed 3 submaximal knee extensions and 5 maximal contractions on an isokinetic dynamometer at 90 degrees/s (CON), 0 degrees/s, and -90 degrees/s on both lower limbs. Statistical analyses for peak and averaged values (best 3 contractions of 5) exhibited good relative reliability (ICCs > .88), except for CON power. Typical error as a coefficient of variation and ratio limits of agreement for peak and averaged score values were larger than desired, with CON power scores demonstrating unacceptable error ranges. Although relative reliability of this 1-session assessment protocol was acceptable, further research is needed to determine whether additional practice trials could enhance absolute reliability.

Aged↗

Reliability and validity of the Outcome Expectations for Exercise Scale-2.

Development of a reliable and valid measure of outcome expectations for exercise for older adults will help establish the relationship between outcome expectations and exercise and facilitate the development of interventions to increase physical activity in older adults. The purpose of this study was to test the reliability and validity of the Outcome Expectations for Exercise-2 Scale (OEE-2), a 13-item measure with two subscales: positive OEE (POEE) and negative OEE (NOEE). The OEE-2 scale was given to 161 residents in a continuing-care retirement community. There was some evidence of validity based on confirmatory factor analysis, Rasch-analysis INFIT and OUTFIT statistics, and convergent validity and test criterion relationships. There was some evidence for reliability of the OEE-2 based on alpha coefficients, person- and item-separation reliability indexes, and R(2)values. Based on analyses, suggested revisions are provided for future use of the OEE-2. Although ongoing reliability and validity testing are needed, the OEE-2 scale can be used to identify older adults with low outcome expectations for exercise, and interventions can then be implemented to strengthen these expectations and improve exercise behavior.

Aged, 80 and over↗

Validity and reliability of acoustic analysis of respiratory sounds in infants.

OBJECTIVE: To investigate the validity and reliability of computerised acoustic analysis in the detection of abnormal respiratory noises in infants. METHODS: Blinded, prospective comparison of acoustic analysis with stethoscope examination. Validity and reliability of acoustic analysis were assessed by calculating the degree of observer agreement using the kappa statistic with 95% confidence intervals (CI). RESULTS: 102 infants under 18 months were recruited. Convergent validity for agreement between stethoscope examination and acoustic analysis was poor for wheeze (kappa = 0.07 (95% CI, -0.13 to 0.26)) and rattles (kappa = 0.11 (-0.05 to 0.27)) and fair for crackles (kappa = 0.36 (0.18 to 0.54)). Both the stethoscope and acoustic analysis distinguished well between sounds (discriminant validity). Agreement between observers for the presence of wheeze was poor for both stethoscope examination and acoustic analysis. Agreement for rattles was moderate for the stethoscope but poor for acoustic analysis. Agreement for crackles was moderate using both techniques. Within-observer reliability for all sounds using acoustic analysis was moderate to good. CONCLUSIONS: The stethoscope is unreliable for assessing respiratory sounds in infants. This has important implications for its use as a diagnostic tool for lung disorders in infants, and confirms that it cannot be used as a gold standard. Because of the unreliability of the stethoscope, the validity of acoustic analysis could not be demonstrated, although it could discriminate between sounds well and showed good within-observer reliability. For acoustic analysis, targeted training and the development of computerised pattern recognition systems may improve reliability so that it can be used in clinical practice.

Acoustics↗

Reliable keratometry with a new hand held surgical keratometer: calibration of the keratoscopic astigmatic ruler.

AIM: Some surgeons consider hand held surgical keratometers unreliable. This may be due to incorrect use through not realising that the distance that the keratometer is held from the cornea influences the shape of the image. When a keratometer is held closer to the astigmatic cornea, the elliptical image will appear more circular, particularly for larger degrees of astigmatism. However, the keratoscopic astigmatic ruler (KAR) has design features that correct the hitherto unrecognised problems with the use of a hand held keratometer. This study assesses the reliability and accuracy of measurement of astigmatism using the KAR. METHODS: The KAR and the Bausch & Lomb keratometer (B&L) were compared using six back surface toric cut contact lens blanks representing 1 to 6 dioptres of astigmatism. Two observers (one experienced in the use of the keratometers, the other a novice) took eight randomly repeated "masked" measurements of each lens blank with the KAR and four measurements with the B&L in a similar fashion. RESULTS: There was no difference between the measurements with either instrument by each of the observers (p = 0.95, ANOVA). The standard error of measurement for the KAR was 0.59 D, for the B&L, 0.31 D. The intraclass correlation coefficient of reliability for the KAR was 0.90 and for the B&L it was 0.97. The coefficient of repeatability for the KAR was plus or minus 0.83 D, and for the B&L plus or minus 0.77 D. The interobserver reliability for the KAR was 0.898, and for the B&L, 0.975. CONCLUSION: These results suggest that the KAR has good reliability and reproducibility and compares favourably with the B&L keratometer. Inexperience with use does not affect reliability.

Analysis of Variance↗

Interobserver and intraobserver reliability of a classification scheme for corneal topographic patterns.

AIMS: To determine the interobserver and the intraobserver reliability of a published classification scheme for corneal topography in normal subjects using the absolute scale. METHOD: A prospective observational study was done in which 195 TMS-1 corneal topography maps in the absolute scale were independently classified twice by three classifiers--a cornea fellow, an ophthalmic technician, and an optometrist. From these observations the interobserver reliability for each category and the intraobserver reliability for each observer were determined in terms of the median weighted kappa statistic for each category and for each observer. RESULTS: For interobserver reliability, the median weighted kappa statistic for each category varied from 0.72 to 0.97 and for intraobserver reliability the range was 0.79 to 0.98. CONCLUSION: This classification scheme is extremely robust and even in the hands of less experienced observers with minimal training it can be relied upon to provide consistent results.

Corneal Topography↗

Reliability of a device measuring triceps surae muscle fatigability.

OBJECTIVE: To examine the test-retest reliability of a protocol using an apparatus designed to standardise the standing heel rise test for the triceps surae muscle. SUBJECTS: 40 healthy subjects volunteered to test short and medium term test-retest reliability (group SM, median age 24 years), and a convenience sample of 38 subjects with a history of unilateral deep vein thrombosis (DVT) volunteered to test long term test-retest reliability (group L, median age 52 years). DESIGN: Subjects carried out 23 heel rises per minute until either the pace or the height could no longer be maintained. Group SM subjects repeated the test 30 minutes later (short term), and again 48 hours later (medium term). Subjects in group L did the test on the unaffected leg, and repeated the test one week later (long term). RESULTS: The median number of heel rises achieved per trial in group SM was 34 (range 16 to 120). The intraclass coefficient (ICC) was 0.93 (SEM 2.1) for both 30 minute and 48 hour test-retest reliability. In group L, the median number of heel rises was 27 (range 9 to 97), with ICC 0.88 and SEM 3.4. CONCLUSIONS: The apparatus is a simple and inexpensive standardised tool that reliably measures triceps surae fatigability in subjects with no current injury. Future research should assess its use in injured patients.

Adolescent↗

Measurement of scapula upward rotation: a reliable clinical procedure.

BACKGROUND: It is important to deal with the scapula when developing rehabilitation strategies for the shoulder complex. This requires clinical measurement tools that are readily available and easy to apply and which provide a reliable evaluation of scapula motion. AIM: To determine the reliability of the Plurimeter-V gravity inclinometer for the measurement of scapular upward rotation positions during humeral elevation in coronal abduction in a group of patients with shoulder pathology. METHOD: Twenty six patients were assessed in two repeat tests within a single testing session. Patients exhibiting a wide spectrum of shoulder pathology were selected. The angle of scapular upward rotation was measured during total shoulder abduction. The measurement protocol was performed twice during a single testing session by a single tester. Results of the two tests were compared and the reliability assessed by intraclass correlation coefficients (ICCs). RESULTS: There was no significant difference in the scapula measurements taken during the two tests at each testing position. Overall, there was very good intrarater reliability (ICC = 0.88). The ICC ranged from 0.81 (at 135 degrees) to 0.94 (at both resting and end of total shoulder abduction range). CONCLUSION: The Plurimeter-V gravity inclinometer can be used effectively and reliably for measuring upward rotation of the scapula in all ranges of shoulder abduction in the coronal plane.

Adult↗

Eccentric and concentric isokinetic knee flexion and extension: a reliability study using the Cybex 6000 dynamometer.

OBJECTIVE: To determine the reliability of the Cybex 6000 isokinetic dynamometer in measuring the knee muscle performance concentrically and eccentrically. METHODS: 18 male and 12 female subjects with no previous knee injuries, who had not previously undergone any isokinetic testing, were studied. The flexor and extensor muscles groups of both knees were tested at 60 degrees s-1 and 120 degrees s-1 with the continuous concentric-eccentric cycle testing protocol. Variables studied included peak torque, total work, and average power. The interclass correlation coefficient (ICC 2,1) was used to determine the reliability with P < 0.05. RESULTS: Peak torque showed significantly greater ICC than the total work and average power, with test-retest reliability ranging from 0.82 to 0.91 for peak torque, from 0.76 to 0.89 for total work, and 0.71 to 0.88 for average power. Average variability for the three variables studied ranged from 9% to 14%. The ICCs for the three variables studied were significantly greater at 120 degrees s-1. The knee extensor muscle group showed greater test-retest reliability, and the results of isokinetic testing in the concentric contraction mode were more reproducible. CONCLUSIONS: The Cybex 6000 isokinetic dynamometer shows high reliability in measuring isokinetic concentric and eccentric variables. Some fluctuation should be allowed when evaluating variations of muscle performance between tests.

Adult↗

The reliability and validity of the physical activity questions in the WHO health behaviour in schoolchildren (HBSC) survey: a population study.

OBJECTIVE: To assess the test-retest reliability and validity of the physical activity questions in the World Health Organisation health behaviour in schoolchildren (WHO HBSC) survey. METHODS: In the validity study, the Multistage Fitness Test was administered to a random sample of year 8 (mean age 13.1 years; n = 1072) and year 10 (mean age 15.1 years; n = 954) high school students from New South Wales (Australia) during February/March 1997. The students completed the self report instruments on the same day. An independent sample of year 8 (n = 121) and year 10 (n = 105) students was used in the reliability study. The questionnaire was administered to the same students on two occasions, two weeks apart, and test-retest reliability was assessed. Students were classified as either active or inadequately active on their combined responses to the questionnaire items. Kappa and percentage agreement were assessed for the questionnaire items and for a two category summary measure. RESULTS: All groups of students (boys and girls in year 8 and year 10) classified as active (regardless of the measure) had significantly higher aerobic fitness than students classified as inadequately active. As a result of highly skewed binomial distributions, values of kappa were much lower than percentage agreement for test-retest reliability of the summary measure. For year 8 boys and girls, percentage agreement was 67% and 70% respectively, and for year 10 boys and girls percentage agreement was 85% and 70% respectively. CONCLUSIONS: These brief self report questions on participation in vigorous intensity physical activity appear to have acceptable reliability and validity. These instruments need to be tested in other cultures to ensure that the findings are not specific to Australian students. Further refinement of the measures should be considered.

Adolescent↗