Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Inter-observer reliability of ten tests used for predicting difficult tracheal intubation.

PURPOSE: To determine inter-observer reliability of ten preoperative airway assessment tests used for predicting difficult tracheal intubation. METHOD: We prospectively assessed 59 patients undergoing elective surgery requiring tracheal intubation at a large metropolitan teaching hospital. Two experienced observers independently conducted the airway assessment tests on the same group of patients. Inter-observer reliability was examined using Kappa (K) and intraclass correlation coefficient (ICC). RESULTS: Two tests--mouth opening (ICC = 0.93) and chin protrusion (ICC = 0.89)--had excellent inter-observer reliability. Seven tests--thyromental distance (ICC - 0.74), subluxation (K = 0.66), atlanto-occipital extension distance (ICC = 0.67) and angle (K = 0.66), profile classification (K = 0.58), ramus length (ICC = 0.53), oropharyngeal best view (K = 0.49)--were moderately reliable. One test--Mallampati technique of assessing oropharyngeal view (K = 0.31)--had poor reliability. CONCLUSION: Many of the preoperative airway tests have only moderate inter-observer reliability. This may provide some insight into why previous research has failed to show that the tests accurately predict difficult tracheal intubation.

Atlanto-Occipital Joint↗

[Reliability and validity of the Anaesthesiological Questionnaire for electively operated patients].

OBJECTIVE: The Anaesthesiological Questionnaire (ANP) is a self-rating method for the assessment of postoperative complaints and patient satisfaction. The questionnaire consists of two parts. Part 1 assesses the intensity of symptoms regarding the postoperative period in the "recovery-room and the first hours on the ward" (19 items) and the "current state" (17 items). Part 2 assesses patient satisfaction with the anaesthetic care as well as the unspecific perioperative care and postoperative convalescence. The questionnaire was designed to fulfill the criteria of reliability and validity and to serve as a practicable means of auditing the quality of routine clinical practice. METHODS: A total of 1,112 patients older than 18 years completed the questionnaire after an elective operation. Additionally data concerning the type of anaesthesia were recorded from the anaesthesia chart. To determine retest-reliability, 94 patients competed the ANP twice postoperatively. RESULTS: The participants of the survey represented 74.6% of the total collective. Out of 19 items 16 had a retest-reliability of r(tt)>0.70, the 3 other items had a reliability of r(tt)>0.50. Reliability (Cronbach's Alpha) of the patient satisfaction scales was between r(tt)=0.76 and r(tt)=0.91. In relation to the period immediately after anaesthesia,women reported more postoperative complaints than men but no differences were found between male and female patients with regard to satisfaction with perioperative care. Younger patients (18-49 years old) described more postoperative complaints than older patients and a lesser degree of satisfaction with perioperative care. There were plausible differences in postoperative complaints between patients who received general vs. regional anaesthesia. Patients reported less postoperative complaints after TIVA than after volatile anaesthetics. The configuration of patient characteristics and anaesthesia gives indications to "risk groups" who predominantly suffer after anaesthesia. CONCLUSIONS: The Anaesthesiological Questionnaire (ANP) is a reliable and valid method for the assessment of postoperative complaints and patient satisfaction.

Adolescent↗

Reliability in the assessment of tendon volume and intratendinous signal of the Achilles tendon on MRI: a methodological description.

The purpose is to introduce a method for accurately and objectively evaluating volume and mean intratendinous signal within the Achilles tendon using MRI. We prospectively studied MRI from 33 patients with chronic Achilles tendinosis (20 males and 13 females) with a median age of 52 years (range 29-70). In all patients, both Achilles tendons were investigated with T1-WI as well as PD-WI MRI. Thus, 66 Achilles tendons were evaluated in the study. Tendon volume and mean intratendinous signal were evaluated using a computerized 3-D seed-growing technique. In general, the computerized 3-D seed-growing technique resulted in an excellent overall observer reliability of the MRI-measurements. The reliability (R) for tendon volume measurements was highest for the T1-WI sequence (R=97.9%). For the mean intratendinous signal, the PD-WI sequence showed the highest reliability (R=88.1%). The same pattern was present when we studied the coefficient of variation (CV). For the CV, lower figures indicate more reliable estimates. CV was 4.9% for tendon volume and 8.9% for mean intratendinous signal. In conclusion, it could be said that a computerized 3-D seed-growing technique to monitor and evaluate the volume of the Achilles tendon and mean intratendinous signal, using MRI, shows an overall excellent reliability regarding inter- as well as intra-observer reliability.

Achilles Tendon↗

Reliability of an instrumental assessment of tardive dyskinesia: results from VA Cooperative Study #394.

Nine VA Medical Centers are participating in a 2-year double-blind placebo controlled study of antioxidant treatment for tardive dyskinesia (TD) conducted by the Department of Veteran Affairs Cooperative Studies Program. One of the principal outcome measures of this study is the score derived from the instrumental assessment of upper extremity dyskinesia. Dyskinetic hand movements are quantified by assessing the variability associated with steady-state isometric force generated by the patient. In the present report, we describe the training procedures and results of a multi-center reliability assessment of this procedure. Data from nine study centers comprising 45 individual patients with six trials each (three from left hand and three from right hand) were reanalyzed by an independent investigator and the results were subjected to reliability assessment. For the statistic of interest (average coefficient of variation over trials 2 and 3 for each hand, then take the larger of these two values), we found very high intraclass correlation coefficients for reliability over all patients across sites (ICC = 0.995). We also calculated the reliability of the measures across trials within patient for each combination of hand (right, left, dominant), rater group (site, control), and trials set (all three, trials 2 and 3). For a given hand and trial set, the reliability of the site raters was similar to that of the control. This study demonstrates that instrumental measures for the assessment of dyskinesia are reliable and can be implemented in multi-center studies with minimal training.

Adult↗

Observer reliability in grading nephrocalcinosis on ultrasound examinations in children.

BACKGROUND: Nephrocalcinosis is often associated with a variety of hypercalcemic conditions. Diagnostic ultrasound is often used for assessing nephrocalcinosis in children, but its reliability has not been proven. OBJECTIVE: To determine the reliability of expert interpretation of sonographic films with a grading scale of severity for nephrocalcinosis. MATERIALS AND METHODS: Fifty-eight ultrasonographic films of 30 children with Williams syndrome and other conditions know to be associated with nephrocalcinosis were assessed. We used a blinded randomized design to assess intra- and interobserver reliability. RESULTS: Grades I, II, and III nephrocalcinosis were noted in 13 %, 19 %, and 27 % of the examinations, respectively. The weighted kappa coefficient was 0.80 (standard error 0.12; 95 % confidence interval 0.68-0.92) for intraobserver agreement and 0.76 (standard error 0.13; 95 % confidence interval 0.63 to 0.89) for interobserver agreement. Reliability in assessing change from one examination to the next, with independently graded films, was fair with an unweighted kappa coefficient of 0.68 (95 % confidence interval 0.38-0.96) and 0.51 (95 % confidence interval 0.21-0.80) for intra- and interobserver reliability, respectively. CONCLUSION: The severity of nephrocalcinosis can be reliably interpreted with an ultrasonography grading scale.

Child↗

Diagnostic interview for genetic studies (DIGS): inter-rater and test-retest reliability of the French version.

The National Institute of Mental Health developed the semi-structured Diagnostic Interview for Genetic Studies (DIGS) for the assessment of major mood and psychotic disorders and their spectrum conditions. The DIGS was translated into French in a collaborative effort of investigators from sites in France and Switzerland. Inter-rater and test-retest reliability of the French version have been established in a clinical sample in Lausanne. Excellent inter-rater reliability was found for schizophrenia, bipolar disorder, major depression, and unipolar schizoaffective disorder while fair inter-rater reliability was demonstrated for bipolar schizoaffective disorder. Using a six-week test-retest interval, reliability for all diagnoses was found to be fair to good with the exception of bipolar schizoaffective disorder. The lower test-retest reliability was the result of a relatively long test-retest interval that favored incomplete symptom recall. In order to increase reliability for lifetime diagnoses in persons not currently affected, best-estimate procedures using additional sources of diagnostic information such as medical records and reports from relatives should supplement DIGS information in family-genetic studies. Within such a procedure, the DIGS appears to be a useful part of data collection for genetic studies on major mood disorders and schizophrenia in French-speaking populations.

Adolescent↗

Static and dynamic standing balance: test-retest reliability and reference values in 9 to 10 year old children.

INTRODUCTION: Based on the literature, reliability reports and normative data for bilateral stance assessments in elementary schoolchildren are limited. The present study was designed to report test-retest reliability and reference values for postural stability in 9 to 10 years old schoolchildren using the Balance Master system. MATERIALS AND METHODS: Twenty children participated in the reproducibility study (mean age 10.1+/-0.7) including test and retest measurement with a one-week interval. The modified clinical test of sensory interaction on balance (mCTSIB) quantified children's static standing balance. The test for the limits of stability (LOS) measured dynamic standing balance. The study sample to determine reference values consisted of 99 children (mean age 9.8+/0.5). RESULTS: The ICCs for inter-item reliability of the four sensory conditions of the mCTSIB showed fair to excellent reliability (ICCs between 0.62 and 0.80). The reproducibility between test and retest was non-significant for the condition 'firm surface with eyes closed' (ICC of 0.37), fair to good for the three other sensory conditions (ICCs between 0.59 and 0.68), and excellent for the composite sway velocity (ICC of 0.77). For all LOS parameters, the significant ICCs showed fair to good reproducibility (ICCs between 0.44 and 0.62), with the exception of the non-significant ICC for the composite reaction time. The ICCs for the separate LOS parameters showed fair to good and excellent reliability for nine parameters (ICCs between 0.46 and 0.81), while 11 separate LOS scores did not demonstrate significant ICCs. DISCUSSION: Analysing reference values, girls performed better on all the composite balance parameters compared to boys, with the exception of reaction time and movement velocity. No differences were found on standing balance scores between 9 and 10 year olds. CONCLUSION: In conclusion, the Balance Master showed fair to good reliability for most postural parameters in 9 to 10 year olds. The current data on postural control in children aged 9 to 10 years are relevant for research in other domains within the clinical field, like obesitas and developmental coordination disorder or in relation to back pain prevalence at early age.

Child↗

Interjudge and intrajudge reliabilities in fiberoptic endoscopic evaluation of swallowing (fees) using the penetration-aspiration scale: a replication study.

This study used Fiberoptic Endoscopic Evaluation of Swallowing (FEES(R)) to assess the reliability of the Penetration-Aspiration Scale (PAS) using 79 swallows and four judges in a replication of a study using videofluoroscopy (VFSS). The swallows were diagnosed using FEES, which allowed for comparison between the two techniques. The findings indicated that all categories of the PAS achieved adequate reliability, both on intrajudge and interjudge assessments. Reliabilities, with the exception of Scale Score 7, were higher in this study than in the original study by Rosenbek and associates. Data analysis indicated that judges were more highly consistent on second ratings compared with their original ratings, indicating a learning curve on the PAS. In addition, findings suggested that the FEES was more reliable on assessing penetration than VFSS, but that VFSS was more reliable on the assessment of the various severities of aspiration. The two techniques were equally effective in discriminating between penetration and aspiration. This study found that FEES was just as reliable as VFSS when using the PAS.

Deglutition Disorders↗

The MISTELS program to measure technical skill in laparoscopic surgery : evidence for reliability.

BACKGROUND: The McGill Inanimate System for Training and Evaluation of Laparoscopic Skills (MISTELS) is a series of five tasks with an objective scoring system. The purpose of this study was to estimate the interrater and test-retest reliability of the MISTELS metrics and to assess their internal consistency. METHODS: To determine interrater reliability, two trained observers scored 10 subjects, either live or on tape. Test-retest reliability was assessed by having 12 subjects perform two tests, the second immediately following the first. Interrater and test-retest reliability were assessed using intraclass correlation coefficients. Internal consistency between tasks was estimated using Cronbach's alpha. RESULTS: The interrater and test-retest reliabilities for the total scores were both excellent at 0.998 [95% confidence interval (CI), 0.985-1.00] and 0.892 (95% CI, 0.665-0.968), respectively. Cronbach's alpha for the first assessment of the test-retest was 0.86. CONCLUSIONS: The MISTELS metrics have excellent reliability, which exceeds the threshold level of 0.8 required for high-stakes evaluations. These findings support the use of MISTELS for evaluation in many different settings, including residency training programs.

Clinical Competence↗

Reliability of motion measurements after total disc replacement: the spike and the fin method.

As motion preservation is one of the main postulated advantages after total disc replacement (TDR) of the lumbar spine, the quantification of the mobility after TDR seems of special clinical interest. Yet, the best method to assess range of motion (ROM) after TDR remains unclear. The aim of the study was the calculation of 95%-confidence intervals (95%-C.I.) for the measurement error accompanying: (1) different methods (2) different observers and (3) different levels of training for radiographic motion analysis after TDR. In 12 patients the level L4-L5 and in another 12 patients level L5-S1 were measured with the Cobb and the superimposition method on flexion-extension X-rays after monosegmental TDR. Both methods were adopted as the landmarks used the spikes of the prosthesis instead the endplates (spike method) and the fin of the prosthesis instead the whole vertebral body (fin method). Measurements were performed by two experienced (O-I and O-III) and one inexperienced observer (O-II). The adopted spike and fin method showed a better reliability compared to the reported results of the original Cobb and superimposition method. The method used was not clinically relevant for the intraobserver reliability in the experienced observer (95%-C.I.: +/-2.0 degrees for the fin and +/-2.1 for the spike method) and for the interobserver reliability for two experienced observers (95%-C.I.: -2.8 degrees /+2.8 degrees for the fin and -2.9 degrees /+3.1 degrees for the spike method). The intraobserver reliability for the inexperienced observer was inferior for both methods compared to the experienced observer but no clinically relevant differences could be observed in interobserver reliability measures. The spike and fin method are reliable methods for study protocols dealing with angular motion after TDR as clinically valid conclusions can be drawn with an accuracy of about +/-2 degrees for the same observer and with an accuracy of about +/-3 degrees for a different observer.

Adult↗

Reliability of the first eye and second eye in frequency doubling technology perimetry.

PURPOSE: To compare the reliability of the perimetry results of the first eye and the second eye with frequency doubling technology (FDT). METHODS: The subjects were 328 residents who underwent the C-20-5 mode of FDT at a city in central Japan. FDT perimetry was always performed first in the right eye and then in the left eye without any time between tests. When more than 33% fixation loss or false-positive error was detected, the result was judged unreliable. RESULTS: Of the 328 subjects, the bilateral perimetry results were reliable in 255 subjects (77.7%), the unilateral results were reliable in 57 (17.4%), and there was not a reliable result in either eye in 16 (4.9%) subjects. Of the 57 subjects whose unilateral result was unreliable, the result of the second eye was unreliable in 50 (88%) subjects, and the result of the first eye was unreliable in 7 (12%). This difference in the reliability between the first eye and second eye was significant (P < 0.001). There were no differences in the age, sex, visual acuity, refractive error, test duration, or number of glaucoma suspects between the two groups whose unilateral result was unreliable in the first eye or the second eye. CONCLUSIONS: The FDT perimetry result of the second eye was less reliable than that of the first eye.

Adult↗

Development, reliability and validity of a new measure of overall health for pre-school children.

BACKGROUND: Few comprehensive systems are available for assessing and reporting the overall health of preschool children. OBJECTIVES: (i) To develop a multi-dimension health status classification system (HSCS) to describe pre-school (PS) children 2.5-5 years of age; (ii) to report reliability and validity of the newly developed measure. DESIGN: Existing systems (Health Utilities Index, Mark 2 and 3) were adapted for application to a pre-school population. The new system was tested for acceptability, validity and reliability. PARTICIPANTS: Three cohorts of children and their parents from Canada and Australia were utilized: Cohort 1 (MAC)-101 3-years old very low birthweight (VLBW, <1500 g) and 50 same age term children from Canada; Cohort 2 (AUS)-150 VLBW 3-years old from Australia; Cohort 3 (OMG)-222 3-years old with cerebral palsy (CP) from Ontario. METHODS: Parental intra-rater reliability was evaluated by completion of the HSCS-PS Parent questionnaire (MAC) at the clinic visit and again 14 days later. Health professionals (MAC) completed the HSCS-PS Clinician questionnaire. Percent agreement and Kappa values were used to assess parent-clinician agreement. Concurrent validity was tested in two populations of VLBW children (MAC and AUS) and a reference group of term children (MAC) by exploring the relationships between dimensions of the HSCS-PS and well-recognized norm-referenced measures: the Bayley Scales of Infant Development (BSID-II), the Vineland Adaptive Behavior Scales (VABS) and the Stanford-Binet (SB). Construct validity was tested by comparing ratings on both the HSCS-PS and the Gross Motor Function classification system (GMFCS) using a population of pre-school children with CP. Analyses were done using chi2, ANOVA and correlations with tau-b statistic. RESULTS: The HSCS-PS has 12 dimensions and 3-5 levels per dimension. Response rate for parental intra-rater reliability was 95%, with percent agreement ranging between 86 and 100%. Kappa values for various dimensions ranged from 0.38 to 1.00. Inter-rater reliability between parents and clinicians showed agreement ranging from 72 to 100%. Kappa values ranged from 0.30 to 1.00. CONCURRENT VALIDITY: There was a statistically significant gradient between HSCS-PS Mobility levels and motor scale scores of the BSID-II and VABS. A significant gradient also occurred when comparing HSCS-PS cognition levels to psychometric scores on the BSID-II and SB, as well as HSCS-PS self-care levels compared to VABS Daily Living scores. DISCRIMINATIVE AND CONSTRUCT VALIDITY: Birthweight category was shown to be a significant determinant of proportion of children with multiple HSCS-PS dimensions affected. In addition, HSCS-PS dimension levels were congruent with GMFCS levels where expected: mobility had excellent correlation; self-care, dexterity, speech and cognitive dimensions had moderate correlations. CONCLUSIONS: The HSCS-PS is readily accepted, quick to complete, widely applicable and provides a multi-dimensional description of health status. Preliminary assessments of reliability and validity are promising. The HSCS-PS can discriminate across populations by birthweight and shows strong relationships with standardized psychometric measures in comparable domains. It can pro- vide a summary profile of functional limitations in various populations of pre-school children in a consistent manner across programs and in different settings.

Australia↗

Reliability of eccentric isokinetic knee flexion and extension measurements.

This study assessed the test-retest reliability of knee isokinetic eccentric muscle performance in subjects with and without a history of tibio-femoral pathology. Nineteen adults were tested at 60 degrees/sec and 180 degrees/sec on three occasions using a standardized protocol that incorporates a same-session learning phase. Results revealed moderate to excellent reliability for average peak torque test-retest ICC (2,1) = .58 to .96, total work ICC = .63 to .93, and power ICC = .67 to .93. Joint angle at peak torque was unreliable (ICC = .01 to .69) for both muscle groups at both angular velocities. Knee flexion reliability was higher than extension reliability at both 60 degrees/sec and 180 degrees/sec. Subjects with tibio-femoral pathologies had ICC values lower than the healthy subjects. Reliable eccentric isokinetic measurements can be obtained for average peak torque, total work, and power. Clinicians should not assume the same degree of reliability in testing patients as in testing healthy subjects.

Adult↗

Isokinetic testing of ankle strength in older adults: assessment of inter-rater reliability and stability of strength over six months.

The study purposes were (1) to estimate the inter-rater reliability of isokinetic strength tests at the ankle in older adults (test-retest interval of three to 7 days), and to determine whether more experienced examiners were more reliable; and (2) to estimate 6 month stability of strength tests. Inter-rater reliability was high for plantar flexion and dorsiflexion tests where average strength was more than about 10 Newton-meters (Nm) (Pearson R = 0.87-0.95). When average strength was less than 10Nm, reliability was less (R = 0.42-0.75). Experienced examiners (physical therapists) and less experienced examiners (research assistants) were equally reliable. Variability in strength over 6 months was no greater than variability over a few days. We conclude that isokinetic tests of ankle strength in older adults are highly reliable and stable when examiners are adequately trained and subjects maintain usual physical activity levels.

Aged↗

Reliability of topographic quantitative EEG amplitude in healthy late-middle-aged and elderly subjects.

Reliabilities of quantitative measures of absolute and relative EEG amplitudes were assessed in healthy older adults under the eyes closed (n = 46) and eyes opened (n = 45) conditions. For the theta, alpha, beta 1, and beta 2 bands, reliabilities of 28 scalp derivations were stable over the 4.5 month test interval. Reliabilities of delta were lower. When appropriate transformations were applied, the reliabilities of absolute EEG amplitude measures tended to exceed those of relative measures. There were not, however, striking differences in reliabilities under the eyes closed, as compared to eyes opened condition. We concluded that when coupled with the criterion of interpretability, the generally higher reliabilities of absolute, as opposed to relative, amplitude measures render them preferable in clinical research.

Aged↗

Theory of reliability, biological systems and aging.

The paper deals with possibilities and prospects of application of reliability theory in biology and particularly in gerontology. The possibilities and limitations of existing reliability methods are analyzed with respect to their eventual use in biology, and new methods suitable for analysis of changes in the reliability of biological systems during their aging process are searched for. A new method is proposed for reliability parameter estimation of elements and subsystems forming biological systems. Application of the method could bring new information regarding the causes of declining reliability of biological systems during their aging. Application of the method is demonstrated by an analysis of reliability of the human organism, and the results are compared with some current theories of aging.

Adolescent↗

Test-retest reliability of the P50 mid-latency auditory evoked response.

Attenuation in mid-latency auditory evoked responses (MLAERs) can be used to study sensory gating. If paired-click stimuli (S1 and S2) are used, lower amplitude in response to S2 vs. S1 (attenuation) is considered evidence for intact sensory gating. However, the need for reliable measurements of MLAER amplitude and attenuation is a recognized problem. Ten normal volunteers were studied six times each. An S1 amplitude test-retest reliability coefficient (r) of 0.585 was obtained when means of two recordings were used vs. reliability coefficients as high as 0.809 for means of six recordings. Averaging a higher number of runs (120 vs. 60) resulted in a reliability coefficient of 0.677/recording. Similar values were obtained for S1 and S2 latencies. Reliability coefficients for S2 attenuation (S2/S1) were not nearly as high (a value of 0.138 when means of all six recordings were used). The S1 amplitude as measured in this study (with 120 averages) appears to be a reliable psychophysiologic measurement, but the S2/S1 attenuation measure is more variable, perhaps reflecting a greater sensitivity of the S2/S1 to uncontrolled variables in this study. Further research to identify such variables is necessary.

Adult↗

The reliability and stability of a quantity-frequency method and a diary method of measuring alcohol consumption.

The study aimed to assess the test-retest reliability of two commonly used measures of alcohol consumption, the quantity-frequency (QF) method and the diary method, as well as the stability of scores on the two measures over time. Two methods of assessing reliability and stability were employed. The first was a traditional method based on calculation of correlation coefficients for agreement between scores on repeated measures over a short retest interval to yield test-retest reliability coefficients, and over a long retest interval to yield stability coefficients. The second method was that devised by Wiley and Wiley (1970) to differentiate the effects of reliability and stability on repeated measures over time. The two methods were applied to a sample of heavy drinkers and to a sample of light drinkers. The results indicated that both the QF and diary measures are reliable in measuring alcohol consumption of light drinkers. Both measures are less reliable for heavy drinkers. The results indicate, in addition, that drinking consumption levels of light drinkers demonstrate a high degree of stability. However, the consumption levels of heavy drinkers demonstrate less stability, especially over a long time period. Heavy drinkers significantly reduced reported levels of alcohol consumption on both measures after the first test, suggesting a regression to the mean effect or the possibility of unintended intervention effects due to repeated measurement of drinking behaviour.

Adult↗