Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

The clubfoot assessment protocol (CAP); description and reliability of a structured multi-level instrument for follow-up.

BACKGROUND: In most clubfoot studies, the outcome instruments used are designed to evaluate classification or long-term cross-sectional results. Variables deal mainly with factors on body function/structure level. Wide scorings intervals and total sum scores increase the risk that important changes and information are not detected. Studies of the reliability, validity and responsiveness of these instruments are sparse. The lack of an instrument for longitudinal follow-up led the investigators to develop the Clubfoot Assessment Protocol (CAP). The aim of this article is to introduce and describe the CAP and evaluate the items inter- and intra reliability in relation to patient age. METHODS: The CAP was created from 22 items divided between body function/structure (three subgroups) and activity (one subgroup) levels according to the International Classification of Function, Disability and Health (ICF). The focus is on item and subgroup development. Two experienced examiners assessed 69 clubfeet in 48 children who had a median age of 2.1 years (range, 0 to 6.7 years). Both treated and untreated feet with different grades of severity were included. Three age groups were constructed for studying the influence of age on reliability. The intra- rater study included 32 feet in 20 children who had a median age of 2.5 years (range, 4 months to 6.8 years). The Unweighted Kappa statistics, percentage observer agreement, and amount of categories defined how reliability was to be interpreted. RESULTS: The inter-rater reliability was assessed as moderate to good for all but one item. Eighteen items had kappa values > 0.40. Three items varied from 0.35 to 0.38. The mean percentage observed agreement was 82% (range, 62 to 95%). Different age groups showed sufficient agreement. Intra- rater; all items had kappa values > 0.40 [range, 0.54 to 1.00] and a mean percentage agreement of 89.5%. Categories varied from 3 to 5. CONCLUSION: The CAP contains more detailed information than previous protocols. It is a multi-dimensional observer administered standardized measurement instrument with the focus on item and subgroup level. It can be used with sufficient reliability, independent of age, during the first seven years of childhood by examiners with good clinical experience.A few items showed low reliability, partly dependent on the child's age and /or varying professional backgrounds between the examiners. These items should be interpreted with caution, until further studies have confirmed the validity and sensitivity of the instrument.

Child↗

Inter-rater reliability of nursing home quality indicators in the U.S.

BACKGROUND: In the US, Quality Indicators (QI's) profiling and comparing the performance of hospitals, health plans, nursing homes and physicians are routinely published for consumer review. We report the results of the largest study of inter-rater reliability done on nursing home assessments which generate the data used to derive publicly reported nursing home quality indicators. METHODS: We sampled nursing homes in 6 states, selecting up to 30 residents per facility who were observed and assessed by research nurses on 100 clinical assessment elements contained in the Minimum Data Set (MDS) and compared these with the most recent assessment in the record done by facility nurses. Kappa statistics were generated for all data items and derived for 22 QI's over the entire sample and for each facility. Finally, facilities with many QI's with poor Kappa levels were compared to those with many QI's with excellent Kappa levels on selected characteristics. RESULTS: A total of 462 facilities in 6 states were approached and 219 agreed to participate, yielding a response rate of 47.4%. A total of 5758 residents were included in the inter-rater reliability analyses, around 27.5 per facility. Patients resembled the traditional nursing home resident, only 43.9% were continent of urine and only 25.2% were rated as likely to be discharged within the next 30 days. Results of resident level comparative analyses reveal high inter-rater reliability levels (most items >.75). Using the research nurses as the "gold standard", we compared composite quality indicators based on their ratings with those based on facility nurses. All but two QI's have adequate Kappa levels and 4 QI's have average Kappa values in excess of.80. We found that 16% of participating facilities performed poorly (Kappa <.4) on more than 6 of the 22 QI's while 18% of facilities performed well (Kappa >.75) on 12 or more QI's. No facility characteristics were related to reliability of the data on which Qis are based. CONCLUSION: While a few QI's being used for public reporting have limited reliability as measured in US nursing homes today, the vast majority of QI's are measured reliably across the majority of nursing facilities. Although information about the average facility is reliable, how the public can identify those facilities whose data can be trusted and whose cannot remains a challenge.

Aged↗

Reliability and validity of the AGREE instrument used by physical therapists in assessment of clinical practice guidelines.

BACKGROUND: The AGREE instrument has been validated for evaluating Clinical Practice Guidelines (CPG) pertaining to medical care. This study evaluated the reliability and validity of physical therapists using the AGREE to assess quality of CPGs relevant to physical therapy practice. METHODS: A total of 69 physical therapists participated and were classified as generalists, specialist or researchers. Pairs of appraisers within each category evaluated independently, a set of 6 CPG selected at random from a pool of 55 CPGs. RESULTS: Reliability between pairs of appraisers indicated low to high reliability depending on the domain and number of appraisers (0.17-0.81 for single appraiser; 0.30-0.96 when score averaged across a pair of appraisers). The highest reliability was achieved for Rigour of Development, which exceeded ICC> 0.79, if scores from pairs of appraisers were pooled. Adding more than 3 appraisers did not consistently improve reliability. Appraiser type did not determine reliability scores. End-users, including study participants and a separate sample of 102 physical therapy students, found the AGREE useful to guide critical appraisal. The construct validity of the AGREE was supported in that expected differences on Rigour of Development domains were observed between expert panels versus those with no/uncertain expertise (differences of 10-21% p = 0.09-0.001). Factor analysis with varimax rotation, produced a 4-factor solution that was similar, although not in exact agreement with the AGREE Domains. Validity was also supported by the correlation observed (Kendall-tao = 0.69) between Overall Assessment and the Rigour of Development domain. CONCLUSION: These findings suggest that the AGREE instrument is reliable and valid when used by physiotherapists to assess the quality of CPG pertaining to physical therapy health services.

Cardiology↗

Validity and reliability of retrospective assessment of disease activity and flare in observational cohorts of lupus patients.

BACKGROUND: The use of validated retrospective tools to assess disease activity in an observational cohort of patients would allow researchers the flexibility to analyze unique exploratory questions. Valid, reliable tools exist for assessing disease activity and flare for Systemic Lupus Erythematosus (SLE) patients. However, these tools have been designed for use in structured settings. Many of these populations under study are subject to strict inclusion, exclusion criteria and disease management protocols. The ability to apply these tools to populations not subject to such control would allow researchers to explore questions about unique populations, new treatment applications or predictors of clinical outcomes. This study sought to establish the reliability and validity of retrospective medical record abstraction of SLE disease activity and flare instruments by two rheumatologists. METHODS: From university rheumatology outpatient clinics, 22 patients were randomly selected to establish intra-rater reliability and 26 patients were selected to establish inter-rater reliability. Two rheumatologists using a Physician Global Assessment (PGA), the Systemic Lupus Erythematosus Disease Activity Index (SLEDAI) and the Safety of Estrogens in Lupus Erythematosus-National Assessment (SELENA) flare tool retrospectively abstracted patient charts. Agreement between the tools was evaluated to assess validity. RESULTS: The mean patient age was 39 y; 96% were female, 54% Caucasian, 27% Hispanic, 19% Asian; median PGA was 1.4 (on a 0-3 scale) and median SLEDAI score was 4 (range 0-27). Intra-rater reliability for PGA, SLEDAI and the SELENA flare tool was 0.88, 0.87 and 0.52, respectively. Inter-rater reliability for PGA, SLEDAI, and SELENA flare was 0.79, 0.75, and 0. 50, respectively. To assess validity, the tools were compared against each other to assess agreement. From the parent study sample (n=54 patients), the disease activity measures, PGA and SLEDAI demonstrated adequate agreement (r=0.60). However, the SELENA flare tool demonstrated poor agreement with either PGA-defined flare or SLEDAI-defined flare (weighted kappa, 0.29 and 0.40 respectively). PGA-defined and SLEDAI-defined flare also demonstrated poor agreement (weighted kappa, 0.35). CONCLUSIONS: These data suggest that investigators can reliably reproduce patient disease activity through retrospective chart abstraction using PGA and SLEDAI. Assessing flare is a more difficult concept. The validity of assessing flare at a specific patient-visit is poor. Retrospective assessment of patient flare risk over a specified time period is conceptually more valid and avoids difficulties assessing timing and duration of flare. As there have been no similar published prospective analyses of validity using the SELENA flare tool, it is not clear if this problem was unique to the method of retrospective chart abstraction, the nature of non-protocol study patient visits, the tool itself or a combination of all three aspects.

Adult↗

An examination of interrater reliability for scoring the Rorschach Comprehensive System in eight data sets.

In this article, we describe interrater reliability for the Comprehensive System (CS; Exner. 1993) in 8 relatively large samples, including (a) students, (b) experienced re- searchers, (c) clinicians, (d) clinicians and then researchers, (e) a composite clinical sample (i.e., a to d), and 3 samples in which randomly generated erroneous scores were substituted for (f) 10%, (g) 20%, or (h) 30% of the original responses. Across samples, 133 to 143 statistically stable CS scores had excellent reliability, with median intraclass correlations of.85, .96, .97, .95, .93, .95, .89, and .82, respectively. We also demonstrate reliability findings from this study closely match the results derived from a synthesis of prior research, CS summary scores are more reliable than scores assigned to individual responses, small samples are more likely to generate unstable and lower reliability estimates, and Meyer's (1997a) procedures for estimating response segment reliability were accurate. The CS can be scored reliably, but because scoring is the result of coder skills clinicians must conscientiously monitor their accuracy.

Adult↗

Interrater reliability of Alzheimer's disease diagnosis.

To determine interrater reliability of dementia diagnosis, 4 physicians experienced in the evaluation of dementia patients applied 3 sets of diagnostic criteria to each of 62 patients, based on a standardized set of medical record information. All patients had undergone similar examinations and follow-up to establish the initial clinical diagnosis (76% had autopsy). Raters were blind to the diagnosis and to follow-up information after the initial evaluation period. This paper presents interrater agreement (kappa values) for a diagnosis of Alzheimer's disease using the American Psychiatric Association diagnostic criteria from the Diagnostic and Statistical Manual (DSM-III), the National Institute of Neurological and Communicative Disorders and Stroke (NINCDS) criteria for the clinical diagnosis of Alzheimer's disease, and the Eisdorfer and Cohen Research Diagnostic Criteria (ECRDC) for primary neuronal degeneration. The NINCDS showed somewhat higher average interrater reliability (kappa = 0.64) than the DSM-III (kappa = 0.55) and considerably higher interrater reliability than the ECRDC (kappa = 0.37). One rater displayed conspicuously lower levels of interrater reliability than the other 3, especially in DSM-III and ECRDC. This study indicates that interrater reliability of DSM-III and NINCDS criteria are comparable. Documentation of interrater reliability and, if necessary, training to improve reliability is an important consideration in research where different observers are diagnosing dementing illnesses.

Aged↗

Reliability of the NINDS Myotatic Reflex Scale.

The assessment of deep tendon reflexes is useful for localization and diagnosis of neurologic disorders, but only a few studies have evaluated their reliability. We assessed the reliability of four neurologists, instructed in two different countries, in using the National Institute of Neurological Disorders and Stroke (NINDS) Myotatic Reflex Scale. To evaluate the role of training in using the scale, the neurologists randomly and blindly evaluated a total of 80 patients, 40 before and 40 after a training session. Inter- and intraobserver reliability were measured with kappa statistics. Our results showed substantial to near-perfect intraobserver reliability, and moderate-to-substantial interobserver reliability of the NINDS Myotatic Reflex Scale. The reproducibility was better for reflexes in the lower than in the upper extremities. Neither educational background nor the training session influenced the reliability of our results. The NINDS Myotatic Reflex Scale has sufficient reliability to be adopted as a universal scale.

Adult↗

Long-term reliability of endoscopic third ventriculostomy.

OBJECTIVE: To describe the short-term operative success and the long-term reliability of endoscopic third ventriculostomy (ETV) for treatment of hydrocephalus and to examine the influence of diagnosis, age, and previous shunt history on these outcomes. METHODS: We retrospectively analyzed 203 consecutive patients from a single institution who had ETV as long as 22.6 years earlier. Patients with hydrocephalus from aqueduct stenosis, myelomeningocele, tumors, arachnoid cysts, previous infection, or hemorrhage were included. RESULTS: The overall probability of successfully performing an ETV was 89% (84-93%). There was support for an association between the surgical success and the individual operating surgeon (odds ratios for success, 0.44-1.47 relative to the mean of 1.0, P = 0.08). We observed infections in 4.9%, transient major complications in 7.2%, and major and permanent complications in 1.1% of 203 procedures. Age was strongly associated with long-term reliability. The longest observed reliability for the 13 patients 0 to 1 month old was 3.5 years. The statistical model predicted the following reliability at 1 year after insertion: at 0 to 1 month of age, 31% (14-53%); at 1 to 6 months of age, 50% (32-68%); at 6 to 24 months of age, 71% (55-85%); and more than 24 months of age, 84% (79-89%). There was no support for an association between reliability and the diagnostic group (n = 181, P = 0.168) or a previous shunt. Sixteen patients had ETV repeated, but only 9 were repeated after at least 6 months. Of these, 4 procedures failed within a few weeks, and 2 patients were available for long-term follow-up. CONCLUSION: Age was the only factor statistically associated with the long-term reliability of ETV. Patients less than 6 months old had poor reliability.

Adolescent↗

Reliability of and correlation between the respiratory therapist written registry and clinical simulation self-assessment examinations.

STUDY OBJECTIVES: The purpose of this study was to determine the reliability of two respiratory therapy self-assessment examinations: the written registry examination (WR), and the clinical simulation examination (CSE). We then used reliability coefficients to test the true correlation between the WR and CSE by employing the Spearman-Brown formula to attenuate for unreliability. DESIGN: This was a nonexperimental correlational study. SETTING: The study was conducted at respiratory therapy education programs located in four states. PARTICIPANTS: Sixty advanced-level respiratory therapy students enrolled in the final semester of their programs. MEASUREMENTS AND RESULTS: Fifty-eight students completed the WR, and 56 students completed the CSE. The reliability coefficient for the WR was 0.79. The reliability coefficient for the CSE when taken as a whole was 0.76. However, the CSE is separated into two sections, information gathering and decision making, which are scored separately. Cronbach alpha computed for the information-gathering section was 0.72, while the alpha coefficient for the decision-making section was only 0.64. The correlation between the WR and CSE was 0.86 after attenuation for reliability. CONCLUSIONS: The estimate of the reliability for the CSE is less than that for the WR, and the two examinations are strongly correlated. This leads us to question whether the CSE adds to the validity or reliability in the testing of respiratory therapists.

Credentialing↗

The Collaborative Longitudinal Personality Disorders Study: reliability of axis I and II diagnoses.

Both the interrater and test-retest-retest reliability of axis I and axis II disorders were assessed using the Structured Clinical Interview for DSM-IV Axis I Disorders (SCID-I) and the Diagnostic Interview for DSM-IV Personality Disorders (DIPD-IV). Fair-good median interrater kappa (.40-.75) were found for all axis II disorders diagnosed five times or more, except antisocial personality disorder (1.0). All of the test-retest kappa for axis II disorders, except for narcissistic personality disorder (1.0) and paranoid personality disorder (.39), were also found to be fair-good. Interrater and test-retest dimensional reliability figures for axis II were generally higher than those for their categorical counterparts; most were in the excellent range (> .75). In terms of axis I, excellent median interrater kappa were found for six of the 10 disorders diagnosed five times or more, whereas fair-good median interrater kappa were found for the other four axis I disorders. In general, test-retest reliability figures for axis I disorders were somewhat lower than the interrater reliability figures. Three test-retest kappa were in the excellent range, six were in the fair-good range, and one (for dysthymia) was in the poor range (.35). Taken together, the results of this study suggest that both axis I and axis II disorders can be diagnosed reliably when using appropriate semistructured interviews. They also suggest that the reliability of axis II disorders is roughly equivalent to that reliability found for most axis I disorders.

Diagnosis, Differential↗

Validity and reliability of self-reported drinking behavior: dealing with the problem of response bias.

This work assesses the validity and reliability of self-reported survey data on drinking behavior. There is evidence to suggest that data are adversely affected by bias from underreporting. This bias affects the validity of measures of consumption of alcohol and can have deleterious effects on the results of some forms of statistical estimation. Data for this study were collected at an isolated military base. The remoteness of this site and the fact that it is a military station made it possible to estimate the actual level of consumption of alcohol for the population by assessing apparent consumption through officially recorded sales of alcohol. The results of eight measures of consumption of alcohol were compared with apparent consumption, as established by documented sales, and the validity and reliability of the various measures were determined using the classical correlational approach. The validity and reliability of the data generated by the self-report survey were also analyzed using LISREL, the measurement model in particular. The results indicate that various instruments used to assess the consumption of alcohol produce very different outcomes in terms of their validity and reliability, some questions being considerably more valid and reliable than others. Two of the more salient characteristics of questions that affect validity and reliability were isolated, namely a question's ability to aid recall and its ability to mitigate the effects of persons providing socially desirable responses. The LISREL results show that these are two underlying factors for the measurement of the consumption of alcohol. It is concluded that questions that produce valid and reliable responses do so for identifiable reasons, and measurement instruments can be improved by incorporating particular features.

Alcohol Drinking↗

Measuring the reliability of observational data: a reactive process.

Reliability of observational data was measured simultaneously by two assessors under two experimental conditions. During overt assessment, observers were told that reliability would be measured by one of the two assessors, thus permitting computation of reliability with an identified and an unidentified assessor. During covert assessment, observers were not informed of the reliability measured. Throughout the study, each of the assessors employed a unique version of a standard observational code. In the overt assessment condition, reliability of observers with the identified assessor was consistently higher than reliability with the unidentified assessor, indicating that observers modified their observational criteria to approximate those of the identified assessor. In the covert assessment condition, reliability with the two assessors was substantially lower than during overt assessment. Further, observers consistently recorded lower frequencies of disruptive behavior than the two assessors during covert assessment.

Journal Article↗

A graphical judgmental aid which summarizes obtained and chance reliability data and helps assess the believability of experimental effects.

Interval by interval reliability has been criticized for "inflating" observer agreement when target behavior rates are very low or very high. Scored interval reliability and its converse, unscored interval reliability, however, vary as target behavior rates vary when observer disagreement rates are constant. These problems, along with the existence of "chance" values of each reliability which also vary as a function of response rate, may cause researchers and consumers difficulty in interpreting observer agreement measures. Because each of these reliabilities essentially compares observer disagreements to a different base, it is suggested that the disagreement rate itself be the first measure of agreement examined, and its magnitude relative to occurrence and to nonoccurrence agreements then be considered. This is easily done via a graphic presentation of the disagreement range as a bandwidth around reported rates of target behavior. Such a graphic presentation summarizes all the information collected during reliability assessments and permits visual determination of each of the three reliabilities. In addition, graphing the "chance" disagreement range around the bandwidth permits easy determination of whether or not true observer agreement has likely been demonstrated. Finally, the limits of the disagreement bandwidth help assess the believability of claimed experimental effects: those leaving no overlap between disagreement ranges are probably believable, others are not.

Journal Article↗

Stulberg classification system for evaluation of Legg-Calvé-Perthes disease: intra-rater and inter-rater reliability.

BACKGROUND: Researchers and clinicians commonly use the classification system of Stulberg et al. as a basis for treatment decisions during the active phase of Legg-Calvé-Perthes disease because of its putative utility as a predictor of long-term outcome. It is generally assumed that this system has an acceptable degree of reliability. This assumption, however, is not convincingly supported by the literature. METHODS: The purpose of the present study was to assess the inter-rater and intra-rater reliability of the classification system of Stulberg et al. with use of a pre-test, post-test design. During the pre-test phase, nine raters independently used the system to evaluate the radiographs of skeletally mature patients who had been managed for Legg-Calvé-Perthes disease. The intervention between the pre-test and post-test phases consisted of a consensus-building session during which all raters jointly arrived at standardized definitions of the various joint structures that are assessed with use of the classification system. The effect of these definitions on reliability then was assessed by reevaluating the radiographs during the post-test phase. RESULTS: The pre-test intra-rater reliability coefficients ranged from 0.709 to 0.915, and the post-test coefficients ranged from 0.568 to 0.874. The pre-test inter-rater reliability coefficients ranged from 0.603 to 0.732, and the post-test coefficients ranged from 0.648 to 0.744. Contributing to the variance was a lack of agreement concerning the assessment of joint structures and the way in which the raters translated these evaluations into a classification according to the system of Stulberg et al. CONCLUSIONS: Although intra-rater reliability was marginally acceptable, the degree of variability between the classifications assigned by different raters even after the intervention - calls into question the reliability of the system of Stulberg et al.; consequently, the validity of any treatment decisions, outcome evaluations, or epidemiological studies based on this system is also in question.

Acetabulum↗

Reliability of three classification systems measuring active motion in brachial plexus birth palsy.

BACKGROUND: Several classification systems for the categorization of function in patients with brachial plexus birth palsy have been proposed. The purpose of this investigation was to determine the intraobserver and interobserver reliability of the modified Mallet Classification, Toronto Test Score, and Hospital for Sick Children Active Movement Scale in the evaluation of these patients. METHODS: Eighty children with brachial plexus birth palsy were evaluated by two trained examiners on two different occasions. Intraobserver and interobserver reliability was determined with use of the kappa statistic. RESULTS: On the basis of the kappa statistic, intraobserver reliability was good to excellent for individual elements of the modified Mallet Classification, Toronto Test Score, and Active Movement Scale in all age-groups. Interobserver reliability for individual elements of these three systems ranged from fair to excellent. When aggregate Toronto Test and modified Mallet scores were assessed, positive intraobserver and interobserver correlations were noted (Pearson r = 0.70 to 0.98, p < 0.001). Internal consistency (test-retest reliability) as determined by the Cronbach alpha for the aggregate Toronto Test and modified Mallet scores was excellent for each age-group (alpha > 0.90, p < 0.001). CONCLUSIONS: The modified Mallet Classification, Toronto Test Score, and Active Movement Scale are reliable instruments for assessing upper-extremity function in patients with brachial plexus birth palsy. The natural history and surgical outcomes of these patients can now be conducted with use of these reliable outcomes instruments.

Adolescent↗

Two measurement techniques for assessing subtalar joint position: a reliability study.

Proper assessment of the subtalar joint is critical for foot and ankle evaluation. Yet, reliability of open kinetic chain goniometric measurements of the subtalar joint has been poor. Two alternative techniques, navicular height and calcaneal position with an inclinometer, have been reported in the literature but lack reliability assessment. The purpose of this study was to determine the intertester and intratester reliability of navicular height and calcaneal position using an inclinometer. Thirty healthy, volunteer subjects (22 females, age 24 +/- 3.6 years; eight males, age 25 +/- 5.1 years) participated in this study. Two testers performed repeated measures on both feet of each subject (N = 60) during two testing sessions. Testers determined the 1) subtalar neutral position, 2) resting position, and 3) difference between these two measurements using an inclinometer for calcaneal position and navicular height. Intratester and intertester reliabilities (ICC 2, 1), standard errors of measurement, and 95% confidence intervals were determined. Intertester and intratester reliability for calcaneal position ranged from .68 to .91 for all measurements. Intertester and intratester reliability for navicular height ranged from .73 to .96 for all measurements. We conclude that these weight-bearing measurement techniques are reliable and acceptable for clinical and research purposes as measured. In addition, we hypothesize that these measurement techniques are simpler than previously described open kinetic chain methods.

Adult↗

The influence of experience on the reliability of goniometric and visual measurement of forefoot position.

Goniometric measurement of forefoot position relative to the rearfoot is a routine procedure used by rehabilitation specialists. This measurement is also frequently made by visual estimation. The influence of tester experience on the reliability of these two techniques at the forefoot is unknown. The purpose of this investigation was to directly examine the reliability of goniometric and visual estimation of forefoot position measurements when experienced and inexperienced testers perform the evaluation. Two clinicians (> or = 10 years experience) and two physical therapy students were recruited as testers. Ten subjects (20-31 years old), free from pathology, were measured. Each foot was evaluated twice with the goniometer and twice with visual estimation by each tester. Intraclass correlation coefficient (ICC) and coefficients of variation method error were used as estimates of reliability. There was no dramatic difference in the intratester or intertester reliability between experienced and inexperienced testers, regardless of the evaluation used. Estimates of intratester reliability (ICC 2,1), when using the goniometer, ranged from 0.08 to 0.78 for the experienced examiners and from 0.16 to 0.65 for the inexperienced examiners. When using visual estimation, ICC (2,1) values ranged from 0.51 to 0.76 for the experienced examiners and 0.53 to 0.57 for the inexperienced examiners. The estimate of intertester reliability [ICC (2,2)] for the goniometer was 0.38 for the experienced examiners and 0.42 for the inexperienced examiners. When using visual estimation, ICC (2,2) values were 0.81 for the experienced examiners and 0.72 for the inexperienced examiners. Although experience does not appear to influence forefoot position measurements, of the two evaluation techniques, visual estimation may be the more reliable.

Adult↗

Reliability of measuring active mandibular excursion using a new tool: the Mandibular Excursiometer.

Measurement tools improve the reliability and validity of measurement. The purpose of this study was to test the intrarater and interrater reliability of a new instrument, the Mandibular Excursiometer, for measuring mandibular excursion on the X and Y axis in the coronal plane during active opening. Two raters measured 12 volunteers. Four ratio, three nominal, and one ordinal scale measurements were analyzed using percent agreement. The Mandibular Excursiometer had high intrarater reliability for vertical opening (100%) and for the categorization of the presence or absence and direction of lateral deviation at the maximum point during opening (92-100%). Overall, moderate intrarater reliability existed for the quantity of lateral deviation at the maximum point during opening (66-83%), presence and direction of deflection (66-83%), presence of deviation or deflection during opening (66-83%), and in which third of opening the maximum point of lateral deviation occurred (66-83%). Moderate interrater reliability existed for vertical opening (75%) and for the classification of presence and direction of lateral deviation at the maximum point during opening (91%). All other measurements had low reliability. The Mandibular Excursiometer had higher intrarater and interrater reliability for measuring deviation and deflection during active mandibular opening than observation alone, based on a comparison with the literature. This measurement can assist in documenting progress while treating patients with TMJ disorders.

Adult↗