Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Inter-rater and intra-rater reliability in the interpretation of MTI Photoscreener photographs of Native American preschool children.

PURPOSE: To evaluate inter- and intra-rater reliability for the interpretation of MTI Photoscreener photographs taken in a population of Native American preschool children with a high prevalence of astigmatism. METHODS: Photographs of 369 children were rated by 11 nonexpert and 3 expert raters. Photographs for each child were scored as pass, refer, or retake. Nonexpert raters scored photos on two separate occasions, permitting analysis of intra-rater reliability. RESULTS: Analyses of pass/refer responses only: inter-rater reliability was moderate to substantial among nonexpert raters and substantial among expert raters. Intra-rater reliability among nonexperts was substantial. Analyses of all responses (pass, refer, and retake): inter-rater reliability for pass and refer scores was moderate among nonexperts and substantial among experts; for retake scores inter-rater reliability was slight for nonexperts and moderate for experts. Intra-rater reliability among nonexperts was substantial for pass and refer scores and moderate for retake scores. CONCLUSIONS: In this population with a high prevalence of astigmatism, whether MTI photoscreening results are interpretable is much more variable among and within raters than whether an interpretable photograph should be scored as pass or refer. The level of agreement among raters in the current study was influenced by the experience of the raters. In addition, nonexpert raters were more likely to deem a photograph uninterpretable than expert raters.

Arizona↗

Interexaminer reliability in physical examination of patients with low back pain.

STUDY DESIGN: Seventy-one patients with low back pain were examined by two physiotherapists (50 patients) and two physicians (21 patients). The two physiotherapists had worked together for many years, but the two physicians had not. The interexaminer reliability of the clinical tests included in the physical examination was evaluated. OBJECTIVES: To evaluate the interexaminer reliability of clinical tests used in the physical examination of patients with low back pain under ideal circumstances, which was the case for the physiotherapists. SUMMARY OF BACKGROUND DATA: Numerous clinical tests are used in the evaluation of patients with low back pain. To reach the correct diagnosis, only tests with an acceptable validity and reliability should be used. Previous studies have mainly shown low reliability. It is important that clinical tests not be rejected because of low reliability caused by differences between examiners in performance of the examination and in their definition of normal results. METHODS: Two examiners, either two physiotherapists or two physicians, independently examined patients with low back pain. RESULTS: In approximately half of the clinical tests studied, an acceptable reliability was demonstrated. CONCLUSION: On the basis of the physiotherapists series, the reliability was acceptable for a number of clinical tests that are used in the evaluation of patients with low back pain. The results suggest that clinical tests should be standardized to a much higher degree than they are today.

Adolescent↗

Apophysial joint degeneration, disc degeneration, and sagittal curve of the cervical spine. Can they be measured reliably on radiographs?

STUDY DESIGN: Interexaminer reliability study. OBJECTIVES: To determine the reliability of grading apophysial joint and disc degenerative changes and the reliability of measuring sagittal curves on lateral cervical spine radiographs. SUMMARY OF BACKGROUND DATA: Several authors have proposed that the presented of degenerative changes and the absence of lordosis in the cervical spine are indicators of poor recovery from neck injuries caused by motor vehicle collisions. The validity of those conclusions is questionable because the reliability of the methods used in their studies to measure the presence of degenerative changes and the absence of lordosis has not been determined. METHODS: Kellgren's classification system for apophysial joint and disc degeneration, as well as the pattern and magnitude of the sagittal curve on 30 lateral cervical spine radiographs were assessed independently by three examiners. RESULTS: Moderate reliability was demonstrated for classifying apophysial joint degeneration with an intraclass correlation coefficient of 0.45 (95% confidence interval, 0.09-0.71). Classifying degenerative disc disease had substantial reliability, with an intraclass correlation coefficient of 0.71 (95% confidence interval, 0.23-0.88). Measuring the magnitude of the sagittal curve from C2 to C7 had excellent interexaminer agreement, with an intraclass correlation coefficient of 0.96 (95% confidence interval, 0.88-0.98) and an interexaminer error of 8.3 degrees. CONCLUSIONS: The classification system for degenerative disc disease proposed by Kellgren et al and the method of measurement of sagittal curves from C2 to C7 demonstrated an acceptable level of reliability and can be used in outcomes research.

Analysis of Variance↗

Multisurgeon assessment of coronal pattern classification systems for adolescent idiopathic scoliosis: reliability and error analysis.

STUDY DESIGN: Three scoliosis surgeons and one orthopedic fellow were presented the anteroposterior radiographs of 70 patients with adolescent idiopathic scoliosis. All the reviewers assigned a type to each curve according to the classification systems of H. A. King and R. W. Coonrad. OBJECTIVES: To compare multisurgeon reliability in applying the classification systems of H. A. King and R. W. Coonrad, and to analyze controversially classified curve patterns. SUMMARY OF BACKGROUND DATA: The system most commonly used to classify adolescent idiopathic scoliosis is King's classification. However, because of poor interobserver reliability, the validity of this system is questioned. In contrast, high interobserver reliability is reported for Coonrad's classification system, which is used less frequently in clinical practice. METHODS: Interobserver agreement and intraobserver reproducibility were tested. Kappa coefficients were used to test reliability. Between the observers, the divergent assignments to curve patterns were analyzed in both quantitative and qualitative terms. An error analysis was performed. RESULTS: Paired comparisons showed a mean interobserver kappa coefficient of 0.45 for King's and 0.38 for Coonrad's classification systems. According to Svanholm et al, these values indicate poor reliability in terms of interobserver agreement. Error analyses for both classification systems showed that the reason for poor reproducibility is disagreement among the observers about structural upper thoracic and structural lumbar curves. CONCLUSIONS: Neither the King nor the Coonrad method appears to have sufficient interobserver reliability. To improve reliability, the authors recommend that the structural stigmas of the upper thoracic and lumbar curves be unequivocally described.

Adolescent↗

Determining the reliability of the Graf classification for hip dysplasia.

We sought to establish the levels of interrater reliability and intrarater reliability of the Graf classification among orthopaedic surgeons in their final training year and who learned the method by instructed teaching or self study. Using standard teaching material developed by Graf, two groups of senior orthopaedic residents at the same training level received structured teaching sessions (Group A, n = 2) or performed self study (Group B, n = 2). Interrater reliability and intrarater reliability were determined (Cohen's weighted kappa). Proportions of correctly rated sonograms were compared between groups, implications of misclassifications were analyzed, and sensitivity analyses were performed. Interrater reliability was 0.59 (95% CI = 0.32-0.85) for Group A, and 0.47 (95% CI = 0.14-0.79) for Group B. Intrarater reliability showed an overall kappa of 0.57 (95% CI = 0.35-0.78) in Group A, and 0.47 (95% CI = 0.19-0.75) in Group B. The proportion of correctly rated sonograms between groups was similar in the original dataset and in the sensitivity analysis. Misclassifications influencing treatment were infrequent; one patient would have received unwarranted treatment and three patients would not have received warranted treatment. The Graf classification showed moderate reliability. Using self study, it can be learned almost, but not quite as effectively as by a structured program.

Clinical Competence↗

Reliability of the visual assessment of cervical and lumbar lordosis: how good are we?

STUDY DESIGN: Blinded test-retest design. OBJECTIVE: To measure the intrarater and interrater reliability of the visual assessment of cervical and lumbar lordosis. SUMMARY OF BACKGROUND DATA: Cervical and lumbar lordoses are frequently evaluated using visual assessment, but little attempt has previously been made to measure the reliability of visual assessment. METHODS: Twenty-eight chiropractors, physical therapists, physiatrists, rheumatologists, and orthopedic surgeons were recruited to evaluate the posture of photographed subjects (with and without back pain). Each clinician rated the lordosis of the cervical and lumbar spines as normal, increased, or decreased. Kappa coefficients (kappa) were calculated to determine intrarater and interrater reliability. RESULTS: Twenty-eight clinicians evaluated photographs of 36 individuals (17 with back pain, 19 without). Mean intrarater reliability was kappa = 0.50 (95% confidence interval 0.02-0.98) and mean interrater reliability was kappa = 0.16 (95% confidence interval 0.00-0.48). No statistically significant difference existed among the five groups of clinicians or between the evaluation of the subjects with and without back pain. CONCLUSION: Intrarater reliability of the visual assessment of cervical and lumbar lordosis was statistically fair, whereas interrater reliability was poor.

Adolescent↗

Reliability of clinical tests in the assessment of patients with neck/shoulder problems-impact of history.

STUDY DESIGN: A clinical trial on patients receiving neck/shoulder physical examinations. OBJECTIVES: To analyze reliability of clinical tests, prevalence of positive findings in the assessment of neck/shoulder problems in primary care patients, and the impact of history, including pain drawing, on these parameters. SUMMARY OF BACKGROUND DATA: Reliability of clinical tests varies, perhaps partly because of the impact of history. To our knowledge, this has not been studied before. METHODS: Two examiners independently assessed 100 patients with a set of 66 clinical tests divided into 9 categories. Half of the patients were examined with and the other half without knowledge of history. Reliability as expressed by percentage agreement, kappa coefficients, and prevalence of positive findings was calculated. RESULTS: Reliability of clinical tests was poor or fair in several categories and did not alter with history. Only a bimanual sensitivity test reached good kappa values. With known history, prevalence of positive findings increased. Bias was apparent in all test categories except sensitivity tests. Four out of five patients were diagnosed to have neurogenic dysfunction in the affected area. CONCLUSIONS: Our sensitivity test was the most reliable and also exempt from bias and should be studied further. Some common tests may not be reliable. History had no impact on reliability of our tests but increased the prevalence of positive findings. Neurogenic dysfunction seems very common in patients with neck and/or shoulder problems and should be screened for.

Adolescent↗

Interobserver and intraobserver reliability in the load sharing classification of the assessment of thoracolumbar burst fractures.

STUDY DESIGN: The Load Sharing Classification of spinal fractures was evaluated by 5 observers on 2 occasions. OBJECTIVE: To evaluate the interobserver and intraobserver reliability of the Load Sharing Classification of spinal fractures in the assessment of thoracolumbar burst fractures. SUMMARY OF BACKGROUND DATA: The Load Sharing Classification of spinal fractures provides a basis for the choice of operative approaches, but the reliability of this classification system has not been established. METHODS: The radiographic and computed tomography scan images of 45 consecutive patients with thoracolumbar burst fractures were reviewed by 5 observers on 2 different occasions 3 months apart. Interobserver reliability was assessed by comparison of the fracture classifications determined by the 5 observers. Intraobserver reliability was evaluated by comparison of the classifications determined by each observer on the first and second sessions. Ten paired interobserver and 5 intraobserver comparisons were then analyzed with use of kappa statistics. RESULTS: All 5 observers agreed on the final classification for 58% and 73% of the fractures on the first and second assessments, respectively. The average kappa coefficient for the 10 paired comparisons among the 5 observers was 0.79 (range 0.73-0.89) for the first assessment and 0.84 (range 0.81-0.95) for the second assessment. Interobserver agreement improved when the 3 components of the classification system were analyzed separately, reaching an almost perfect interobserver reliability with the average kappa values of 0.90 (range 0.82-0.97) for the first assessment and 0.92 (range 0.83-1) for the second assessment. The kappa values for the 5 intraobserver comparisons ranged from 0.73 to 0.87 (average 0.78), expressing at least substantial agreement; 2 observers showed almost perfect intraobserver reliability. For the 3 components of the classification system, all observers reached almost perfect intraobserver agreement with the kappa values of 0.83 to 0.97 (average, 0.89). CONCLUSIONS: Kappa statistics showed high levels of agreement when the Load Sharing Classification was used to assess thoracolumbar burst fractures. This system can be applied with excellent reliability.

Humans↗

Measurement of fracture kyphosis with the Oxford Cobbometer: intra- and interobserver reliabilities and comparison with other techniques.

STUDY DESIGN: Statistical analysis of 3 techniques for measuring thoracolumbar kyphosis secondary to fracture. OBJECTIVES: To determine the reliability of using an Oxford Cobbometer and assess the most reliable measurement technique. SUMMARY OF BACKGROUND DATA: The reproducibility of Cobb angles for the assessment of saggital plane deformity on spine radiographs has been shown to have significant variability in both intra- and interobserver error. METHODS: Twenty-four lateral spine radiographs of patients with thoracic and lumbar vertebral fractures were measured on 2 separate occasions, in random order, by 4 blinded observers using the same Oxford Cobbometer and ruler. RESULTS: Method 2, the angle from the inferior endplate of the vertebra above the fractured vertebra to the superior endplate of the vertebra below the fractured vertebra, had the greatest intraobserver and interobserver reliabilities (rho = 0.856-0.976 and rho = 0.95, respectively). The other 2 methods had lower reliabilities; however, all 3 methods were well above the statistically acceptable threshold of >0.8, and the intraobserver reliabilities with each observer was 99% overall. These reliabilities supersede results reported previously using the conventional Cobb technique. The absolute mean difference between readings and 95% limit of agreement also improves on previous data, 2 degrees and +/- 5.8 degrees , respectively. CONCLUSIONS: Highest intraclass correlation coefficients were obtained using method 2. Using the Oxford Cobbometer to measure fracture kyphosis has higher reliability than the standard Cobb angle technique. It is easy and quick to use in a clinical setting.

Diagnostic Techniques and Procedures↗

Reliability of end, neutral, and stable vertebrae identification in adolescent idiopathic scoliosis.

STUDY DESIGN: Analysis of radiographic interpretation and vertebral level identification. OBJECTIVES: To assess the intra- and interobserver reliability by observer training level used for selecting the end vertebra (EV), neutral vertebra (NV), and stable vertebra (SV) in adolescent idiopathic scoliosis patients. SUMMARY OF BACKGROUND DATA: Various radiographic and clinical factors are important in surgical planning. For adolescent idiopathic scoliosis, an analysis of the end, neutral, and stable vertebrae are of paramount importance for understanding spinal deformity management and determining the distal fusion level. Additionally, the development and comparison of optimal surgical techniques requires reliable, reproducible radiographic parameters. METHODS: One hundred consecutive radiographs of operative cases of adolescent idiopathic scoliosis were evaluated on three separate occasions by three surgeons (2700 data points) at various levels of training (fellowship-trained spine surgeon, fellow in-training, orthopedic surgery resident). For each iteration, the observers attempted to identify the distal structural Cobb curve EV, NV, and SV. The radiographs included preselected Lenke type 1, 3, and 5 curves in random order. The average main thoracic curve was 53 degrees (range, 30-82 degrees) with a T8-T9 average apex, whereas the average thoracolumbar curve was 33 degrees (range, 18-65 degrees). Intra- and interobserver reliability was assessed by means of Cohen's Kappa correlation coefficient, and raw percentages of agreement were recorded. RESULTS: Intraobserver reliability was good to excellent for determining the EV (kappaa = 0.69-0.88), good for determining the NV (kappaa = 0.65-0.73), and good to excellent for determining the SV (kappaa = 0.74-0.91) with 83.5, 72.2, and 85.6% intraobserver agreement, respectively. A trend was noted towards greater intraobserver reliability with increasing levels of observer experience. Interobserver reliability was poor (kappaa = 0.26-0.39) for each vertebral level, with interobserver agreement for only 48.7% of EV, 41.7% of NV, and 51.0% of SV. However, interobserver agreement increased significantly when concurrence within one vertebral level was assessed, with 91, 73, and 76% agreement for identifying the EV, NV, and SV, respectively. CONCLUSIONS: Radiographic determination of the EV, NV, and SV demonstrated good to excellent intraobserver, but poor interobserver, reliability. Interobserver agreement was fair to good when concurrence within one adjacent level was assessed. Observer experience level may be a factor. The difficulties in identifying these vertebral levels represent a potential obstacle to reproducible patient-specific fusion level determination and to the optimization and uniformity of patient care.

Adolescent↗

Interobserver and intraobserver reliability of maximum canal compromise and spinal cord compression for evaluation of acute traumatic cervical spinal cord injury.

STUDY DESIGN: Prospective, blinded validation study of an objective, quantitative measure to assess maximum canal compromise (MCC) and maximum spinal cord compression (MSCC) in individuals with acute cervical spinal cord injury (SCI). OBJECTIVE: To examine the intraobserver and interobserver reliability of MCC and MSCC in individuals with acute traumatic cervical SCI. SUMMARY OF BACKGROUND DATA: To date, few quantitative reliable radiologic methods for assessing the extent of spinal cord compression in the setting of acute SCI have been reported. MCC and MSCC, as assessed on mid-sagittal CT and T2-weighted MR images, respectively, appear to have potential clinical and prognostic value. To date, the validation of these assessment tools has been limited to a small number of observers at a single institution. However, to date no study has focused on the reliability of these radiologic parameters among a large cohort of spine surgeons from North America and abroad. This type of validation is critical to allow the broader use of these outcome measures in research studies and in clinical practice. METHODS: Mid-sagittal MRI and CT images of cervical spine were selected from 10 individuals with acute traumatic cervical SCI. A total of 28 spine surgeons independently estimated CT MCC, T1-weighted MRI MCC, and T2-weighted MRI MSCC on two occasions using a calibrated ruler. In the first round of measurements, the observers estimated the radiologic parameters using only written instructions. The second measurement set was obtained after an interactive teaching session on the methodology. The order of the images was altered for the second set of measurements. RESULTS: Analysis using parametric and nonparametric statistics indicated high intraobserver reliability for CT MCC, T1-weighted MRI MCC, and T2-weighted MSCC with interclass correlation coefficients (ICCs) of 0.92, 0.95, and 0.97, respectively. The interobserver reliability for all three radiologic parameters was considered moderate with ICCs ranging from 0.35 to 0.56. CONCLUSION: Our results indicate that the intraobserver reliability for the MCC and MSCC was high. Although the interobserver reliability for all three radiologic parameters in the present study was below 0.75, the observed differences were small and largely accounted for by the limitations in the precision of the calibrated ruler. For cases with minimal cord compression, the measurement of canal stenosis (MCC) proved more accurate. In contrast, in cases with severe cord compression, the assessment of MSCC was more accurate. It is anticipated that the use of digital imaging technologies will further enhance the precision of these outcome measures.

Acute Disease↗

Intraobserver and interobserver reliability indices for drawing scanning laser ophthalmoscope optic disc contour lines with and without the aid of optic disc photographs.

PURPOSE: To investigate the effects on topographic optic disc analysis of defining regions of interest (ROIs) by drawing contour lines with and without photographic aid. PATIENTS AND METHODS: Forty-five patients had optic disc imaging by stereoscopic photography and confocal scanning laser ophthalmoscopy using the Heidelberg Retinal Tomograph (HRT). Two experienced observers defined ROIs with a stereoscopic optic disc photograph, with a non-stereoscopic photograph and without any photographic guide. Intraclass coefficients and 95% tolerance limits for change were calculated for each HRT optic disc parameter for the following situations: (1) Intraobserver reliability for ROIs defined with non-stereoscopic photographs and no photograph compared against ROIs defined using stereoscopic photographs; (2) interobserver reliability for ROIs defined without the aid of a photograph, with stereoscopic photographs and with non-stereoscopic photographs; and (3) intraobserver reliability for ROIs defined twice for 23 patients using non-stereoscopic photographs. RESULTS: Intraclass correlation coefficients ranged from 0.63 to 1.00 (ie, 'substantial' to 'perfect' agreement). The 95% tolerance limits for change ranged from 3% to 34%. In general, intraobserver reliability indices were higher than interobserver reliability indices but no method of ROI definition appeared superior. The least reliable measurements were for rim area, rim volume, and disc area. CONCLUSIONS: Non-stereoscopic optic disc photographs are an acceptable alternative to stereoscopic photographs when defining ROIs for HRT analysis. Observers with impaired stereopsis may be able to define reliable ROIs. For HRT algorithms designed to determine a 'diagnosis' of glaucoma, rim area, rim volume, and disc area data should be used with caution.

Adult↗

Vastus medialis H-reflex reliability during standing.

Vastus medialis H-reflex is a valid measure to examine quadriceps muscle voluntary activation and inhibition after knee injury. Its reliability during repeated sessions has not been established. The purpose of this study was to establish the intrasession and intersession reliability of vastus medialis H-reflex amplitude recordings during standing with varied knee flexion angles (0, 30, 45, and 60 degrees). Electromyography unit was used to elicit and record the vastus medialis H-reflex from the right leg of five healthy subjects. The femoral nerve was stimulated using 0.5-millisecond pulses at 0.2 pps of H-maximum. Four recordings of the vastus medialis H-reflex amplitude were recorded in three trials for each knee flexion angle within each session for two consecutive days. Reliability was calculated using intraclass correlation coefficients (ICC). Intrasession reliability during standing with varied knee angles was high (ICC [2, 4] range from 0.76 to 0.98), and intersession reliability during standing with varied knee angles was moderate to high (ICC [2, 1] range from 0.51 to 0.84). Recording four traces of vastus medialis H-reflex amplitude per trial was reliable. Vastus medialis H-reflex amplitude recordings while standing during varied knee flexion are reliable within and between sessions.

Adult↗

Reliability of the ICD-10 classification of adverse familial and environmental factors.

BACKGROUND: The tenth revision of the International Classification of Diseases (ICD-10) contains a number of categories and guidelines for coding different types of adverse familial and environmental situations. These categories have been selected on the basis of empirical evidence that the adverse situations might represent important psychiatric risk factors. Prior studies, using written case vignettes or video- or audiotaped semi-structured interviews, showed that the interrater reliability of these classifications is satisfactory. METHOD: We have tested the interrater reliability of the categories in a multicentre study by using an 'in vivo' design. Classifications were performed in day-to-day practice, using information that is normally available. Three hundred children from 0 to 13 years were involved in the study. These cases had been admitted to various institutes for children with psychiatric disorders, developmental delays and/or adverse psychosocial circumstances. Two clinicians and a student classified the psychosocial situation of each child. In total, 51 clinicians and 6 undergraduate students were involved in the study. RESULTS: It was found that, with the exception of more or less objective categories, the reliability is not satisfactory. Post hoc analyses showed that the insufficient reliability is only partly due to confounding factors: no clear indications were found for information variance, but observation variance may have played a part. Having additional information from two family questionnaires hardly contributes to a better reliability. CONCLUSIONS: The reliability of the psychosocial axis of the ICD-10 is not satisfactory if tested in day-to-day practice. In comparing this result with other reliability studies, it seems that the absence of adequate information to code the psychosocial axis may be a fundamental problem in obtaining sufficient interrater agreement. Apparently, the classification of this axis requires information that is often not available in common practice.

Adolescent↗

Reliability of the clinical teaching effectiveness instrument.

INTRODUCTION: The Clinical Teaching Effectiveness Instrument (CTEI) was developed to evaluate the quality of the clinical teaching of educators. Its authors reported evidence supporting content and criterion validity and found favourable reliability findings. We tested the validity and reliability of this instrument in a European context and investigated its reliability as an instrument to evaluate the quality of clinical teaching at group level rather than at the level of the individual teacher. METHODS: Students participating in a surgical clerkship were asked to fill in a questionnaire reflecting a student-teacher encounter with a staff member or a resident. We calculated variance components using the urgenova program. For individual score interpretation of the quality of clinical teaching the standard error of estimate was calculated. For group interpretation we calculated the root mean square error. RESULTS: The results did not differ statistically between staff and residents. The average score was 3.42. The largest variance component was associated with rater variance. For individual score interpretation a reliability of > 0.80 was reached with 7 ratings or more. To reach reliable outcomes at group level, 15 educators or more were needed with a single rater per educator. DISCUSSION: The required sample size for appraisal of individual teaching is easily achievable. Reliable findings can also be obtained at group level with a feasible sample size. The results provide additional evidence of the reliability of the CTEI in undergraduate medical education in a European setting. The results also showed that the instrument can be used to measure the quality of teaching at group level.

Clinical Clerkship↗

Conditional reliability of admissions interview ratings: extreme ratings are the most informative.

CONTEXT: Admissions interviews are unreliable and have poor predictive validity, yet are the sole measures of non-cognitive skills used by most medical school admissions departments. The low reliability may be due in part to variation in conditional reliability across the rating scale. OBJECTIVES: To describe an empirically derived estimate of conditional reliability and use it to improve the predictive validity of interview ratings. METHODS: A set of medical school interview ratings was compared to a Monte Carlo simulated set to estimate conditional reliability controlling for range restriction, response scale bias and other artefacts. This estimate was used as a weighting function to improve the predictive validity of a second set of interview ratings for predicting non-cognitive measures (USMLE Step II residuals from Step I scores). RESULTS: Compared with the simulated set, both observed sets showed more reliability at low and high rating levels than at moderate levels. Raw interview scores did not predict USMLE Step II scores after controlling for Step I performance (additional r2 = 0.001, not significant). Weighting interview ratings by estimated conditional reliability improved predictive validity (additional r2 = 0.121, P < 0.01). CONCLUSIONS: Conditional reliability is important for understanding the psychometric properties of subjective rating scales. Weighting these measures during the admissions process would improve admissions decisions.

Education, Medical, Undergraduate↗

The reliability and validity of three dimensional ultrasound volumetric measurements using an in vitro balloon and in vivo uterine model.

OBJECTIVE: To evaluate the reliability and validity of two and three dimensional ultrasound volumetric measurements using balloon and uterine models. DESIGN: Prospetive observational study. SETTING: Obstetric ultrasound department at a university teaching hospital. METHOD: Two and three dimensional ultrasound volumetric measurements (with 5, 10 and 15 ultrasonic slices) were performed on 30 different sets of ultrasound images obtained from 15 water filled balloons with volumes ranging from 19 to 697mL. The measurements were performed independently by two observers who were blinded to the true volumes of the balloons. For the uterine model, only three dimensional ultrasonic volume measurements were performed independently on 16 uteri by two observers who were again unaware of the definitive uterine volumes. OUTCOME MEASURE: For the assessment of intra-and inter-rater reliability, the intraclass correlation coefficient was used. The index of concordance between the ultrasonic volumes and those obtained by the reference standard (validity) was assessed with the conventional Pearson's correlation coefficient, limits of agreement method and the intra-class correlation coefficient. RESULTS: High levels of reliability and validity were obtained for both two and three dimensional ultrasound balloon volume measurements. For two dimensional ultrasonic volume measurements, the intra-class correlation coefficient ranged from 0.992 to 0.998 for reliability and validity whereas the Pearson's correlation coefficient for validity was 0.996. With three dimensional ultrasonic volume measurements, the intra-class correlation coefficient ranged from 0.991 to 0.999 for reliability and validity whereas the Pearson's correlation coefficient for validity was 0.999. Both two and three dimensional ultrasonic measurements tended to underestimate the true balloon volume with the largest observed mean difference obtained with three dimensional ultrasound measurements using five ultrasonic slices and the smallest value obtained with three dimensional ultrasound measurements employing 15 ultrasonic slices. The mean difference in volume measurement for two dimensional ultrasound was intermediate between these two values. However, two dimensional ultrasound volume measurement generated the largest range between the limits of agreement whereas the smallest range was obtained with three dimensional ultrasound using 10 ultrasonic slices. The intra-class correlation coefficient for reliability and validity with three dimensional ultrasonic uterine volume estimation ranged from 0.956 to 0.996 whereas the Pearson's correlation coefficient for validity ranged from 0.993 to 0.999). The use of three dimensional ultrasound also consistently under-estimated the actual uterine volumes. The larger the number of ultrasonic slices employed for three dimensional ultrasound, the smaller was the mean difference between the ultrasonic and true uterine volume measurements and the smaller the limits of agreement. CONCLUSIONS: The reliability and validity of balloon and uterine volume measurement by three dimensional ultrasound is high. This allows further research on three dimensional ultrasound for measuring pelvic organ volumes in the prediction of pelvic pathology.

Female↗

Methods of failure and reliability assessment for mechanical heart pumps.

Artificial blood pumps are today's most promising bridge-to-recovery (BTR), bridge-to-transplant (BTT), and destination therapy solutions for patients suffering from intractable congestive heart failure (CHF). Due to an increased need for effective, reliable, and safe long-term artificial blood pumps, each new design must undergo failure and reliability testing, an important step prior to approval from the United States Food and Drug Administration (FDA), for clinical testing and commercial use. The FDA has established no specific standards or protocols for these testing procedures and there are only limited recommendations provided by the scientific community when testing an overall blood pump system and individual system components. Product development of any medical device must follow a systematic and logical approach. As the most critical aspects of the design phase, failure and reliability assessments aid in the successful evaluation and preparation of medical devices prior to clinical application. The extent of testing, associated costs, and lengthy time durations to execute these experiments justify the need for an early evaluation of failure and reliability. During the design stages of blood pump development, a failure modes and effects analysis (FMEA) should be completed to provide a concise evaluation of the occurrence and frequency of failures and their effects on the overall support system. Following this analysis, testing of any pump typically involves four sequential processes: performance and reliability testing in simple hydraulic or mock circulatory loops, acute and chronic animal experiments, human error analysis, and ultimately, clinical testing. This article presents recommendations for failure and reliability testing based on the National Institutes of Health (NIH), Society for Thoracic Surgeons (STS) and American Society for Artificial Internal Organs (ASAIO), American National Standards Institute (ANSI), the Association for Advancement of Medical Instrumentation (AAMI), and the Bethesda Conference. It further discusses studies that evaluate the failure, reliability, and safety of artificial blood pumps including in vitro and in vivo testing. A descriptive summary of mechanical and human error studies and methods of artificial blood pumps is detailed.

Animals↗