Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Feasibility and reliability of an in-training assessment programme in an undergraduate clerkship.

INTRODUCTION: Structured assessment, embedded in a training programme, with systematic observation, feedback and appropriate documentation may improve the reliability of clinical assessment. This type of assessment format is referred to as in-training assessment (ITA). The feasibility and reliability of an ITA programme in an internal medicine clerkship were evaluated. The programme comprised 4 ward-based test formats and 1 outpatient clinic-based test format. Of the 4 ward-based test formats, 3 were single-sample tests, consisting of 1 student-patient encounter, 1 critical appraisal session and 1 case presentation. The other ward-based test and the outpatient-based test were multiple sample tests, consisting of 12 ward-based case write-ups and 4 long cases in the outpatient clinic. In all the ITA programme consisted of 19 assessments. METHODS: During 41 months, data were collected from 119 clerks. Feasibility was defined as over two thirds of the students obtaining 19 assessments. Reliability was estimated by performing generalisability analyses with 19 assessments as items and 5 test formats as items. RESULTS: A total of 73 students (69%) completed 19 assessments. Reliability expressed by the generalisability coefficients was 0.81 for 19 assessments and 0.55 for 5 test formats. CONCLUSIONS: The ITA programme proved to be feasible. Feasibility may be improved by scheduling protected time for assessment for both students and staff. Reliability may be improved by more frequent use of some of the test formats.

Clinical Clerkship↗

The reliability and validity of a matrix to assess the completed reflective personal development plans of general practitioners.

INTRODUCTION: We wished to determine whether assessors could make reliable and valid judgements about the quality of completed reflective personal development plans (PDPs) for the purpose of accrediting UK general practitioners (GPs) for a postgraduate education allowance using a marking matrix, and secondly, to plan a feasible model of PDP assessment in the context of forthcoming GP appraisal/revalidation that would overcome the main sources of error identified from this study. METHODS: Within generalisability theory, a variance components analysis on PDP scores estimated reliability and the effect on them of varying, for example, the number of assessors. We investigated the construct validity of the matrix through its internal consistency and detection of differences in the quality of PDPs. RESULTS: For a single PDP and one assessor, 37.6% of the variance in scores was due to true differences in the quality of the PDP. Between 5 and 7 PDP assessors are needed to achieve summative reliability of greater than 0.8. While increasing the number of judges is important, reliability could also be improved by addressing assessor subjectivity. Construct validity was demonstrated, as the matrix distinguished between good, satisfactory and poor PDPs, and it had good internal consistency. CONCLUSION: PDP assessment has reasonable summative characteristics for the purpose of assessing GPs' reflective continuing professional development. If doctors could include their PDPs within their revalidation folders as evidence of their reflections on pursuing better clinical performance, we have described a reliable, valid and feasible method of external assessment.

Accreditation↗

Reliability of the functional independence measure for children in normal Thai children.

BACKGROUND: The Functional Independence Measure for children (WeeFIM) is a new instrument for evaluating functionality in disabled children aged 9-100 months. It was developed to determine a child's functional capacity and performance. With no baseline information about Thai children, it is difficult to assess whether a patient is initially high or low with respect to function. METHODS: The aim of this study, therefore, was to evaluate the interrater, intrarater reliability and appropriateness of the use of the WeeFIM and to establish a normative data profile suitable for Thai children. The WeeFIM is an instrument used to assess independence in self-care, sphincter control, transfer, locomotion, communication, and social cognition. RESULTS: Direct interviews were conducted in the communities for 569 normal Thai children (289 girls and 280 boys) aged 6-100 months. The interrater and intrarater reliability scores were examined. The WeeFIM total and domain scores increased progressively with age. Intraclass correlation coefficients for the reliability for the WeeFIM domain score ranged from 0.90 to 0.99. Total WeeFIM intraclass correlation coefficients values were greater than 0.97 for all analyses. The authors classified the 18 items into six groups according to the degree of correlation with age. Most items were highly correlated with age as indicated by a Spearman's correlation coefficient greater than 0.8. The interrater and intrarater reliability of the WeeFIM subscores was high. CONCLUSION: The current study demonstrated that WeeFIM could be employed as a useful and reliable instrument for assessing functional independence for Thai children. Therefore, usage of WeeFIM with different age criteria for achieving independence should be adopted. Normative functional independence measures for a large group of Thai children will enhance the knowledge base about their development measurement and provide a database for future investigations on clinical population in Thailand.

Adolescent↗

Reliability and validity of women's recall of mammographic screening.

OBJECTIVE: To assess the reliability and validity of self-reported attendance for mammographic screening. METHODS: To assess reliability of recall of attendance for a screening mammogram, 100 women selected at random were interviewed twice (approximately one week apart). To assess validity, 127 women who reported having a mammogram within the national breast screening program (BreastScreen Australia) consented to having their reports verified by the national program. RESULTS: Test-retest reliability for the question "Have you ever had a mammogram?" was perfect (agreement 100%, kappa 1). Validity was also high. About one-quarter of women (24.4%) recalled the exact date of their last mammogram and a further third (39.4%) correctly reported the month in which the mammogram was done. Almost all (91.3%) women reported the mammogram date accurately to within 12 months of the recorded date. CONCLUSIONS: These data suggest that Australian women provide reliable and valid information in relation to mammographic screening attendance. IMPLICATIONS: Self-reported data about attendance for mammographic screening are likely to provide reliable and valid estimates for research and health services evaluation purposes.

Adult↗

An empirical look at the Defense Mechanism Test (DMT): reliability and construct validity.

Although the Defense Mechanism Test (DMT) has been in use for almost half a century, there are still quite contradictory views about whether it is a reliable instrument, and if so, what it really measures. Thus, based on data from 39 female students, we first examined DMT inter-coder reliability by analyzing the agreement among trained judges in their coding of the same DMT protocols. Second, we constructed a "parallel" photographic picture that retained all structural characteristic of the original and analyzed DMT parallel-test reliability. Third, we examined the construct validity of the DMT by (a) employing three self-report defense-mechanism inventories and analyzing the intercorrelations between DMT defense scores and corresponding defenses in these instruments, (b) studying the relationships between DMT responses and scores on trait and state anxiety, and (c) relating DMT-defense scores to measures of self-esteem. The main results showed that the DMT can be coded with high reliability by trained coders, that the parallel-test reliability is unsatisfactory compared to traditional psychometric standards, that there is a certain generalizability in the number of perceptual distortions that people display from one picture to another, and that the construct validation provided meager empirical evidence for the conclusion that the DMT measures what it purports to measure, that is, psychological defense mechanisms.

Adult↗

The Diagnostic Interview Schedule for Children (DISC-2.1) in Spanish: reliability in a Hispanic population.

The reliability across time, informants and interviewers of the Spanish translation of the DISC-2.1 was tested on a Puerto Rican Hispanic sample using a test-retest design. Levels of reliability between clinic and community samples and between younger and older children were compared to explore the sources of low reliability for certain psychiatric disorders. Parents' reports tended to be more reliable than those of their children, although the difference was less obvious with older children. Reliability was generally higher for the externalizing disorders and when the second interviewer was a psychiatrist rather than a lay interviewer.

Adolescent↗

Reliability of goniometric measurements of hip motion in spastic cerebral palsy.

Reliability of goniometric measurements of ranges of motion in the right hip of four children with mild or moderate spastic diplegia was studied by comparing results when physical therapists used specific and non-specific measurement instructions. Measurements of hip extension, abduction and external rotation were made. The use of specific measurement instructions improved inter-rater reliability only in the case of external rotation. Inter-session reliability was not improved by the use of specific measurement instructions. It is concluded that goniometric measurements have such a low level of reliability that they can only be used to assist clinical judgment, and that they are not sufficiently reliable to be used in studied of cerebral palsy.

Adolescent↗

Some aspects of the reliability of Touwen's examination of the child with minor neurological dysfunction.

Three studies concerning the inter-rater and test-retest reliability of the Touwen examination are presented. The results of the first study showed that it was not possible to achieve acceptable levels of reliability using the manual as the only reference for instruction. Although the reliability estimates of the total scores were good, inter-rater reliability for the nine groups of items and the individual tasks within them was poor. When methodology and interpretation of performance was agreed between observers, a second study showed that these disagreements diminished. The third study demonstrated that the short-term stability of total scores is good, but reliability for group and individual item scores remained poor.

Child↗

The reliability of ERP components in the auditory oddball paradigm.

Nineteen adolescents (average age 15 years, 3 months) were tested and retested using a standard 40 target, auditory oddball ERP paradigm across an interval of 1 year, 10 months to determine reliability of the ERP components, both in terms of intersubject stability and score agreement and in terms of trait (between-session reliability) versus state (within-session reliability). Significant trait stability was found for the N100, P200, and P300 latencies (r = .48, .51, and .74, respectively), and for P300 amplitude (r = .62), supporting the P300 as a reliable measure, with the stability required for group research but not necessarily for clinical applications. Discussion and examples illustrate the application of reliability information to the planning and evaluation of ERP paradigms.

Adolescent↗

Reliability of the Body Awareness Scale-Health.

In physical therapy the clinical assessment Body Awareness Scale-Health (BAS-H) focusing on the quality of movements and movement behaviour has previously been studied for validity. The aim of this study was to address the inter-rater reliability and test-retest reliability in three groups. The groups assessed were: patients in psychiatric care with eating disorders (n = 26), patients in rehabilitation of prolonged musculoskeletal pain (n = 22) and healthy individuals (n = 22). Results revealed inter-rater reliability (n = 70) of the BAS-H total to be 79.9 % with acceptable agreement (accepting one scale-step of difference) and 48.7% with perfect agreement. Weighted Kappa ranged between 0.34 and 0.92. Test-retest reliability (n = 54) as a mean for both raters were found to be 90.5% for the BAS-H total with acceptable agreement and 60.4% with perfect agreement. Weighted Kappa ranged between 0.65 and 0.92. The BAS-H seems to be a reliable assessment in the rehabilitation of patient with prolonged pain, psychiatric disorders and healthy controls when used according to the manual. The authors, however, suggest some revisions.

Activities of Daily Living↗

The reliability of survey assessments of characteristics of medical clinics.

OBJECTIVE: To assess the reliability of survey measures of organizational characteristics based on reports of single and multiple informants. DATA SOURCE: Survey of 330 informants in 91 medical clinics providing care to HIV-infected persons under Title III of the Ryan White CARE Act. STUDY DESIGN: Cross-sectional survey. DATA COLLECTION METHODS: Surveys of clinicians and medical directors measured the implementation of quality improvement initiatives, priorities assigned to aspects of HIV care, barriers to providing high-quality HIV care, and quality improvement activities. Reliability of measures was assessed using generalizability coefficients. Components of variance and clinician-director differences were estimated using hierarchical regression models with survey items and informants nested within organizations. PRINCIPAL FINDINGS: There is substantial item- and informant-related variability in clinic assessments that results in modest or low clinic-level reliability for many measures. Directors occasionally gave more optimistic assessments of clinics than did clinicians. CONCLUSIONS: For most measures studied, obtaining adequate reliability requires multiple informants. Using multiple-item scales or multiple informants can improve the psychometric performance of measures of organizational characteristics. Studies of such characteristics should report the organizational level reliability of the measures used.

Adult↗

Creating high reliability in health care organizations.

OBJECTIVE: The objective of this paper was to present a comprehensive approach to help health care organizations reliably deliver effective interventions. CONTEXT: Reliability in healthcare translates into using valid rate-based measures. Yet high reliability organizations have proven that the context in which care is delivered, called organizational culture, also has important influences on patient safety. MODEL FOR IMPROVEMENT: Our model to improve reliability, which also includes interventions to improve culture, focuses on valid rate-based measures. This model includes (1) identifying evidence-based interventions that improve the outcome, (2) selecting interventions with the most impact on outcomes and converting to behaviors, (3) developing measures to evaluate reliability, (4) measuring baseline performance, and (5) ensuring patients receive the evidence-based interventions. The comprehensive unit-based safety program (CUSP) is used to improve culture and guide organizations in learning from mistakes that are important, but cannot be measured as rates. CONCLUSIONS: We present how this model was used in over 100 intensive care units in Michigan to improve culture and eliminate catheter-related blood stream infections--both were accomplished. Our model differs from existing models in that it incorporates efforts to improve a vital component for system redesign--culture, it targets 3 important groups--senior leaders, team leaders, and front line staff, and facilitates change management-engage, educate, execute, and evaluate for planned interventions.

Health Facilities↗

Using adapted short MASTs for assessing parental alcoholism: reliability and validity.

In previous research adapted versions of the Short Michigan Alcoholism Screening Test (SMAST) have been employed to assess an individual's father's (F-SMAST) and mother's alcohol abuse (M-SMAST). However, to date psychometric information on these forms has been limited. In order to more broadly assess the psychometric properties of these forms, several critical issues in five related studies were addressed. The samples for the five studies were drawn from a college population at a large midwestern university. Overall, the reliability and validity of the adapted SMASTs appears to be quite good. The F-SMAST demonstrated high reliability (from the standpoint of internal consistency, temporal stability, and reliability across siblings) as well as validity (both in respect to convergence with an interview measure and with father's own report on a parallel instrument). Furthermore, shortening both of these instruments to nine-item versions appears to improve their reliability and validity. For researchers and clinicians interested in assessing parental history of alcoholism, the F-SMAST and M-SMAST would appear to be a reliable and valid paper-and-pencil measure.

Alcoholism↗

An evaluation of the reliability and validity of the Functional Assessment Inventory.

The reliability of the Functional Assessment Inventory (FAI) was evaluated using a sample of VA domiciliary and nursing home patients. The interobserver and interrater reliability coefficients of the summary rating scales, based on a single assessment, tended to be higher than their test-retest reliability coefficients, based on two independent assessments separated by a modal four-week interval. Validity coefficients, using the OARS instrument ratings as criteria, also based on two independent assessments several weeks apart, were, on the average, as high as the test-retest reliability coefficients. More specifically, the mental health, physical health, and activities of daily living rating scales, along with the objectively scored Short Portable Mental Status Questionnaire and Short Psychiatric Evaluation Schedule, tended to yield relatively similar scores with repeated measurement, while the social resources and economic resources scales were somewhat less stable, a discrepancy possibly explained by the homogeneous nature of the social and economic status of most of the patients (institutionalized veterans). Thus the reliability and validity of the FAI are satisfactory, but the stability of some of its scales requires further investigation.

Activities of Daily Living↗

Assessing diagnostic approaches to depression in medically ill older adults: how reliably can mental health professionals make judgments about the cause of symptoms?

OBJECTIVE: To evaluate the reliability of the DSM-IV approach and five other schemes for counting symptoms toward the diagnosis of depression in hospitalized medically ill older patients and to examine whether mental health professionals can reliably make judgments about the etiology (medical or psychological) of depressive symptoms. METHOD: A sample of 38 patients aged 60 years or older admitted to the general medicine, cardiology, or neurology services at Duke University Medical Center were evaluated for depression using a structured psychiatric interview and the Hamilton Depression Scale. Interrater reliability for the diagnostic schemes, for unstructured clinical diagnoses, and for determinations of the causes of individual depressive symptoms was assessed by three pairs of mental health professionals. RESULTS: Agreement between raters for structured diagnoses was high regardless of diagnostic strategy, with the DSM-IV approach being only slightly less reliable than the strict inclusive approach (Kappa 0.88 vs Kappa 1.0, respectively). For all diagnostic approaches, there was perfect agreement between raters for eight cases of major depression. Agreement for unstructured clinical diagnoses of depression (K = 0.50) was much lower than for the structured diagnoses. Agreement between raters on the etiology of individual depression criterion symptoms assessed by structured interview was greater than 80% for 14 of 19 symptoms. Correlation between raters' depression severity ratings on the Hamilton Scale using the DSM-IV etiologic approach was equivalent to that using the strict inclusive approach (0.98 vs 0.95, respectively). CONCLUSIONS: Mental health professionals can be trained to make judgments reliably about the cause (medical or psychological) of symptoms in hospitalized older medical patients. The "strict inclusive" and other diagnostic schemes for counting symptoms toward the diagnosis of depression have only marginal, if any, benefit compared with the current DSM-IV approach.

Age Factors↗

Interrater reliability of the Clinical Dementia Rating in a multicenter trial.

OBJECTIVE: To test the interrater reliability of the Clinical Dementia Rating (CDR) in a multicenter clinical trial. DESIGN: Observational study. SETTING: Training session for a multicenter trial of milameline, a direct muscarinic agonist, in the treatment of Alzheimer's disease. PARTICIPANTS: Twenty-four raters (physicians and nurses) familiar with drug trials and expert in the care of patients with Alzheimer's disease. METHODS: Independent scoring of the CDR using four videotaped CDR interviews. OUTCOME MEASURE: Interrater reliability, as tested by the Kappa statistic RESULTS: The overall interrater reliability was 0.62. Within the CDR domains, the global kappas ranged from 0.33 +/- 0.06 to 0.88 +/- 0.06. CONCLUSIONS: The data support moderate to high overall interrater reliability but show important difficulties in the reliable assessment of early dementia.

Dementia↗

Reliability and validity testing of three breastfeeding assessment tools.

OBJECTIVE: This study examined validity and reliability of three clinical instruments that assess feedings at the breast. DESIGN: A descriptive correlational design testing the validity and interrater and test-retest reliability of instruments. SETTING: Hospital rooms and the participants' homes. SUBJECTS: Eleven breastfeeding women and their neonates were videotaped in 23 breastfeeding observations. INTERVENTIONS: The Infant Breastfeeding Assessment Tool (IBFAT), the Mother Baby Assessment Tool (MBA), and the LATCH assessment tool were scored by three nurse raters using videotapes of breastfeedings. Instruments were completed twice by each rater with a 6-month period between administration. MAIN OUTCOME MEASURES: To test validity, test-retest, and interrater reliability, Spearman correlation coefficients among raters' breastfeeding assessment scores, among scores of each instrument, and between test and retest scores of raters. Percent of agreement among raters for each of the items in the three tools. RESULTS: Reliability coefficients for all three assessment tools are below acceptable levels for clinical decisions. Spearman rank coefficients of pairwise interrater correlations were .57, .27, and .69 for the IBFAT: .66, .64, and .33 for the MBA; and .11, .46, and .48 for the LATCH assessment tool. Spearman rank coefficients among instrument scores were .69, .78, and .68. Test-retest correlations were .88, .78, and .64. Percent of agreement among raters for each of the items in the three tools was highly variable, ranging from 37.0 to 97.2. CONCLUSION: The IBFAT, MBA, and LATCH as tools to measure breastfeeding effectiveness are not sufficiently reliable at this stage in their development; thus, these tools cannot be valid for clinical use. These tools need to be revised and retested before use in clinical practice to identify breastfeeding mother-infant pairs who need intervention.

Adult↗

Reliability of length measurements in full-term neonates.

OBJECTIVE: To describe and compare the intra- and interexaminer reliability of four techniques for measuring length in full-term newborns and to determine whether the different techniques yield significantly different measurements. DESIGN: A descriptive study, describing the intra- and interexaminer reliability of four length measurement techniques: crown-heel, supine, paper barrier, and Neo-infantometer. The nurses were blind to their own and to the other nurse's measurements. The order of the nurses and the order in which the measurements were obtained was randomized. SETTING: Mothers' rooms in a university hospital. PARTICIPANTS: Thirty-two healthy full-term newborns. INTERVENTIONS: Length measurements using four different length techniques were obtained twice each by two experienced neonatal nurses. MAIN OUTCOME MEASURES: To measure the intra- and interexaminer reliability, the following statistics were calculated: mean absolute differences, standard deviations, technical error of measurement; percentage less than .5 and 1.0 cm, and percentage of error. RESULTS: Intra- and interexaminer differences were significantly larger when examiners used the crown-heel measurement technique. Although the intra- and interexaminer reliability of length measurements obtained with the supine, paper barrier, and Neo-infantometer techniques did not differ significantly, the amount of error in these measurements was large. CONCLUSIONS: Measurements obtained using the crown-heel technique are significantly less reliable than measurements obtained using the supine, paper barrier, or Neo-infantometer techniques.

Anthropometry↗