Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

Reliability of normalisation methods for EMG analysis of neck muscles.

Acceptable reliability of normalisation contractions in electromyography (EMG) is paramount for testing conducted over a number of days or if normal laboratory strength testing equipment is unavailable. This study examined the reliability of maximal voluntary isometric contractions (MVIC) and sub-maximal (60%) isometric contractions for use in neck muscle EMG studies. Surface EMG was recorded bilaterally from eight sites around the neck at C4/5 level from five healthy male subjects. Subjects performed MVIC and sub-maximal normalisation contractions using an isokinetic dynamometer (ID) and a portable cable dynamometer with attached strain gauge (PCD) in addition to a MVIC against a manual resistance (MR). Subjects were tested in flexion, extension, left and right lateral bending and were retested by the same tester within a two-week period. Intra class correlation co-efficients (ICC) were calculated for each testing method and contraction direction and a mean ICC was calculated across all contraction directions. All normalisation methods produced excellent within-day reliability (mean ICC >0.80) but only the MVICs using the ID and PCD had acceptable reliability when assessed between-days. This study confirmed the validity of using MVICs elicited using the ID and PCD as reliable reference contractions for the normalisation of neck EMG.

Adult↗

Validity and reliability study of the Thai version of WHO SCAN: somatoform and dissociative symptoms section.

OBJECTIVES: To determine the validity and reliability of the Thai version of the WHO Somatoform and Dissociative Symptoms Section of the Schedules for Clinical Assessment in Neuropsychiatry (SCAN) Version 2.1 MATERIAL AND METHOD: The SCAN interview version 2.1 Somatoform and Dissociative Symptoms Section was translated into Thai. The content validity of the translation was verified by comparing a back-translation (to English) of the Thai version to the English original. Whenever inconsistencies were encountered, the Thai version was adapted so that it correctly conveyed the meaning of the original English version. The revised Thai version was then field-tested nationwide for the comprehensibility of the relatively technical language. Between October 2003 and August 2004, 30 persons were recruited for the reliability study (16 males; 14 females) Fifteen subjects had somatoform disorders and 15 were normal. The number of years of formal education varied widely and occupations were diverse. Subjects were interviewed by a psychiatrist competent in using the Thai version of SCAN. The interviews were recorded on video so that the material could be rerated. RESULTS: Based on the response from Thai subjects and consultations with competent psychiatrists, the content validity was established. The time taken to interview a somatoform patient averaged 57.1 +/- 12.1 minutes while it was 42.1 +/- 13.9 minutes for a normal subject. The inter-rater reliability (kappa) of the 113 Items were: 0.81-1.0, 0.61-0.80 and 0. 00-0.20 in 49.6, 30.0 and 8.9 percent, respectively. Kappas could not be calculated for 11.5% of the Items. The intra-rater reliabilities were. 0.81-1.0, 0.61-0.80 and 0.00-0.20 in 54.9, 26.5 and 2.7 percent, respectively. Kappas could not be calculated for 15.9% of the Items. CONCLUSION: The Thai version of the Somatoform and Dissociative Symptoms Section of SCAN version 2.1 proved to be a valid and reliable tool for assessing somatoform and dissociative symptoms among Thai speakers.

Adult↗

[Comment on "the reliability and validity of the effort-reward imbalance - the Chinese version"].

OBJECTIVE: To evaluate reliability and validity of the effort-reward imbalance (ERI) in the Chinese version. METHODS: A cross-sectional survey was conducted comprising a large sample of 4782 subjects in China, using ERI in the Chinese version. This scale contained 23 scaled items while the questionnaire including questions on the effort and reward at work, over-commitment, the full CES-D scale of depression and a range of other characteristics. Reliability analysis was applied to evaluate reliability of the ERI scale in the Chinese version and factor analysis was applied to analyze validity of the scale. RESULTS: Theoretical hypothesis on the ERI model was supported by the data derived in this study. Reliability and validity of the effort sub-scale, the reward sub-scale of the ERI scale in the Chinese version seemed to be better, but reliability and validity of the over-commitment sub-scale were not perfect. CONCLUSION: The results of the study showed that the effort sub-scale, the reward sub-scale of the ERI in the Chinese version was applicable to the Chinese population but the scaled items of the over-commitment sub-scale should be further modified.

China↗

Reliability of Enderlein's darkfield analysis of live blood.

CONTEXT: In 1925, the German zoologist Günther Enderlein, PhD, published a concept of microbial life cycles. His observations of live blood using darkfield microscopy revealed structures and phenomena that had not yet been described. Although very little research has been conducted to explain the phenomena Dr. Enderlein observed, the diagnostic test is still used in complementary and alternative medicine. OBJECTIVE: To test the interobserver reliability and test-retest reliability of 2 experienced darkfield specialists who had undergone comparable training in Enderlein blood analysis. SETTING: Inpatient clinic for internal medicine and geriatrics. METHODS: Both observers assessed 48 capillary blood samples from 24 patients with diabetes. The observers were mutually blind and assessed their findings according to a specific item randomization list that allowed observers to specify whether Enderlein structures were visible or not. RESULTS: The interobserver reliability for the visibility of various structures was kappa = .35 (95% CI: .27-.43), the test-retest reliability was kappa = .44 (95% CI: .36-.53). CONCLUSIONS: This pilot study indicates that Enderlein darkfield analysis is very difficult to standardize and that the reliability of the diagnostic test is low.

Adult↗

Reliability and validity of data for 2 newly developed shuttle run tests in children with cerebral palsy.

BACKGROUND AND PURPOSE: The purpose of this study was to examine the reliability and validity of data obtained with 2 newly developed shuttle run tests (SRT-I and SRT-II) to measure aerobic power in children with cerebral palsy (CP) who were classified at level I or II on the Gross Motor Function Classification System (GMFCS). The SRT-I was developed for children at GMFCS level I, and the SRT-II was developed for children at GMFCS level II. SUBJECTS: Twenty-five children and adolescents with CP (10 female, 15 male; mean age = 11.9 years, SD = 2.9), classified at GMFCS level I (n = 14) or level II (n = 11), participated in the study. METHODS: To assess test-retest reliability of data for the 10-m shuttle run tests, the subjects performed the same test within 2 weeks. To examine validity, the shuttle run tests were compared with a GMFCS level-based treadmill test designed to measure peak oxygen uptake. RESULTS: Statistical analyses revealed test-retest reliability for exercise time (number of levels completed) (intraclass correlation coefficients of .97 for the SRT-I and .99 for the SRT-II) and reliability for peak heart rate attained during the final level (intraclass correlation coefficients of .87 for the SRT-I and .94 for the SRT-II). High correlations were found for the relationship between data for both shuttle run tests and data for the treadmill test (r = .96 for both). DISCUSSION AND CONCLUSION: The results suggest that both 10-m shuttle run tests yield reliable and valid data. Moreover, the shuttle run tests have advantages over a treadmill test for children with CP who are able to walk and run (GMFCS level I or II).

Adolescent↗

[Validity and reliability of the Japanese version of the Neuropsychiatric Inventory Caregiver Distress Scale (NPI D) and the Neuropsychiatric Inventory Brief Questionnaire Form (NPI-Q)].

OBJECTIVE: Neuropsychiatric disturbances are common and burdensome symptoms of dementia. Assessment and measurement of neuropsychiatric disturbances are indispensable to the management of patients with dementia. Neuropsychiatric Inventory (NPI) is a comprehensive assessment tool that evaluates psychiatric symptoms in dementia. We translated the NPI-Caregiver Distress Scale part of NPI (NPI-D) and NPI-Brief Questionnaire Form (NPI-Q) into Japanese and examined their validity and reliability. SUBJECTS AND METHODS: The subjects were 152 demented patients and the caregivers who lived with them. These patients consisted of 76 women and 76 men; their mean age was 73.9 +/- 7.8 (S.D.; range: 49 to 93) years. Their caregivers consisted of 46 men and 106 women; their mean age was 65.0 +/- 11.4 (S.D.; range: 35 to 90) years. The Mini-Mental State Examination (MMSE) was conducted with all patients and NPI-Q, NPI, NPI-D, and the Zarit caregiver burden interview (ZBI) were conducted with all caregivers. We examined validity of NPI-D by comparing its score with the MMSE and ZBI scores, and the validity of NPI-Q by comparing its score with the NPI and NPI-D scores. In order to evaluate test-retest reliability, NPI-D was re-adopted to 30 randomly selected caregivers by a different examiner one month later and NPI-Q was re-executed by 27 randomly selected caregivers one day later. RESULTS: Total NPI-D score was significantly correlated with ZBI (rs = 0.59, p < 0.01). Test-retest reliability of NPI-D was adequate (ri = 0.47, p < 0.01). Total NPI-Q severity score and distress score were strongly correlated with NPI (r = 0.77, p < 0.01) and NPI-D (r = 0.80, p < 0.01) scores, respectively. Test-retest reliability of the scores of NPI-Q was acceptably high (the severity score; ri = 0.81, p < 0.01, the distress score; ri = 0.80, p < 0.01). CONCLUSION: The Japanese version of NPI-D and NPI-Q demonstrated sufficient validity and reliability as well as the original version of them. These are useful tools for evaluating psychiatric symptoms in demented patients and their caregivers' distress attributable to these symptoms.

Caregivers↗

Reliability of the capsaicin cough reflex sensitivity test in healthy children.

Testing cough reflex sensitivity (CRS) in children requires suitable methodology. A CRS test performed under control of inspiratory flow rate (IFR) shows excellent reliability in children, but it is difficult to perform, especially in younger children. The aim of the present study was to find whether the capsaicin CRS test performed without direct control of constant IFR in healthy children is reliable enough for practical use. The CRS test was performed in 27 healthy children, aged 7-17 yr three times within 8 days. Cough was induced by inhalation of capsaicin aerosol in doubling concentrations (0.61-1250 micromol/l) for 400 ms each. CRS was defined as the lowest capsaicin concentration that evoked 2 or more coughs (C2). Although the intraclass correlation coefficient values showed good to excellent reliability of this test, the within-subject standard deviation values revealed lower reliability of this method compared to the CRS test performed under control of IFR. From the results obtained it is reasonable to conclude that the method using uncontrolled IFR in CRS testing provides acceptable precision only when a bigger sample size is used or more tests are performed. Good to excellent reliability of this method was found in children with higher values of C2 and in those aged 13-17 yr.

Adolescent↗

Validity and reliability of active shape models for the estimation of Cobb angles in adolescent idiopathic scoliosis.

Choosing the most suitable treatment for scoliosis relies heavily on accurate and reproducible Cobb angle measurement from successive radiographs. The objective is to reduce variability of Cobb angle measurement by reducing user intervention and bias. Custom software to automate Cobb angle measurement from posteroanterior radiographs was developed using active shape models. Validity and reliability of the automated system against a manual and semi-automated measurement method was conducted by two examiners each performing measurements on 3 occasions from a test set (N=22). A training set (N=47) of radiographs representative of curves seen in a scoliosis clinic was used to train the software to recognize vertebrae from T4 to L4. Images with a maximum Cobb angle between 20 degrees and 50 degrees, excluding surgical cases, were selected for training and test sets. Automated Cobb angles were calculated using best-fit slopes of the detected vertebrae endplates. Intra-class correlation coefficient (ICC) and standard error of measurement (SEM) showed high intra-examiner (ICC > 0.90, SEM 2-3 degrees) and inter-examiner (ICC > 0.82, SEM 2-4 degrees), but poor inter-method reliability (ICC=0.30, SEM 8-9 degrees). The automated method underestimated large curves. The reliability improved (ICC = 0.70, SEM 4-5 degrees) with exclusion of the 4 largest curves (>40 degrees) in the test set. The automated method was reliable for moderate sized curves, but did not properly detect vertebrae in larger curves. Optimization of constraints on scaling, rotation, translation, and iteration may improve reliability with larger curves.

Adolescent↗

3D back shape in healthy young adults: an inter-rater and intra-rater reliability study.

UNLABELLED: Whilst postural evaluation of spinal dysfunction is routine in physiotherapy practice, objective measurements are rarely undertaken due to the scarcity of reliable low cost assessment tools. Warren et al [12] found high intrarater reliability but inter-rater reliability has not as yet been evaluated. The purpose of this study was to assess both the intra-rater and inter-rater reliability of the Middlesbrough Integrated Digital Assessment system (MIDAS). METHODS: A convenience sample of twenty-five healthy University of Teesside students was recruited for the study. One rater palpated fifteen key landmarks on each subjects back. Each of three raters took two measurements on each subject in a standardized upright posture. Rater order was randomized to minimize data recording bias. X (medio-lateral), Y (antero-posterior) and Z (height) landmark positions were recorded via a computer interface. DATA ANALYSIS: Intraclass Correlation Coefficients (ICC 2,1) were used to analyse data using SPSS v13. RESULTS: Both intra-rater agreement (mean ICCs - rater 1 r= 0.970, rater 2 r= 0.965 and rater 3 r= 0.965, p<0.001) and inter-rater agreement (mean ICCs r = 0.967, p<0.001 was very high between repeated measures and between markers. Error values for the z-axis (height) were lowest. CONCLUSIONS: The system demonstrated both high inter-rater and intra-rater reliability. Before the system is used on a clinical population, data output needs to be converted from raw format to a clinically applicable format. Work is currently being undertaken to develop an interactive visual display and control for postural sway.

Adult↗

Cross-cultural adaptation of modified Oswestry Low Back Pain Disability Questionnaire to Thai and its reliability.

OBJECTIVE: The present study aimed to cross-culturally adapt the modified Oswestry Low Back Pain Disability Questionnaire (ODQ) into Thai. MATERIAL AND METHOD: The process comprised of an initial forward translations from English to Thai, synthesis of the translations, back translation, and back translation approval. The approved version of Thai ODQ was then calculated for test-retest reliability. Forty patients with LBP, aged 40.1+/-10.7 years, were recruited into a test-retest reliability study. RESULTS: The test-retest reliability, calculated by intraclass correlation coefficient, was assessed on two occasions separated by a time interval of 20-30 minutes. The values of test-retest reliability of items ranged from 0.80-1.00. The value of total score was 0.98. CONCLUSION: This finding indicated good reliability of the Thai version modified ODQ.

Adult↗

Inter- and intra-rater reliability of the Thai version of SCAN: Use of Alcohol and Use of Tobacco Section.

OBJECTIVE: To assess inter and intra rater reliability of the Thai version of Use of Alcohol and Use of Tobacco Section of the WHO Schedules for Clinical Assessment in Neuropsychiatry (SCAN). MATERIAL AND METHOD: Fifteen alcohol and/or tobacco dependence patients and fifteen controls with ages of > or = 18 years were recruited from October 2003 to August 2004 at Srinagarind Hospital. One psychiatrist interviewed and first rated under video-recording, then re-rated two weeks later. Another psychiatrist rated independently by looking at the videotapes for inter-rater reliability testing. RESULTS: The intra-rater kappa was 'excellent' for both sections. The inter-rater kappa for all items of "tobacco use" were 'excellent' (mean kappa = 0.84), and 'good' (mean kappa = 0.66) for the "alcohol use". Items of dependence had 'good' kappa (mean kappa = 0.73-0.95), except items of 'activities limitation (interest neglect)' and 'use despite knowledge of psychological/physical problems (continued use)' was 'fair' (mean kappa 062-0.49). The poor kappa (mean kappa = 0.17-0.38) was found in the item of 'physical and mental health problems due to drinking'. CONCLUSION: Thai SCAN provided reliable inter-rater diagnostic information for alcohol and tobacco dependence, but fair reliability for alcohol abuse. Improve understanding of the item concepts, better validity in questioning; giving examples of the focus symptoms, and frequent discussion about the respondent's answers, might improve overall validity and reliability.

Adult↗

Intratester and intertester reliability of the cervical range of motion device.

The purpose of this study was to investigate the intratester and intertester reliability of the Cervical Range of Motion instrument (CROM) for measuring cervical flexion, extension, lateral flexion, and rotation. Twenty able-bodied subjects were tested by two testers on two different occasions. Pearson product-moment correlations for intratester reliability ranged from .63 to .90 for tester one and from .62 to .91 for tester two. Intertester reliability was good. Coefficients ranged from .80 to .87 for session one and .74 to .85 for session two. Paired data t-tests showed that there were no significant differences between testers or sessions (p = .01). The results suggest that the CROM has acceptable intratester and intertester reliability. The CROM has many benefits including ease of application and reliability. More research is needed on patients with cervical dysfunction.

Adult↗

Statistical methodology for reliability studies.

Although a considerable number of reliability studies have appeared in the chiropractic literature in the last 10 yr, there is still no consensus on the definition of reliability, appropriate statistics for reliability testing or interpretation of results. This paper offers suggestions for the standardization of experimental design, analysis and evaluation for reliability studies. The primary focus is on appropriate concordance statistics, i.e., Kappa and intraclass correlation. These are presented in detail, together with a discussion of their limitations. Less suitable, but often used indices such as Pearson's r are also discussed. A measure of precision, the interexaminer measurement error, is introduced and a methodology for evaluating precision is presented. Clinical issues which must be considered in reliability studies are discussed; specifically, those issues relevant to studies of diagnostic indicators for spinal manipulation.

Analysis of Variance↗

Maximal exercise testing of mentally retarded adolescents and adults: reliability study.

Few data are available regarding maximal exercise testing of mentally retarded individuals. No data are available on the reliability of maximal exercise testing of mentally retarded individuals. The purpose of this study was to determine the reliability of graded exercise testing of mentally retarded adolescents and adults. The testing was conducted at two geographically different centers. At Center A, 14 mentally retarded adolescents (11 boys, three girls) with Down syndrome, who were educable or trainable, were recruited from a nonresidential school. The subjects completed two Balke-Ware treadmill protocols until exhaustion. The treadmill time and heart rate (HR) were recorded. The time between tests was approximately one week. At Center B, 21 mentally retarded adults (14 women, seven men means IQ = 56) were recruited from local workshops and group homes. These subjects completed a treadmill walking protocol, with metabolic measurements, until exhaustion. The time between tests varied from one to four months. At Center A, the subjects achieved a mean treadmill time of 8.72min on test one and 8.84min on test two (means HR = 174 and 175bpm, respectively). The reliability coefficient between the two tests was .94. At Center B, the subjects achieved a mean V0(2)max of 27.2mL.kg-1.min-1 on test one and 26.9mL.kg-1.min-1 on test two. The reliability coefficient was .93. These data show that maximal exercise testing is reliable for these populations of mentally retarded individuals, exhibiting similar values to their nonretarded peers.

Adolescent↗

Reliability and validity of an objective structured clinical examination for assessing the clinical performance of residents.

Clinical performance of residents should be assessed as reliably and validly as possible. This study investigated the reliability and validity of an objective structured clinical examination (OSCE) for assessing clinical performance of internal medicine residents. Residents were required to take a 17-patient OSCE in their first and second year. Reliability of the OSCE was 0.40. Validity studies indicated second-year students were significantly better than third-year students for five of six OSCE skill scores; first-year students were significantly better for three scores. Resident's scores for diagnosis, plan, and total significantly increased on their second OSCE. Generally faculty overall ratings of residents' clinical performance did not correlate with OSCE scores. American Board of Internal Medicine certifying examination scores were consistently positively correlated only with diagnosis. This 17-case OSCE is a feasible method for obtaining moderately reliable, valid data not available from other sources about the clinical performance of residents. More cases should be added to increase its reliability.

Clinical Competence↗

[Evaluation of the reliability of anthropometric measures].

As part of the Early Malnutrition and Its Effects on Youth Project, a study of the reliability of anthropometric measures was carried out through the re-measurement of 226 adolescents of the rural area of Guatemala. In all anthropometric variables, the intra-measure coefficients of reliability were higher than 0.96 and the inter-measure coefficients of reliability higher than 0.91. A significant effect of the measure was found, suggesting the existence of systematic differences among measures. This information will permit corrections by measure effect in the analysis of data. The effect of different grades of "data editing" and "validity, reliability and agreement check", of measures was evaluated. The conclusion was that even though the persons in charge of measurements were trained and supervised, the deletion of values outside ranks did not result in an appreciable decrease of the technical error of measurement. As expected, data editing and validity carried out through the comparison of two measurements taken from each subject and eliminating those whose differences were classified as error according to statistical criteria, decreased the technical error of measurement, in most of the variables measured. Nevertheless, with the exception of one of the variables, this decrease was not marked. This indicates that the technical error of measurement obtained from all the subjects, many of whom were not re-measured, was within acceptable ranks. Results of the study demonstrate the importance of including re-measurement of a certain percentage of subjects when designing anthropometric studies. This will allow evaluation of the reliability of measurements, identification of bias due to measurements, and their correction.

Adolescent↗

Reliability of flexible fiberoptic nasopharyngoscopy for evaluation of velopharyngeal function in a clinical population.

Flexible fiberoptic nasopharyngoscopy (FFN) has become a popular clinical tool for evaluating velopharyngeal function. The literature contains numerous reports of FFN methodology and findings. However, there are few published reports that address clinical rating and reporting schemes or that evaluate viewer reliability. This study was designed to evaluate the reliability of visual perceptual ratings of FFN video images for assessing velopharyngeal structure and function in a clinical population. Ninety-five videotaped clinical evaluations were presented to and judged by three expert raters and nine novice raters from the fields of speech pathology, otolaryngology, and plastic surgery, using a standard rating form. A clinical rating scheme was used for quantifying perceptual judgments of velopharyngeal activity. Results suggest that videotaped FFN evaluations may be rated reliably, that expert raters working as a group are more reliable than novice raters working individually, and that the 125 evaluations presented without feedback are insufficient to improve the novices' reliability. The combined auditory and visual perceptual evaluation inherent in FFN may be its most significant asset for both clinical and research applications.

Adolescent↗

The reliability of estimates of interhemispheric transmission times derived from unimanual and verbal response latencies.

In this study, the reliability of estimates of interhemispheric transmission times derived from verbal and unimanual responses was assessed. Responses were made to presentations of lateralized light flashes in a series of 20 experimental sessions, with responses being made to stimuli presented at three different retinal eccentricities. The reliability of estimates of interhemispheric transmission across 20 experimental sessions was generally poorer for verbal compared to unimanual responses. This was particularly true for stimuli presented close to the centre of the visual field. In addition, for unimanual response conditions, estimates of interhemispheric transmission times for the female subjects were longer in duration and exhibited poorer reliability than estimates found for the male subjects. In the verbal response conditions, there were no differences between the mean durations of the estimates found for the male and female subjects. The results suggest that for centrally located stimulus presentations, verbal responses do not produce reliable estimates of interhemispheric transmission time. This relatively poor reliability across experimental sessions may account for the mixed results seen in previous studies which have used verbal reaction times to estimate interhemispheric transmission time. The collection of unimanual responses to lateralized visual stimulus presentations is likely the best method of estimating interhemispheric transmission time in normal man.

Adult↗