Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Reliability of motor cortex transcranial magnetic stimulation in four muscle representations.

OBJECTIVE: Motor cortex plasticity may underlie motor recovery after stroke. Numerous studies have used transcranial magnetic stimulation (TMS) to investigate motor system plasticity. However, research on the reliability of TMS measures of motor cortex organization and excitability is limited. We sought to test the reliability of these TMS measurements. METHODS: Twenty healthy volunteers were tested twice over a two-week period using TMS to determine motor threshold, map topography, and stimulus-response curves for first dorsal interosseous (FDI), abductor pollicis brevis (APB), extensor digitorum communis (EDC), and flexor carpi radialis (FCR) muscles. RESULTS: We found moderate to good test-retest reliability TMS measurements of motor threshold (ICC=0.90-0.97), map area (ICC=0.63-0.86) and location (ICC=0.69-0.86), and stimulus-response curves (ICC=0.60-0.83). CONCLUSIONS: TMS assessments of motor representation size, location, and excitability are generally reliable measures, although their reliability may vary according to the muscle under investigation. SIGNIFICANCE: These results suggest that TMS measurements of motor cortex function are reliable enough to be potentially useful in investigation of motor system plasticity.

Adult↗

The Japanese version of the Depressive Experiences Questionnaire: its reliability and validity for lifetime depression in a working population.

We have developed a Japanese version of the Depressive Experiences Questionnaire (DEQ), devised by Blatt et al., for assessing depression-prone personality and examined the questionnaire's reliability (test-retest reliability and internal consistency) and validity. To examine the questionnaire's validity, we evaluated its factorial validity and discriminant power for depression (i.e., construct validity). To test the construct validity of the DEQ with and without depression proneness, the scores on the DEQ subscales were compared between subjects with and without a lifetime history of major depressive disorder (MDD). The Inventory to Diagnose Depression, Lifetime version (IDDL), was used to identify lifetime depression. The reliability tests showed that the Japanese version has reliability almost similar to that of the original version. While the self-criticism has good reliability, the dependency appears to have only modest reliability. In the comparisons between subjects with and without lifetime histories of major depression, the former had significantly higher scores on the self-criticism dimension of the DEQ than did the latter, suggesting that the Japanese version of the DEQ, especially the self-criticism, may have the ability to distinguish individuals with lifetime depression from normal controls. We conclude that the DEQ is an acceptable instrument for assessing the depression-prone personality.

Adult↗

The interrater reliability of the Structured Interview for DSM-IV Personality.

We examined the joint interview interrater reliability of the Structured Interview for Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition, Personality Disorders (SIDP-IV) in 433 non-treatment-seeking military recruits. Reliability was computed for the diagnosis of a specific personality disorder (PD) and for the number of PD criteria present, and computed using a dimensional score. Reliability increased when PDs were computed using dimensional scores rather than categorical scores. Avoidant and dependent PDs demonstrated the highest interrater reliability, whereas schizoid and schizotypal showed the lowest. This large sample allowed us to perform item-level analyses of the SIDP-IV. Interrater reliability for each of the PD criteria was generally more than 0.70, with the notable exception of criteria scored through observation only. Overall, the SIDP-IV demonstrated good reliability in a non-treatment-seeking population.

Diagnostic and Statistical Manual of Mental Disord↗

The Cannabis Problems Questionnaire: factor structure, reliability, and validity.

AIM: To develop a multi-dimensional valid and reliable measure of cannabis-related problems. METHOD: The Cannabis Problems Questionnaire (CPQ) was developed from the Alcohol Problems Questionnaire to measure cannabis treatment outcome. The CPQ was administered on two occasions 1 week apart to a stratified sample of adults who had used cannabis at least once in the previous 3 months. Exploratory factor analyses were conducted and the relationship between items of the CPQ and measures of daily use and dependence assessed. The reliability of the CPQ was also assessed using a test-retest and inter-rater reliability methodology. RESULTS: Exploratory factor analyses revealed a three factor solution best described the data accounting for 57% of the variance in the larger item set. The CPQ is highly reliable with test-retest tetrachoric correlations of between 0.92 and 1.00 and inter-rater reliability correlations between 0.74 and 1.00. The total CPQ score classified DSM-IV cannabis dependence with 84% specificity and sensitivity and daily cannabis use with 83% specificity and 55% sensitivity. CONCLUSIONS: The 22-item CPQ is a valid, reliable and sensitive measure of cannabis-related problems for use with clinical and research populations of current cannabis users.

Adolescent↗

Diagnostic Interview for Genetic Studies (DIGS): inter-rater and test-retest reliability and validity in a Spanish population.

OBJECTIVE: To test the reliability and validity of the DIGS in Spanish population. METHODS: Inter-rater and test-retest reliability of the Spanish version of DIGS was tested in 95 inpatients and outpatients. The resultant diagnoses were compared with diagnoses obtained by the LEAD (Longitudinal Expert All Data) procedure as "gold standard". The kappa statistic was used to measure concordance between blind inter-raters and between the diagnoses obtained by LEAD procedure and through the DIGS. RESULTS: Overall kappa coefficient for inter-rater reliability was 0.956. The kappa value for individual diagnosis varied from major depression=0.877 to schizophrenia=1. Test-retest reliability was 0.926. Kappa for all individual target diagnoses ranged from 0.776 (major depression) to 1. Kappa between LEAD procedure and DIGS ranged from 0.704 (major depression) to 0.825 (bipolar I disorder). CONCLUSION: Most of the DSM-IV major psychiatric disorders can be assessed with acceptable to excellent reliability with the Spanish version of the DIGS interview. The Spanish version of DIGS showed an acceptable to excellent concurrent validity. Giving the good reliability and validity of Spanish version of DIGS it should be considered to identify psychiatric phenotypes for genetics studies.

Adult↗

The standard gamble demonstrated lower reliability than the feeling thermometer.

BACKGROUND AND OBJECTIVE: Participants rated clinical marker states (CMS) to make respondents familiar with the task of preference instruments, ground their ratings in relation to other health states, and help investigators interpret patient ratings. The objective was to assess the reliability of CMS using appropriate reliability statistics. STUDY DESIGN AND SETTING: Eighty-one patients rated CMSs for mild, moderate, and severe chronic respiratory disease using the feeling thermometer (FT) and the standard gamble (SG) before and after a 3-month respiratory rehabilitation program. To assess reliability we used (a) intraclass correlation coefficients (ICC) with the variance between CMSs as signal and the variance between raters, the variance within raters, and the signal as noise; (b) scatter plots; and (c) Bland-Altman plots. RESULTS: ICCs were 0.47 for the FT and 0.37 for the SG. Scatter and Bland-Altman plots showed large between- and within-person variability; 64.2% and 11.3% of the CMSs ratings were in the correct order on both occasions on the FT and SG, respectively. CONCLUSION: Our results suggest moderate reliability of CMSs ratings for the FT and poor reliability for the SG, which may explain their lack of improving the SG's measurement properties. Investigators should use appropriate reliability statistics when addressing related issues.

Aged↗

Intra-rater reliability of electromyographic recordings and subjective evaluation of neck muscle fatigue among helicopter pilots.

UNLABELLED: The aim was to evaluate the reliability of a method of measuring neck muscle fatigue among helicopter pilots. METHOD: Surface EMG from three areas in the neck region, bilaterally, was recorded among 10 male helicopter pilots while they were performing isometric contractions in flexion and extension for 45 s, sustaining a force representing 75% of maximum strength in a seated position. Perceived fatigue was rated using the Borg CR-10 scale. The test was repeated twice the first day and then two additional times with one-week intervals. Variables analyzed were the slope of the median frequency change, the normalized slope, and the ratings after 15, 30 and 45 s; and also the initial median frequency (IMDF). The intra-class correlation (ICC) and the measurement error (S(w)), intra- and inter-day were calculated statistically. RESULTS: The best reliability for the slope was found for the 45 s intra-day analysis taking all measurements into account (ICC 0.65-0.83). The reliability after 30 s was poorer but still acceptable (ICC 0.52-0.71). For the subjective ratings, the highest reliability was found after 30 s inter-day (ICC 0.86-0.88). IMDF showed generally high reliability for the intra-day analyses (ICC 0.63-0.80). CONCLUSION: The method is reliable for use in further research. Since performing a contraction of 75% of maximum was quite strenuous, we recommend that the protocol be shortened to 30 s.

Adult↗

Reliability of techniques to assess human neuromuscular function in vivo.

The purpose of this study was to comprehensively evaluate the reliability of a large number of commonly utilized experimental tests of in vivo human neuromuscular function separated by 4-weeks. Numerous electrophysiological parameters (i.e., voluntary and evoked electromyogram [EMG] signals), contractile properties (i.e., evoked forces and rates of force development and relaxation), muscle morphology (i.e., MRI-derived cross-sectional area [CSA]) and performance tasks (i.e., steadiness and time to task failure) were assessed from the plantarflexor muscle group in 17 subjects before and following 4-weeks where they maintained their normal lifestyle. The reliability of the measured variables had wide-ranging levels of consistency, with coefficient of variations (CV) ranging from approximately 2% to 20%, and intraclass correlation coefficients (ICC) between 0.53 and 0.99. Overall, we observed moderate to high-levels of reliability in the vast majority of the variables we assessed (24 out of the 29 had ICC>0.70 and CV<15%). The variables demonstrating the highest reliability were: CSA (ICC=0.93-0.98), strength (ICC=0.97), an index of nerve conduction velocity (ICC=0.95), and H-reflex amplitude (ICC=0.93). Conversely, the variables demonstrating the lowest reliability were: the amplitude of voluntary EMG signal (ICC=0.53-0.88), and the time to task failure of a sustained submaximal contraction (ICC=0.64). Additionally, relatively little systematic bias (calculated through the limits of agreement) was observed in these measures over the repeat sessions. In conclusion, while the reliability differed between the various measures, in general it was rather high even when the testing sessions are separated by a relatively long duration of time.

Adult↗

Intrarater reliability in the ultrasound diagnosis of medial and lateral orbital wall fractures with a curved array transducer.

PURPOSE: The aims of the study were to document the effectiveness of ultrasound (US) in diagnosing orbital wall fractures when compared with computed tomography (CT) and to measure the intraobserver reliability of US using a curved array transducer. MATERIALS AND METHODS: From December 2003 to March 2004, 13 patients with the clinical diagnosis of an orbital trauma were investigated prospectively by CT (reference) and 2 US investigators. Both orbits were investigated. Sensitivity, specificity, accuracy, and positive and negative predictive value were calculated. The statistical difference between the 2 US investigators was calculated by a chi-square test. The interrater reliability was calculated using the lambda coefficient. Values below 0.4 represent poor reliability, between 0.4 and 0.75 represent fair to good reliability, and a score > 0.75 is graded as excellent reliability. RESULTS: The comparison of the results of the 2 US investigators by the chi-square test showed P values of .385 for the medial orbital wall and .638 for the lateral orbital wall, which shows no significant difference. The lambda-value for the investigation of the medial orbital wall reached 0.429, 0.714, and 0.750. The lambda-value for the investigation of the lateral orbital wall yielded 0.647, 0.750, and 0.882. These values show a good and excellent inter-rater reliability. CONCLUSION: The US investigation does not yet reach the diagnostic quality of CT. US could be a helpful diagnostic imaging tool in cases with clear clinical symptoms. The results of the current study and the previously published results imply that US has the potential to reach the same diagnostic quality as CT in the future, but further studies must be performed to improve the diagnostic quality of the method.

Adolescent↗

Reliability in perceptual analysis of voice quality.

This study focuses on speaking voice quality in male teachers (n = 35) and male actors (n = 36), who represent untrained and trained voice users, because we wanted to investigate normal and supranormal voices. In this study, both substantial and methodologic aspects were considered. It includes a method for perceptual voice evaluation, and a basic issue was rater reliability. A listening group of 10 listeners, 7 experienced speech-language therapists, and 3 speech-language therapist students evaluated the voices by 15 vocal characteristics using VA scales. Two sets of voice signals were investigated: text reading (2 loudness levels) and sustained vowel (3 levels). The results indicated a high interrater reliability for most perceptual characteristics. Connected speech was evaluated more reliably, especially at the normal level, but both types of voice signals were evaluated reliably, although the reliability for connected speech was somewhat higher than for vowels. Experienced listeners tended to be more consistent in their ratings than did the student raters. Some vocal characteristics achieved acceptable reliability even with a smaller panel of listeners. The perceptual characteristics grouped in 4 factors reflected perceptual dimensions.

Adult↗

Inter-examiner reliability of passive assessment of intervertebral motion in the cervical and lumbar spine: a systematic review.

A systematic review was conducted to determine inter-examiner reliability of passive assessment of segmental intervertebral motion in the cervical and lumbar spine as well as to explore sources of heterogeneity. Passive assessment of motion is used to decide on treatments for neck and low-back pain patients. Inter-examiner reliability has been a matter of debate, resulting in questions about professional credibility and accountability. A structured search for relevant studies in MEDLINE and CINAHL was followed by extensive reference tracing and hand searching. Studies presenting estimates of reliability for individual motion segments were included. No language restrictions were imposed. Study quality was assessed using criteria derived from the Standards for Reporting of Diagnostic Accuracy (STARD) statement and a quality assessment tool for studies of diagnostic accuracy included in systematic reviews (QUADAS). Study selection, quality assessment, and data extraction were performed by two reviewers independently. Qualitative analyses and additional subgroup analyses were conducted. Nineteen studies were included. Two studies satisfied criteria for external and internal validity, of which one found fair to moderate reliability. Assessment of motion segments C1-C2 and C2-C3 almost consistently reached at least fair reliability. Overall, inter-examiner reliability was poor to fair. However, most studies were found to be of poor methodological quality. We propose explicit recommendations for the conduct and reporting of future research.

Cervical Vertebrae↗

The reliability of selected motion- and pain provocation tests for the sacroiliac joint.

The objective of the study was to assess inter-rater reliability of one palpation and six pain provocation tests for pain of sacroiliac origin. The sacroiliac joint (SIJ) is a potential source of low back and pelvic girdle pain. Diagnosis is made primarily by physical examination using palpation and pain provocation tests. Previous studies on the reliability of such tests have reported inconclusive and conflicting results. Fifty-six women and five men aged 18-50 years old were included in the study. Fifteen patients had ankylosing spondylitis; 30 women had post partum pelvic girdle pain for more than 6 weeks; and 16 people had no low back or pelvic girdle pain. All participants were examined twice on the same day by experienced manual therapists. Percentage agreement and kappa statistic were used to evaluate the tests reliability. Results showed percentage agreement and kappa values ranged from 67% to 97% and 0.43 to 0.84 for the pain provocation tests. For the palpation test the percent agreement was 48% and the kappa value was -0.06. Clusters of pain provocation tests were found to have good percentage agreement, and kappa values ranged from 0.51 to 0.75. In conclusion this study has shown the reliability of the pain provocation tests employed were moderate to good, and for the palpation test, reliability was poor. Clusters out of three and five pain provocation tests were found to be reliable. The cluster of tests should now be validated for assessment of diagnostic power.

Adolescent↗

Using standardized video cases for assessment of medical communication skills: reliability of an objective structured video examination by computer.

OBJECTIVE: Using standardized video cases in a computerized objective structured video examination (OSVE) aims to measure cognitive scripts underlying overt communication behavior by questions on knowledge, understanding and performance. In this study the reliability of the OSVE assessment is analyzed using the generalizability theory. METHODS: Third year undergraduate medical students from the Academic Medical Center of the University of Amsterdam answered short-essay questions on three video cases, respectively about history taking, breaking bad news, and decision making. Of 200 participants, 116 completed all three video cases. Students were assessed in three shifts, each using a set of parallel case editions. About half of all available exams were scored independently by two raters using a detailed rating manual derived from the other half. Analyzed were the reliability of the assessment, the inter-rater reliability, and interrelatedness of the three types of video cases and their parallel editions, by computing a generalizability coefficient G. RESULTS: The test score showed a normal distribution. The students performed relatively well on the history taking type of video cases, and relatively poor on decision making and did relatively poor on the understanding ('knows why/when') type of questions. The reliability of the assessment was acceptable (G = 0.66). It can be improved by including up to seven cases in the OSVE. The inter-rater reliability was very good (G = 0.93). The parallel editions of the video cases appeared to be more alike (G = 0.60) than the three case types (G = 0.47). DISCUSSION: The additional value of an OSVE is the differential picture that is obtained about covert cognitive scripts underlying overt communication behavior in different types of consultations, indicated by the differing levels of knowledge, understanding and performance. The validation of the OSVE score requires more research. CONCLUSION AND PRACTICE IMPLICATIONS: A computerized OSVE has been successfully applied with third year undergraduate medical students. The test score meets psychometric criteria, enabling a proper discrimination between adequately and poorly performing students. The high inter-rater reliability indicates that a single rater is permitted.

Communication↗

An in vitro evaluation of the reliability and validity of an electronic pantograph by testing with five different articulators.

STATEMENT OF PROBLEM: The degree of reliability and validity of a new electronic pantographic instrument, the Cadiax Compact, has not been established. PURPOSE: The purpose of this study was to test the reliability and validity of the electronic pantograph in calculating condylar settings for 5 different articulators (Denar D5A, Denar Mark II, Whip Mix 8500, Hanau Modular, and Panadent PCH). MATERIAL AND METHODS: Pantograph sensors were mounted to each articulator with custom-made mounting devices. Border movements were made on each articulator with known condylar settings to produce readings at 3-, 5-, and 10-mm condylotrack distances. The condylar settings investigated included horizontal condylar inclination (HCI), immediate mandibular lateral translation (IMLT), progressive mandibular lateral translation (PMLT), top wall, and rear wall. Reliability was assessed by the relative size of standard deviations. Validity was tested with 1-way analysis of variance and the Tukey HSD Test (alpha=.05). RESULTS: The reliability readings for the condylar settings at the 10-mm condylotrack distance, for the most part, were more consistent than those at the 5-mm distance, which, in turn, were more consistent than those at the 3-mm distance. When there were exceptions to this trend, the differences in standard deviations were very small (0.006 mm for IMLT, 0.03 and 0.12 degrees for HCI). When comparing validity between the condylotrack distances, the smallest deviations, in general, were found at the 10-mm distance, the next smallest at the 5-mm distance, and the largest, at the 3-mm distance for HCI. For PMLT, in general, the 10-mm and 5-mm deviations were smaller than the 3-mm deviation. For IMLT, there was no significant difference between the 3-, 5-, and 10-mm deviations. In analyzing validity between the articulators, the smallest deviations from the preset values, in general, were found with the Denar Mark II. CONCLUSION: The standard deviations for assessing reliability and the mean deviations for assessing validity were both relatively small in comparison to the average values of the condylar determinants. Therefore, the electronic pantograph was determined to be both reliable and valid.

Analysis of Variance↗

Development and reliability of the HAM-D/MADRS interview: an integrated depression symptom rating scale.

The Hamilton Rating Scale for Depression (HAM-D) and the Montgomery-Asberg Depression Rating Scale (MADRS), two widely used depression scales, each have unique advantages and limitations for research. The HAM-D's limited sensitivity and multidimensionality have been criticized, despite the scale's popularity. The MADRS, designed to be sensitive to treatment changes, is briefer and more uniform. A limitation of the MADRS is the lack of a structured interview, which may affect reliability. The HAM-D and the MADRS are often used conjointly as endpoints in depression trials. We designed a hybrid questionnaire that allows administration of MADRS and 31 HAM-D items simultaneously. Seventy mood disorder patients (60 bipolar I, 10 major depressive disorder) were administered the HAM-D/MADRS Interview (HMI) as part of a larger study. Interrater reliability for 50 patients was excellent for the HAM-D and the MADRS (ICC=0.97-0.98). MADRS item reliabilities (ICC=0.86-0.97) were higher than obtained in studies that did not use a structured interview. Reliability coefficients for seven HAM-D(31) 'atypical' symptoms ranged from 0.77 to 0.95. HMI was highly correlated with the Global Clinical Impressions Scale. This is the first study we know of to investigate the reliability of a structured interview of either the MADRS or of the HAM-D(31). The HMI provides an easily administered, reliable method of rating depression severity which may improve consistency and validity of study findings.

Adolescent↗

An assessment tool for brachial plexus regional anesthesia performance: establishing construct validity and reliability.

BACKGROUND AND OBJECTIVES: Technical proficiency in regional anesthesia is often determined subjectively through in-training evaluations. Objective assessment tools improve these evaluations by providing criteria for measurement. However, any evaluation instrument needs to be valid and reliable before it is adopted into a curriculum. The purpose of this study is to determine the validity and reliability of a devised assessment of residents performing an interscalene brachial plexus block (ISB). METHODS: In this prospective study, 10 junior trainees and 10 senior trainees were videotaped performing an ISB. Junior trainees were defined as in their first year of anesthetic training and had performed less than 10 ISBs independently. Senior trainees had completed at least 1 year of anesthesia training and had performed greater than 10 ISBs independently. Two blinded expert raters independently evaluated the performance of the ISB using a checklist and global rating scale. Construct validity was established if the assessments were able to reliably discriminate between different levels of training. RESULTS: Senior trainees performed an ISB significantly better than junior trainees when assessed using the global rating scale (P < .05) and checklist (P < .001). The overall interrater reliability for the global rating scores was excellent (r = 0.85, P < .05) and was good for the checklist scores (r = 0.74, P < .05). CONCLUSIONS: Both assessment modalities were valid, in that they reliably discriminated between different levels of training. Objective measures of technical skills are feasible, timely, and improve the validity and reliability of competency assessments.

Anesthesia, Conduction↗

Internal consistency and test-retest reliability of the Chinese version of the self-report health-related quality of life measure for children and adolescents with epilepsy.

PURPOSE: The aim of this validation study was to evaluate the internal consistency (internal reliability) and test-retest reliability (external reliability) of the Chinese version of the self-report health-related quality of life measure for children and adolescents with epilepsy. METHODS: Children and adolescents with epilepsy between the ages of 8 and 18 years were conveniently sampled in two regional hospitals in Hong Kong. They were requested to complete the 25-item questionnaire twice, with a test-retest interval of 10 to 14 days. Internal consistency and test-retest reliability were measured with Cronbach's alpha coefficient and the intraclass correlation coefficient, respectively. RESULTS: A sample of 50 patients completed the first questionnaire. Internal consistency was adequate on four of five subscales and marginal in the remaining one. Forty-two subjects repeated the questionnaire. The test-retest reliability was acceptable for all five subscales. CONCLUSIONS: The Chinese version of the health-related quality of life measure for children and adolescents with epilepsy demonstrated acceptable internal consistency and test-retest reliability. Further studies are required to study other psychometric properties such as construct validity and factor analysis.

Adolescent↗

Test-retest reliability of colorectal testing questions on the Massachusetts Behavioral Risk Factor Surveillance System (BRFSS).

BACKGROUND: Information on use of colorectal cancer tests, particularly for the purpose of population surveillance, is often obtained through self-report. The Behavioral Risk Factor Surveillance System (BRFSS) is a major source for population-based estimates and is used by health professionals, public health organizations, and researchers to identify and quantify self-reported utilization of screening procedures. METHODS: We provide estimates of the reliability of responses among persons age > or = 50 to questions on the 1999 BRFSS questionnaire addressing two colorectal cancer testing procedures, fecal occult blood test (FOBT), and sigmoidoscopy or colonoscopy (endoscopy), based on responses of 868 persons who responded to a callback survey. RESULTS: We found moderate reliability for questions addressing ever having an FOBT exam, (Kappa [K] = 0.55, 95% confidence interval [95% CI]: 0.49-0.61) and good reliability for questions addressing ever having an endoscopy exam (K = 0.69, 95% CI: 0.65-0.74). Questions addressing the timing of the most recent exam were only slightly less reliable (K = 0.49, 95% CI: 0.43-0.55 and K = 0.62, 95% CI: 0.57-0.67, respectively). We observed comparable reliability across levels of most demographic and risk factor characteristics for both ever having and recency of exam. CONCLUSION: Our results suggest that colorectal cancer testing questions on the BRFSS display a reasonable level of test-retest reliability.

Age Distribution↗