Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Reliability of stabilised commercial dynamometers for measuring hip abduction strength: a pilot study.

BACKGROUND: Reliable quantification of hip abductor strength in a clinical setting is challenging. OBJECTIVES: To examine the intrarater and interrater reliability of three commonly used commercial dynamometers in the measurement of hip abduction. METHODS: Supine gravity minimised measures of unilateral hip abduction strength were recorded in 10 women (mean (SD) age 23.5 (1.9) years) using three different commercially available dynameters. Measurements were repeated over a three day period with a different device used on each day. RESULTS: Intrarater reliability ranged from 0.880 to 0.958 across the three devices, and measures of interrater reliability ranged from 0.899 to 0.948. CONCLUSION: Commercially available dynamometers can be used to quantify hip abduction strength with good to excellent reliability. A previously undescribed method of quantifying hip abduction strength in a clinical setting using readily available instrumentation is presented.

Adult↗

An inexpensive and reliable new haemoglobin colour scale for assessing anaemia.

AIM: To describe a new inexpensive method (the WHO Colour Scale) for estimating haemoglobin concentration from a drop of blood by means of a colour scale, and to compare its reliability with a standard laboratory method of measuring haemoglobin, and its clinical usefulness in field trials. METHODS: The new colour scale method was used to measure haemoglobin concentration in 1213 random venous blood samples from routine work in four laboratories (one each in the UK, South Africa, Thailand, and Switzerland). Limited field trials of the method for assessing clinical usefulness were done in a rural hospital (in South Africa) staffed by nurses, at two blood donor sessions (one each in South Africa and Thailand), and by nonlaboratory personnel in malaria clinics (in Thailand), following training and a short practice session. RESULTS: In the laboratory based comparability study the presence of anaemia was reliably detected using the new method with 91% sensitivity and 86% specificity. Clinically relevant levels of anaemia (mild to moderate, pronounced, and severe) were graded and serious anaemia (< 8 g/dl) was identified with an efficiency of 89%. The clinical trials showed the ease and reliability with which the colour scale could be used by non-laboratory persons after brief training. The blood donor trials showed it to be at least as reliable as the copper sulphate method with the advantage of being more convenient. CONCLUSIONS: The preliminary studies have shown that the WHO Colour Scale is a reliable screening method for detecting anaemia, especially for diagnosing serious anaemia. Following a brief training session health workers found it simple to use and, at a cost of about 1/10th that for traditional photometric analysis, it should be of value in "countries in need" for primary health centres, obstetrical management, paediatric clinics, tropical disease control programmes, blood transfusion donor selection, as well as for industrial health checks and epidemiological studies.

Anemia↗

Reliability and validity of the short form of the child health questionnaire for parents (CHQ-PF28) in large random school based and general population samples.

STUDY OBJECTIVES: This study assessed the feasibility, reliability, and validity of the 28 item short child health questionnaire parent form (CHQ-PF28) containing the same 13 scales, but only a subset of the items in the widely used 50 item CHQ-PF50. DESIGN: Questionnaires were sent to a random regional sample of 2040 parents of schoolchildren (4-13 years); in a random subgroup test-retest reliability was assessed (n = 234). Additionally, the study assessed CHQ-PF28 score distributions and internal consistencies in a nationwide general population sample of (parents of) children aged 4-11 (n = 2474) from Statistics Netherlands. MAIN RESULTS: Response was 70%. In the school and general population samples seven scales showed ceiling effects. Both CHQ summary measures and one multi-item scale showed adequate internal consistency in both samples (Cronbach's alpha>0.70). One summary measure and one scale showed excellent test-retest reliability (intraclass correlation coefficient >0.70); seven scales showed moderate test-retest reliability (intraclass correlation coefficient 0.50-0.70). The CHQ could discriminate between a subgroup with no parent reported chronic conditions (n = 954) and subgroups with asthma (n = 134), frequent headaches (n = 42), and with problems with hearing (n = 38) (Cohen's effect sizes 0.12-0.92; p<0.05 for 39 of 42 comparisons). CONCLUSIONS: This study showed that the CHQ-PF28 resulted in score distributions, and discriminative validity that are comparable to its longer counterpart, but that the internal consistency of most individual scales was low. In community health applications, the CHQ-PF28 may be an acceptable alternative for the longer CHQ-PF50 if the summary measures suffice and reliable estimates of each separate CHQ scale are not required.

Adolescent↗

Fatigue and daytime sleepiness rating scales in myotonic dystrophy: a study of reliability.

OBJECTIVES: To assess the reliability of the Epworth Sleepiness Scale (ESS), Daytime Sleepiness Scale (DSS), Chalder Fatigue Scale (CFS), and Krupp's Fatigue Severity Scale (KFSS) in patients with myotonic dystrophy type 1 (DM1). METHODS: In total, 27 patients with DM1 were administered the questionnaires on two occasions, with a 2 week interval. Internal consistency and test retest reliability were measured using intraclass correlation coefficients (ICCs), and Cronbach's alpha, Cohen's kappa, and Goodman-Kruskal's gamma coefficients. RESULTS: Internal consistency of the CFS and KFSS were adequate (alpha > 0.70) but that of the ESS was weak (alpha = 0.24). Both daytime sleepiness and fatigue rating scales showed significant test retest reliability. Test retest reliability for individual items revealed inconsistencies for some ESS and CFS items. CONCLUSIONS: Reliability of the CFS, DSS, and KFSS was high, allowing their use for individual patients with DM1, but that of the ESS was lower, rendering its current usage in DM1 questionable. Fatigue rating scales such as the KFSS, which are based on the behavioural consequences of fatigue, may constitute a more accurate and comprehensive measure of fatigue severity in the DM1 population.

Adult↗

Psychometric properties of the Need for Recovery after work scale: test-retest reliability and sensitivity to detect change.

BACKGROUND: Monitoring worker health and evaluating occupational healthcare interventions requires sensitive instruments that are reliable over time. The Need for Recovery scale (NFR), which quantifies workers' difficulties in recovering from work related exertions, may be a relevant instrument in this respect. OBJECTIVES: To examine (1) the NFR's test-retest reliability and (2) the NFR's sensitivity to detect the effect of a fatigue inducing change, namely an increase in working hours. METHODS: Two year longitudinal data of 526 truck drivers and 144 nurses were used. Two week, one year, and two year test-retest reliability was examined in both stable and unstable work environments by calculating intraclass correlations (ICCs). Work environmental (in)stability was quantified by four events that might have occurred during the follow up period: (1) a reorganisation or merge (0 = yes, 1 = no), (2) a change of supervisor or management (0 = yes, 1 = no), (3) a change in working hours or work schedules (0 = yes, 1 = no), and (4) a change in work activities, position, or duties (0 = yes, 1 = no). The four scores constituted a work (in)stability index ranging from 0 to 4. The NFR's sensitivity to detect the effect of the increase in working hours was assessed indirectly by comparing it with an alternative scale, namely the Checklist Individual Strength. RESULTS: Test-retest reliability over a two year interval was good to excellent when applied in stable work environments (ICCs 0.68 to 0.80) but, as expected, poor to fair when applied in unstable work environments (ICCs 0.30 to 0.55). The NFR was sensitive in detecting an increase in work related fatigue due to the increase in working hours (effect size 0.40). CONCLUSIONS: The NFR's test-retest reliability and sensitivity to detect change are favourable. This implicates that the NFR may form a valuable part of health surveys and may be a useful tool for evaluating occupational healthcare interventions.

Adult↗

An analysis of the reliability of self reported work histories from a cohort of workers exposed to polychlorinated biphenyls.

An investigation was conducted to examine the reliability (reproducibility) of self reported occupational histories obtained from a cohort of 326 capacitor manufacturing workers who had participated in an epidemiological study relating health abnormalities to exposure to polychlorinated biphenyls (PCBs). For a subsample of the cohort (n = 164) in which occupational histories were obtained twice (in 1976 and 1979), reliability of cumulative exposure to PCBs ranged from 93.6% for the early PCB period (1947-70) to 95.7% for the late PCB period (1971-6). These respective reliabilities were lower, however, for workers who changed jobs often. Workers above the median value of a weighted job change index had early and late reliabilities of 89.9% and 83.6% respectively. Reliability is a relevant factor when calculating power or sample size during the planning stage of epidemiological studies, for interpretation or adjustment of estimates in the analysis stage, or for determination of study feasibility.

Adult↗

The reliability of a structured examination protocol and self administered vaginal swabs: a pilot study of gynaecological outpatients in Goa, India.

OBJECTIVES: Low participation rates for gynaecological examination and low reliability of clinical reporting of gynaecological examination findings are problems in community studies of gynaecological morbidity in India. This pilot study aimed to describe the reliability of a new examination protocol for recording the findings of gynaecological examination and the reliability and acceptability of the use of self administered vaginal swabs for the diagnosis of reproductive tract infections. METHOD: 75 women attending a gynaecology outpatient clinic were purposively sampled. Each woman was examined by two gynaecologists independently who recorded findings on the new examination protocol. Two swabs were collected from each woman, one by the gynaecologist and one by the woman. Swabs were smeared on separate slides which were stained and read for bacterial vaginosis and candidiasis by laboratory technicians blind to the mode of collection of the slides. RESULTS: The study showed a high inter-rater reliability for most of the items of the examination protocol. The interslide agreement for the diagnosis of the two RTIs was high. One third of women preferred the self administered swab. CONCLUSIONS: The examination protocol is a reliable method of recording gynaecological examination findings, and self administered swabs a useful way of obtaining vaginal specimens from women who did not wish to undergo gynaecological examinations in studies in the Indian setting.

Adolescent↗

Human brain: reliability and reproducibility of pulsed arterial spin-labeling perfusion MR imaging.

The Committee of Human Research of the University of California San Francisco approved this study, and all volunteers provided written informed consent. The goal of this study was to prospectively determine the global and regional reliability and reproducibility of noninvasive brain perfusion measurements obtained with different pulsed arterial spin-labeling (ASL) magnetic resonance (MR) imaging methods and to determine the extent to which within-subject variability and random noise limit reliability and reproducibility. Thirteen healthy volunteers were examined twice within 2 hours. The pulsed ASL methods compared in this study differ mainly with regard to magnetization transfer and eddy current effects. There were two main results: (a) Pulsed ASL MR imaging consistently had high measurement reliability (intraclass correlation coefficients greater than 0.75) and reproducibility (coefficients of variation less than 8.5%), and (b) random noise rather than within-subject variability limited reliability and reproducibility. It was concluded that low signal-to-noise ratios substantially limit the reliability and reproducibility of perfusion measurements.

Adult↗

Reliability of the Karolinska Psychodynamic Profile (KAPP) among patients with and without psychoactive substance abuse disorders.

BACKGROUND: The Karolinska Psychodynamic Profile (KAPP) is a rating instrument, based on psychoanalytic theory, that assesses different aspects of character from clinical interviews. The aim of the present study was to examine interrater reliability of the KAPP in a sample of patients with and without psychoactive substance abuse disorders, using interviewers and a reliability judge who had not been trained by the developers of the instrument. METHODS: The sample comprised 47 consecutive patients with and without psychoactive substance abuse disorders, who were referred to an outpatient psychotherapy unit specializing in the treatment of substance abuse and dependence. The two interviewers and the reliability judge had not been trained by the developers of the KAPP, and they worked outside the setting where it was constructed. RESULTS: The intraclass correlations were satisfactory for the total sample (mean 0.84, median 0.89, range 0.62-0.95), as well as for various sub-samples, such as males and females, and patients with and without substance abuse disorders. CONCLUSIONS: The results show that interviewers and a reliability judge who had not been trained by the developers of the KAPP can attain high interrater reliability in a sample of patients with substance abuse disorders. Some recommendations for conducting KAPP interviews are given.

Adult↗

Reliability study on the Japanese version of the Clinician's Interview-Based Impression of Change. Analysis of subscale items and 'clinician's impression'.

BACKGROUND/AIMS: The Japanese version of the Clinician's Interview-Based Impression of Change plus Caregiver Input (CIBIC-plus J) consists of 3 subscales: Disability Assessment of Dementia Scale (DAD), Behavioral Pathology in Alzheimer's Disease Rating Scale (Behave-AD), and Mental Function Impairment Scale (MENFIS), as well as the Clinician's Global Impression of Change (CGIC). While the interrater reliability of CGIC has already been reported, that of the 3 subscales has not. The aim of the present report was to examine the reliabilities of the subscale items and investigate their relationships with CGIC. METHODS: Eleven raters who were clinical physicians watched videotapes of 20 patients with Alzheimer's disease, completed the CIBIC-plus J assessment form, and assigned a CGIC score to the patients. Reliability was assessed using the kappa coefficient. RESULTS: The kappa coefficient of the subscale items was in most instances higher than that of CGIC (0.453) and substantial reliability was observed. The Spearman rank correlation that was calculated between CGIC and the total score change of items was very high for MENFIS (0.990) and DAD (0.910), and moderate for Behave-AD items (0.576). The incidence of comments by the raters was highest for MENFIS (89%), followed by DAD (70%). The incidence was low for Behave-AD items (48%). CONCLUSION: Based on the results, it is concluded that DAD, Behave-AD, and MENFIS are necessary constituents of CIBIC-plus J, and indispensable for the reliability of CGIC.

Aged↗

Modified National Institutes of Health Stroke Scale for use in stroke clinical trials: prospective reliability and validity.

BACKGROUND AND PURPOSE: To prospectively evaluate the reliability and validity of this previously developed stroke scale in an independently collected cohort. The National Institutes of Health Stroke Scale (NIHSS) has been criticized for its complexity and variability. Prior formal clinimetric analyses were used to obtain a modified version of NIHSS (mNIHSS), which retrospectively demonstrated improved reliability and validity. We sought to prospectively measure the reliability and validity of the mNIHSS. METHODS: Forty-five patients with a history of stroke or intracerebral hemorrhage were evaluated at the University of California, San Diego, Stroke Center from September 2000 through March 2001. Each patient was tested by 2 NIHSS-certified neurologists using the NIHSS, mNIHSS, Barthel Index, and Modified Rankin scales. RESULTS: There were a large percentage of high kappa values using the mNIHSS. Only 10 (66.67%) of 15 NIHSS kappa scores showed excellent agreement, whereas 10 (90.91%) of 11 mNIHSS kappa scores showed excellent agreement. As predicted, the mNIHSS was more reliable than the NIHSS because of the exclusion of items with low kappa values. With the use of correlation coefficient analysis, the mNIHSS was as valid as the NIHSS. CONCLUSIONS: This prospective study found high reliability and continued validity by using a previously developed mNIHSS. Items found to have low kappa values were consistent with the previously derived retrospective mNIHSS. The resulting mNIHSS scale has much higher kappa values. The mNIHSS showed improved agreement between examiners and was also easier to administer, having fewer and simpler items. Further prospective evaluation should assess whether the mNIHSS could be used in lieu of the NIHSS.

Academic Medical Centers↗

Diagnostic reliability of the percutaneous ultrasonic Doppler technique for vertebral arterial occlusive diseases.

There is little data on the diagnostic reliability of the ultrasonic Doppler technique for vertebral arterial occlusive lesions. Percutaneous vertebral Doppler examination and the vertebral angiograms were compared to determine the diagnostic reliability of this technique in 64 vertebral arteries of 53 patients with cerebrovascular disease. The percutaneous vertebral Doppler findings were quantitatively analyzed using a sound spectrograph and were classified into three types: no flow signal type, poor flow type and normal flow type. In nine patients with the no flow type the angiograms revealed vertebral occlusion or a missing vertebral artery in six, giving a diagnostic reliability of 67%. In 17 patients with poor flow type the angiograms revealed vertebral occlusion or a missing vertebral artery in five, terminal narrowing of the artery in nine, and hypoplasia in two giving a diagnostic reliability of 94%. For all vertebral arteries examined with this technique, including normal ones, the diagnostic reliability was 92% (59/64). Percutaneous vertebral Doppler examination has clinical usefulness as a screening test for occlusive vertebral arterial diseases.

Adult↗

Ion channel stochasticity may be critical in determining the reliability and precision of spike timing.

The firing reliability and precision of an isopotential membrane patch consisting of a realistically large number of ion channels is investigated using a stochastic Hodgkin-Huxley (HH) model. In sharp contrast to the deterministic HH model, the biophysically inspired stochastic model reproduces qualitatively the different reliability and precision characteristics of spike firing in response to DC and fluctuating current input in neocortical neurons, as reported by Mainen & Sejnowski (1995). For DC inputs, spike timing is highly unreliable; the reliability and precision are significantly increased for fluctuating current input. This behavior is critically determined by the relatively small number of excitable channels that are opened near threshold for spike firing rather than by the total number of channels that exist in the membrane patch. Channel fluctuations, together with the inherent bistability in the HH equations, give rise to three additional experimentally observed phenomena: subthreshold oscillations in the membrane voltage for DC input, "spontaneous" spikes for subthreshold inputs, and "missing" spikes for suprathreshold inputs. We suggest that the noise inherent in the operation of ion channels enables neurons to act as "smart" encoders. Slowly varying, uncorrelated inputs are coded with low reliability and accuracy and, hence, the information about such inputs is encoded almost exclusively by the spike rate. On the other hand, correlated presynaptic activity produces sharp fluctuations in the input to the postsynaptic cell, which are then encoded with high reliability and accuracy. In this case, information about the input exists in the exact timing of the spikes. We conclude that channel stochasticity should be considered in realistic models of neurons.

Action Potentials↗

The Richmond Agitation-Sedation Scale: validity and reliability in adult intensive care unit patients.

Sedative medications are widely used in intensive care unit (ICU) patients. Structured assessment of sedation and agitation is useful to titrate sedative medications and to evaluate agitated behavior, yet existing sedation scales have limitations. We measured inter-rater reliability and validity of a new 10-level (+4 "combative" to -5 "unarousable") scale, the Richmond Agitation-Sedation Scale (RASS), in two phases. In phase 1, we demonstrated excellent (r = 0.956, lower 90% confidence limit = 0.948; kappa = 0.73, 95% confidence interval = 0.71, 0.75) inter-rater reliability among five investigators (two physicians, two nurses, and one pharmacist) in adult ICU patient encounters (n = 192). Robust inter-rater reliability (r = 0.922-0.983) (kappa = 0.64-0.82) was demonstrated for patients from medical, surgical, cardiac surgery, coronary, and neuroscience ICUs, patients with and without mechanical ventilation, and patients with and without sedative medications. In validity testing, RASS correlated highly (r = 0.93) with a visual analog scale anchored by "combative" and "unresponsive," including all patient subgroups (r = 0.84-0.98). In the second phase, after implementation of RASS in our medical ICU, inter-rater reliability between a nurse educator and 27 RASS-trained bedside nurses in 101 patient encounters was high (r = 0.964, lower 90% confidence limit = 0.950; kappa = 0.80, 95% confidence interval = 0.69, 0.90) and very good for all subgroups (r = 0.773-0.970, kappa = 0.66-0.89). Correlations between RASS and the Ramsay sedation scale (r = -0.78) and the Sedation Agitation Scale (r = 0.78) confirmed validity. Our nurses described RASS as logical, easy to administer, and readily recalled. RASS has high reliability and validity in medical and surgical, ventilated and nonventilated, and sedated and nonsedated adult ICU patients.

Adult↗

The Brown Assessment of Beliefs Scale: reliability and validity.

OBJECTIVE: The authors developed and evaluated the reliability and validity of the Brown Assessment of Beliefs Scale, a clinician-administered seven-item scale designed to assess delusions across a wide range of psychiatric disorders. METHOD: The authors developed the scale after reviewing the literature on the assessment of delusions. Four raters administered the scale to 20 patients with obsessive-compulsive disorder (OCD), 20 patients with body dysmorphic disorder, and 10 patients with mood disorder with psychotic features. Audiotaped interviews of scale administration conducted by one rater were independently scored by the other raters to evaluate interrater reliability. The scale was administered to 27 patients twice to determine test-retest reliability. Other insight instruments as well as scales that assess symptom severity were administered to assess convergent and discriminant validity. Sensitivity to change was assessed in a multicenter treatment study of sertraline for OCD. RESULTS: Interrater and test-retest reliability for the total score and individual item scores was excellent, with a high degree of internal consistency. One factor was obtained that accounted for 56% of the variance. Scores on the Brown Assessment of Beliefs Scale were not correlated with symptom severity but were correlated with other measures of insight. The scale was sensitive to change in insight in OCD but was not identical to improvement in severity. CONCLUSIONS: The Brown Assessment of Beliefs Scale is a reliable and valid instrument for assessing delusionality in a number of psychiatric disorders. This scale may help clarify whether delusional and nondelusional variants of disorders constitute the same disorder as well as whether delusionality affects treatment outcome and prognosis.

1-Naphthylamine↗

Reliability of self-reports about sexual risk behavior for HIV among homeless men with severe mental illness.

The reliability of self-reports of sexual behaviors related to HIV transmission was examined in a study of homeless men with severe mental illness. Thirty-nine patients of a New York City shelter psychiatric program were interviewed about their sexual behaviors in the past six months. The same interview was administered twice, with a one- to two-week interval between interviews. Test-retest reliability was assessed using kappa and intraclass correlation coefficients. Reliability estimates ranged from.49 to.93 for overall sexual activity, number of partners, and specific behaviors other than receptive anal sex. Reliability was lower for condom use. The authors conclude that reliable self-reports about sexual behavior can be obtained from homeless men with severe mental illness.

Adult↗

Reliability of a craniomandibular index.

The Craniomandibular Index (CMI) was developed to provide a standardized measure of severity of problems in mandibular movement, TMJ noise, and muscle and joint tenderness for use in epidemiological and clinical outcome studies. The instrument was designed to have clearly defined objective criteria, simple clinical methods, and ease in scoring; it is divided into the Dysfunction Index and the Palpation Index. Inter-rater reliability (three raters) and intra-rater reliability (19 patients examined twice by one rater) were tested to determine whether the instrument has operational definitions sufficiently precise to allow for consistency in use between different raters and with one rater over time. Intraclass Correlation Coefficient for inter-rater reliability was 0.84 for the Dysfunction Index, 0.87 for the Palpation Index, and 0.95 for the CMI. Correlation for intra-rater reliability was 0.92 for the Dysfunction Index, 0.86 for the Palpation Index, and 0.96 for the CMI. These results support the reliability of the CMI for use in epidemiological and clinical studies. Users are cautioned about the subjectivity of numerous items within the CMI and the strict methodological guidelines that must be followed in order to assure accuracy and reproducibility of results.

Humans↗

Can quality of movement be measured? Rasch analysis and inter-rater reliability of the Motor Evaluation Scale for Upper Extremity in Stroke Patients (MESUPES).

OBJECTIVE: Clinical scales evaluating arm function after stroke are weak at detecting quality of movement. Therefore a new scale, the Motor Evaluation Scale for Upper Extremity in Stroke Patients (MESUPES), was developed, comprising 22 items pertaining to arm and hand performance. The scale was investigated for validity and unidimensionality using the Rasch measurement model, and for inter-rater reliability. SETTING: Twelve hospitals and rehabilitation centres in Belgium, Germany and Switzerland. PATIENTS: There were 396 patients (average age 63.38+/-12.89 years) in the Rasch study and 56 patients (average age 65.68+/-12.75 years) in the reliability study. MAIN MEASURES: The scale was examined on its fit to the Rasch model, thereby evaluating the scale's unidimensionality and validity. Differential item functioning was performed to test the stability of item hierarchy on several variables. Inter-rater reliability was examined with kappa values, weighted percentage agreement and intraclass correlation coefficients (ICC). RESULTS: Based on Rasch analysis, five items were removed. The MESUPES was divided in two tests: the MESUPES-arm test (8 items) and MESUPES-hand test (9 items). Both scales fitted the Rasch model. All items were stable among the subgroups of the sample. ICCs were 0.95 (95% confidence interval (CI) 0.91 -0.97) and 0.97 (95% CI 0.95-0.98) for the total score on arm and hand test respectively. The scale was also reliable at item level (weighted kappa 0.62 -0.79, weighted percentage agreement 85.71 -98.21). CONCLUSION: The MESUPES-arm and MESUPES-hand meet the statistical properties of reliability, validity and unidimensionality. Both tests provide a useful clinical and research tool to qualitatively evaluate arm and hand function during recovery after stroke.

Aged↗