Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

[Reliability of basic geriatric assessment].

This study investigates the Geriatric Basis Assessment (GBA) in terms of its reliability. Data from 1037 patients were collected. The reliability was estimated relating to the lambda 2 coefficient. It is necessary to define the items in different categories: the first variable means valuation 1 of each item and not 2, 3, 4; the second variable means valuation 1, 2 against 3, 4; the third variable means the valuation 1, 2, 3 and not 4. The table shows only little difference concerning the lambda 2 coefficients. In conclusion, 80% of the variability of the GBA items can be explained by differences in the patients themselves, while 20% is due to the inaccurate assessment system. For 343 patients, data for both Barthel index and GBA were available. As presumed, correlations between Barthel and connected GBA items were observed. However, the correlations were too weak to predict the Barthel scores from the corresponding GBA item accurately enough. The Barthel index appears to include similar, but not exactly the same aspects as the GBA. The reliability of the Barthel index (lambda 2 = 0.89 for the first variable) is slightly higher compared to the GBA but it is not suitable as a criterion of validity. Both the validity of the GBA and the Barthel index can not be determined lacking an external measure. As an example, a suitable criterion of validity could be the reintegration into the familiar surroundings preceding the hospital stay. When developing the GBA, it was not assumed that geriatric patients could be correctly diagnosed on the basis of an overall score alone or to allocate them to adequate care using that score as a sole indicator. Crucial for these purposes is the test profile as a whole, including the impairments, disabilities handicaps, and last but not least the diseases of the individual patient. Furthermore, the depiction of the GBA profile at admission and discharge allows one to identify those items, on which therapy has a significant influence and those which remain more or less stable. As presumed, items with minor initial deficits (e.g., motivation, eyesight, hearing, depression, capability of verbal expression, situative adaptability, understanding) showed only small differences between admission and discharge. On the other hand, items strongly influenced by geriatric treatment were, e.g., mobility (walking, transfer), functions of internal medicine, and domestic care. Prognostically significant are those items which are crucial for reintegration and describe a deficiency but cannot be altered reliably. Such items are the person, to whom the patient relates most closely, situative adaptability, motivation, orientation, capability of verbal expression, and possibly depression. All of these parameters are more difficult to influence than the activities of daily living assessed by the Barthel index. Further investigations should clarify whether the GBA can be a reliable tool for allocating a patient to adequate care. However, the requirement for such a criterion of validity is that this allocation is truly optimal for the patient.

Activities of Daily Living↗

Reliability of peak VO(2) and maximal cardiac output assessed using thoracic bioimpedance in children.

The purpose of this study was to evaluate the reliability of a thoracic electrical bioimpedance based device (PhysioFlow) for the determination of cardiac output and stroke volume during exercise at peak oxygen uptake (peak VO(2) in children. The reliability of peak VO(2) is also reported. Eleven boys and nine girls aged 10-11 years completed a cycle ergometer test to voluntary exhaustion on three occasions each 1 week apart. Peak VO(2) was determined and cardiac output and stroke volume at peak VO(2) were measured using a thoracic bioelectrical impedance device (PhysioFlow). The reliability of peak VO(2) cardiac output and stroke volume were determined initially from pairwise comparisons and subsequently across all three trials analysed together through calculation of typical error and intraclass correlation. The pairwise comparisons revealed no consistent bias across tests for all three measures and there was no evidence of non-uniform errors (heteroscedasticity). When three trials were analysed together typical error expressed as a coefficient of variation was 4.1% for peak VO(2) 9.3% for cardiac output and 9.3% for stroke volume. Results analysed by sex revealed no consistent differences. The PhysioFlow method allows non-invasive, beat-to-beat determination of cardiac output and stroke volume which is feasible for measurements during maximal exercise in children. The reliability of the PhysioFlow falls between that demonstrated for Doppler echocardiography (5%) and CO(2) rebreathing (12%) at maximal exercise but combines the significant advantages of portability, lower expense and requires less technical expertise to obtain reliable results.

Cardiac Output↗

Accuracy and reliability of the ParvoMedics TrueOne 2400 and MedGraphics VO2000 metabolic systems.

This study examined the accuracy and reliability of the MedGraphics VO2000 (VO2000) portable metabolic system and the ParvoMedics TrueOne 2400 (TrueOne 2400) metabolic cart against the criterion Douglas bag (DB) method. Ten healthy males (age 20 +/- 1.7 years) had their gas exchange variables measured at rest and during cycling at 50, 100, 150, 200, and 250 W. Each stage was 10-12 min. For half of the stage gas exchange was measured with the DB and TrueOne 2400 simultaneously and for the other half of the stage gas exchange was measured with the VO2000. The testing was performed on two separate days and the order in which the equipment was used in each stage was randomized. Reliability between days for V (E) (CV 7.3-8.8%) was similar among devices, however, for VO2, and VCO2 the VO2000 (CV 14.2-15.8%) was less reliable compared to the DB (CV 5.3-6.0%) and TrueOne 2400 (CV 4.7-5.7%). The TrueOne 2400 was not significantly different from the DB at rest or any work rate for V (E), VO2, or VCO2 (P > or = 0.05). The VO2000 was significantly different from the DB for V (E) at 50-100 W, VO2 at rest and 100-250 W, and VCO2 at rest and 200-250 W (all, P < 0.05). The TrueOne 2400 provides accurate and reliable results for the measurement of gas exchange variables. The VO2000 portable metabolic system was less reliable for measuring VO2 and VCO2 and generally overestimates VO2 at most cycling work rates. Further research is needed to confirm the results found with the VO2000.

Adult↗

Reliability of pain threshold measurement in young adults.

The objective was to examine reliability of pressure and thermal (cold) pain threshold assessment in persons less than 25 years of age, using intra-class correlation (ICC) and coefficients of repeatability and variability. We measured thresholds to pain from pressure algometry and ice placed at the hand and head in 10 healthy volunteers aged 18-25. Intra-rater reliability was examined with ICC. Coefficients of repeatability (CR) and variability (CV) were estimated. Reliability of repeat assessments was high as assessed by ICC, although coefficients of repeatability and variation indicated considerable inter-individual variation in repeat measurements. Pressure algometry and strategically placed ice appear to be reliable techniques for assessing pain processing in young adults. Reliability studies employing ICC may benefit from complementary estimation of CR and CV.

Adolescent↗

Heritability and reliability of P300, P50 and duration mismatch negativity.

BACKGROUND: Event-related potentials (ERPs) have been suggested as possible endophenotypes of schizophrenia. We investigated the test-retest reliabilities and heritabilities of three ERP components in healthy monozygotic and dizygotic twin pairs. METHODS: ERP components (P300, P50 and MMN) were recorded using a 19-channel electroencephalogram (EEG) in 40 healthy monozygotic twin pairs, 19 of them on two separate occasions, and 30 dizygotic twin pairs. Zygosity was determined using DNA genotyping. RESULTS: High reliabilities were found for the P300 amplitude and its latency, MMN amplitude, and P50 suppression ratio components. ICC=0.86 and 0.88 for the P300 amplitude and P300 latency respectively. Reliability of MMN peak amplitude and mean amplitude were 0.67 and 0.66 respectively. P50 T/C ratio reliability was 0.66. Model fitting analyses indicated a substantial heritability or familial component of variance for these ERP measures. Heritability estimates were 63 and 68% for MMN peak amplitude and mean amplitude respectively. For P50 T/C ratio, 68% heritability was estimated. P300 amplitude heritability was estimated at 69%, and while a significant familiality effect was found for P300 latency there was insufficient power to distinguish between shared environment and genetic factors. CONCLUSIONS: The high reliability and heritability of the P300 amplitude, MMN amplitude, and P50 suppression ratio components supports their use as candidate endophenotypes for psychiatric research.

Adult↗

Reliability of lumbar paravertebral EMG assessment in chronic low back pain.

The reliability of lumbar paravertebral EMG assessment was investigated in a sample of 70 patients with chronic low back pain, (CLBP). Dual-site EMG monitoring was employed during both static postures and movements. Flexion and rotation indices were divided to assess the reliability of patterning of paravertebral EMG during movement. Within-session reliabilities computed for the full sample ranged from 0.66 to 0.97, and between-session reliabilities, computed on a subset of 29 patients retested after varying intervals, ranged from 0.26 to 0.92. Average EMG levels, flexion, and rotation indices showed no statistically significant differences between surgical (n = 40) and nonsurgical patients (n = 30), although EMG variability was consistently greater for surgical patients across the postures and movements. These results indicate that lumbar paravertebral EMG can be reliably measured and therefore has potential utility as an assessment and treatment variable in CLBP.

Adult↗

Test-retest reliability of assessing psychiatrically ill patients in a multi-center design.

In a test-retest reliability study involving 25 psychiatric patients and 5 professional raters we demonstrate that research clinicians from collaborating institutions are able to achieve good reliability for most areas of the SADS and RDC when assessing psychiatrically ill patients under interview conditions that provide even less data than ideally obtained in the practice of clinical research. We expect greater reliability in the actual use of the SADS/RDC on most items and diagnoses since the SADS is intended to be used in conjunction with information obtained from relatives, friends, and treatment staff to confirm and clarify the judgements made by the raters on the patient interviews. Moreover, we are reassured that the diagnosis of schizo-affective disorders and schizophrenia is protected from the item unreliability found with specific delusions and hallucinations. Similarly, the difficulties in determining the episodic and chronic nature of the present episode does not substantially interfere with making an RDC diagnosis of the current condition. A complex diagnostic interview system such as the SADS and RDC requires multiple complementary techniques to determine reliability. We find that establishing explicit procedures for raters to discuss and categorize the reasons for their disagreements on individual items and diagnoses provides valuable data for understanding reliability problems. This has helped us to identify specific areas of the interview and criteria that require further clarification and more intensive rater training to improve ratings made by interviewers.

Affective Disorders, Psychotic↗

Reliability of reaction time measurements in brain-damaged patients.

The higher sensitivity of chronometric analysis leads to an increase of their use in the evaluation of brain-damaged patients. However, patients have usually higher reaction time (RT) variability, suggesting a possible decrease in the reliability of RT measurement. This reliability has never been evaluated in patients. This study assessed the reliability of RT estimates in 2 populations of brain-damaged patients and controls, using the Kendall coefficient of concordance W. Concordance coefficients were computed for simple detection and choice tests. W values were good and significant. This study suggests that RT measured in simple detection and choice tasks represent a reliable measurement and supports the idea that the higher sensitivity associated with the use of chronometric analysis is not obtained at the cost of a lower reliability. Moreover, these results provide information for the elaboration of RT tests.

Adult↗

Hamilton Anxiety Rating Scale Interview guide: joint interview and test-retest methods for interrater reliability.

The Hamilton Anxiety Rating Scale (HARS) is the most widely used semistructured assessment scale in treatment outcome studies of anxiety. Interrater reliability coefficients for the HARS have been previously reported. However, differences in the way clinicians assess symptom severity may reduce reliability. A structured interview guide--The Hamilton Anxiety Rating Scale Interview Guide (HARS-IG)--was developed to standardize clinical probe questions and to minimize interrater variance. Joint-interview and test-retest methods of interrater reliability assessment were used in a group of 30 inpatients. Intraclass coefficient calculations revealed improved interrater agreement with the HARS-IG versus the HARS. The findings of this study demonstrate that the HARS-IG is a more reliable assessment instrument than the semistructured HARS and that it meets established standards of reliability assessment.

Adult↗

A new measure of quality of life in depression: testing the reliability and construct validity of the QLDS.

Our previous paper described the development of a new quality of life scale for use with people suffering from depression; the Quality of Life in Depression Scale (QLDS). This paper reports on the testing of the scale for reliability and construct validity. Reliability was assessed by giving the questionnaire to the same set of patients on two occasions 2 weeks apart. This test-retest technique yielded a correlation of 0.94, with high internal consistency at both time 1 and time 2. A test of split-half reliability also indicated very high reliability. Construct validity was measured by comparing scores on the QLDS with those on an established scale of well-being from the same group of patients. The results gave a correlation between the two measures of 0.79, giving a satisfactory validity. It is concluded that the QLDS is a reliable and valid measure which is easy to use and acceptable to patients. Further tests of discriminative, concurrent and criterion validity are planned.

Depression↗

Reliability of the Spanish version of the Nottingham Health Profile in patients with stable end-stage renal disease.

OBJECTIVE: Since reproducibility of results is a basic prerequisite of health status measures for its use in prospective and evaluative studies, the reliability of the Spanish version of the Nottingham Health Profile (NHP), a multi-dimensional perceived health status measure, was assessed in a sample of stable end-stage renal disease (ESRD) patients. METHODS: The NHP was administered on two occasions four weeks apart to a group of hospital hemodialysis program patients who were clinically stable according to their physicians. Correlations of scores and agreement of first and second administrations were assessed together with internal consistency. Afterwards, analyses were repeated taking into account the time (before, during or after the dialysis) and the method of administration (self vs interviewer), and the interviewer. RESULTS: Spearman correlation coefficients (rs) between responses to the first and to the second administration were > 0.6 for all of the six dimensions of the NHP (range = 0.69-0.85) and in every sub-group analyzed (P < 0.01). Agreement percent (AP) between items was > 0.4 (0.48-0.65). Internal consistency was 0.91 for the whole profile and > 0.5 (0.58-0.86) when analyzed by individual dimensions. Reliability did not vary significantly with the time nor the method of administration (self or interviewer). CONCLUSIONS: Overall, results suggest that the Spanish version of NHP is sufficiently reliable to be used in ESRD patients. While a higher reliability would have been achieved by a shorter retest period, the study provides a realistic approximation to the reliability of the questionnaire in actual research and clinical applications.

Female↗

A reliable and valid method for evaluating cardiopulmonary resuscitation training outcomes.

In order to compare the quality of CPR performance after various training methods, training outcome assessment must provide meaningful data and do it in a way that is reliable. Few studies have provided details of their assessment procedures, and even fewer report on whether the measures to evaluate performance are reliable (yielding information consistently over multiple trials), or valid (measuring the outcome intended). Few studies have attempted to replicate assessment methods used by other authors. Conventional skill sheets have not been shown to assess compressions and ventilations reliably and validly. When using an instrumented manikin, skill checklists can be simplified by eliminating qualitative assessment of compressions and ventilations. Using a sample of 171 CPR trainees rated by trained evaluators, we provide details of agreement between two evaluators and use an established statistic (Cronbach's alpha) to assess the reliability of a 14-item simplified CPR checklist. The level of agreement between two raters was high (Pearson product-moment correlation = 0.87) as was the reliability estimate obtained by Cronbach's alpha (0.89). As criterion-related evidence of the validity of the CPR checklist to assess CPR performance, a correlation with a five-point subjective overall rating of CPR was estimated (Spearman correlation = 0.92). We urge standardized reporting of CPR training outcomes in order to achieve comparability across studies.

Adult↗

Reliability of spontaneous electrodermal activity in the cat as a function of waking and sleep stages.

This study was designed to examine the reliability of spontaneous electrodermal activity (EDA) as a function of the stages of sleep and waking in freely moving cats. Five adult cats were observed during waking--sleep sessions. The results show that: (a) for all stages, the reliability of EDA is slightly higher for amplitude of SSPRs than for frequency; (b) during drowsiness, a maximum of reliability is observed, as is a slight decrease during slow wave sleep; during paradoxical sleep, reliability decreases greatly to below that of the waking level; (c) the reliability of spontaneous EDA appears to be higher in waking cats than that quoted for human subjects. These results are discussed with reference to individual characteristics and state variables.

Analysis of Variance↗

Reliability of spontaneous electrodermal activity in humans as a function of sleep stages.

This study was designed to examine the reliability of spontaneous electrodermal activity (EDA) as a function of the sleep stages in human subjects. Recordings were made from 10 volunteer paid male students during four complete nights. The results show that: (a) reliability of EDA varies as a function of sleep stages for both frequency and amplitude parameters although in a different way: frequency reliability shows a U-shape curve whereas amplitude reliability grows monotically with the depth of sleep and; (b) paradoxical sleep appears to be the most reliable stage for both frequency and amplitude variables. These results are compared to those obtained in waking human subjects and in sleeping cats.

Animals↗

IASP taxonomy of chronic pain syndromes: preliminary assessment of reliability.

Communication and consequently advancement of knowledge in understanding and treatment of chronic pain has been hindered by the absence of a taxonomy of chronic pain syndromes. Recently the IASP Subcommittee on Taxonomy proposed a classification method based on a multiaxial system. In the present study the interjudge reliability of 2 of the 5 axes, body location and presumed etiology are evaluated. Overall, axis I demonstrated good reliability, however, the reliability of several categories contained within this axis were low enough to suggest minor changes to this axis may increase its clinical utility. Axis V was found to have only fair reliability and many of the categories comprising this axis were demonstrated to have reliabilities that are not clinically acceptable. The implications of these results for future development and refinement of the IASP taxonomy are discussed.

Chronic Disease↗

Pressure pain thresholds, clinical assessment, and differential diagnosis: reliability and validity in patients with myogenic pain.

Four studies are presented testing the validity and reliability of pressure pain thresholds (PPTs) and of examination parameters believed to be important in the clinical assessment of sites commonly used for such measures in patient samples. Forty-five patients with a myogenous temporomandibular disorder were examined clinically prior to PPT measures. Criteria for history and examination included functional aspects of the pain, tissue quality of the pain site, and the type of pain elicited from palpation. Control sites within the same muscle and in the contralateral muscle were also examined. PPTs were measured as an index of tenderness using a strain gauge algometer at these sites. The data from the 5 male subjects were excluded from subsequent analyses due to the higher PPT in the males and to their unequal distribution among the various factorial conditions. The first study demonstrated strong validity in PPT measures between patients (using pain sites replicating the patients' pain) and matched controls (n = 11). The PPT was not significantly different between the primary pain site (referred pain and non-referred pain collapsed) and the no-pain control site in the same muscle (n = 16). The PPT was significantly lower at the pain site compared to the no-pain control site in the contralateral muscle (n = 13). The second study indicated adequate reliability in patient samples of the PPT measures. In the third study, the PPT was significantly lower at sites producing referred pain on palpation compared to sites producing localized pain on palpation. The PPT findings from the control sites were inconsistent on this factor. The fourth study presented preliminary evidence that palpable bands and nodular areas in muscle were most commonly associated with muscle regions that produce pain; such muscle findings were not specific, however, for regions that produce pain. Further, the intraexaminer reliability in reassessing these pain sites qualitatively was only fair. Referred pain had a poor association with the pain pattern and physical findings, which may suggest a need to reevaluate part of the theory regarding referred muscle pain. The reliability of PPT measures was better overall than the reliability of the signs and site-specific symptoms, suggesting that pressure pain thresholds may be an important tool in clinical studies of pain. PPT measures demonstrate a high within-subject variability in pain patient subjects as well as non-pain subjects.(ABSTRACT TRUNCATED AT 400 WORDS)

Adult↗

Preclinical evaluation of the reliability of a 50 MeV racetrack microtron.

PURPOSE: A 50 MeV racetrack microtron has been installed and tested at Memorial Sloan-Kettering Cancer Center. It is designed to execute multi-segment conformal therapy automatically under computer control using scanned X ray and electron beams from 10 to 50 MeV. Prior to acceptance of the machine from the manufacturer, formal reliability testing was carried out. Only in this way could confidence be gained in its usefulness for routine 3D computer-controlled conformal therapy. MATERIALS AND METHODS: To assess reliability, a set of 25 multi-segment test cases, each consisting of 10 to 17 fixed segments, was developed. The field arrangements and modalities for some of the test cases were identical to 3D conformal treatments that were being delivered with multiple static fields on conventional linear accelerators at our institution. Other cases were designed to explore reliability under more complex sets of conditions. These cases were "treated" repeatedly during a total period of 45 hours, over 5 days. During the treatments, ion chambers attached to the head of the machine provided dosimetric data for each field. Data from sensors connected to every set-up parameter (for example, couch positions, gantry angle, collimator leaf positions, etc.) were recorded and verified by an external computer. RESULTS: While preliminary tests indicated an interlock rate of 5%, final reliability test results demonstrated an interlock fault rate of approximately 0.5%. The reproducibility of dosimetric data and geometric setup parameters was within specifications. As an example, leaf position reproducibility in the patient plane was within 0.5 mm for 97% of the setups. The times required to carry out treatments were recorded and compared with the times to carry out identical treatments on a conventional linear accelerator with cerrobend blocks. Areas where additional time savings can be achieved were identified. CONCLUSIONS: As an integral part of acceptance testing, the Scanditronix MM50 was rigorously tested for reliability. The machine successfully passed these tests, providing increased confidence in its usefulness for routine 3D conformal therapy.

Humans↗

Reliability of prehospital rating scales for case severity and status change.

The purpose of this report is to determine the reliability and sensibility of the currently available prehospital rating scales by a prospective evaluation of ambulance call reports using generalizability methodology. Sequential samples of emergency call data from the Hamilton Base Hospital Paramedic program database were used to sample calls randomly for a two-phase study. Phase I and II used blinded ambulance call report forms presented to six rates during three sessions in each phase. Generalizability (reliability) coefficients were then generated to determine the degree of reliability for the scales in both phases of the study. The generalizability coefficients for all scales are substantial or excellent using the standards commonly applied to agreement statistics. The conclusion of the study is that all ambulance officers can use the prehospital scales reliably. The reliability of these general measures is one of the parameters that will allow us to evaluate where basic and advanced prehospital care have an impact on overall patient outcome.

Emergency Medical Technicians↗