Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Teaching DSM-III to clinicians. Some problems of the DSM-III system reducing reliability, using the diagnosis and classification of depressive disorders as an example.

Experiences from teaching DSM-III to more than three hundred Norwegian psychiatrists and clinical psychologists suggest that reliable DSM-III diagnoses can be achieved within a few hours training with reference to the decision trees and the diagnostic criteria only. The diagnoses provided are more reliable than the corresponding ICD diagnoses which the participants were more familiar with. The three main sources of reduced reliability of the DSM-III diagnoses are related to: poor knowledge of the criteria which often is connected with failure of obtaining diagnostic key information during the clinical interview; unfamiliar concepts and vague or ambiguous criteria. The two first issues are related to the quality of the teaching of DSM-III. The third source of reduced reliability reflects unsolved validity issues. By using the classification of five affective case stories as examples, these sources of diagnostic pitfalls, reducing reliability and ways to overcome these problems when teaching the DSM-III system, are discussed. It is concluded that the DSM-III system of classification is easy to teach and that the system is superior to other classification systems available from a reliability point of view. The current version of the DSM-III system, however, partly owes a high degree of reliability to broad and heterogeneous diagnostic categories like the concept major depression, which may have questionable validity. Thus, the future revisions of the DSM-III system should, above all, address the issue of validity.

Depressive Disorder↗

Reliability reporting practices in rape myth research.

A number of school-based programs address sexual violence by focusing on adolescents' attitudes about rape or acceptance of rape myths. However, many problems exist in the literature regarding measurement of rape myth acceptance, including issues of reliability and validity. This paper addresses measurement reliability issues and reviews reliability reporting practices of studies using the Burt Rape Myth Acceptance Scale. Less than one-half of the 68 articles examined reported reliability coefficients for the data collected. Almost one-third of the studies did not mention reliability. Examples of acceptable reliability reporting are provided. It is argued that reliability coefficients for the data actually analyzed should always be assessed and reported when interpreting program results. Implicationsfor school health research and practice are discussed.

Adolescent↗

Reliability of an individually molded shank shell for measuring tibial transverse rotations during the stance phase of walking.

Use of a shank shell has been shown to estimate tibial transverse rotations better than skin-mounted markers. However, the day-to-day reliability of the transverse tibial rotations using an individually molded shank shell has not been previously investigated. This study examined the between-tests and trials reliability of an individually molded shank shell for measuring peak tibial internal and external rotations, time of peak values, and tibia range of motion during 5 walking trials. The trial-to-trial reliability of tibial transverse rotations was measured in 14 healthy individuals while the test-retest reliability was measured in 10 persons on two occasions. Trial-to-trial reliability for peak transverse rotations, time of peak values, and tibia range of motion ranged from ICC (3,1) 0.59-0.95. The PCA between trials showed that 88-99 % of values were within 3 degrees of agreement. Test-retest reliability for peak rotations, tibia range of motion, and time of peak values ranged from ICC (3,1) 0.70-0.89 with SEM 1.6-2.21 degrees , 0.021 %, and 0.034 %, respectively. The PCA between tests showed that 70-100 % of values were within 3 degrees of agreement. The use of an individually molded shell and the close attachment of the shank shell to the individual's shank resulted in reliable test-retest and trial-to-trial data.

Adult↗

Validity, reliability, and applicability of seven definitions of hip osteoarthritis used in epidemiological studies: a systematic appraisal.

OBJECTIVE: To summarise and review articles addressing the quality (validity, reliability, applicability) of seven commonly used definitions of hip osteoarthritis (OA) for epidemiological studies in order to use it primarily as a classification criterion. METHODS: Medline and Embase were searched and articles studying the validity, reliability, or applicability of the definitions of hip OA were selected. Two reviewers independently extracted data on the quality of the seven definitions. RESULTS: Review of the literature showed the validity of the various definitions of hip OA, in particular, has barely been investigated. Minimal joint space (MJS) demonstrated the highest (intra- and interrater) reliability, and showed the highest association with hip pain and restricted internal rotation compared with the other definitions of hip OA. The reliability of the Kellgren and Lawrence grade and the index according to Lane is comparable with that of the MJS, but the construct validity should be investigated more thoroughly. The reliability and validity according to the Croft grade were inferior to the MJS, the Kellgren and Lawrence grade, and the index according to Lane. Despite precise and extensive development, the ACR criteria showed poor reliability and poor cross-validity (agreement between three ACR criteria sets) in a primary care setting. CONCLUSIONS: The reliabilities of MJS, Kellgren and Lawrence, and the index according to Lane were comparable, but the MJS had the highest relationship with hip pain in a male population. Considering how often definitions of hip OA are used, it is surprising that the validity has been so poorly investigated, and the validity needs to be studied more thoroughly.

Adult↗

Measuring the effectiveness of cataract surgery: the reliability and validity of a visual function outcomes instrument.

AIMS: To assess test-retest reliability and validity of the "TyPE" patient self assessed visual function questionnaire, as part of a study in two hospitals measuring the effectiveness of cataract surgery. The American TyPE questionnaire had minor adaptations made for use in Britain. METHODS: Test-retest reliability was assessed on 63 out of 378 adult cataract surgery patients in the study, using Spearman correlation coefficients and kappa coefficients of agreement. "Construct" validity was evaluated by comparing the association between changes in visual function questionnaire scores after surgery, with patients' perception of change in visual function obtained by independent interview of 24 patients. RESULTS: The TyPE questionnaire items showed very good test-retest reliability. Average Spearman and kappa coefficients for 39 patients from hospital 1 were 0.93 and 0.84 respectively. Spearman and kappa coefficients of 0.9 and 0.81 were obtained for those nine patients in hospital 2 where both the test and retest questionnaires were filled in by the same people. However, for the 15 patients from hospital 2, where the questionnaire was filled in by different people in the retest, reliability was less good: the Spearman coefficients were still high, average 0.72, but the kappa coefficients were poor, 0.27. Good construct validity was exhibited, with a correlation of 0.79 between change in distance vision score from the questionnaires and the independent interview. CONCLUSIONS: The adapted TyPE questionnaire is both very reliable and has good construct validity. The kappa coefficient should be used wherever possible to evaluate reliability. The test-retest reliability and validity and practicability of other visual function questionnaires have not been assessed adequately, and further development should be carried out of all such questionnaires, so that they may be introduced into routine clinical care.

Activities of Daily Living↗

Retest reliability of surveillance questions on health related quality of life.

STUDY OBJECTIVES: Health related quality of life (HRQoL) is an important surveillance measure for monitoring the health of populations, as proposed in the American public health plan, Healthy People 2010. The authors investigated the retest reliability of four HRQoL questions from the US Behavioral Risk Factor Surveillance System (BRFSS). DESIGN: Randomly sampled BRFSS respondents from the state of Missouri were re-contacted for a retest of the HRQoL questions. Reliability was estimated by kappa statistics for categorical questions and intraclass correlation coefficients for continuous questions. SETTING: Missouri, United States. PARTICIPANTS: 868 respondents were re-interviewed by telephone about two weeks after the initial interview (mean 13.5 days). Participants represented the adult, non-institutionalised population of Missouri: 59.1% women; mean age 49.5 years; 93.2% white race. MAIN RESULTS: Retest reliability was excellent (0.75 or higher) for Self-Reported Health and Healthy Days measures, and moderate (0.58 to 0.71) for other measures. Reliability was lower for older adults. Other demographic subgroups (for example, gender) showed no regular pattern of differing reliability and there was very little change in reliability by the time interval between the first and second interview. CONCLUSIONS: Retest reliability of the HRQoL Core is moderate to excellent. Scaling options will require future attention, as will research into appropriate metrics for what constitutes important population group differences and change in HRQoL.

Adolescent↗

SF 36 health survey questionnaire: I. Reliability in two patient based studies.

OBJECTIVE: To assess the reliability of the SF 36 health survey questionnaire in two patient populations. DESIGN: Postal questionnaire followed up, if necessary, by two reminders at two week intervals. Retest questionnaires were administered postally at two weeks in the first study and at one week in the second study. SETTING: Outpatient clinics and four training general practices in Grampian region in the north east of Scotland (study 1); a gastroenterology outpatient clinic in Aberdeen Royal Hospitals Trust (study 2). PATIENTS: 1787 patients presenting with one of four conditions: low back pain, menorrhagia, suspected peptic ulcer, and varicose veins and identified between March and June 1991 (study 1) and 573 patients attending a gastroenterology clinic in April 1993. MAIN MEASURES: Assessment of internal consistency reliability with Cronbach's alpha coefficient and of test-retest reliability with the Pearson correlation coefficient and confidence interval analysis. RESULTS: In study 1, 1317 of 1746 (75.4%) correctly identified patients entered the study and in study 2, 549 of 573 (95.8%). Both methods of assessing reliability produced similar results for most of the SF 36 scales. The most conservative estimates of reliability gave 95% confidence intervals for an individual patient's score difference ranging from -19 to 19 for the scales measuring physical functioning and general health perceptions, to -65.7 to 65.7 for the scale measuring role limitations attributable to emotional problems. In a controlled clinical trial with sample sizes of 65 patients in each group, statistically significant differences of 20 points can be detected on all eight SF 36 scales. CONCLUSIONS: All eight scales of the SF 36 questionnaire show high reliability when used to monitor health in groups of patients, and at least four scales possess adequate reliability for use in managing individual patients. Further studies are required to test the feasibility of implementing the SF 36 and other outcome measures in routine clinical practice within the health service.

Adolescent↗

Action potential propagation through embryonic dorsal root ganglion cells in culture. II. Decrease of conduction reliability during repetitive stimulation.

1. The reliability of the propagation of action potentials (AP) through dorsal root ganglion (DRG) cells in embryonic slice cultures was investigated during repetitive stimulation at 1-20 Hz. Membrane potentials of DRG cells were recorded intracellularly while the axons were stimulated by an extracellular electrode. 2. In analogy to the double-pulse experiments reported previously, either one or two types of propagation failures were recorded during repetitive stimulation, depending on the cell morphology. In contrast to the double-pulse experiments, the failures appeared at longer interpulse intervals and usually only after several tens of stimuli with reliable propagation. 3. In the period with reliable propagation before the failures, a decrease in the conduction velocity and in the amplitude of the afterhyperpolarization (AHP), an increase in the total membrane conductance, and the disappearance of the action potential "shoulder" were observed. 4. The reliability of conduction during repetitive stimulation was improved by lowering the extracellular calcium concentration or by replacing the extracellular calcium by strontium. The reliability of conduction decreased by the application of cadmium, a calcium channel blocker, 4-amino pyridine, a fast potassium channel blocker, or apamin or muscarine, the blockers of calcium-dependent potassium channels. The reliability of conduction was not effected by blocking the sodium potassium pump with ouabain or by replacing extracellular sodium with lithium. 5. In the period with reliable propagation cadmium, apamin, and muscarine reduced the amplitude of the AHP. The shoulder of the action potential was more pronounced and not sensitive to repetitive stimulation when extracellular calcium was replaced by strontium. It disappeared when cadmium was applied. 6. In DRG somata changes of the intracellular Ca2+ concentration were monitored by measuring the fluorescence of the Ca2+ indicator Fluo-3 with a laser-scanning confocal microscope. During repetitive stimulation, an accumulation of intracellular calcium occurred that recovered very slowly (tens of seconds) after the AP trains. 7. Computer model simulations performed in analogy to the experimental protocols produced conduction failures during repetitive stimulation only when the calcium currents during the APs were reduced. 8. From these findings it is concluded that conduction failures during repetitive stimulation are dependent on an accumulation of intracellular calcium leading to an inactivation of calcium currents, combined with small contributions of an accumulation of extracellular potassium and a summation of slow potassium conductances.

Acetylcholine↗

Measurement of ventricular size: reliability of the frontal and occipital horn ratio compared to subjective assessment.

INTRODUCTION: The frontal and occipital horn ration (FOR) has recently been described as a simple, linear measurement of ventricular size that correlates very well with ventricular volume. This study further characterizes the measurement properties of the FOR by investigating its interobserver reliability and comparing it to a subjective assessment of ventricular size. METHODS: Axial images (CT and MR) of children with hydrocephalus taken before and after third ventriculostomy were reviewed by 4 independent observers. Two observers were blinded to patient identity and clinical status and 2 observers were nonblinded. Each observer independently recorded linear measurements from which the FOR was calculated for each image. Each reviewer also made a separate subjective assessment of the degree of hydrocephalus on a 9-point adjectival scale. Reliability was calculated using a repeated-measures analysis of variance (ANOVA) and an intraclass correlation coefficient (ICC) with random image and observer effects. RESULTS: There were 120 separate observations (4 observers, 30 images). The FOR ranged from 0.33 to 0.75 (mean 0.55, standard deviation 0.11). The reliability coefficient was 0.93 (95% confidence interval, CI 0.80-0.97) between the 2 blinded observers and 0.98 (95% CI, 0.95-0.99) between the 2 nonblinded observer. The overall interobserver reliability for all 4 observers was 0.95 (95% CI 0.92-0.98). The mean FOR for each observer was very similar, regardless of the observer's blinding status. However, the reliability of the observers' subjective assessment of the hydrocephalus was much lower (ICC = 0.77, 95% CI 0. 60-0.88). CONCLUSIONS: The FOR demonstrates excellent interobserver reliability (>0.9) and was superior to subjective assessments of hydrocephalus. In this study, excellent reliability was maintained regardless of the blinding status of the observers. This further demonstrates the properties of the FOR as a simple and reproducible measure of ventricular size. It is suitable for use in clinical studies, possibly even in situations in which observer blinding is not possible.

Cerebral Ventricles↗

Inter-rater reliability of delirium rating scales.

Delirium continues to be under-recognized despite use of rating scales with apparently high inter-rater reliability. We analyzed the inter-reliability data of published rating scales for delirium using a standard questionnaire to evaluate if the inter-rater reliability was assessed rigorously. Most studies employed a heterogeneous group of cognitively disordered elderly, however other aspects of inter-rater reliability estimation were less than rigorous. This suggests that the reported reliability may be spuriously high, which may have implications on the ability of clinicians to discriminate delirium from other causes of cognitive impairment in practice. The methodology of assessing inter-rater reliability of delirium scales needs to improve and reliability should be evaluated when the settings of administration change substantially.

Delirium↗

A modified National Institutes of Health Stroke Scale for use in stroke clinical trials: preliminary reliability and validity.

BACKGROUND AND PURPOSE: The National Institutes of Health Stroke Scale (NIHSS) is accepted widely for measuring acute stroke deficits in clinical trials, but it contains items that exhibit poor reliability or do not contribute meaningful information. To improve the scale for use in clinical research, we used formal clinimetric analyses to derive a modified version, the mNIHSS. We then sought to demonstrate the validity and reliability of the new mNIHSS. METHODS: The mNIHSS was derived from our prior clinimetric studies of the NIHSS by deleting poorly reproducible or redundant items (level of consciousness, face weakness, ataxia, dysarthria) and collapsing the sensory item into 2 responses. Reliability of the mNIHSS was assessed with the certification data originally collected to assess the reliability of investigators in the National Institute of Neurological Disorders and Stroke (NINDS) rtPA (recombinant tissue plasminogen activator) Stroke TRIAL: Validity of the mNIHSS was assessed with the outcome results of the NINDS rtPA Stroke Trial: RESULTS: Reliability was improved with the mNIHSS: the number of scale items with poor kappa coefficients on either of the certification tapes decreased from 8 (20%) to 3 (14%) with the mNIHSS. With the use of factor analysis, the structure underlying the mNIHSS was found identical to the original scale. On serial use of the scale, goodness of fit coefficients were higher with the mNIHSS. With data from part I of the trial data, the proportion of patients who improved >/=4 points within 24 hours after treatment was statistically significantly increased by tPA (odds ratio, 1.3; 95% confidence limits, 1.0, 1.8; P=0.05). Likewise, the odds ratio for complete/nearly complete resolution of stroke symptoms 3 months after treatment was 1.7 (95% confidence limits, 1.2, 2.6) with the mNIHSS. Other outcomes showed the same agreement when the mNIHSS was compared with the original scale. The mNIHSS showed good responsiveness, ie, was useful in differentiating patients likely to hemorrhage or have a good outcome after stroke. CONCLUSIONS: The mNIHSS appears to be identical clinimetrically to the original NIHSS when the same data are used for validation and reliability. Power appears to be greater with the mNIHSS with the use of 24-hour end points, suggesting the need for fewer patients in trials designed to detect treatment effects comparable to rtPA. The mNIHSS contains fewer items and might be simpler to use in clinical research trials. Prospective analysis of reliability and validity, with the use of an independently collected cohort, must be obtained before the mNIHSS is used in a research setting.

Clinical Trials as Topic↗

An intervention to improve the reliability of manuscript reviews for the Journal of the American Academy of Child and Adolescent Psychiatry.

OBJECTIVE: The effects of methods used to improve the interrater reliability of reviewers' ratings of manuscripts submitted to the Journal of the American Academy of Child and Adolescent Psychiatry were studied. METHOD: Reviewers' ratings of consecutive manuscripts submitted over approximately 1 year were first analyzed; 296 pairs of ratings were studied. Intraclass correlations and confidence intervals for the correlations were computed for the two main ratings by which reviewers quantified the quality of the article: a 1-10 overall quality rating and a recommendation for acceptance or rejection with four possibilities along that continuum. Modifications were then introduced, including a multi-item rating scale and two training manuals to accompany it. Over the next year, 272 more articles were rated, and reliabilities were computed for the new scale and for the scales previously used. RESULTS: The intraclass correlation of the most reliable rating before the intervention was 0.27; the reliability of the new rating procedure was 0.43. The difference between these two was significant. The reliability for the new rating scale was in the fair to good range, and it became even better when the ratings of the two reviewers were averaged and the reliability stepped up by the Spearman-Brown formula. The new rating scale had excellent internal consistency and correlated highly with other quality ratings. CONCLUSIONS: The data confirm that the reliability of ratings of scientific articles may be improved by increasing the number of rating scale points, eliciting ratings of separate, concrete items rather than a global judgment, using training manuals, and averaging the scores of multiple reviewers.

Algorithms↗

Two-unit reliability analysis of questionnaires used in a regulatory system.

An interjudge reliability test was conducted to evaluate the questionnaires used in the surveillance of residential care institutions. Because the reliability test was carried out as part of the routine surveillance program and not as part of a controlled experiment, it was subject to deviations from the optimal reliability test model. However, this nonpure design provided an opportunity to not only examine the reliability of the items in the surveillance tool, but also to gain a better understanding of the use of a reliability test in an "imperfect" field setting. Two different surveyor teams administered the 257 questions on the questionnaires to a representative sample of 32 institutions on two separate occasions. In order to explain the variance in the reliability scores, a multivariate analysis was conducted for two units of analysis: the surveillance questions and the institutions. Based on the results of the reliability test, changes were introduced to improve the questionnaires and their administration.

Aged↗

A reliability study of the universal goniometer, fluid goniometer, and electrogoniometer for the measurement of ankle dorsiflexion.

This study investigated the reliability of three goniometers, the universal, fluid, and electro-goniometers, in the measurement of ankle dorsiflexion. Intra- and interobserver reliability were assessed using 10 healthy volunteers and five observers. A standardized ankle position was used to measure full range of active dorsiflexion. Intraobserver reliability was assessed using one observer over two successive occasions. Interobserver reliability was assessed among five observers over five separate occasions. A one-factor analysis of variance to examine intraobserver reliability demonstrated no significant difference between each of the devices on the two occasions. A multifactorial analysis of variance demonstrated significant differences among observers and again among devices (P < 0.001). Secondary analysis for interdevice reliability demonstrated significant differences among the three devices (p < 0.1). The study suggests that each device cannot be used reliably among observers or be used interchangeably, and clinical judgment based on angular changes of less than 10 degrees are invalid if rigid protocols are not followed.

Adult↗

Reliability of F-scan in-shoe measurements of plantar pressure.

Research by our group and others indicates that many amputations of the lower limb occur after foot ulceration in patients with diabetes. It has been proposed that diabetic foot ulcers are mainly caused by repetitive trauma in areas of high plantar pressure during walking. Recent technology permits in-shoe measurement of plantar pressure. We assessed the reliability of the F-Scan in-shoe system for measurement of plantar pressure (Tekscan Inc., Boston, MA) in 51 subjects from a cohort of 977 diabetic veterans enrolled in a prospective study of risk factors for foot ulceration and amputation (the Seattle Diabetic Foot Study). Subjects were tested twice, wearing their own shoes. We used the coefficient of variation (CV) and the intra-class correlation coefficient (ICC) to estimate the reliability of F-Scan measurements of pressure. Peak pressure over the metatarsal heads proved to have the best indices of reliability, with CVs of 0.150 and 0.155, and ICCs of 0.755 and 0.751. Coefficients of variation for the heel, whole foot, and hallux ranged from 0.148 to 0.240, with ICCs ranging from 0.493 to 0.832. By published standards, peak pressures over the metatarsal heads and right hallux met the criteria for excellent reliability. Our ICCs for high pressures under the foot, heel, metatarsal heads, and hallux, and for peak pressures under the heel and left hallux represented fair-to-good reliability. No F-Scan plantar measurements could be judged by these criteria as having poor reliability. This clinical study found that for elderly patients with diabetes who were wearing their own shoes and were tested on two different days with different insoles, the F-Scan insole system was generally reliable for measurements of high pressure and peak pressure.

Adult↗

Field reliability of comprehensive system scoring in an adolescent inpatient sample.

The extent to which the Comprehensive System for the Rorschach is reliably scored has been a topic of some controversy. Although several studies have concluded it can be scored reliably in research settings, little is known about its reliability in field settings. This study evaluated the reliability of both response-level codes and protocol-level scores among 84 adolescent psychiatric inpatients in a clinical setting. Rorschachs were originally administered and scored for clinical purposes. Among response codes, 87% demonstrated acceptable reliability(> .60), and most coefficients exceeded .80. Results were similar for protocol-level scores, with only one score demonstrating less than adequate reliability. The findings are consistent with previous evidence, indicating reliable scoring is possible even in field settings.

Adolescent↗

Test-retest reliability of psychological and neurobehavioral tests self-administered by computer.

A series of 12 psychological and 7 neurobehavioral performance tests were administered twice to a nonclinical normative sample with 1 week between administrations. The tests were presented in a self-administered computerized format. One week test-retest reliabilities were comparable to conventional administration formats. The results suggest that individual test reliability is not affected when tests are administered as part of an extensive multi-measure battery. Computer administered test reliability coefficients also were compared to a Mixed Format (computer-conventional) administration with mixed format reliabilities generally similar to the reliabilities of published conventional tests but also generally lower than same format testing. Compared to psychological test reliability, neurobehavioral test reliability appeared more vulnerable to decreases with mixed format testing. These conclusions should not be generalized to all computer implemented tests as the qualities of the test implementation will affect the outcome.

Adult↗

Assessing reliability of a measure of self-rated health.

The test-retest reliability of self-rated health is analysed and compared with the reliability of health questions phrased more as well as less precisely. Differences in reliability between men and women and between age groups are also assessed. The study is based on 204 and 409 re-interviews from the 1991 Swedish Level of Living Survey and the 1989 Survey of Living Conditions respectively. The results show that the reliability of self-rated health is as good as or even better than that of most of the more specific questions. Only an indicator of high blood pressure showed significantly higher reliability. The reliability of self-rated health is good in all subgroups studied, and is even excellent among older men. It is concluded that the good overall reliability of self-rated health found in this study is in line with previous results concerning the validity of people's assessments of their general health as well as results concerning the basis upon which they make these judgements.

Adult↗