Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

The IPCS Collaborative Study on Neurobehavioral Screening Methods: VI. Agreement and reliability of the data.

The IPCS Collaborative Study on Neurobehavioral Screening Methods was undertaken to determine the intra- and inter-laboratory reliability of a functional observational battery (FOB) and an automated assessment of motor activity in eight laboratories world-wide. The effects of seven chemicals (acrylamide, bis-acrylamide, p,p'-DDT, lead acetate, parathion, toluene, and triethyl tin) were studied during two dosing regimens: single-dose and four-week repeated dosing. All participating laboratories generally could detect and characterize the effects of known neurotoxicants, even though there were some differences in outcome on specific endpoints. The results were further evaluated to assess the agreement across laboratories in the dose-response data at the expected times of maximal effect (time of peak effect for the single-dose studies, and during or at the end of dosing for repeated-exposure studies). Percent agreement was calculated as the percentage of laboratories agreeing on an outcome (whether it be a significant dose effect or not). As an alternative approach, slopes of the dose-response functions were calculated, and reliability of those slope estimates across laboratories and chemicals was determined. Reliability was defined as the degree of agreement across laboratories (intraclass correlation coefficient) of the dose-response slopes within and between chemicals. These reliability estimates were calculated for each domain and for each endpoint. Relative reliability of the endpoints was evaluated, and hypotheses concerning the influence of outlying data were tested. The data clearly showed that reliability was not influenced by the objectivity or subjectivity of the test measure. Thus these data provide additional information regarding the reliability and robustness of the tests across the participating laboratories.

Animals↗

[Reliability of the evaluation of regional cerebral blood flow using SPECT in chronic schizophrenia and bipolar disorder].

OBJECTS: Demonstrate the reliability of cerebral SPECT using 99 mTc-HMPAO. METHODS & MATERIALS: Evaluation of cerebral blood flow using SPECT in 24 patients with schizophrenia, 24 patients with bipolar disorder and 20 controls. In the study we have reliability between observers and intraobserver. In both cases kappa statistic has been applied for measuring reliability. RESULTS: reliability between observers represents a kappa coefficient of 0.71. Intraobserver reliability, with a medium grade concordance slightly superior, shows a medium kappa coefficient of 0.74. CONCLUSIONS: Visual evaluation of SPECT images using 99mTc--HMPAO is a trustworthy technique to document the different patterns of regional cerebral blood flow. Reliability is determinate by the improvement, during visual analysis of reliability between observers (kappa: 0.71) and intraobservers (kappa: 0.74).

Algorithms↗

Active lateral neck flexion range of motion measurements obtained with a modified goniometer: reliability and estimates of normal.

OBJECTIVE: To describe a new method for measuring lateral neck flexion range of motion (ROM), document the reliability of the method and present estimates of normal. SUBJECTS: One hundred thirty-five subjects ranging in age from 14-95 yr. Two physical therapists with 13 and 2 yr of experience, respectively, served as testers. INTERVENTION: Measurement of active lateral neck flexion ROM using a universal goniometer modified by the placement of a portion of a small paper clip through the axis. The goniometer arms were aligned with the subject's nose, and the free-swinging paper clip (pendulum) was used as a marker. The more experienced therapist measured lateral flexion of 100 subjects to establish intratester reliability and estimates of normal. Both therapists measured 35 subjects to determine intertester reliability. MAIN OUTCOME MEASURE: Degrees of lateral neck flexion. RESULTS: Intraclass correlation coefficients for intratester reliability exceeded 0.90. Coefficients for intertester reliability were 0.86 and 0.65. ROM decreased with increasing age. CONCLUSION: The modified goniometer is inexpensive, easy to use and can yield high intratester reliability and satisfactory intertester reliability. The estimates of normal provide preliminary values with which a patient's lateral neck flexion ROM can be compared.

Adolescent↗

Further analysis of the reliability of the posterior tangent lateral lumbar radiographic mensuration procedure: concurrent validity of computer-aided X-ray digitization.

OBJECTIVE: To investigate the reliability of a specific method of radiographic analysis of the geometric configuration of the lumbopelvic spine in the sagittal plane, and to investigate the concurrent validity of a computer-aided digitization procedure designed to replace the more tedious and time-consuming manual measurement process. DESIGN: A blind, repeated-measures design was used. The results of radiographic measures derived through the traditional manual marking method were compared with measures derived by computer-aided digitization of lateral lumbopelvic radiographs. SETTING: Private chiropractic clinic. MAIN OUTCOME MEASURES: Pearson's product-moment correlation coefficients, paired sample t tests and intraclass correlation co-efficients (ICC) were used to examine intraexaminer reliability, and repeated measures of analysis of variance were used to examine interexaminer reliability for relative rotation angles for T12-L1, L1-L2, L2-L3, L3-L4, L4-L5, L5-S1, overall lordosis measurement [absolute rotation angle (ARA)] from L1-L5 and Cobb angle of overall lordosis measured from the inferior surface of T12 to the superior surface of S1, Ferguson's sacral base angle to horizontal, angle of pelvic tilt (arcuate angle) to horizontal and anteroposterior thoracic translation (Sz) in millimeters. RESULTS: ICC estimates for intraexaminer reliability were in the range of 0.96-0.98 for the L1-L5 ARA, a range of 0.87-0.99 for the arcuate angle measurement, 0.83-0.94 for the Ferguson's angle measurement, 0.88-0.95 for the Cobb angle measurement from the inferior surface of T12 compared with the superior surface of S1 and 0.98-1.00 for the translation measurement of the lower thoracic spine to S1 (Sz). The intersegmental measurement's (T12-L1, L1-L2, L2-L3, L3-L4, L4-L5, L5-S1) correlations ranged from a low of 0.55 to a high of 0.97. Examination of these findings suggests that the reliability for the three doctors is acceptable with only the T12-L1 intersegmental measure falling below 0.70 for the least experienced examiner. Average ICC of interexaminer reliability for manual and computer-aided digitizing examiners were the following: 0.96 for the L1-L5 ARA; 0.84 for the arcuate angle measurement; 0.82 for the Ferguson's angle measurement; 0.88 for the Cobb angle measurement; 1.00 for the Sz translation measurement; and values of 0.65, 0.73, 0.74, 0.75, 0.89 and 0.81 for relative rotation angle measurements T12-L1, L1-L2, L2-L3, L3-L4, L4-L5 and L5-S1, respectively. CONCLUSION: The data tend to support the reliability of this method of radiographic analysis of the geometric configuration of the lumbopelvic spine as viewed on lateral lumbopelvic radiographs. The additional data presented here tend to support the concurrent validity of the computer-aided digitization method of analysis inasmuch as the measures determined by the digitizing examiners are essentially identical to those determined by the manual method plus or minus the average standard error of measure of each value.

Humans↗

Reliability of reported age at menopause.

Age at menopause is an important epidemiologic characteristic whose reliability of reporting in the US population is not known. The authors examined four hypotheses about the reliability of reported age at menopause in the United States: 1) women with hysterectomy-induced menopause more reliably report their age at menopause than women who have undergone natural menopause; 2) reliability declines with time since menopause; 3) reliability declines with age; and 4) women with higher educational levels report their age at menopause more reliably than women with less education. The authors used linear regression models among 2,545 women in the First National Health and Nutrition Examination Survey and Followup Study (1971-1984) and compared responses at first and follow-up interviews. Among women who had undergone a natural menopause, 44% reported their age at menopause within one year from the first to second interviews; among women who had undergone a hysterectomy-induced menopause, 59% reported their age at menopause within one year from first to follow-up interviews. Only hysterectomy status and years from menopause to follow-up interview were significantly associated with the absolute difference between age at menopause reported at first and follow-up interviews. The authors conclude that caution in studies involving age at menopause may enhance our understanding of this critical event in the lives of women.

Adult↗

AIDS knowledge and attitudes among injection drug users: the issue of reliability.

Among injection drug users (IDUs), AIDS-related knowledge and attitudes have not consistently predicted AIDS risk behavior. This may be due in part to the limited reliability of indexes used to measure drug users' AIDS knowledge and attitudes. In addition, the substantive interpretation of findings is confounded if index reliability is lower for particular demographic groups (e.g., ethnic populations and women). This report is based on 8 measures of AIDS-related knowledge and attitudes in a sample of 332 injection drug users in Los Angeles. The reliability of knowledge and attitude indexes for the overall sample is generally acceptable for the purpose of group comparison (average alpha = .60). But reliability is consistently lower for respondents who are Hispanic (average alpha = .49) and respondents with less formal education (alpha = .56). The reliability of 2 measures of sex-related attitudes is lower for female respondents. It is therefore important that the reliability of knowledge and attitude indexes be assessed not just for drug-user samples as a whole, but also within demographic groups of substantive interest.

Acquired Immunodeficiency Syndrome↗

Interrater and intrarater reliability in the evaluation of velopharyngeal insufficiency within a single institution.

OBJECTIVE: To explore the interrater and intrarater reliability in nasoendoscopic assessment of velopharyngeal (VP) function using the standardized reporting method described by Golding-Kushner within a single institution. DESIGN: Prospective blinded study. SETTING: Academic, tertiary care, pediatric hospital. PARTICIPANTS: Six health care providers (2 pediatric otolaryngology faculty members, 2 pediatric otolaryngology fellows, and 2 speech pathologists) independently rated 50 videotaped nasoendoscopy segments twice. The segments on the videotape were obtained in a clinical setting. MAIN OUTCOME MEASURES: The Golding-Kushner rating system was used to rate VP function. Raters described VP closure quantitatively by rating palatal and lateral pharyngeal wall movement for each segment. They also qualitatively described characteristics of the VP gap, rated gap size as none, small, medium, or large, and estimated the percentage gap size relative to the resting position. Reliability coefficients were calculated for the data sets. RESULTS: Fairly good interrater and intrarater reliability was seen in the quantitative measures. Faculty otolaryngologists rated segments more similarly to each other than did pediatric otolaryngology fellows, but intrarater reliability was similar for both the experienced and less experienced otolaryngologists. Less consistency was seen in the ratings of the speech pathologists. Raters tended to rate with less consistency when describing qualitative characteristics of the VP gap than when making quantitative measurements. CONCLUSIONS: The Golding-Kushner scale is a reasonably reliable tool for reporting nasoendoscopic findings at our institution. However, these data also indicate that there exists room for improvement and that rater training may increase reliability.

Adolescent↗

Reliability and validity in binary ratings: areas of common misunderstanding in diagnosis and symptom ratings.

Confusion may exist between the reliability of a binary rating (for example, schizophrenia versus not-schizophrenia) and its implications for validity. High reliability does not guarantee validity, but paradoxically, low reliability does not imply poor validity in all contexts. Changes in the base rate or in experimental design may indicate high validity even when the reliability was thought to be low. Attempts to improve the psychiatric nomenclature by increasing only reliability run the risk of the "attenuation paradox" where further increases in reliability will make the ratings less valid. Finally, the assumption of random error in making diagnoses does not always hold, so that statistical analyses must be adjusted accordingly. New statistical methods are needed to index only false-positive or false-negative rates in order to quantify the error that will reduce some validity coefficients.

Bipolar Disorder↗

Reliability of lifetime diagnosis. A multicenter collaborative perspective.

It is important to determine the reliability of lifetime diagnosis in a nonpatient population, for this type of diagnostic data and this type of sample are used in many genetic, epidemiological, and nosological studies. We examined the reliability of lifetime diagnosis when the Schedule for Affective Disorders and Schizophrenia-Lifetime Version and Research Diagnostic Criteria were used to interview ill and well relatives of probands in the National Institute of Mental Health Collaborative Study of the Psychobiology of Depression. Subjects were interviewed three times, so data are available concerning both short- and long-interval test-retest reliability. Short-interval test-retest reliability was excellent for both diagnoses and symptoms. Reliability was also quite high in the long-interval test-retest study. We conclude that it is possible to make lifetime diagnoses reliably in a nonpatient population.

Anxiety Disorders↗

The assessment of affective disorders in children and adolescents by semistructured interview. Test-retest reliability of the schedule for affective disorders and schizophrenia for school-age children, present episode version.

The reliability of assessment of Research Diagnostic Criteria and DSM-III axis I affective disorders in children and adolescents was studied using a semistructured diagnostic interview. The Schedule for Affective Disorders and Schizophrenia (SADS) for School-Age Children (Kiddie SADS) Present Episode Version, an adaptation of the adult SADS for children was used. Fifty-two subjects, aged 6 through 17 years, were interviewed in a test-retest format by one of three pairs of interviewers. Assessment of symptoms and composite scales of the depressive syndrome were determined to have acceptable reliability, as were three depressive diagnoses. Conduct disorder was assessed with high reliability. Four anxiety disorders and their composite symptoms were assessed with unacceptable reliability; only separation anxiety was assessed with acceptable reliability. The results of this study showed generally lower reliability of symptoms, scales, and diagnoses than did two studies of adults using the SADS.

Adolescent↗

Reliability of DSM-III-R anxiety disorder categories. Using the Anxiety Disorders Interview Schedule-Revised (ADIS-R).

A large reliability study of DSM-III-R anxiety disorders is reported in which outpatients (n = 267) received two independent structured interviews (Anxiety Disorders Interview Schedule-Revised). It is the only reliability study to date in which the final DSM-III-R criteria are used throughout the study. Reliability was assessed for each diagnosis when it was assigned as a principal diagnosis and when it was assigned as either a principal or an additional diagnosis. Excellent reliability was obtained for current principal diagnoses of simple phobia, social phobia, and obsessive-compulsive disorder. Agreement was good for panic disorder when all severity levels of agoraphobic avoidance were combined. Reliability was fair for generalized anxiety disorder. Remaining diagnostic difficulties, particularly in identifying levels of agoraphobic avoidance and in reliably diagnosing generalized anxiety disorder, are discussed in the context of changes in diagnostic criteria that are under consideration for DSM-IV.

Adult↗

Monitoring sedation status over time in ICU patients: reliability and validity of the Richmond Agitation-Sedation Scale (RASS).

CONTEXT: Goal-directed delivery of sedative and analgesic medications is recommended as standard care in intensive care units (ICUs) because of the impact these medications have on ventilator weaning and ICU length of stay, but few of the available sedation scales have been appropriately tested for reliability and validity. OBJECTIVE: To test the reliability and validity of the Richmond Agitation-Sedation Scale (RASS). DESIGN: Prospective cohort study. SETTING: Adult medical and coronary ICUs of a university-based medical center. PARTICIPANTS: Thirty-eight medical ICU patients enrolled for reliability testing (46% receiving mechanical ventilation) from July 21, 1999, to September 7, 1999, and an independent cohort of 275 patients receiving mechanical ventilation were enrolled for validity testing from February 1, 2000, to May 3, 2001. MAIN OUTCOME MEASURES: Interrater reliability of the RASS, Glasgow Coma Scale (GCS), and Ramsay Scale (RS); validity of the RASS correlated with reference standard ratings, assessments of content of consciousness, GCS scores, doses of sedatives and analgesics, and bispectral electroencephalography. RESULTS: In 290-paired observations by nurses, results of both the RASS and RS demonstrated excellent interrater reliability (weighted kappa, 0.91 and 0.94, respectively), which were both superior to the GCS (weighted kappa, 0.64; P<.001 for both comparisons). Criterion validity was tested in 411-paired observations in the first 96 patients of the validation cohort, in whom the RASS showed significant differences between levels of consciousness (P<.001 for all) and correctly identified fluctuations within patients over time (P<.001). In addition, 5 methods were used to test the construct validity of the RASS, including correlation with an attention screening examination (r = 0.78, P<.001), GCS scores (r = 0.91, P<.001), quantity of different psychoactive medication dosages 8 hours prior to assessment (eg, lorazepam: r = - 0.31, P<.001), successful extubation (P =.07), and bispectral electroencephalography (r = 0.63, P<.001). Face validity was demonstrated via a survey of 26 critical care nurses, which the results showed that 92% agreed or strongly agreed with the RASS scoring scheme, and 81% agreed or strongly agreed that the instrument provided a consensus for goal-directed delivery of medications. CONCLUSIONS: The RASS demonstrated excellent interrater reliability and criterion, construct, and face validity. This is the first sedation scale to be validated for its ability to detect changes in sedation status over consecutive days of ICU care, against constructs of level of consciousness and delirium, and correlated with the administered dose of sedative and analgesic medications.

Aged↗

Pacemaker and ICD generator reliability: meta-analysis of device registries.

CONTEXT: Despite there being millions of pacemaker and implantable cardioverter-defibrillator (ICD) generator implants worldwide, little is known about device reliability. OBJECTIVES: To perform a meta-analysis of prospective pacemaker and ICD registries to determine annual rates of pacemaker and ICD malfunction and to identify trends in these rates. DATA SOURCES: MEDLINE (January 1966 to April 2005), the Cochrane Central Register of Controlled Trials (through second quarter 2005), the Cochrane Database of Systematic Reviews (through second quarter 2005), and a bibliographic review of secondary sources. Search terms included pacemaker, artificial; defibrillators, implantable; registries; performance; and malfunction. STUDY SELECTION: Eligible registries were prospective; reported the number of patients with pacemaker, ICD, or both; and allowed determination of the annual number of device malfunctions. Of 1007 references screened, 3 registries meeting the selection criteria were identified and included 2.1 million pacemaker person-years and 14,821 ICD person-years of observation. DATA EXTRACTION: Included the annual number of patients with pacemakers and ICDs at risk of device failure and the annual number of generator malfunctions (1983-2004 for pacemakers, 1988-2004 for ICDs). A device malfunction was defined as an integral component failure that required device explantation prior to reaching elective replacement. Failures of pacemaker and ICD electrodes were not included in the study. DATA SYNTHESIS: There were 2981 pacemaker and 384 ICD generator malfunctions. Pacemaker reliability improved markedly during the 1980s (P for trend <.001) and the pacemaker malfunction rate remained low during the remainder of the study. Implantable cardioverter-defibrillator reliability improved during the first 10 study years (P for trend <.001). From 1998-2002, however, the ICD malfunction rate increased more than 4-fold (P for trend <.001), before decreasing substantially in the latter 2 years of the study. Overall, the mean (SE) annual ICD malfunction rate was about 20-fold higher than the pacemaker malfunction rate (26.5 [3.8] vs 1.3 [0.1] malfunctions per 1000 person-years, P<.001). Battery malfunctions were the most common cause of device failure. CONCLUSIONS: Pacemaker reliability has improved markedly. In contrast, after more than a decade of improving device reliability, the ICD malfunction rate transiently increased before experiencing substantial reductions in the latter 2 study years. Whether increasing device sophistication accounts for the observed decrease in reliability is not known. Continued monitoring of pacemaker and ICD performance is required.

Defibrillators, Implantable↗

Reliability of cognitive tests used in Alzheimer's disease.

Assessment of cognitive status is a key component of monitoring Alzheimer's patients during the course of their illness. The reliability of a cognitive test is a measure of its reproducibility under replicate conditions. In the classical setting, reliability is defined in three ways: the ratio of the variance of the true scores to the variance of the observed scores; the correlation of observed scores on two parallel forms of the test, and the square of the correlation between the observed score and the true score. In the classical case of independence of true scores and measurement errors, the three definitions are equivalent. Estimation of reliability through analysis of variance techniques and construction of confidence intervals is accomplished when the true scores and errors are normally distributed. This paper examines a non-parametric, probabilistic estimate of reliability as the probability that, given a parallel test, the second set of scores has the same ranking as the first set. In the classical case there is a monotonic relationship between this measure and the reliability. This measure is also linked to Kendall's tau. The performance of the probabilistic measure is compared with the traditional measures in a variety of models, including those with bounded scales, and those with skewed distributions. The ideas are extended to the case of the reliability of change scores and to biased estimators of true scores. In this context truncation models and Bayes estimates of true scores are considered.

Aged↗

Person reliability of psychiatric patients' responses to a psychopathology inventory.

Person-reliability indices can assist clinicians in determining the interpretability of a patient's responses to the Basic Personality Inventory (BPI). Using an initial sample of 65 psychiatric patients, we found that: (1) different person-reliability indices showed modest evidence of psychometric adequacy and tended not to be confounded with general psychopathology; (2) a content consistency index of person reliability was predictably related to other item change variables, whereas within-session profile stability was related to across-session measures of profile stability: and (3) evidence for the ability of person-reliability indices to moderate the validity of clinical criteria was weak. Results provide cautious support for a multidimensional conceptualization of the person reliability construct on the BPI but demand further evaluation of the clinical utility of person reliability indices.

Adolescent↗

Inter-rater reliability for function and strength measurements in the acute care hospital after elective hip and knee arthroplasty.

OBJECTIVE: To determine the inter-rater reliability of function and strength measurements in patients undergoing elective hip and knee arthroplasty in an acute care setting. METHOD: Forty-four patients underwent either total hip or knee arthroplasty. Patients were rated by 4 occupational therapists and 7 physical therapists on their performance of 5 functional tasks: lower extremity dressing, toilet transfer, supine-to-sit transfer, sit-to-stand transfer, and ambulation to 100 feet. Strength measurements of the quadriceps femoris muscle were measured quantitatively with a Microfet hand-held dynamometer. Data were analyzed to determine the interrater reliability using the Kappa statistic (K) for the functional tasks and the intra-class correlation coefficient (ICC) for the strength measurements. RESULTS: A high level of inter-rater reliability was achieved for lower extremity dressing, toilet transfer, supine-to-sit transfer, sit-to-stand transfer, and ambulation to 100 feet, as evidenced by K values between 0.75 and 0.99. Reliability was also excellent for quantitative strength measurements using the dynamometer, with an ICC of 0.94. CONCLUSION: This study demonstrated excellent interrater reliability with measurements of function and strength post-operatively after elective hip and knee arthroplasty. The practical implication is that by using a standardized measurement tool in the acute care setting, the treatment team can more reliably assess patients' progress, which may aid clinical decision making.

Activities of Daily Living↗

Analyzing reliability of change in depression among persons with rheumatoid arthritis.

OBJECTIVE: To examine several methods of determining reliability of change constructs in depressive symptoms in patients with rheumatoid arthritis (RA) and to demonstrate the strengths, weaknesses, and uses of each method. METHODS: Data were analyzed from a cohort of 54 persons with RA who participated in a combined behavioral/pharmacologic intervention of 15 months duration. These longitudinal data were used to examine 3 methodologies for assessing the reliability of change for various measures of depression. The specific methodologies involved the calculations of reliable change, sensitivity to change, and reliability of the change score. RESULTS: The analyses demonstrated differences in reliability of change performance across the various depression measures, which suggest that no single measure of depression for persons with RA should be considered superior in all contexts. CONCLUSION: The findings highlight the value of utilizing reliability of change constructs when examining changes in depressive symptoms over time.

Arthritis, Rheumatoid↗

Functional index-2: Validity and reliability of a disease-specific measure of impairment in patients with polymyositis and dermatomyositis.

OBJECTIVE: To revise the content of the Functional Index in myositis (FI) and to evaluate measurement properties of a revised FI. METHODS: Previously performed FI (n = 287) were analyzed for internal redundancy and consistency, and ceiling and floor effects. Content was evaluated and a preliminary revised FI was developed. To evaluate the construct validity of the preliminary revised FI, it was compared with isokinetic measurements of muscular strength and endurance, the Myositis Activities Profile, disease impact on general wellbeing, and creatine phosphokinase levels. Minor adjustments were made and the revised FI was investigated for interrater reliability and intrarater reliability over a 1-week period. After this, some minor, additional adjustments were made leading to the final version, FI-2. RESULTS: Five tasks were removed from the original FI due to ceiling effects. Performance pace and number of repetitions were modified for the remaining tasks. A moderate correlation (r(s) = 0.58) was found between the shoulder flexion task of the preliminary revised FI and isokinetic measurements of shoulder flexion endurance. Intraclass correlation coefficient (ICC) for interrater reliability of the revised FI varied from 0.86-0.99 with no systematic differences. ICC for intrarater reliability varied from 0.56-0.99 with systematic differences (P < 0.05) between test and retest in 3 of the tasks. The sit-up task was excluded due to low intrarater reliability resulting in the final 7-item FI-2. There was a good correlation between tasks on the right and left side suggesting that the FI-2 could be performed unilaterally. CONCLUSION: The FI-2 is a valid and reliable outcome measure of impairment for patients with polymyositis or dermatomyositis. It is well tolerated and the unilateral FI-2 requires a maximum of 20 minutes to perform. Further evaluation of sensitivity to change and testing in healthy individuals needs to be conducted.

Activities of Daily Living↗