Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Portable mood mapping: the validity and reliability of analog scale displays for mood assessment via hand-held computer.

The long-term natural time course of mood change remains poorly understood, and improved methods that assay multiple mood symptoms quickly and reliably are crucial to further progress. This study describes the reliability and validity of the new visual analog scale (VAS) display method for a recently developed 19-item VAS-based mood questionnaire, the VMQ, administered via hand-held computer (HHC). The effect of the smaller HHC screen size on accuracy and precision of VAS completion was investigated in 28 subjects using 4- and 10-cm paper-based VASs to indicate six specified dates within the year. The influence of digital vs. paper medium was then tested in 39 subjects who completed the same task, using 10-cm paper and 4-cm HHC-based VASs. Test-retest reliability was evaluated in 29 subjects who completed the questionnaire on a HHC twice, 10 min apart. Since the HHC presents VMQ scales with text anchor orientation set randomly, we also considered whether subjects might inadvertently transpose responses on the HHC. We found that reducing VAS size produced no significant loss of response precision or accuracy in subject response. Moreover, there was no significant loss of accuracy or precision between 10-cm paper and 4-cm HHC-based versions of the VAS. HHC-based items also demonstrated excellent test-retest reliability, with excellent values of Cronbach's alpha. The transposition error rate was negligible (0.27%). Our study provides initial evidence that the HHC-based VAS display used in the VMQ is a reliable and valid tool for comprehensive collection of analog mood scale data.

Adult↗

Reliability and accuracy of dermatologists' clinic-based and digital image consultations.

BACKGROUND: Telemedicine technology holds great promise for dermatologic health care delivery. However, the clinical outcomes of digital image consultations (teledermatology) must be compared with traditional clinic-based consultations. OBJECTIVE: Our purpose was to assess and compare the reliability and accuracy of dermatologists' diagnoses and management recommendations for clinic-based and digital image consultations. METHODS: One hundred sixty-eight lesions found among 129 patients were independently examined by 2 clinic-based dermatologists and 3 different digital image dermatologist consultants. The reliability and accuracy of the examiners' diagnoses and the reliability of their management recommendations were compared. RESULTS: Proportion agreement among clinic-based examiners for their single most likely diagnosis was 0. 54 (95% confidence interval [CI], 0.46-0.61) and was 0.92 (95% CI, 0. 88-0.96) when ratings included differential diagnoses. Digital image consultants provided diagnoses that were comparably reliable to the clinic-based examiners. Agreement on management recommendations was variable. Digital image and clinic-based consultants displayed similar diagnostic accuracy. CONCLUSION: Digital image consultations result in reliable and accurate diagnostic outcomes when compared with traditional clinic-based consultations.

Adult↗

Potentialities of theoretical and experimental prediction of Life Support Systems reliability.

To develop and design Life Support Systems it is necessary to evaluate their reliability. However direct experiments take much time, are very expensive, and therefore are practically impossible. Promising way is to use approximate estimates of reliability, which need essentially fewer amounts of experimental data. Two types of estimates of Life Support System reliability--additive and multiplicative ones are considered in the paper. Additive estimate is based on the assumption that total system failure probability is low and therefore it can be considered as the sum of failure probability of separate units. Additive approach allows obtaining near lower-bounded estimate of failure probability. Multiplicative estimate allows evaluating the possibility of system catastrophe due to simultaneous effect of several factors when each of them separately is not dangerous. Evaluation shows that the possible error of reliability forecast increases with the increasing of number of external factors faster than exponential function. An illustration of the ecological similarity approach as promising tool for providing estimation of full-scale system reliability by means the set of small similar experimental models.

Biomass↗

The comparability and reliability of five health-state valuation methods.

The objective of the study was to consider five methods for valuing health states with respect to their comparability (convergent validity, value functions) and reliability. Valuation tasks were performed by 104 student volunteers using five frequently used valuation methods: standard gamble (SG), time trade-off (TTO), rating scale (RS), willingness-to-pay (WTP) and the paired comparisons method (PC). Throughout the study, the EuroQol classification system was used to construct 13 health-state descriptions. Validity was investigated using the multitrait-multimethod (MTMM) methodology. The extent to which results of one method could be predicted by another was examined by transformations. Reliability of the methods was studied parametrically with Generalisability Theory (an ANOVA extension), as well as non-parametrically. Mean values for SG were slightly higher than TTO values. The RS could be distinguished from the other methods. After a simple power transformation, the RS values were found to be close to SG and TTO. Mean values of WTP were linearly related to SG and TTO, except at the extremes of the scale. However, the reliability of WTP was low and the number of inconsistencies substantial. Valuations made by the RS proved to be the most reliable. Paired comparisons did not provide stable results. In conclusion, the results of the parametric transformation function between RS and SG/TTO provide evidence to justify the current use of RS (with transformations) not only for reasons of feasibility and reliability but also for reasons of comparability. A definite judgement on PC requires data of a complete design. Due to the specific structure of the correlation matrix which is inherent in valuing health states, we believe that full MTMM is not applicable for the standard analysis of health-state valuations.

Adult↗

Interrater reliability in myofascial trigger point examination.

The myofascial trigger point (MTrP) is the hallmark physical finding of the myofascial pain syndrome (MPS). The MTrP itself is characterized by distinctive physical features that include a tender point in a taut band of muscle, a local twitch response (LTR) to mechanical stimulation, a pain referral pattern characteristic of trigger points of specific areas in each muscle, and the reproduction of the patient's usual pain. No prior study has demonstrated that these physical features are reproducible among different examiners, thereby establishing the reliability of the physical examination in the diagnosis of the MPS. This paper reports an initial attempt to establish the interrater reliability of the trigger point examination that failed, and a second study by the same examiners that included a training period and that successfully established interrater reliability in the diagnosis of the MTrP. The study also showed that the interrater reliability of different features varies, the LTR being the most difficult, and that the interrater reliability of the identification of MTrP features among different muscles also varies.

Adult↗

Diagnostic interview for genetic studies (DIGS): inter-rater and test-retest reliability of alcohol and drug diagnoses.

The semi-structured diagnostic interview for genetic studies (DIGS) was developed to assess major mood and psychotic disorders and their spectrum manifestations in genetic studies. Our research group developed a French version of the DIGS and tested its inter-rater and test-retest reliability in psychiatric patients. In this article, we present estimates of the reliability of substance use and antisocial personality disorders. High kappa coefficients for inter-rater reliability were found for drug and alcohol as well as antisocial personality diagnoses and slightly lower kappas for test-retest reliability. Combined with evidence of the reliability of major mood and psychotic disorders, these findings support the suitability of the DIGS for studies of familial aggregation and comorbidity of psychiatric disorders including substance use and antisocial personality disorders.

Adolescent↗

Assessing reliability of categorical substance use measures with latent class analysis.

This article illustrates the use of the latent class model to identify classes of individuals and to assess the psychometric reliability of categorical items. The latent class model is a categorical latent variable model used to identify homogeneous classes of respondents such that class membership accounts for item responses. The assessment of measurement reliability comes directly from the estimates of the model. Although not based on classical test theory, the reliability assessment procedures described here answer the same question-that is, how consistent or dependable is measurement? The goal is to identify reliable indicators of a characteristic by examining measurement error and the inter-relatedness of the items. Methods for estimating the reliability of individual items as well as sets of items are presented. These methods are illustrated with data on cigarette smoking from a national sample of adolescents. By using the procedures described here, researchers are able to determine: (1). which classes of people are measured well and which are not; (2). which items perform well and which do not; and (3). whether items need to be altered or added in order to measure and identify particular classes better.

Adolescent↗

The Alcohol Use Disorder and Associated Disabilities Interview Schedule-IV (AUDADIS-IV): reliability of alcohol consumption, tobacco use, family history of depression and psychiatric diagnostic modules in a general population sample.

BACKGROUND: the purpose of this study was to assess the test-retest reliability of newly introduced or revised modules of the Alcohol Use Disorder and Associated Disabilities Interview Schedule-IV (AUDADIS-IV), including alcohol consumption, tobacco use, family history of depression, and selected DSM-IV axis I and II psychiatric disorders. METHODS: kappa and intraclass correlation coefficients were calculated for the AUDADIS-IV modules using a test-retest design among a total of 2657 respondents, in subsets of approximately 400, randomly drawn from the National Epidemiologic Survey on Alcohol and Related Conditions (NESARC). RESULTS: reliabilities for alcohol consumption, tobacco use and family history of major depression measures were good to excellent, while reliabilities for selected DSM-IV axis I and II disorders were fair to good. The reliabilities of dimensional symptom scales of DSM-IV axis I and axis II disorders exceeded those of their dichotomous diagnostic counterparts and were generally in the good to excellent range. CONCLUSIONS: the high reliability of alcohol consumption, tobacco use, family history of depression and psychiatric disorder modules found in this study suggests that the AUDADIS-IV can be a useful tool in various research settings, particularly in studies of the general population, the target population for which it was designed.

Adolescent↗

The alcohol use disorder and associated disabilities interview schedule (AUDADIS): reliability of alcohol and drug modules in a clinical sample.

The alcohol use disorder and associated disabilities interview schedule (AUDADIS), was designed for use in the general population, and was previously shown to have good reliability in a sample of household residents. However, measurement problems are different in clinical samples. Thus, a test-retest study was conducted of the AUDADIS in a clinical sample of 296 substance-using patients from substance- and psychiatrically-identified treatment settings. Reliability for current drug-specific AUDADIS dependence diagnoses was good to excellent for high-prevalence as well as low-prevalence drug categories. Reliability for abuse diagnoses was not as good, although this was due to the hierarchical nature of the abuse diagnosis itself, rather than its defining criteria. Demographic and other factors were investigated for their potential effects on the reliability of alcohol and cocaine diagnoses; low severity was the only consistent predictor of unreliability for both of these categories. Reliability of consumption variables was generally good, with a few notable exceptions. Results suggest that the AUDADIS can be used in research comparing treated to community samples of individuals with alcohol and drug diagnoses.

Adolescent↗

[Linear factor analytic models for reliability analysis of composite variables].

BACKGROUND: Confirmatory factor analysis allows testing whether a composite variable may be considered as a reliable measure of a psychological attribute which is defined within a population. METHODS: Models for parallel, tau equivalent and congeneric measurements are presented along with their reliability coefficients. RESULTS: When the variables are not tau equivalent, the averaged inter-item correlation should be preferred to coefficient alpha, which is not an estimator of the reliability of the corresponding data. As a rule, interpretation of a coefficient as a reliability coefficient requires that the corresponding structural model of the composite be known. Beyond unidimensionality, simultaneous analysis of several congeneric variables through the use of cross-sectional or longitudinal hierarchical models entails fragmenting the theoretical variables. CONCLUSION: Interpreting a composite variable whose theoretical structure is corroborated by a hierarchical model may raise some difficulties because of its multidimensionality. Reliability formulae which account for this fragmentation are detailed.

Factor Analysis, Statistical↗

Assessment of medical students' communicative behaviour and attitudes: estimating the reliability of the use of the Amsterdam attitudes and communication scale through generalisability coefficients.

It is widely accepted that adequate attitudes and communicative skills are among the essential objectives in medical education. The Amsterdam attitude and communication scale (AACS) was developed to assess communicative skills and professional attitudes of medical students. More specifically, it was designed to evaluate the clinical behaviour of clerks to establish their suitability for the medical profession. The AACS covers nine dimensions. Moreover, an overall judgement of the student's performance is included. The present paper reports first results on the reliability of the use of the AACS. Data were collected in the course of an AACS training programme for future judges: senior medical and nursing staff members (N=98). Participants judged three videotapes of clerks interviewing patients at the bedside. For the assessment of videotapes, the first four dimensions of the AACS and the overall judgement are relevant. By applying Generalisability Theory to the training data we can forecast the reliability of the AACS in practice and gain insight in the number of raters that is needed to achieve sufficient reliability in clinical practice. If clerk behaviour is rated by six judges, summative assessment is sufficiently precise, i.e. <0.25. When using the full AACS, covering 10 items, the same number of judges is needed. Scores on individual AACS items are not sufficiently reliable. In conclusion, the results indicate that students' behaviour can be evaluated in a reliable manner using the AACS as long as enough judges and items are involved.

Attitude of Health Personnel↗

Development of a reliable and valid questionnaire to test the prostate cancer knowledge of men with the disease.

Several prostate cancer knowledge questionnaires exist but none have demonstrated both reliability and validity when used by men with the disease. This study aims to develop a reliable and valid knowledge questionnaire for men with prostate cancer. After developing a 40-item Prostate Cancer Knowledge Questionnaire (PCKQ-40) in phase I, it was piloted in phase II with 391 medical students. This resulted in the reduction of the tool by removing items of poor discriminatory value. A balance between true and false items and content domains was also maintained in the reduced scale. The PCKQ-12 had a moderate internal consistency. Phase III of the study assessed the reliability and construct validity of the tool by measuring the prostate cancer knowledge levels of men with prostate cancer (n=28), men with other cancers (n=10) and men without cancer (n=84). Men with prostate cancer achieved significantly higher PCKQ-12 scores compared with men without cancer, supporting the construct validity of the tool. The tool's reliability was also confirmed with a moderate internal consistency. This study provides some evidence for the reliability and validity of the PCKQ-12 and supports its use with men with prostate cancer in further research.

Adolescent↗

The Prosthetic Upper Extremity Functional Index: development and reliability testing of a new functional status questionnaire for children who use upper extremity prostheses.

The Prosthetic Upper Extremity Functional Index (PUFI) was developed by the authors' clinical research group to evaluate the extent to which a child actually uses a prosthetic limb for daily activities, the comparative ease of task performance with and without the prosthesis, and its perceived usefulness. The PUFI's test-retest and interrater reliability were evaluated with 24 children. Intraclass coefficients (ICCs) were calculated for each of four subscales of the PUFI--specifically, method of performance, ease of prosthetic use, usefulness of the prosthesis, and ease of performance without the prosthesis. The ICCs were greater than 0.65, indicating good test-retest reliability for the older-child respondents (n = 10) and fair to good reliability (ICCs, 0.40 to 0.84) for the parent respondents overall (n= 21). Interrater (child-parent) reliability was lower, with ICCs from 0.30 to 0.77. This finding was not unexpected, since a child and parent may rate in the context of different functional environments. The prosthesis was used 53% of the time by older children and more than 75% of the time by younger children. The results provide evidence that the PUFI has good test-retest reliability overall as a measure of a child's ability to perform upper extremity activities with a prosthesis.

Activities of Daily Living↗

Manual muscle strength testing: intraobserver and interobserver reliabilities for the intrinsic muscles of the hand.

The reliability of manual muscle strength testing of the intrinsic muscles of the hand is reported. The muscle strengths of 28 patients who had neuropathies of the ulnar nerve or the ulnar and median nerves were graded by two physiotherapists to determine intraobserver and interobserver reliabilities. Muscle strength was graded using the numeric scale developed by the Medical Research Council (grades 0 to 5). Reliabilities were established for nine muscles or muscle groups. Intraobserver reliabilities ranged from 0.71 to 0.96 and interobserver reliabilities from 0.72 to 0.93. It is difficult to isolate, and hence grade, most of the intrinsic muscles of the hand. Therefore, it is suggested that specific movements be tested and graded when assessing and evaluating muscle or nerve function.

Adult↗

The case for comprehensive quality indicator reliability assessment.

To demonstrate the importance of evaluating overall quality indicator reliability, in addition to component or variable level reliability, a comparison of interrater agreement on four chart-abstracted pneumonia-related processes of care was conducted. The hospital medical records of 356 Medicare patients' recent discharges for pneumonia were independently abstracted by different abstractors. Kappa, prevalence and bias-adjusted kappa, P(pos), P(neg), and the Bias Index were used to assess reliability of composite quality indicators and their components. The adjusted kappas for the data elements used to determine eligibility to receive as well as to derive the pneumonia-related processes of care ranged from 0.68 to 1.0. The adjusted kappa associated with overall eligibility to receive the pneumonia-related processes of care was 0.63. The kappa statistics for determining if processes of care were provided ranged from 0.56 to 0.83 and increased to 0.65 and 0.85 upon adjustment for the prevalence effect. Kappas for the composite quality indicators were lower, but improved with adjustment for the prevalence effect. The composite quality indicator with the highest adjusted kappa value was oxygenation assessment (0.93); the composite quality indicator with the lowest adjusted kappa value was antibiotic administration within 8 hours of hospital arrival (0.74). This study establishes the reliability of pneumonia indicators and underscores the need for reliability assessment at the quality indicator level, as well as at the component level.

Humans↗

The structured interview for anorexic and bulimic disorders for DSM-IV and ICD-10 (SIAB-EX): reliability and validity.

OBJECTIVE: For reliable and valid assessment and diagnostic categorization of eating disorders, self-report measures have considerable limitations. A semi-structured interview - the SIAB-EX - was developed for a more reliable and valid assessment of eating disorders. METHODS: One study (videotapes of 31 inpatients, seven raters) was made to establish inter-rater reliability; in another study with 80 patients the SIAB-EX was compared to another semi-structured interview designed for comparable purposes (EDE). In a third study data was obtained on 377 eating disorder patients seeking treatment to explore discriminant and convergent (construct) validity using the following self-rating scales: EDI, TFEQ, SCL-90, BDI, and the PERI Demoralization Scale. RESULTS: Inter-rater reliability of dichotomous ratings was good with mean kappa values of.81 (current) and.85 (past). Comparison of the SIAB-EX with the EDE generally showed quite similar results and higher intercorrelation of the total scale (.77). There are, however, a number of differences between the two scales, which are discussed in detail. Construct validity of the SIAB-EX was established. CONCLUSION: Inter-rater reliability was good. Convergent and discriminant (construct) validity of the SIAB-EX was demonstrated. The constructs assessed by the SIAB and its subscales and items are discussed in the context of their correlations with other well-known scales.

Adaptation, Psychological↗

Reliability and validity of a multi-pad pressure evaluator for pressure ulcer management.

It is often helpful to assess the pressures exerted upon the bony prominences when monitoring the likely outcome of pressure ulcer prevention or treatment. However, in the clinical setting, hard pressure sensors may damage the skin and operational difficulties may influence their reliability and validity. The authors have developed a multi-pad pressure sensor and tested its clinical reliability and validity. The inter-rater and intra-rater reliability were calculated using the coefficient of variation data from 10 patients. After a comparison analysis, the multi-pad was more reliable than a single-pad type pressure sensor. A validation test was conducted in 79 elderly patients. The mean interface pressures recorded among patients who had erythema or stage I pressure ulcers at the sacrum were significantly higher than were the contact pressures measured in patients with no pressure damage. The pressure sensor exhibited satisfactory clinical reliability and validity. Furthermore, it may be that for Japanese elderly patients the maximum pressure that can be tolerated by the tissues around the sacrum may be 40-50 mmHg.

Aged↗

Reliability of Chalmers' scale to assess quality in meta-analyses on pharmacological treatments for osteoporosis.

PURPOSE: This study estimates the inter-rater and test-retest reliability of Chalmers' quality score scale in the context of bone mass loss and fracture rate in postmenopausal women. METHODS: An exhaustive literature search was performed on Medline to locate clinical trials studying the effect of medication use on bone mass loss and fracture rate in postmenopausal women. Twenty articles were randomly selected and four raters independently assessed the quality of each article with Chalmers' scale. Among the 20 articles, 10 were blinded on authors' names, journal, year of publication and source of funding. Raters were also asked to assess all 20 articles one more time, two months after the first evaluation. Intraclass (ICC) and test-retest correlation coefficients were calculated. RESULTS: The overall inter-rater ICC was 0.66 [0.55, 0.79](95%). The overall test-retest reliability of Chalmers' scale was 0.81 [0.67, 0. 98](95%). When ratings were stratified according to articles' blinding status, blinded assessments generated a smaller inter-rater ICC than non-blinded assessments: 0.30 [0.17, 0.53](95%) vs. 0.80 [0. 71, 0.90](95%). In addition, analyzing sub-scales separately generated different estimates of reliability. CONCLUSIONS: This study shows that the reliability of the quality scale developed by Chalmers substantially varies between sub-scales, and is highly dependent on articles' blinding status. The possibility of bias in rating non-blinded articles can not be ruled out. The reliability of the scale can also be dependent on the outcome studied.

Adult↗