Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Reliability of the 7-point subjective global assessment scale in assessing nutritional status of dialysis patients.

Subjective global assessment (SGA) is a method to score nutritional status in a standardized way. The original 3-point scale has been replaced by a 7-point scale. The reliability of the latter scale has never been tested. We therefore assessed inter-observer and intra-observer reliability. Furthermore, we examined the relationship of SGA with other objective nutritional parameters. In 13 hemodialysis and 9 peritoneal dialysis patients, two nurses assessed SGA. They re-examined the same patients two weeks later. Anthropometric measurements and blood samples were taken at the first assessment. According to SGA, 2 patients (9%) were classified as severely malnourished, 6 (27%) as mildly malnourished, and 14 (64%) as well nourished. The 7-point SGA scale showed fair inter-observer reliability [intraclass correlation (ICC) = 0.72] and good intra-observer reliability (ICC = 0.88). A strong correlation was present between the 7-point SGA scale and body mass index (BMI) (r = 0.79, p < 0.001), % fat (r = 0.77, p < 0.001), and mid arm circumference (r = 0.71, p < 0.001). Lower correlations were found with mid arm muscle circumference and serum albumin. With respect to biochemical markers, the strongest relationship was found with prealbumin (r = 0.60, p = 0.004). We conclude that the 7-point SGA scale is a valid and reliable tool to assess nutritional status among end-stage renal disease patients. We suggest that one observer or a select group of observers perform the assessments to gain maximum benefit from the reliability of the SGA instrument.

Anthropometry↗

Reliability of the quadriceps angle measurement.

The quadriceps angle (Q-angle) is used to determine patellofemoral alignment. Although this measurement has been used to evaluate and treat patellofemoral joint pathology, few studies have examined its reliability. This study evaluated the interobserver and intraobserver reliability of the Q-angle measurement. To investigate the interobserver reliability of the Q-angle, 25 individuals of varying levels of training served as observers and participants as each measured the other 24 participants. To investigate the intraobserver reliability of the Q-angle, 3 of the observers measured 13 of the participants an additional 2 times. Additionally, clinically derived Q-angle measurements were compared with radiographically derived measurements. The reliability analysis was performed using intraclass correlation coefficients. For interobserver measurements, the intraclass correlation coefficients ranged from 0.17-0.29 for the four variables evaluated (right and left, extension and flexion). For intraobserver measurements, the intraclass correlation coefficients ranged from 0.14-0.37. The average intraclass correlation coefficient between the clinically and radiographically derived measurements ranged from 0.13-0.32. This study demonstrates poor interobserver and intraobserver reliability of Q-angle measurement and poor correlation between clinically and radiographically derived Q-angles.

Adult↗

[The role of observer for the reliability of Dutch version of the Discomfort Scale-Dementia of Alzheimer Type (DS-DAT)].

The role of the observer in the reliability of the Dutch Discomfort Scale-Dementia of Alzheimer Type (DS-DAT). The Discomfort Scale of Dementia of the Alzheimer Type (DS-DAT) is an instrument to assess discomfort in severely demented patients. No data on the reliability of assessment using a Dutch translation were available. In this paper, we analyse the role of the observer in the reliability of rating. This is of importance for studies in which many physicians perform multiple assessments. Twenty-eight nursing home physicians in training rated the DS-DAT in five nursing home patients with dementia presented on videotape. This was repeated after five months. All the physicians were previously trained in the use of the instrument. The results were statistically analysed using random effects analysis of variance. The Intra-class Correlation Coefficient (ICC) was 0.74 for inter-observer reliability and 0.97 for intra-observer reliability. Variance between subsequent assessments was small, but physicians appeared to differ somewhat among themselves in the way they rated the videotaped patients. A future complete reliability assessment of rating the DS-DAT in clinical practice would involve patient variation as well, scoring patients in clinical practice.

Adult↗

Health survey reliability in an active duty population.

We examined the reliability of data collected from the Health Enrollment Assessment Review (HEAR) survey, a self-report instrument administered by the Department of Defense styled after the Behavioral Risk Factor Surveillance System (BRFSS). Survey responses from a convenience sample of active duty service members who completed a HEAR survey on two occasions were examined. We measured test-retest reliability by comparing individuals' responses to several of the survey items and compared HEAR reliability patterns with the reliability patterns of similar studies conducted with the BRFSS. The majority of estimates reflected fair to excellent reliability. We found substantial agreement between the results of this investigation and similar BRFSS studies. Our findings support the reliability of responses from the HEAR survey for active duty military members. Our results were generally consistent with those of other studies despite differences in survey administration, respondent characteristics, and privacy guarantees.

Adult↗

Reliability of the Severin classification in the assessment of developmental dysplasia of the hip.

This study was undertaken to investigate the reliability of Severin's classification at various ages, and to determine whether the reliability is improved by careful measurement rather than subjective assessment. The radiographs taken at ages 6, 10 and 16 years of 20 randomly selected patients treated for developmental dysplasia of the hip were graded by six observers on two separate occasions using parameters measured according to Severin's criteria. In addition, four of these observers regraded the same radiographs using subjective assessment without measurements being made on two other separate occasions. Agreements between and within observers were evaluated using the weighted Cohen's kappa statistic for each age group. Intraobserver reliability was good, there being a close association between the measured and the subjectively observed ratings. This accords with the subjective nature of this classification. The interobserver reliability was found to be poor although it improved when direct measurements were made. Overall agreement between observers improves as patient age increases. It is concluded that comparisons between different observers using the Severin classification system are not reliable. However, a single investigator comparing treatment modalities in the same study allowing for individual bias in assessing deformity and subluxation would produce reliable results.

Acetabulum↗

[Reliability of diagnostic instruments: investigating the psychiatric DSM-III checklist applied to community samples].

This study focused on the reliability of the DSM-III inventory of psychiatric symptoms in representative general population samples in three Brazilian cities. Reliability was assessed through two different designs: inter-rater reliability and internal consistency. Diagnosis of lifetime (k = 0.46) and same-year generalized anxiety (k = 1.00), lifetime depression (k = 0.77), and lifetime alcohol abuse and dependence (k = 1.00) was consistently reliable in the two methods. Lifetime diagnosis of agoraphobia (k = 1.00), simple phobia (k = 0.77), non-schizophrenic psychosis (k = 1.00), and psychological factors affecting physical health (1.00) showed excellent reliability as measured by the kappa coefficient. The main reliability problem in general population studies is the low prevalence of certain diagnoses, resulting in small variability in positive answers and hindering kappa estimation. Therefore it was only possible to examine 11 of 39 diagnoses in the inventory. We recommend test and re-test methods and a short time interval between interviews to decrease the errors due to such variations.

Humans↗

Reliability of biomechanical variables during wheelchair ergometry testing.

Wheelchair ergometer testing is used to characterize wheelchair propulsion mechanics. The reliability of kinematic and kinetic measures has not been investigated for wheelchair ergometer testing. In this study, test-retest reliability of biomechanical measurements on a wheelchair ergometer was determined during a submaximal endurance test. Ten nondisabled subjects (seven male, three female), inexperienced in wheelchair propulsion, completed three separate submaximal fatigue tests. An instrumented wheelchair ergometer was used to measure handrim kinetics while three-dimensional kinematic data were collected. Analysis of variance was used to determine if measurement differences existed across the tests. Intraclass correlation coefficients (ICC) were calculated to determine the reliability of the measurements. The majority of handrim and temporal variables were found to be reliable. Joint kinematic variables were less reliable, especially those involving wrist movements in the fatigued state. It was concluded that most biomechanical variables obtained during wheelchair ergometry were reliable.

Adult↗

Reliability of the mini nutritional assessment (MNA) in institutionalized elderly people.

OBJECTIVES: To measure the reliability of the Mini Nutritional Assessment (MNA) in institutionalized elderly people. DESIGN: 12 day interobserver reliability study. PARTICIPANTS AND SETTING: All subjects admitted to two long term geriatric units in Mataró (Barcelona, Spain) over 4 months during 1996 (n=67). MEASUREMENTS: in each center, different trained nurses independently administered the MNA on two separate occasions. RESULTS: Mean (standard deviation) scores for the two assessments of the MNA were 20.8 (5.4) and 21.3 (4.6) respectively. Internal consistency, estimated by the Cronbach's Alpha, were 0.83 and 0.74 for the first and second assessment respectively. Test-retest reliability, according to the intraclass correlation coefficient (ICC), was 0.89 for the total MNA score and higher than 0.89 for its continuous items. According to the Kappa index, test-retest reliability for the stratified total MNA was substantial (0.78); for the 18 ordinal or nominal items of the MNA it was 'almost perfect' or 'substantial' in 12 items, 5 were 'moderate' to 'fair' and in I item it was 'slight'. Subjective health evaluation, the number of glasses of liquids per day, and brachial circumference (this former with an ICC=0.91) were the items with the lowest Kappa indices. CONCLUSION: The MNA test has good levels of reliability, according to its internal consistency and its test-retest reproducibility. Some improvements can still be introduced by refining the categorization and content of some items with low reliability.

Aged↗

Reliability in rehabilitation measurement.

The article examines reliability as a key component of the psychometric properties of tests and inventories. What is reliability, types of reliability, why is reliability important, and other reliability issues are addressed in a manner that will help Work readers to better understand and apply psychometric considerations in the context of rehabilitation research. The authors contend that reliability is an essential condition for the psychometric soundness of measurement instruments that are used in rehabilitation research and practice.

Journal Article↗

The reliability of determining effort level of lifting and carrying in a functional capacity evaluation.

OBJECTIVES: To establish inter- and intra-rater reliability of observations in a functional capacity evaluation. BACKGROUND: Functional capacity evaluations are used to assess a person's functional capacity as it relates to work. Lifting and carrying are important aspects of a functional capacity evaluation. An evaluator determines the patient's levels of effort through standardized observations. Questions remain with regards to the reliability of these observations. METHODS: Four healthy subjects were videotaped while performing two lifts and four carries with progressive loads. The videotape was scrambled randomly and viewed twice by 3 physical therapists and 2 occupational therapists. The evaluators determined the amount of effort it required (light, medium, heavy, and maximum). The inter- and intra-rater reliability of the observations was expressed by means of percentage agreement. RESULTS: Inter-rater reliability ranged 87-96%, intra-rater reliability ranged 93-97%. CONCLUSION: The results indicate that by means of standardized observations, therapists can reliably determine effort level during lifting and carrying in healthy subjects, and thus affirm the findings of other studies of similar design.

Adult↗

Assessment of the reliability of protein-protein interactions and protein function prediction.

As more and more high-throughput protein-protein interaction data are collected, the task of estimating the reliability of different data sets becomes increasingly important. In this paper, we present our study of two groups of protein-protein interaction data, the physical interaction data and the protein complex data, and estimate the reliability of these data sets using three different measurements: (1) the distribution of gene expression correlation coefficients, (2) the reliability based on gene expression correlation coefficients, and (3) the accuracy of protein function predictions. We develop a maximum likelihood method to estimate the reliability of protein interaction data sets according to the distribution of correlation coefficients of gene expression profiles of putative interacting protein pairs. The results of the three measurements are consistent with each other. The MIPS protein complex data have the highest mean gene expression correlation coefficients (0.256) and the highest accuracy in predicting protein functions (70% sensitivity and specificity), while Ito's Yeast two-hybrid data have the lowest mean (0.041) and the lowest accuracy (15% sensitivity and specificity). Uetz's data are more reliable than Ito's data in all three measurements, and the TAP protein complex data are more reliable than the HMS-PCI data in all three measurements as well. The complex data sets generally perform better in function predictions than do the physical interaction data sets. Proteins in complexes are shown to be more highly correlated in gene expression. The results confirm that the components of a protein complex can be assigned to functions that the complex carries out within a cell. There are three interaction data sets different from the above two groups: the genetic interaction data, the in-silico data and the syn-express data. Their capability of predicting protein functions generally falls between that of the Y2H data and that of the MIPS protein complex data. The supplementary information is available at the following Web site: http://www-hto.usc.edu/-msms/AssessInteraction/.

Computational Biology↗

Reliability of the PEDro scale for rating quality of randomized controlled trials.

BACKGROUND AND PURPOSE: Assessment of the quality of randomized controlled trials (RCTs) is common practice in systematic reviews. However, the reliability of data obtained with most quality assessment scales has not been established. This report describes 2 studies designed to investigate the reliability of data obtained with the Physiotherapy Evidence Database (PEDro) scale developed to rate the quality of RCTs evaluating physical therapist interventions. METHOD: In the first study, 11 raters independently rated 25 RCTs randomly selected from the PEDro database. In the second study, 2 raters rated 120 RCTs randomly selected from the PEDro database, and disagreements were resolved by a third rater; this generated a set of individual rater and consensus ratings. The process was repeated by independent raters to create a second set of individual and consensus ratings. Reliability of ratings of PEDro scale items was calculated using multirater kappas, and reliability of the total (summed) score was calculated using intraclass correlation coefficients (ICC [1,1]). RESULTS: The kappa value for each of the 11 items ranged from.36 to.80 for individual assessors and from.50 to.79 for consensus ratings generated by groups of 2 or 3 raters. The ICC for the total score was.56 (95% confidence interval=.47-.65) for ratings by individuals, and the ICC for consensus ratings was.68 (95% confidence interval=.57-.76). DISCUSSION AND CONCLUSION: The reliability of ratings of PEDro scale items varied from "fair" to "substantial," and the reliability of the total PEDro score was "fair" to "good."

Databases, Bibliographic↗

Letournel classification for acetabular fractures. Assessment of interobserver and intraobserver reliability.

BACKGROUND: A fracture classification system enables communication among surgeons and provides guidelines for treatment as well as some estimate of prognosis. Thus, the system should be anatomically meaningful and reliable. The purpose of this study was to assess the interobserver and intraobserver reliability of Letournel's acetabular fracture classification and the effect of computed tomography on its reliability. METHODS: Plain radiographs (anteroposterior and Judet views) and axial computed tomography scans were randomly chosen from an acetabular fracture database, with at least five cases of each fracture type and eight of the most common types. The study involved three groups of three orthopaedic surgeons: (1) surgeons who had studied under Letournel, (2) surgeons who specialized in acetabular fracture surgery, and (3) general trauma surgeons. Each observer read the radiographs twice, and at each session the fractures were classified first on the basis of the radiographs only and then in combination with the computed tomography scan. Observer agreement was then assessed with the unweighted kappa coefficient (kappa). We also calculated the frequency with which the observers agreed with the diagnosis made intraoperatively by the treating orthopaedic surgeon. RESULTS: The interobserver reliability without and with computed tomography during the first session was 0.70 and 0.74, respectively, for group 1, 0.71 and 0.69 for group 2, and 0.51 and 0.51 for group 3. The results of the second session were similar. When the two sessions were compared, intraobserver reliability without and with computed tomography was 0.80 and 0.83 for group 1, 0.80 and 0.80 for group 2, and 0.64 and 0.69 for group 3. The overall agreement of the radiographic observation with the fracture pattern observed at surgery was 74%. CONCLUSIONS: Letournel's acetabular classification with use of plain radiographs with or without supplemental computed tomography scans has substantial reliability (kappa > 0.7) when used by surgeons who have been taught how to interpret the images or by those who treat acetabular fractures on a regular basis. The value of computed tomography scans in the evaluation of acetabular fractures has been well established for the identification of loose bodies and articular impaction; however, they do not appear to be essential for the classification of acetabular fractures.

Acetabulum↗

Reliability and validity of a histologic score as a marker for skin cancer chemoprevention studies.

OBJECTIVE: To develop a reliable and valid scoring system for grading skin biopsies from actinic keratosis (AK) and sun-damaged skin for use in evaluating the efficacy of skin cancer chemopreventive agents. STUDY DESIGN: A panel of dermatopathologists developed histologic criteria and diagnostic definitions for the progression of lesions from early AK to AK. The criteria were then applied to a sample of 335 histologic slides from an ongoing chemoprevention study. A 10% sample of 35 slides was reread in order to assess intrarater reliability. RESULTS: Six of the 7 criteria demonstrated high reliability (> 85%). The total histologic score, calculated using the 6 criteria, was found to significantly differentiate between (blinded) biopsy location (normal, pre-AK, AK and adjacent to squamous cell carcinoma) and histologic diagnosis (normal, pre- or early AK, AK and squamous cell carcinoma). CONCLUSION: The total histologic score, having demonstrated reliability on repeated readings and validity in its association with biopsy location and histologic score, is a reliable and valid end point for judging the efficacy of agents in skin cancer chemoprevention studies. Additional interrater reliability tests utilizing larger test sets and a rigorous statistical design should be undertaken to establish its portability.

Biopsy↗

Intrarater and Interrater Reliability of the Beighton and Horan Joint Mobility Index.

OBJECTIVE: Clinicians may benefit from using a joint mobility index to screen for individuals on the high end of the spectrum of joint laxity (ie, those with generalized joint laxity), which may be associated with musculoskeletal complaints. Reliability of the Beighton and Horan Joint Mobility Index (BHJMI) has not been reported in the literature. Our purpose was to determine intrarater and interrater reliability of (1) composite BHJMI scores (the overall score from 0 to 9), and (2) categorized scores, the BHJMI scores in 3 categories (0 to 2, 3 to 4, and 5 to 9) DESIGN AND SETTING: This was an intrarater and interrater reliability study. Data were collected in an academic physical therapy department and in a high school. SUBJECTS: Forty-two (intrarater) and 36 (interrater) female volunteers, aged 15 to 45 years. MEASUREMENTS: Subjects were screened using the BHJMI. Percentage agreement and the Spearman rho were used to analyze BHJMI composite and category scores. RESULTS: The percentage agreement and the Spearman rho for intrarater and interrater reliability of BHJMI composite scores were 69% and.86 and 51% and.87, respectively. The percentage agreement and the Spearman rho for intrarater and interrater reliability of the category scores were 81% and.81 and 89% and.75, respectively. CONCLUSIONS: Reliability of the BHJMI was good to excellent in screening for generalized joint laxity in females aged 15 to 45 years.

Journal Article↗

Reliability of Joint Position Sense and Force-Reproduction Measures During Internal and External Rotation of the Shoulder.

OBJECTIVE: To determine the reliability of 2 common measures of proprioception. DESIGN AND SETTING: Joint position sense (JPS) and force reproduction (FR) were measured in the dominant shoulder using internal-rotation (IR) and external-rotation (ER) target angles on 2 consecutive days. SUBJECTS: Thirty-one healthy subjects (age = 22.0 +/- 2.8 years, height = 171.8 +/- 9.2 cm, mass = 69.5 +/- 15.9 kg) who did not regularly compete in overhand sports volunteered to participate in the study. MEASUREMENTS: Error scores were calculated at 2 target angles by averaging the absolute difference of 3 trials of JPS and FR. Reliability was determined by comparing the error scores obtained on 2 consecutive days. RESULTS: The inclinometer was found to be a reliable instrument as both intertester (.999) and intratester (.999) intraclass correlation coefficients were high. The JPS and FR measurements were also found to be reliable, with intraclass correlation coefficients ranging from.978 to.984. No differences were observed between trials for either measure. CONCLUSIONS: The inclinometer was a reliable instrument and can provide an affordable and accurate measure of range of motion and JPS. Both JPS and FR were also reliable measures of proprioception in the shoulder. Further research is needed to identify the specific mechanism of proprioception during these tasks.

Journal Article↗

Number of stimuli as a reliability parameter in perimetry.

Catch trials test patient performance during automated, static perimetry, but their adequacy to estimate reliability is uncertain even though up to 10% of the test time is reserved for catch trials. The 308 visual fields (program G1, all 3 phases, Octopus 201) of 308 eyes of 308 glaucoma, suspected glaucoma, and normal subjects were studied. The 108 visual fields (mean sensitivity > 10 dB; corrected loss variance < 50 dB2) without false responses to catch trials were considered reliable. A multiple linear regression analysis of these 108 fields was performed and revealed the following result (r2 = 0.751): Number of stimuli = 480 + (40.short-term fluctuation) + (8.8.the square root of the index corrected loss variance) - (2.2.mean sensitivity). This equation was used to estimate the number of stimuli required of a reliable subject to complete an examination. Excess stimuli would thus be a sign of reduced reliability. The difference between the estimated and the actual number of stimuli was called the 'stimulus discrepancy'. In 169 fields with false-positive and 58 fields with false-negative responses, the false-positive and false-negative responses correlated with the 'stimulus discrepancy' (r = 0.19, P = 0.014; r = 0.29, P < 0.026, respectively). The number of stimuli depends not only on reliability but also on the software and hardware of the perimeter. 'Stimulus discrepancy' may be an additional useful perimetric reliability parameter which does not require extra testing time.

Adult↗

[Safety and reliability verification for manned spacecraft crew support facilities].

OBJECTIVE: To verify reliability and safety of the crew support facilities on board manned spacecraft. METHOD: Several comprehensive qualitative and quantitative projection verification technique, such as analysis, check, demonstration, tests and reliability assessment, were used. RESULT: Work-items that were specified in the reliability and safety program were realized. Assurance measures for safety and reliability critical items were available. FMEA and SHA and FTA were brought in all around. Safety and reliability on equipment levels and system levels was fully carried out. CONCLUSION: The facility safety and reliability achieved design specification and met the requirements of spaceflight tests for manned spacecraft.

Ergonomics↗